Elepsy region electroencephalogram recognition method and system based on interpretable codebook and multi-task pre-training

By adopting an interpretable codebook and multi-task pre-training method in electroencephalogram recognition in the tuberculosis area, the shortcomings of existing time series analysis models in terms of interpretability and universality are solved, cross-domain effective representation learning and classification tasks are achieved, and an in-depth understanding of the model decision-making process is provided.

CN120180246APending Publication Date: 2025-06-20SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510198083.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing time series analysis models have shortcomings in interpretability and versatility, which are difficult to provide a deep understanding of the model decision process, and are difficult to migrate to different domains and datasets.

Method used

EEG recognition method for epilepsy-induced area based on interpretable codebooks and multi-task pre-training is adopted. A globally perceived potential embedding is generated through a time series encoder, attribute tuples are extracted and mapped to discrete codebooks, and time series subsequences are reconstructed in combination with attribute tuples, reconstruction, quantification and untangling losses are optimized, and multi-task pre-training is completed.

Benefits of technology

It realizes the extraction of interpretable abstract shapes from time series data as markers, effectively representing learning and classification tasks across different fields and data sets, while providing an in-depth understanding of the model decision process, improving the interpretability and universality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180246A_ABST
    Figure CN120180246A_ABST
Patent Text Reader

Abstract

The invention relates to an epilepsy region electroencephalogram recognition method and system based on an interpretable codebook and multi-task pre-training, and the method comprises the following steps: (1) obtaining electroencephalogram signal data, and carrying out the preprocessing, including band-pass filtering, noise removal and standardization; (2) dividing the univariate time sequence into non-overlapping fragments; (3) extracting an attribute tuple from the potential embedding through an attribute decoder; (4) performing vector quantization on the potential embedding; (5) reconstructing a time sequence sub-sequence by combining a shape decoder with the attribute tuple; and (6) extracting potential space marks and codebook histogram features based on the pre-training model. The method has the advantages that interpretable abstract shapes are extracted from the time series data to serve as marks, effective representation learning and classification tasks across different fields and data sets are achieved, deep understanding of the model decision process is provided, and researchers and practitioners are assisted in better analyzing and processing the time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for epileptogenic zone EEG recognition based on an interpretable codebook and multi-task pre-training, belonging to the technical field of time series analysis. Background Art

[0002] 1. Processing Challenges in Epileptogenic Zone EEG Recognition Epileptogenic Zone Identification (EZI) is an important direction in epilepsy research, aiming to locate the epileptogenic zone of epilepsy by analyzing electroencephalogram (EEG) signals. The epileptogenic zone refers to the area of the cerebral cortex that causes epileptic seizures. Precise localization of these areas is crucial for the clinical diagnosis and treatment of epilepsy (such as surgical treatment). Traditionally, the identification of the epileptogenic zone mainly relies on the visual analysis of electroencephalogram (EEG) signals by clinicians. By analyzing the long-term EEG monitoring data of epilepsy patients, doctors can observe the activity characteristics of different brain regions before and after seizures, thereby inferring the location of the epileptogenic zone. However, this method has limitations, mainly relying on the experience and subjective judgment of operators, and it is difficult to accurately identify complex epileptic seizures. Especially in the EEG activities recorded by scalp or intracranial electrodes, characteristic waveforms such as epileptiform spikes often have a short duration, irregular shape, and are easily confused with artifacts, increasing the difficulty of visual identification and prone to misjudgment or unstable diagnostic accuracy.

[0003] 2. Limitations of Existing Deep Learning Models Black-box property: Although deep learning has made progress in time series analysis in recent years, such as models based on the Transformer structure (such as TST, TimeGPT-1, MOMENT, etc.) can be used for various downstream tasks after pre-training, most of these models are like black boxes and cannot provide human-understandable representations. They are usually pre-trained by predicting the next or masked timestamp, time window, or patch, lacking the concept of discrete tokens in natural language processing, and it is difficult to explain how the model learns and makes decisions from time series data.

[0004] Lack of interpretability and generality: Some models (such as TOTEM) use VQ-VAE to obtain a codebook and reconstruct time series, but the tokens in the codebook are only latent vector representations and lack physical meaning. In interpretable time series modeling, although there are shapelets that can convert time series into low-dimensional representations for classification, this kind of shape feature based on a predefined length has poor flexibility, and the shape features optimized for specific datasets are difficult to transfer to different fields and cannot meet the needs of cross-domain general modeling. Therefore, developing a time series classification method that is both highly interpretable and can achieve cross-domain generalization has become a research hotspot. Summary of the Invention

[0005] To overcome the defects of the prior art, the present invention aims to provide an epileptogenic zone EEG recognition system based on an interpretable codebook and multi-task pre-training, so as to solve the deficiencies of existing time series analysis models in terms of interpretability and generality. This model can extract interpretable abstract shapes from time series data as markers, achieve effective representation learning and classification tasks across different domains and datasets, and at the same time provide an in-depth understanding of the model's decision-making process to assist researchers and practitioners in better analyzing and processing time series data.

[0006] The present invention provides an epileptogenic zone EEG recognition method and system based on an interpretable codebook and multi-task pre-training. The technical solution of the present invention is as follows: An epileptogenic zone EEG recognition method based on an interpretable codebook and multi-task pre-training, characterized by comprising the following steps: (1) Obtain EEG signal data and perform preprocessing, including band-pass filtering, noise removal, and normalization; (2) Divide the univariate time series into non-overlapping segments, and generate globally perceptive latent embeddings through a Transformer-based time series encoder; (3) Extract attribute tuples from the latent embeddings through an attribute decoder, where the attribute tuples include abstract shape encodings, offsets, scale factors, relative start positions, and relative lengths; (4) Perform vector quantization on the latent embeddings, and map them to the abstract shape encodings in the discrete codebook, where the codebook contains multiple low-dimensional vectors to represent basic time series morphological features; (5) Reconstruct the time series subsequences through a shape decoder in combination with the attribute tuples, optimize the reconstruction loss, vector quantization loss, and shape disentanglement loss, and complete multi-task pre-training; (6) Extract latent space markers and codebook histogram features based on the pre-trained model, and input them into a classifier to achieve the classification and recognition of epileptogenic zone EEG signals.

[0007] In the step (2), the implementation of the time series encoder includes: Evenly divide the univariate time series into K non-overlapping segments of fixed length, generate initial feature embeddings through linear projection and positional embeddings, input them into a multi-layer Transformer encoder for global-to-local feature fusion, and output high-dimensional latent embeddings.

[0008] In the step (4), the specific steps of the vector quantization are as follows: Compress the continuous latent embeddings into a low-dimensional space, calculate their Euclidean distances from all vectors in the codebook, select the nearest discrete codebook vector as the quantization result, and optimize the codebook distribution through gradient clipping and entropy incentives.

[0009] In the step (5), the reconstruction loss includes: a global reconstruction loss, which constrains the reconstruction accuracy of the complete time series through the mean square error; a local reconstruction loss, which constrains the shape reconstruction accuracy through the mean square error of the subsequence.

[0010] The codebook histogram is generated by statistically counting the occurrence frequencies of each codebook index in the latent embedding, and is used to represent the abstract shape distribution characteristics of the time series.

[0011] The classifier adopts a support vector machine (SVM), and makes a multi-modal classification decision by combining the latent space markers and the codebook histogram features.

[0012] An epileptogenic zone EEG recognition system based on an interpretable codebook and multi-task pre-training, comprising: a data preprocessing module, which is used to filter, denoise and standardize the original EEG signals; a time series encoding module, which generates latent embeddings based on the Transformer architecture; an attribute decoding and vector quantization module, which extracts attribute tuples and maps them to a discrete codebook; a shape decoding and reconstruction module, which reconstructs the time series subsequences in combination with the attribute tuples; a multi-task pre-training module, which optimizes the reconstruction, quantization and disentanglement losses; a feature extraction and classification module, which realizes the identification of the epileptogenic zone based on the latent space markers and the codebook histogram.

[0013] The codebook of the vector quantization module has a dimension of 8 and a capacity of 1024, and the codebook vectors are discretely encoded through low-dimensional projection and Euclidean distance matching.

[0014] The number of Transformer layers of the time series encoding module is 8, the number of decoder layers is 2, the embedding dimension is 512, and the number of shards is 64.

[0015] The classification module supports zero-shot generalization and performs cross-domain classification on unseen data sets based on the pre-trained codebook.

[0016] The advantages of the present invention are: (1) Improve interpretability Abstract shape representation: By decomposing the time series into abstract shapes and their attributes, it provides an intuitive and interpretable way to understand the time series data. For example, in the visualization, the abstract shapes learned by the model can be clearly seen, and these shapes can capture the key patterns in the time series, providing researchers with in-depth insights into the internal structure of the data.

[0017] Discriminability of the codebook histogram: The codebook histogram has clear interpretability in classification tasks, capable of showing the differences in code frequencies among samples of different classes, thus helping to understand how the model makes classification decisions based on shape features. For example, when differentiating time series data of different gesture classes, it can be intuitively seen from the histogram which abstract shapes are more prominent in specific classes, providing a clear discriminative basis for classification.

[0018] (2)Enhanced generality Cross-dataset performance: Pre-trained and tested on multiple datasets from the UEA Multivariate Time Series Classification Archive, VQShape performs excellently in the cross-domain generalization ability of unseen samples. Compared with other models, it achieves comparable or even better performance on most datasets, demonstrating its ability to learn general feature representations applicable to different domains and datasets.

[0019] Zero-shot learning ability: VQShape and its codebook can generalize on datasets and domains not included in the pre-training process, demonstrating its potential in zero-shot learning scenarios and providing the possibility to handle new and unseen time series data.

[0020] (3)Efficient model performance Accuracy comparable to advanced models: In the multivariate time series classification task, VQShape shows comparable performance compared to current advanced baseline models (such as TimesNet, T-Rep, TS2Vec, MOMENT, etc.). In the test on 29 UEA datasets, its average accuracy, median accuracy and other metrics are at the same level as other excellent models, indicating that VQShape does not sacrifice model accuracy while ensuring interpretability.

[0021] Flexible model configuration and optimization: The configuration parameters of the model (such as codebook size, shape loss weight, etc.) can be adjusted according to specific tasks. By experimenting, the impact of different parameter settings on model performance is discovered, thus providing flexible optimization strategies for different application scenarios. For example, in the study of codebook size, it is found that a good balance is achieved between abstractness and expressiveness, which is suitable for pre-training on UEA datasets and can produce the best histogram representation, providing guidance for the optimization of the model in practical applications. Description of the drawings

[0022] Figure 1 is a schematic flowchart of the present invention.

[0023] Figure 2 is a visualization schematic diagram of the codebook.

[0024] Figure 3It is a comparison chart of the evaluation indicators of the results of the present invention, the results of the present invention, and other comparison methods. Detailed implementation manners

[0025] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, these embodiments are merely exemplary and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that modifications or substitutions can be made to the details and forms of the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications and substitutions fall within the protection scope of the present invention.

[0026] Refer to Figures 1 to 3 , the present invention relates to an epileptogenic zone EEG recognition method based on an interpretable codebook and multi-task pre-training, including the following steps: (1) Obtain EEG signal data and perform preprocessing, including band-pass filtering, noise removal, and normalization; (2) Divide the univariate time series into non-overlapping slices, and generate globally perceived latent embeddings through a Transformer-based time series encoder; (3) Extract attribute tuples from the latent embeddings through an attribute decoder, where the attribute tuples include abstract shape encoding, offset, scale factor, relative start position, and relative length; (4) Perform vector quantization on the latent embeddings and map them to the abstract shape encoding in a discrete codebook, where the codebook contains multiple low-dimensional vectors to represent basic temporal morphological features; (5) Reconstruct the time series subsequences through a shape decoder in combination with the attribute tuples, optimize the reconstruction loss, vector quantization loss, and shape disentanglement loss, and complete multi-task pre-training; (6) Extract latent space markers and codebook histogram features based on the pre-trained model, and input them into a classifier to realize the classification and recognition of epileptogenic zone EEG signals.

[0027] In the step (2), the implementation of the time series encoder includes: Uniformly divide the univariate time series into K non-overlapping slices of fixed length, generate initial feature embeddings through linear projection and positional embedding, input them into a multi-layer Transformer encoder for global-to-local feature fusion, and output high-dimensional latent embeddings.

[0028] In the step (4), the specific steps of the vector quantization are: Compress the continuous latent embeddings into a low-dimensional space, calculate their Euclidean distances from all vectors in the codebook, select the nearest discrete codebook vector as the quantization result, and optimize the codebook distribution through gradient clipping and entropy incentive.

[0029] In the step (5) described above, the reconstruction loss includes: a global reconstruction loss that constrains the reconstruction accuracy of the complete time series through the mean square error; a local reconstruction loss that constrains the shape reconstruction accuracy through the mean square error of the subsequences.

[0030] The codebook histogram is generated by statistically counting the occurrence frequencies of each codebook index in the latent embedding, and is used to represent the abstract shape distribution characteristics of the time series.

[0031] The classifier uses a support vector machine (SVM) and jointly performs multi-modal classification decision-making with the latent space markers and the codebook histogram features.

[0032] The present invention can extract interpretable abstract shapes from time series data as markers, realize effective representation learning and classification tasks across different fields and datasets, and at the same time provide an in-depth understanding of the model decision-making process, assisting researchers and practitioners to better analyze and process time series data.

[0033] 1. Database Introduction The Bern_Barcelona database is provided by the Center for Neurology of the University of Bern and is mainly used for epilepsy research. This database contains electroencephalogram (EEG) data of 5 patients with long-term drug-tolerant temporal lobe epilepsy. The data is collected using the international 10-20 electrode system, and the reference electrodes are placed at Fz and Pz outside the skull. The sampling frequency is 512 Hz or 1024 Hz, depending on the number of leads (1024 Hz is used when the number of leads exceeds 64, and 512 Hz is used when it is less than 64). The sampling time of each EEG signal data is 20 seconds, with a total of 10,240 data points. The EEG signals in the database are processed by a fourth-order Butterworth filter, and the band-pass filtering range is 0.5 to 150 Hz, and forward and reverse filtering are adopted to reduce phase distortion.

[0034] In this database, experienced experts classify the signals into epileptogenic zone EEG signals and non-epileptogenic zone EEG signals, including 3,750 signals in the epileptogenic zone and 3,750 signals in the non-epileptogenic zone, that is, there are a total of 7,500 EEG signals, and each signal contains data of two channels. All EEG signals have been visually inspected before being officially included in the database, and signals with obvious measurement interference have been removed. The original data in the database is stored in.txt file format, and each file contains the EEG signal data samples of two channels recorded in the same time period. This database provides a reliable data source for epilepsy research and EEG signal identification and detection, and is particularly suitable for the localization of epileptic seizure areas and the training and verification of related detection algorithms.

[0035] 2. Shape-Level Representation In the time series classification task, the dataset is defined as , where each sample is a multivariate time series containing M variables, corresponding to class labels . To deeply analyze the time series characteristics, each multivariate sample can be decomposed into M independent univariate sequences, denoted as .

[0036] For any univariate time series (n is the total length), its subsequence is defined as a finite set of consecutive observations. Let represent the local sequence starting from position i with length k, satisfying the constraint condition . This sliding window-based interception method can effectively capture the local pattern features in the time series.

[0037] To achieve refined modeling, this paper proposes to use a combination of five attributes to describe the subsequence . The specific information is shown in Table 1. Set the minimum length threshold . Through the pre-trained Transformer model, the univariate time series X can be transformed into a structured feature set , where each feature tuple contains multi-dimensional information such as shape, position, and scale. In addition, the model also learns a reusable abstract shape encoding dictionary , which is applicable to datasets from different fields.

[0038]

[0039] VQShape model architecture Time series encoder: This part adopts a patch-based time series feature encoding method based on Transformer. This method transforms the original time series into a globally aware latent representation through three-stage processing. The time series is patched and each patch is encoded into an embedding vector using linear projection and positional embedding, and then the latent representations of these patches are generated through the Transformer encoder. These latent embeddings represent not only the information of a single patch, but also the information of all patches, helping the model capture global time series features. The time series encoder is expressed as: , specifically as follows: First, the univariate time series is evenly divided into K fixed-length, non-overlapping patches, and each patch can be regarded as a fixed-length time window. The number of patches K and the dimension are dynamically determined by the total length L of the sequence and the model parameters. Each patch carries the local features within a specific time window. Subsequently, each patch is mapped to In the n-dimensional space, an initial feature embedding is formed. To maintain temporal dependence, a trainable position encoding matrix is used to mark the positions of the slices, and the temporal order information is incorporated into the feature representation through vector superposition. The embedding matrix generated by this process contains both local feature details and encodes global position relationships.

[0040] Finally, the embedding matrix is input into a multi-layer Transformer encoder for feature optimization. Through the multi-head self-attention mechanism, the Transformer dynamically establishes feature associations between different slices, and the representation of each slice fuses the context information of other slices during the iterative update process. After N layers of processing, the Transformer outputs a higher-dimensional latent embedding , and this global-local feature fusion mechanism enables the final representation to retain both fine-grained temporal features and macroscopic pattern perception capabilities.

[0041] Attribute decoder: The attribute decoder receives the latent embedding and extracts the attribute tuple . Formally, performs: , where each decoding function is implemented by a multi-layer perceptron (MLP) with one hidden layer and a ReLU activation function. represents the attribute tuple before quantization. This decoder can restore the attribute information of the time series from the latent representation, providing a basis for subsequent quantization and reconstruction.

[0042] Codebook and vector quantization: This section proposes a method for modeling discrete latent spaces based on vector quantization (VQ). The core goal is to map continuous features into abstract shape representations through low-dimensional coding constraints. First, a discrete codebook is defined, where each vector represents a basic shape pattern in the latent space. The coding dimension is restricted to 8 to force the model to ignore detailed noise (such as high-frequency fluctuations or local perturbations) and focus on globally significant temporal morphological features (such as trends, periods, etc.).

[0043] The continuous latent vector output by the encoder is compressed into a low-dimensional space through linear projection, and then vector quantization is performed: ; this process replaces with the discrete vector that is the closest in Euclidean distance in the codebook., thus realizing the conversion from continuous features to discrete symbols. This operation improves the model performance through dual constraints: firstly, the low-dimensional projection forms an information bottleneck, forcing the features to only retain the abstract shapes shared among different samples (such as basic patterns like smooth rise and step decline); secondly, the discretization process naturally filters out redundant information, making the decoder tend to generate waveforms with regular structures during reconstruction.

[0044] In the reconstruction stage, the discrete code is remapped to the high-dimensional space as the decoder input. Since the quantization process only transmits the information of shape primitives, the signal output by the decoder will automatically align to the standard patterns defined in the codebook, thus suppressing the risk of overfitting. This method is inherited from the VQ-VAE framework, but further strengthens the feature abstraction ability through the strong constraints of low-dimensional coding.

[0045] Shape decoder: The core goal of the shape decoder is to restore the latent code into a time series signal aligned with the shape of the input subsequence. Given the original time series X, a subsequence is intercepted from the position with a length of . Its abstract shape code , offset and scale factor constitute the decoding input.

[0046] First, is normalized to a sequence with a fixed length through bilinear interpolation, eliminating the influence of the original length : The shape decoder generates a new sequence with the same length as the interpolated target subsequence based on the abstract shape code : This output is an abstract waveform without scale information and offset. Therefore, the final output needs to multiply by and add to restore this information, specifically expressed as:

[0047] This process rescales the normalized waveform output by the decoder to the amplitude range of the original data and adjusts it to the target length through the inverse interpolation transformation, ensuring that the reconstructed signal is aligned with in terms of shape, scale, and position.

[0048] Attribute coding and reconstruction: The quantized attribute tuple is projected linearly through a learnable Mapping to a Low-Dimensional Space: This projection is carried out through a linear layer in a neural network, so it is learnable and can be adaptively adjusted during training to better capture the underlying structure of the input data. Then, the time series decoder receives all of the and outputs the reconstructed time series based on this information . The goal of the decoder is to recover the original time series data from these encoded representations. Among them, represents different time steps.

[0049] 4. Pre-training Strategy Training Objective Reconstruction Loss: Time series reconstruction requires the model to minimize the difference between the original input and the reconstruction result when reconstructing the entire time series, making the time series output by the model as close as possible to the original input and extracting useful representations in the latent space. To establish a latent representation with physical meaning, the model constrains the reconstruction process through a dual supervision mechanism: Time Series Global Reconstruction Loss: Constrains the reconstruction quality of the complete time series through mean square error ; Subsequence Local Reconstruction Loss: Constrains the shape reconstruction accuracy of the subsequence ; Among them is the target reconstruction value of this subsequence, is the reconstructed signal, and the loss function calculates the average of the mean square errors of all subsequences, aiming to enable the model to accurately recover each subsequence in the time series, thereby improving the interpretability and accuracy of the model. In this way, the model can capture the important features in the time series data and convert them into an efficient representation that is convenient for subsequent tasks.

[0050] Vector Quantization Loss: This method defines a vector quantization objective function based on VQ-VAE for jointly training the encoder and the codebook C. In addition, this paper introduces an additional entropy term to encourage the diverse use of the codebook. Experiments show that this method can improve the pre-training stability and avoid the problem of codebook degradation. The complete codebook learning objective function is defined as:

[0051] Among them is the gradient clipping operator, is the entropy function of discrete variables. The first half aims to minimize the distance between the vector output by the encoder and the closest code C selected from the codebook. The distance metric function takes the encoded vector Convert the Euclidean distance to the category distribution for all vectors in the codebook, which represents the most likely code in the codebook selected by the encoder output. All vectors in The Euclidean distance of is transformed into a category distribution, which represents the Most likely which code in the codebook is selected by the encoder output.

[0052] The formula consists of three parts, and its physical meaning is as follows:

[0053] Shape disentanglement loss: Attribute tuple The optimization goal of is to accurately reconstruct the subsequence. It should be noted that since Defines the position of the target shape Therefore, its learning process is crucial for shape abstraction and codebook construction. Therefore, a diversity regularization term is introduced to force the latent space markers to capture shape-level information with different positions and scales. This regularization term is defined as:

[0054] where the coordinate transformation function is defined as:

[0055] and Provide with the position encoding of, Perform scale transformation on to control its distribution in the latent space.

[0056] Hyperparameter Sets the minimum allowable distance threshold in the transformation space for determining whether two samples are diverse enough. The overall pre-training objective is to minimize the weighted loss:

[0057] where is a hyperparameter that defines the weights of each part. In the experiment, is set, . The weighted combination of these loss terms works together to help the model balance different aspects of the objective during the learning process, so as to better decouple and capture diverse details when processing shape information, as follows.

[0058]

[0059] Model configuration and training details: In this paper, all input univariate time series are interpolated to length , and divide it into small pieces, and the dimension of each small piece is . The Transformer layer in the encoder and the decoder has 8 heads, the embedding dimension , and the size of the feed-forward layer is 2048. This paper adopts an asymmetric structure, the encoder has 8 layers, the decoder has 2 layers, the encoder is responsible for extracting information, and the decoder generates the final output. The codebook C contains codes, and the dimension of each code is . The subsequence and the decoded sequence have a length of . We set the minimum shape length . Under these settings, the model has 37.1 million parameters.

[0060] In the pre-training stage, we train the model using the AdamW optimizer, with weight decay , , , the gradient clipping value is 1.0, and the effective batch size is 2048. We adopt the cosine learning rate scheduler, with the initial learning rate of , the final learning rate of , and perform 1 cycle of linear warm-up. The pre-training dataset contains univariate time series extracted from the training sets of 29 datasets in the UEA multivariate time series classification archive, totaling 1,387,642 univariate time series. This paper trains on this dataset for 50 cycles, using bfloat-16 mixed precision.

[0061] 5. Downstream task representation Latent space tokens: Tokens are a way to represent input data, similar to the latent space feature maps in other VQ (vector quantization) methods (such as VQ-VAE and VQ-GAN). Tokens are a way to transform input data into a latent space representation. These tokens contain multiple information dimensions such as the quantized codebook vectors, the mean and standard deviation of the subsequence, the starting position of the subsequence, and the relative length, etc. For the input univariate time series x , the token representation form is defined as: ; where is the number of time series segments, represents the codebook vector dimension. This token representation performs multi-dimensional information fusion and is suitable for complex tasks, but its interpretability in classification tasks is relatively weak.

[0062] Codebook Histogram: Inspired by the Concept Bottleneck Models (CBMs) in computer vision, each latent vector in the codebook in this paper is regarded as a concept in time series data. The model determines the code index corresponding to each latent vector by minimizing the distance between the latent vector and a certain code : :

[0063] Once all latent vectors are mapped to their corresponding code indices, the histogram representation can be calculated. The histogram representation r is a vector, where each element represents the frequency of the corresponding code appearing in all latent vectors, defined as:

[0064] Each element here represents the number of times the index appears in the index list . The code histogram is a method of representing time series data by statistically counting the occurrence frequencies of each code index. It can provide a useful and interpretable representation of time series data, especially suitable for classification tasks because it can generate rule-like predictions that are easy to understand and interpret.

[0065] 6. Feature Extraction and Classification The trained time series classification model based on abstract shapes outputs two types of feature representations, namely latent space tokens and codebook histogram. The two feature representations are respectively fed into a classifier for processing to perform subsequent classification tasks. The following is a detailed description of this process. First, the time series feature extraction and classification model based on abstract shapes maps the input data to the latent space and generates two types of feature representations. The latent space tokens represent the discretized representation of the input data in the latent space, while the codebook histogram represents the distribution of each input sample in the codebook space. Both of these feature representations contain important information of the input data and have different forms of expression. The latent space tokens and the codebook histogram are saved separately and fed into the classifier for classification respectively. Each feature representation is processed by an independent classifier to perform a separate classification task. Specifically, the feature vectors generated by the latent space tokens and the codebook histogram respectively are used as inputs, and through the decision-making process of the classifier, the corresponding classification results are output.

[0066] In this way, we can not only capture the multi-level information of the input data by using the latent space markers and codebook histograms of the model, but also achieve more accurate classification through the powerful discrimination ability of the classifier. By combining multi-modal features with an SVM classifier, this method effectively improves the performance and robustness of the model in classification tasks.

[0067] The present invention also relates to an epileptogenic zone EEG recognition system based on an interpretable codebook and multi-task pre-training, including: A data preprocessing module for filtering, denoising, and normalizing the original EEG signals; A time series encoding module for generating latent embeddings based on the Transformer architecture; An attribute decoding and vector quantization module for extracting attribute tuples and mapping them to a discrete codebook; A shape decoding and reconstruction module for reconstructing time series subsequences in combination with attribute tuples; A multi-task pre-training module for optimizing reconstruction, quantization, and disentanglement losses; A feature extraction and classification module for realizing epileptogenic zone recognition based on latent space markers and codebook histograms.

[0068] The codebook of the vector quantization module has a dimension of 8, a codebook capacity of 1024, and the codebook vectors are discretely encoded through low-dimensional projection and Euclidean distance matching.

[0069] The Transformer of the time series encoding module has 8 layers, the decoder has 2 layers, the embedding dimension is 512, and the number of shards is 64.

[0070] The classification module supports zero-shot generalization and performs cross-domain classification on unseen data sets based on the pre-trained codebook.

[0071] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for identifying epileptogenic zones based on interpretable codebook and multi-task pre-training, characterized in that: The following steps are involved: (1) Obtain EEG signal data and perform preprocessing, including bandpass filtering, noise removal, and standardization; (2) Divide the univariate time series into non-overlapping segments and generate globally aware latent embeddings via a Transformer-based time series encoder; (3) extracting an attribute tuple from the latent embedding through an attribute decoder, wherein the attribute tuple includes an abstract shape code, an offset, a scale factor, a relative starting position, and a relative length; (4) vector quantizing the latent embedding and mapping it to an abstract shape code in a discrete codebook, wherein the codebook contains multiple low-dimensional vectors to represent the underlying temporal morphological features; (5) Reconstruct the time series subsequence by combining the shape decoder with the attribute tuple, optimize the reconstruction loss, vector quantization loss and shape disentanglement loss, and complete multi-task pre-training; (6) Based on the pre-trained model, latent space markers and codebook histogram features are extracted and input into the classifier to achieve classification and recognition of EEG signals in the epileptogenic zone.

2. The method for identifying epileptogenic zones based on an interpretable codebook and multi-task pre-training according to claim 1, characterized in that: In the step (2), the implementation of the time series encoder includes: The univariate time series is evenly divided into K non-overlapping slices of fixed length, and the initial feature embedding is generated by linear projection and position embedding. The multi-layer Transformer encoder is input for global to local feature fusion, and high-dimensional latent embedding is output.

3. The method for identifying epileptogenic zones based on interpretable codebook and multi-task pre-training according to claim 1, characterized in that: In the step (4), the specific steps of the vector quantization are: The continuous latent embedding is compressed into a low-dimensional space, its Euclidean distance to all vectors in the codebook is calculated, the nearest neighbor discrete codebook vector is selected as the quantization result, and the codebook distribution is optimized through gradient truncation and entropy excitation.

4. The method for identifying epileptogenic zones based on interpretable codebook and multi-task pre-training according to claim 1, characterized in that: In the step (5), the reconstruction loss includes: a global reconstruction loss, which constrains the reconstruction accuracy of the complete time series through a mean square error; The local reconstruction loss constrains the shape reconstruction accuracy through the mean square error of subsequences.

5. The method for identifying epileptogenic zones based on interpretable codebook and multi-task pre-training according to claim 1, characterized in that: The codebook histogram is generated by counting the occurrence frequency of each codebook index in the potential embedding, and is used to represent the abstract shape distribution characteristics of the time series.

6. The method for identifying epileptogenic zones based on interpretable codebook and multi-task pre-training according to claim 1, characterized in that: The classifier adopts support vector machine (SVM) and combines latent space labeling and codebook histogram features to make multimodal classification decisions.

7. An EEG recognition system for epileptogenic zones based on an interpretable codebook and multi-task pre-training according to any one of claims 1 to 6, characterized in that: include: Data preprocessing module, used to filter, denoise and standardize the raw EEG signals; A time series encoding module that generates latent embeddings based on the Transformer architecture; Attribute decoding and vector quantization module, extracts attribute tuples and maps them to discrete codebooks; The shape decoding and reconstruction module combines the attribute tuples to reconstruct the time series subsequence; Multi-task pre-training module to optimize reconstruction, quantization and disentanglement losses; The feature extraction and classification module realizes the identification of epileptogenic zone based on latent space labeling and codebook histogram.

8. The system according to claim 7, characterized in that The codebook dimension of the vector quantization module is 8, the codebook capacity is 1024, and the codebook vector is discretized by low-dimensional projection and Euclidean distance matching.

9. The system according to claim 7, characterized in that The time series encoding module has 8 Transformer layers, 2 decoder layers, 512 embedding dimensions, and 64 slices.

10. The system according to claim 7, characterized in that The classification module supports zero-shot generalization and performs cross-domain classification of unseen datasets based on a pre-trained codebook.