EEG characterization method based on wavelet neural quantization training and semantic alignment
By employing wavelet neural quantization training and a parallel feature extraction architecture, the problems of codebook bias and heterogeneous data unification in EEG signal processing were solved, achieving efficient representation and cross-modal alignment of EEG signals and improving the accuracy of EEG signal classification.
Patent Information
- Application Number
- CN202510864949.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-07
AI Technical Summary
Existing EEG signal processing methods suffer from codebook bias, challenges in unifying heterogeneous data from different domains, and a lack of effective cross-modal feature alignment strategies, resulting in poor self-supervised reconstruction performance and low efficiency in multimodal fusion.
A training method based on wavelet neural quantization is adopted, which combines multi-resolution time-frequency analysis and neural codebook learning to construct a parallel feature extraction architecture. Feature extraction and reconstruction are performed through wavelet quantization encoder and decoder, and cross-modal alignment framework is used to associate EEG spatiotemporal features with clinical diagnostic descriptions to improve signal representation capabilities.
This study aims to achieve efficient compression and reconstruction of the spatiotemporal-spectral features of EEG signals in scenarios with few samples, thereby improving the ability to uniformly represent heterogeneous data, enhancing cross-modal feature alignment capabilities, and improving the accuracy of EEG signal classification.
Smart Images

Figure CN120910552A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to the field of electroencephalogram (EEG) processing, and more particularly to an EEG representation method based on wavelet neural quantization training and semantic alignment. BACKGROUND
[0002] Electroencephalogram (EEG), as a non-invasive technique for recording brain electrical activity, plays a crucial role in brain-computer interfaces and neurological disease diagnosis. The high temporal resolution and low cost make EEG a core data source for studying brain function. This has led to various applications based on EEG signals, including human emotion recognition, body movement imagination, automatic sleep stage classification, seizure detection, and fatigue detection. However, the development of EEG representation learning is affected by problems such as high collection cost, large individual differences, and label scarcity. In recent years, the general artificial intelligence framework represented by large language models (LLMs) has provided a new approach to solving complex data representation problems. In particular, self-supervised pre-training has significant potential by learning general features from a large amount of unlabeled data. Building a basic model that can effectively capture the temporal-spatial and spectral features of electroencephalogram and adapt to various downstream tasks through the "pre-training and fine-tuning" paradigm has become a research hotspot in the field of electroencephalogram signal processing. Despite this, the high noise of EEG data, multi-source heterogeneity (such as different devices, channel configurations, and preprocessing programs), and the complexity of cross-modal feature alignment pose serious challenges to existing methods. All these problems emphasize the need for more effective representation learning and modal alignment methods for EEG signals.
[0003] Although some progress has been made in deep learning-based EEG representation methods, they still face three major problems in practical applications:
[0004] 1) Codebook bias problem in self-supervised EEG representation: In the process of quantizing EEG with a codebook, the non-stationary adaptation of traditional K-means clustering to time-frequency-space features leads to the accumulation of codebook representation deviation from the true signal distribution, which in turn weakens the effect of self-supervised reconstruction EEG representation learning.
[0005] 2) Difficulty in unifying heterogeneous data from different domains: Multi-source EEG data differs significantly in channel number, sampling rate, and noise distribution, and existing methods have difficulty in balancing the preservation of the original signal's temporal-spatial pattern and the robustness of device / preprocessing differences in unified processing.
[0006] 3) Lack of effective cross-modal feature alignment strategy: The lack of effective cross-modal mapping mechanism leads to low efficiency of feature coordination in multi-modal fusion when the semantic association between EEG signals and text (such as clinical diagnosis description) is not fully exploited.
[0007] Therefore, there is a need to improve the prior art to solve at least one of the above problems.
[0008] It should be noted that the background art is only intended to introduce the relevant information of the present application, so as to help understand the technical solutions of the present application, but does not mean that the relevant information must be prior art. The relevant information is submitted and disclosed together with the present application scheme, and in the absence of evidence that the relevant information has been disclosed before the filing date of the present application, the relevant information should not be regarded as prior art. SUMMARY
[0009] Therefore, the purpose of the present application is to overcome the defects of the above prior art, and to provide a training method based on wavelet neural quantization and an EEG signal classification method.
[0010] The purpose of the present application is achieved by the following technical solutions:
[0011] According to the first aspect of the present application, a training method based on wavelet neural quantization is provided, comprising: S1, obtaining training data, which includes a plurality of EEG signal samples, each sample including a plurality of EEG sub-signals, each EEG sub-signal being a part of a single channel signal; S2, training a wavelet quantization neural marker and its neural codebook using the training data, comprising: using the vector quantization encoder of the wavelet quantization neural marker to perform discrete wavelet transform on the input sample, generating a feature block corresponding to each sub-signal of each sample obtained by discrete wavelet transform, and extracting the latent representation of each feature block; obtaining the closest discrete vector matched from the neural codebook according to the latent representation, and using the neural decoder based on wavelet inverse transform to reconstruct the EEG sub-signal corresponding to each feature block according to the discrete vector matched by each feature block; and training and updating the parameters of the neural marker and the discrete vectors of its neural codebook according to the first total loss function, wherein the first total loss function includes the weighted sum of the codebook matching loss sub-function, the reconstruction loss sub-function of the EEG sub-signal and the time-frequency domain structure constraint loss sub-function. The scheme can at least achieve the following beneficial technical effects: the scheme proposes a vector quantization framework based on discrete wavelet transform, for the first time combining multi-resolution time-frequency analysis with neural codebook learning, realizing efficient compression and reconstruction of EEG signal space-frequency spectrum features in a small sample scene, and significantly improving the unified representation ability of heterogeneous data.
[0012] Optionally, each feature block corresponding to each sub-signal of each sample is generated in the following manner: the entire sample is subjected to discrete wavelet transform to obtain time-frequency sub-band sequences of multiple resolutions including approximation coefficients and detail coefficients; each frequency sub-band sequence is segmented to obtain multiple sub-blocks; and the sub-blocks related to each sub-signal are spliced in a predetermined order to obtain the feature block corresponding to the sub-signal obtained by the discrete wavelet transform, and the feature blocks corresponding to different sub-signals have the same length. This scheme can achieve at least the following beneficial technical effects: in this way, the feature blocks corresponding to different sub-signals are all fixed-length vectors, which can be directly used as the input of the vector quantization encoder, thus preserving the frequency characteristics of each sub-band and realizing the structured processing of the features corresponding to the input sample, facilitating subsequent encoding and analysis.
[0013] Optionally, the vector quantization encoder comprises: a channel-level Transformer module, configured to independently map each input feature block to generate query, key, and value vectors, capture the dependency relationship between different channels according to the query, key, and value vectors corresponding to each feature block by means of a cross-channel attention mechanism, obtain attention output, and output a first feature matrix corresponding to each feature block by performing nonlinear transformation on the attention output using a feedforward network; a time-space convolution module, configured to sequentially perform one-dimensional time convolution to extract time features and spatial depth separable convolution to extract spatial features on each first feature matrix, and perform weighted fusion and filtering of invalid information on the time features and spatial features by using gate weights generated by a gating mechanism to obtain spatio-temporal features; and a position encoding module, configured to add position information to the spatio-temporal features by sinusoidal time encoding to obtain a second feature matrix containing spatio-temporal position information, and obtain a latent representation corresponding to the feature block by linear projection of the second feature matrix. This scheme can achieve at least the following beneficial technical effects: the channel-level Transformer of this scheme is composed of a multi-head self-attention layer and a feedforward network, which effectively captures the collaborative mode of activities in different brain regions and realizes global channel interaction by channel-independent mapping and cross-channel attention modeling. Subsequently, the spatio-temporal convolution adopts a cascaded structure of one-dimensional time convolution (including dilated convolution) and spatial depth separable convolution, and combines the gating mechanism to fuse the spatio-temporal features, thereby preserving the transient waveform and spatial distribution details of the electroencephalogram signal. Finally, the position encoding integrates the channel position of the sinusoidal time encoding, which more accurately models the context and analyzes the spatial dependency of the electroencephalogram signal sequence, thereby improving the representation ability of the EEG signal.
[0014] Optionally, the first total loss function is:
[0015]
[0016] wherein, is a codebook matching loss sub-function for constraining the latent representation to be as close as possible to the discrete vectors in the neural codebook, a reconstruction loss sub-function for calculating a point-wise amplitude error and a waveform similarity error weighted sum between the reconstructed EEG sub-signals and the EEG sub-signals of the samples, a time-frequency domain structure constraint loss sub-function for calculating a weighted sum of the approximation coefficient structure loss and the detail coefficient structure loss in the reconstruction process. This scheme can at least achieve the following beneficial technical effects: this scheme uses the codebook matching loss sub-function to make the discrete vectors in the neural codebook as accurately represent the relevant latent representation as possible, so as to more accurately compress and extract effective features; the reconstruction loss sub-function extracts rich reconstruction loss terms from both the point-wise amplitude error and the waveform similarity error, so as to improve the accuracy of the features in the discrete vectors that express the amplitude and waveform related information; the time-frequency domain structure constraint loss sub-function is designed based on the multi-resolution analysis angle of discrete wavelet transform, aiming at the low-frequency trend and high-frequency details after signal decomposition, to ensure the hierarchical preservation of time-frequency domain features; through the weighting of multiple sub-losses, the multi-aspect ability and accuracy of the features expressed by the discrete vectors in the neural codebook are improved.
[0017] Optionally, the method further comprises training the parallel feature encoder, which includes: obtaining the final matched discrete vector of each EEG sub-signal of each sample by using the S2 trained wavelet quantization neural marker and its neural codebook; performing random mask processing on the EEG sub-signals in each sample to obtain a mask-containing sample containing partially masked EEG sub-signals; after using the parallel feature encoder to extract time, space and semantic features from the mask-containing sample in parallel, using the superimposed Transformer block to fuse the time, space and semantic features to obtain the parallel features of each EEG sub-signal, and using the parallel features to determine the predicted discrete vector corresponding to each EEG sub-signal; updating the parameters of the parallel feature encoder according to the error between the predicted discrete vector corresponding to each EEG sub-signal and its final matched discrete vector. This scheme can at least achieve the following beneficial technical effects: this scheme constructs an encoding architecture containing time dynamics, space interaction and semantic association, and through the collaborative design of the superimposed Transformer and the spatio-temporal convolution, it reduces the computational complexity while preserving the physiological specificity of the signal, providing a more universal feature extraction scheme for high-noise and sparse data scenarios.
[0018] Optionally, the parallel feature encoder comprises: a time encoder, which adopts a pre-trained time encoder to extract time features of each EEG sub-signal from the input masked samples; a spatial encoder, which is configured to maintain local spatial encoding parameters corresponding to each brain region group according to a predefined brain region grouping rule, capture spatial correlation features between channels within the corresponding brain region group from the masked samples by using the local spatial encoding parameters corresponding to each brain region group, obtain local spatial features corresponding to each EEG sub-signal, and perform pooling operation on all local spatial features of the same sample after attention mechanism-based processing, to obtain global spatial features, and aggregate the local spatial features and the global spatial features corresponding to each EEG sub-signal to obtain spatial features of the EEG sub-signal; a semantic encoder, which is configured to extract semantic features of each EEG sub-signal from the masked samples; and a plurality of stacked Transformer blocks, which are configured to fuse the time features, the spatial features and the semantic features of each EEG sub-signal according to the attention mechanism to obtain parallel features of each EEG sub-signal. The scheme can achieve at least the following beneficial technical effects: the scheme extracts features from multiple angles by using the time encoder, the spatial encoder and the semantic encoder, and fuses the features by using the plurality of Transformer blocks, which can improve the knowledge richness implied by the parallel features and effectively improve the performance of the parallel feature encoder in extracting features; moreover, the spatial encoder uses predetermined local spatial encoding parameters according to the brain region grouping rule to extract local spatial features of each group, so as to more accurately express the implied information of each brain region group; the advantages of local and global spatial features are fully utilized, the local spatial features and the global spatial features are aggregated to form richer spatial feature representation, which not only retains the spatial specificity of local brain regions, but also integrates the spatial correlation between global channels, so as to improve the spatial expression ability of the final features and provide better support for the prediction ability of subsequent tasks.
[0019] Optionally, the training of the parallel feature encoder further comprises: obtaining a clinical diagnosis description text corresponding to each EEG signal sample, extracting a text embedding from the clinical diagnosis description text by using a preset text encoder; fusing the parallel features of each EEG sub-signal extracted from the mask-containing sample corresponding to each EEG signal sample to obtain an EEG embedding; and updating the parameters of the parallel feature encoder according to a second total loss function, wherein the second total loss function comprises a vector reconstruction sub-loss function for calculating the error between the predicted discrete vector corresponding to each EEG sub-signal and the final matched discrete vector, and an alignment loss sub-function for calculating the deviation between the text embedding and the EEG embedding. The scheme can at least achieve the following beneficial technical effects: the scheme proposes a cross-modal alignment framework based on dynamic semantic coupling, correlates the EEG spatial-temporal features (such as frequency band energy) and the semantic of the clinical diagnosis description text through a contrast loss, and synchronously learns the signal recovery and semantic logic alignment by using a mask reconstruction mechanism, so as to improve the accuracy of the clinical semantics implied by the parallel features, and improve the accuracy of the downstream tasks.
[0020] According to a second aspect of the present application, a method for classifying EEG signals is provided, which comprises: obtaining an EEG signal of a target, inputting the EEG signal of the target into the wavelet quantization neural marker trained according to the method of the first aspect to extract the latent representation of each EEG sub-signal and match the closest discrete vector from the neural codebook according to the latent representation; fusing the closest discrete vectors corresponding to all EEG sub-signals of the target EEG signal, inputting the EEG embedding into a pre-trained classifier, and determining the classification result corresponding to the target.
[0021] According to a third aspect of the present application, a method for classifying EEG signals is provided, which comprises: obtaining an EEG signal of a target, inputting the EEG signal of the target into the parallel feature encoder trained according to the method of the first aspect to extract the parallel features of each EEG sub-signal; fusing the parallel features of all EEG sub-signals of the target EEG signal to obtain an EEG embedding, inputting the EEG embedding into a pre-trained classifier, and determining the classification result corresponding to the target. BRIEF DESCRIPTION OF DRAWINGS
[0022] The embodiments of the present application will be further described below with reference to the accompanying drawings, in which:
[0023] Figure 1 A modular schematic diagram of the training method according to the embodiments of the present application. DETAILED DESCRIPTION
[0024] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application.
[0025] As mentioned in the background section, although some progress has been made in deep learning-based EEG representation methods, they still face three major problems. Among them, for the codebook bias problem in self-supervised EEG representation, the present invention proposes a wavelet transform-based vector quantization framework, which first combines multi-resolution time-frequency analysis with neural codebook learning, improves the non-stationary adaptation of time-frequency-space features, and realizes efficient compression and reconstruction of EEG signal space-spectrum features in a small sample scenario, significantly improving the unified representation ability of heterogeneous data. For the problem of unified heterogeneous data across domains, the present invention constructs a three-dimensional parallel feature extraction paradigm: an encoding architecture is constructed to include time dynamics, spatial interactions, and semantic associations, and through the collaborative design of spatial convolution and channel-level Transformer blocks, the physiological specificity of the signal is preserved while reducing computational complexity, providing a universal feature extraction solution for high-noise and sparse data scenarios. For the problem of lacking effective cross-modal feature alignment strategies, the present invention proposes a dynamic semantic coupling-based cross-modal alignment framework, which correlates EEG space-time features (such as frequency band energy) and clinical text semantics through a contrastive loss, and combines a mask reconstruction mechanism to simultaneously learn signal recovery and semantic logic alignment, improving the cross-modal semantic alignment ability of related encoders.
[0026] Embodiment 1
[0027] According to an embodiment of the present invention, referring to Figure 1 , a wavelet neural quantization-based training method is provided, comprising
[0028] S1, obtaining training data, which includes a plurality of EEG signal samples, each sample including a plurality of EEG sub-signals, each EEG sub-signal being a part of a single channel signal;
[0029] S2, performing first-stage training, training a wavelet quantization neural marker and its neural codebook using the training data, comprising:
[0030] S21, using the vector quantization encoder of the wavelet quantization neural marker to perform discrete wavelet transform on the input sample, generating a feature block corresponding to each sub-signal of each sample obtained by discrete wavelet transform, and extracting the latent representation of each feature block;
[0031] S22, obtaining the closest discrete vector matched from the plurality of discrete vectors in the neural codebook according to the latent representation, and using the neural decoder based on the inverse wavelet transform to reconstruct the EEG sub-signal corresponding to each feature block according to the discrete vector matched by each feature block;
[0032] S23, training and updating the parameters of the neural marker and the discrete vectors of the neural codebook thereof according to the first total loss function, wherein the first total loss function comprises a weighted sum of a codebook matching loss sub-function, a reconstruction loss sub-function of the EEG sub-signals and a time-frequency domain structure constraint loss sub-function;
[0033] S3, performing second-stage training to train the parallel feature encoder, which comprises:
[0034] S31, obtaining the final matching discrete vector of each EEG sub-signal of each sample by using the wavelet quantization neural marker and the neural codebook thereof trained in the first stage;
[0035] S32, performing random mask processing on the EEG sub-signals in each sample to obtain a mask-containing sample containing partially masked EEG sub-signals;
[0036] S33, after extracting time, space and semantic features from the mask-containing sample in parallel by using the parallel feature encoder, performing fusion processing on the time, space and semantic features by using the stacked Transformer block to obtain the parallel features of each EEG sub-signal, and determining the predicted discrete vector corresponding to each EEG sub-signal using the parallel features;
[0037] S34, updating the parameters of the parallel feature encoder according to the error between the predicted discrete vector corresponding to each EEG sub-signal and the final matching discrete vector thereof.
[0038] In order to better understand the present application, each step will be described in detail below in combination with specific embodiments.
[0039] Step S1: obtaining training data, which comprises a plurality of EEG signal samples, each sample comprising a plurality of EEG sub-signals, each EEG sub-signal being a part of a single channel signal.
[0040] According to an embodiment of the present application, the EEG signal samples of the training data comprise multi-channel EEG signal samples collected from electrodes deployed on the brain of a subject. Different electrodes correspond to different data channels. Different EEG signal samples can be collected from different subjects. Each EEG signal can be divided into a plurality of EEG sub-signals according to different channels c (such as FPz, AF7, F5, FC1, Cz, CP6, P8, O1, etc. in the EEG10-20 system) and different time instants t, so as to obtain a plurality of EEG sub-signals. Figure 1 wherein the EEG sub-signal represents the EEG sub-signal of channel c at time instant t.
[0041] Step S2: training the wavelet quantization neural marker and the neural codebook thereof using the training data.
[0042] Before introducing step S2 (the training of the first stage), the wavelet quantization neural marker and its neural codebook are described.
[0043] According to an embodiment of the present application, the wavelet quantization neural marker comprises a vector quantization encoder, a neural decoder and a neural codebook. Wherein:
[0044] 1) Vector quantization encoder
[0045] The vector quantization encoder comprises a channel-level Transformer module, a time-space convolution module and a position encoding module. Unlike previous work, our design focuses more on channel relationships and more comprehensively explores the spatio-temporal correlation of signals, which is more advantageous in capturing long-distance dependencies or dynamic position relationships.
[0046] The channel-level Transformer module is used to independently map each feature block of the input to generate query, key and value vectors, capture the dependency relationship between different channels according to the query, key and value vectors corresponding to each feature block through the cross-channel attention mechanism, obtain the attention output, and use the feedforward network to perform nonlinear transformation on the attention output to output the first feature matrix corresponding to each feature block. The channel-level Transformer module is composed of a multi-head self-attention layer and a feedforward network, which effectively captures the collaborative mode of activities in different brain regions through channel-independent mapping and cross-channel attention modeling, and realizes global channel interaction.
[0047] The time-space convolution module is used to sequentially perform one-dimensional time convolution to extract time features and spatial depth separable convolution to extract spatial features for each first feature matrix, and uses the gating weight generated by the gating mechanism to weight and fuse the time features and spatial features and filter invalid information to obtain spatio-temporal features. The time-space convolution module adopts a cascaded structure of one-dimensional time convolution (including dilated convolution) and spatial depth separable convolution, combines the gating mechanism to fuse spatio-temporal features, and preserves the transient waveform and spatial distribution details of the electroencephalogram signal.
[0048] The position encoding module is used to add position information to the spatio-temporal features through sinusoidal time encoding to obtain a second feature matrix containing spatio-temporal position information, and project the second feature matrix through linear projection to obtain the latent representation of the corresponding feature block. The position encoding module integrates the channel position of the sinusoidal time encoding, which more accurately models the context and analyzes the spatial dependence of the electroencephalogram signal sequence. The feature block of the original EEG sub-signal of the signal at time t located in channel c is finally mapped into a latent representation by the VQ Encoder. .
[0049] 2) Neural codebook
[0050] The neural codebook includes a plurality of discrete vectors which are learnable. The neural codebook is defined where K' is the number of discrete vectors in the neural codebook, and D' is the dimension of the vectors. The dimension of the discrete vectors is consistent with the dimension of the latent representation. represents the i-th discrete vector, and Figure 1 v1-v12 represent the 1st-12th discrete vectors. The discrete vectors in the codebook are essentially the cluster centers of the latent representations of the EEG sub-signals, representing the typical patterns in the high-dimensional space (e.g. different brain rhythms, cognitive states corresponding to feature clusters). The vector values are determined by the distribution of the latent representations; for example, the codebook vectors corresponding to low-frequency stable rhythms can have a smooth numerical distribution, while the vectors corresponding to high-frequency transient activities can contain dramatic fluctuations in numerical values.
[0051] The role of the neural codebook is to quantize the latent representation, calculate the distance (Mahalanobis distance or Euclidean distance) between the latent representation and each discrete vector, and find the most similar latent codebook , so that quantization is completed.
[0052] Illustratively, taking the Mahalanobis distance as an example, the distance between the latent representation and each discrete vector is calculated as:
[0053]
[0054] where represents the transpose, is the feature covariance matrix, which is used to eliminate the correlation between the latent representation and the discrete vector and to standardize the scale, so that the distance measurement is more consistent with the statistical properties of the EEG signal.
[0055] Illustratively, the initial neural codebook can be generated based on pre-training or random initialization to generate initial discrete vectors. For example, K vectors can be randomly sampled from the latent representations of the sub-signals corresponding to a large amount of EEG data as initial cluster centers. Then, the values of the discrete vectors can be dynamically updated as the training progresses.
[0056] 3) Neural decoder
[0057] A neural decoder based on inverse wavelet transform (IDWT) is constructed, and the neural decoder is a neural network. The neural decoder includes a coefficient reconstruction layer and a sub-signal reconstruction layer. The quantized discrete vectors are input as input, and the coefficient reconstruction layer including multiple convolution layers is used for feature transformation to obtain estimated wavelet coefficients and , and then the sub-signal reconstruction layer is used to integrate the estimated wavelet coefficients into the reconstructed EEG sub-signal according to the IDWT It should be understood that, represents an estimated approximation coefficient, represents an estimated detail coefficient.
[0058] Step S2 includes sub-steps: S21-S23, which are respectively schematically described below.
[0059] Step S21: using the vector quantization encoder of the wavelet quantization neural marker to perform discrete wavelet transform on the input sample, to generate a feature block corresponding to each sub-signal of each sample obtained by discrete wavelet transform, and to extract a latent representation of each feature block.
[0060] According to an embodiment of the present application, the feature block corresponding to each sub-signal of each sample obtained by discrete wavelet transform is generated in the following manner: the entire sample is subjected to discrete wavelet transform to obtain time-frequency sub-band sequences of multiple resolutions including approximation coefficients and detail coefficients; each frequency sub-band sequence is segmented to obtain multiple sub-blocks; and the sub-blocks related to each sub-signal are spliced in a predetermined order to obtain the feature block corresponding to the sub-signal obtained by discrete wavelet transform, and the feature blocks corresponding to different sub-signals are of the same length.
[0061] For example, a single EEG signal sample containing multiple segments is taken as input. The original EEG signal is first subjected to discrete wavelet transform (DWT) to obtain approximation coefficients (A) and detail coefficients (D1, D2, D3) of different levels of detail. For example, after 3-level discrete wavelet transform (DWT) is performed on an EEG signal of length 1024, 4 sub-bands are obtained, which are:
[0062] 1) Approximation coefficients (cA3)
[0063] This is the low-frequency component retained after 3-level decomposition, which reflects the overall trend and low-frequency information of the signal. Length: 1024 / 23=128.
[0064] 2) Detail coefficients (cD3)
[0065] This is the high-frequency component obtained by 3-level decomposition, which contains the local detail changes of the signal at the 3rd level, and the length is: 1024 / 23=128.
[0066] 3) Detail coefficients (cD2)
[0067] This is the high-frequency component obtained by 2-level decomposition, which reflects the local detail changes of the signal at the 2nd level. Length: 1024 / 22=256.
[0068] 4) Detail coefficients (cD1)
[0069] This is the high frequency component obtained by the first level decomposition, which contains the local detail changes of the signal at the first level. Length: 1024 / 21 = 512.
[0070] Each coefficient can be divided into multiple sub-blocks, and then grouped and merged according to the correspondence between the sub-blocks and the EEG sub-signals to obtain the feature blocks corresponding to the EEG sub-signals obtained by the discrete wavelet transform. Subsequently, the feature blocks Input vector quantization encoder to extract latent representation .
[0071] Step S22: Obtain the closest discrete vector matched from the plurality of discrete vectors in the neural codebook according to the latent representation, and reconstruct the EEG sub-signal corresponding to each feature block according to the discrete vector matched by the feature block using the neural decoder based on the inverse wavelet transform.
[0072] According to an embodiment of the present application, the latent representation matches the nearest discrete vector from the neural codebook , and the EEG sub-signal corresponding to the feature block is reconstructed using the quantized discrete vector .
[0073] Step S23: Update the parameters of the neural marker and the discrete vectors of its neural codebook according to the first total loss function, wherein the first total loss function includes the weighted sum of the codebook matching loss sub-function, the reconstruction loss sub-function of the EEG sub-signal and the time-frequency domain structure constraint loss sub-function.
[0074] According to an embodiment of the present application, the first total loss function is:
[0075] wherein, is the codebook matching loss sub-function for constraining the latent representation to be as close as possible to the discrete vectors in the neural codebook, is the reconstruction loss sub-function for calculating the weighted sum of the point-by-point amplitude error and the waveform similarity error between the reconstructed EEG sub-signal and the sample EEG sub-signal, is the time-frequency domain structure constraint loss sub-function for calculating the weighted sum of the approximation coefficient structure loss and the detail coefficient structure loss in the reconstruction process.
[0076] The sub-losses of the first total loss function will be described in detail below.
[0077] 1) Codebook matching loss sub-function
[0078] According to an optional embodiment of the present application, which can be defined as:
[0079]
[0080] in, yes The corresponding weights It is the covariance matrix. Indicates transpose. Indicates and The closest discrete vector. During wavelet quantization neural labeler training, this formula is based on weighted K-means clustering and Mahalanobis distance, allowing the latent representation to... The values should be as close as possible to the neural codebook. Discrete vectors in similar.
[0081] According to another optional embodiment of the invention, a regularization term can be introduced to constrain codebook diversity. The previous formula only focuses on "the matching between the current latent representation and discrete vectors in the codebook," and by adding a regularization term, excessive concentration of codebook vectors can be avoided.
[0082]
[0083] in, It is a similarity penalty term between discrete vectors in the neural codebook. The higher the similarity with discrete vectors, the better. The larger the similarity, the greater the penalty term. In other words, if the similarity between discrete vectors is too similar, the larger the penalty term, the more diverse the codebook becomes, ensuring that there are as many differentiated exclusive codebook vectors as possible between different EEG signal feature patterns, thus improving the quantitative representation capability.
[0084] According to another optional embodiment of the present invention, the formula in the above embodiments uses fixed weights, which makes it impossible to focus on learning important physiological patterns. It can be changed to dynamically calculate the weights based on the local complexity of the signal, adapting to the signal complexity through dynamic weighting.
[0085]
[0086] in, Indicates dynamic weights, , This represents the rate of change of amplitude in the neighborhood of an EEG sub-signal (e.g., the neighborhood of channel c in an EEG signal at time t). This represents an exponential function. Through dynamic weighting, training focuses on optimizing codebook matching for key signals, such as complex EEG events (e.g., spikes, sharp waves). Larger, corresponding to dynamic weight This allows the model to focus more on important neurophysiological patterns, thereby improving the neural codebook's ability to represent these patterns and enhancing the performance of subsequent downstream tasks.
[0087] 2) Reconstruction loss function
[0088] According to an optional embodiment of the present invention, the reconstruction loss function It can be simply the point-by-point amplitude error (mse). The point-by-point amplitude error (mse) directly measures the difference in amplitude between the original signal and the reconstructed signal at each point, preserving the basic accuracy constraints.
[0089]
[0090] in, This represents the total number of channels. This refers to the time period (or number of time points, number of sampling points) of a single EEG signal. This represents the EEG sub-signal of channel c at time t. This represents the EEG sub-signal of channel c at time t during reconstruction.
[0091] According to another optional embodiment of the invention, the reconstruction loss function It can be solely for waveform similarity loss. Waveform similarity loss is calculated using normalized cross-correlation (NCC) to measure the overall waveform similarity:
[0092]
[0093]
[0094] in, Represents the original EEG sub-signal The mean, This represents the EEG sub-signal of channel c at time t. Indicates reconstruction signal The mean, Represents the original signal With reconstruction signal The normalized cross-correlation is calculated. The loss value ranges from 0 to 2, with smaller values indicating higher similarity.
[0095] According to yet another embodiment of the present invention, reconstruction loss It can be composed of point-by-point amplitude error (mse) and waveform similarity loss. Therefore, the reconstruction loss linearly combines the point-by-point amplitude error and waveform similarity loss, requiring only two weight hyperparameters:
[0096]
[0097] in, express The weighting coefficients, express weighting coefficients of
[0098] 3) Time-frequency domain structure constraint loss sub-function
[0099] To ensure the accuracy and reversibility of the feature representation in the quantization process, while preserving the time-frequency domain structure information of the EEG signal, the time-frequency domain structure constraint loss is based on the multi-resolution analysis of discrete wavelet transform, and is designed for the low-frequency trend and high-frequency details after signal decomposition, to ensure the hierarchical preservation of time-frequency domain features. Preferably, including the approximation coefficient structure loss and the detail coefficient structure loss :
[0100]
[0101]
[0102]
[0103] wherein, , are respectively , weighting coefficients of . The approximation coefficient structure loss ensures the accuracy of the approximation coefficient in reconstruction, maintaining the main low-frequency component of the signal. The detail coefficient structure loss ensures the accuracy of the high-frequency details in reconstruction.
[0104] The first total number function is used to guide the optimization of the encoder (E), the decoder (D) and the neural codebook (C). In the early stage of training, the adaptability of the codebook to the data distribution can be increased (i.e. the weighting coefficient of is greater than the weighting coefficient of ). In the later stage, the reconstruction effect is emphasized (i.e. the weighting coefficient of is greater than the weighting coefficient of ). At the same time, when calculating , a feature semantic similarity measure can also be introduced to iteratively update the model parameters.
[0105] Step S3: training the parallel feature encoder.
[0106] After the training of the neural codebook is completed, the input EEG signal is subjected to probabilistic mask processing, and then input into the parallel feature encoder to extract feature embeddings from time, space and semantics in parallel. Then the model restores the masked area using context information and codebook prior knowledge. Thus, a universal parallel feature encoder can be trained.
[0107] For ease of understanding, the structure of the parallel feature encoder is first introduced schematically.
[0108] According to an embodiment of the present application, the parallel feature encoder comprises a temporal encoder, a spatial encoder, a semantic encoder and a plurality of stacked Transformer blocks. Wherein:
[0109] 1) Temporal encoder
[0110] The temporal encoder can adopt a pre-trained temporal encoder for extracting temporal features of each EEG sub-signal from the input mask samples.
[0111] In order to save training overhead, the temporal encoder can use the temporal encoder that has been trained in the existing large model. For example, the architecture of the temporal encoder in the LaBram model is used, and the parameters of this part are frozen when building the parallel feature encoder. The temporal encoder in the LaBram model is pre-trained with multi-task EEG data and has excellent temporal feature extraction capability. It is composed of multiple 1-D convolution blocks, each of which contains a convolution layer, a group normalization layer and a GELU activation function. For the input EEG signal , the encoding generates an embedding is a d-dimensional vector. With this pre-trained and parameter-frozen temporal encoder, temporal features are efficiently extracted while reducing computational overhead, allowing the model to benefit from the rich temporal representations learned during the pre-training phase. It should be understood that other temporal encoders can also be used. Alternatively, without considering the impact of time overhead, the implementer can design the temporal encoder structure by himself using convolution layers, pooling layers and activation functions, etc.
[0112] 2) Spatial encoder
[0113] According to an embodiment of the present application, the spatial encoder is used to maintain local spatial encoding parameters corresponding to each brain region group according to a pre-defined brain region grouping rule, capture spatial correlation features between channels within the corresponding brain region group from the mask samples using the local spatial encoding parameters corresponding to each brain region group, obtain local spatial features corresponding to each EEG sub-signal, and perform pooling operation after processing all local spatial features of the same sample based on attention mechanism respectively to obtain global spatial features. The local spatial features and global spatial features corresponding to each EEG sub-signal are aggregated to obtain the spatial features of the EEG sub-signal. According to the physiological function and anatomical structure of the brain region, the channels are divided according to the pre-defined grouping rule. An illustrative grouping rule is shown in the following table:
[0114]
[0115] For each group of channels, in order to capture the spatial correlation features between the channels within the local brain region, the following method is used for processing:
[0116]
[0117] wherein, represents a learnable weight matrix for local spatial encoding, represents local channel data obtained by dimension reordering of channel-related data of the same group, represents a bias term, represents local spatial features obtained by first linear transformation, then bias addition, and finally ReLU activation on the local channel data . Focus on the spatial structure within the local brain region, i.e., the local dependency pattern between channels.
[0118] Then, all local spatial features of the same sample can be respectively processed based on an attention mechanism and then globally pooled to obtain global spatial features . Alternatively, all local spatial features of the same sample can be directly averaged and pooled to obtain global spatial features. The global feature vector integrates the spatial information of all channels and captures the global spatial dependency and interaction pattern across brain regions. To fully utilize the advantages of local and global spatial features, the local spatial features and the global spatial features are aggregated to form richer spatial feature representations which not only retain the spatial specificity of local brain regions but also integrate the spatial correlation between global channels.
[0119] 3) Semantic encoder
[0120] According to an embodiment of the present application, the semantic encoder is used to extract semantic features of each EEG sub-signal from the mask-containing sample.
[0121] According to an optional embodiment of the present application, the input signal (original EEG sub-signal of channel c and time point t) is first frequency-decomposed using fast Fourier transform (FFT):
[0122]
[0123] The power spectrum feature is calculated, representing the energy distribution of channel c in frequency band b. The semantic information of the EEG signal depends on the association between the frequency band (such as δ, θ, α, β, γ wave) and the brain region function (for example, α wave is related to the resting state, and β wave is related to the motor cortex activation). Through channel attention and spectral attention mechanisms, key spectral components are adaptively selected:
[0124] Channel attention focuses on task-critical brain regions (such as central region channels in motor tasks) through weight , and spectral attention focuses on key frequency components through weight Capture characteristic frequency bands (e.g. epileptic abnormal high-frequency oscillations):
[0125]
[0126]
[0127] Synthesize weights , locate key semantic features in the "channel-frequency band" joint space.
[0128] Adjust complex Fourier coefficients , recover time-domain signals using inverse FFT:
[0129]
[0130] Combine residual connections , while screening spectral features, retain the time-domain dynamic characteristics of the original signal (e.g. time localization of spiky and sharp waves).
[0131] Perform channel-by-channel layer normalization, and project normalized features to a shared semantic space through a multi-layer perception:
[0132]
[0133]
[0134] where, represents the signal feature value of channel c at time t after residual connection, represents the signal feature value of channel c at other time t', represents the signal feature value of channel c at time t", represents the minimum value (e.g. 10 -8 order of magnitude), set the purpose is to prevent the denominator from being 0, flatten features into vectors, is a nonlinear activation function, , represents the bias. This process converts the original signal into a distributed semantic representation with neurophysiological significance, so that signals with similar neurophysiological mechanisms are distributed adjacent in the embedding space (e.g. different samples of alpha wave activity form clusters), thus achieving explicit encoding of semantic information.
[0135] According to another optional embodiment of the present application, the normalization of the previous embodiment is "independent operation within the channel", which cannot highlight the cooperative mode between channels. Inter-channel attention can be added to allow the normalization of brain electrical signal strongly related channels (e.g. adjacent channels of motor cortex) to influence each other, highlighting the cooperative mode:
[0136]
[0137]
[0138]
[0139] First, E n Do inter-channel attention ( It is attention weight), strengthens the brainwave channel coordination features (such as multi-channel synchronization abnormalities during epileptic discharge), and then performs normalization and MLP mapping based on the weighted features to make the semantic encoding more in line with the neural mechanism of brain region collaboration.
[0140] According to another optional embodiment of the present invention, the normalization method can be replaced with instance normalization. If it is desired to focus more on the independence of "channel-time" instances, layer normalization can be replaced with instance normalization (normalizing each instance corresponding to each channel and time separately). After modification:
[0141]
[0142] in, (Time mean of channel c) (Time variance of channel c), subsequent MLP mapping remains unchanged:
[0143]
[0144] This instance normalization is a more rigorous "channel-time instance" standardization, which is suitable for highlighting the characteristics of single-channel transient events (such as spikes) in EEG signals, avoiding the influence of "signals at different times within a channel being flattened" in layer normalization, and preserving more fine-grained temporal dynamics.
[0145] 4) Multiple stacked Transformer blocks
[0146] According to one embodiment of the present invention, multiple superimposed Transformer blocks are used to fuse the temporal features, spatial features and semantic features of each EEG sub-signal according to an attention mechanism to obtain the parallel features of each EEG sub-signal.
[0147] Step S3 is the second stage of training, which includes S31-S33. The following is an illustrative explanation of each sub-step.
[0148] Step S31: Use the wavelet quantization neural labeler trained in S2 and its neural codebook to obtain the discrete vector of the final matching of each EEG sub-signal for each sample.
[0149] According to an embodiment of the present application, after the wavelet quantization neural marker and its neural codebook are trained, the EEG sub-signals of each sample are input into the wavelet quantization neural marker to extract latent representations, and then the distance between each discrete vector in the neural codebook and the latent representations is calculated to find the discrete vector with the minimum distance, thereby obtaining the final matching discrete vector of each EEG sub-signal.
[0150] Step S32: Random mask processing is performed on the EEG sub-signals in each sample to obtain a masked sample containing partially masked EEG sub-signals.
[0151] According to an embodiment of the present application, some EEG sub-signals in the sample are randomly set to 0 to obtain a masked sample containing partially masked EEG sub-signals.
[0152] Step S33: After the time, space and semantic features are extracted from the masked sample in parallel by the parallel feature encoder, the time, space and semantic features are fused by using the stacked Transformer block to obtain the parallel features of each EEG sub-signal, and the parallel features are used to determine the predicted discrete vector corresponding to each EEG sub-signal.
[0153] According to an embodiment of the present application, the masked sample is input into the parallel feature encoder of the above embodiment to extract the parallel features of each EEG sub-signal, and then the prediction head is used to determine the predicted discrete vector according to the parallel features. For example, the discrete vector closest to the parallel features is calculated as the predicted discrete vector:
[0154]
[0155] wherein, represents the discrete vector closest to the parallel features, represents the parallel features and the distance between the parallel features and the discrete vector .
[0156] Step S34: The parameters of the parallel feature encoder are updated according to the error between the predicted discrete vector corresponding to each EEG sub-signal and the final matching discrete vector.
[0157] According to an optional embodiment of the present application, when the parallel feature encoder is trained, the optimization target can be set to minimize the error between the predicted discrete vector corresponding to each EEG sub-signal and the final matching discrete vector, so as to update the parameters of the parallel feature encoder.
[0158] According to another optional embodiment of the present application, in order to strengthen feature understanding, data implicit patterns are mined, and the model learns the internal structure and spatiotemporal semantics of the EEG signal, such as enabling the model to learn the semantic information shared in the cross-modal data, which can be further improved. Preferably, the training of the parallel feature encoder further includes: obtaining the clinical diagnosis description text corresponding to each EEG signal sample, and extracting text embedding from the clinical diagnosis description text by using a preset text encoder; obtaining EEG embedding according to the parallel features of each EEG sub-signal extracted from the mask sample corresponding to each EEG signal sample; and updating the parameters of the parallel feature encoder according to a second total loss function, wherein the second total loss function includes a vector reconstruction sub-loss function for calculating the error between the predicted discrete vector of each EEG sub-signal and the final matched discrete vector thereof, and an alignment loss sub-function for calculating the deviation between the text embedding and the EEG embedding.
[0159] wherein the parallel features (EEG embedding Z E ) obtained after parallel feature encoding are focused on feature alignment in the semantic dimension to capture the correlation between the textual description and the EEG signal change pattern.
[0160] Suppose there is a batch of EEG signal samples and the corresponding clinical diagnosis description texts , wherein N is the number of samples. After the foregoing processing, the EEG signal has been subjected to the mask operation and obtained the preliminary feature representation (C is the number of channels, and T is the time length), and the i-th clinical diagnosis description text is encoded by using the advanced model multilingual-e5-large-instruct in the mteb list, so as to map to a 1024-dimensional vector. It should be understood that other large models, such as Ling-Embed-Mistral or Qwen3-Embedding-0.6b, can also be used to extract text embedding from the clinical diagnosis description text.
[0161] To realize the alignment of the text and the EEG embedding, a contrastive loss function is constructed to maximize the mutual information of the two types of embedding. For the batch samples and , first, a cosine similarity matrix is calculated, and the loss function forces the same type of sample embedding to be close in the feature space and the different types of sample embedding to be far away by comparing the similarity of the positive sample pair and the negative sample pair . Illustratively, the alignment loss sub-function of the deviation between the EEG embeddings is defined as:
[0162]
[0163] wherein, denotes the batch size, is a temperature parameter for adjusting the sharpness of the similarity distribution, denotes the similarity between the EEG embedding of an EEG signal and the text embedding corresponding to the EEG signal, denotes the similarity between the EEG embedding of an EEG signal and the text embedding corresponding to other EEG signals. The loss function enables the model to learn the semantic information shared across the modal data (e.g. similar texts and EEG signal change patterns form tight clusters in the embedding space).
[0164] In the above embodiment, it is divided into two stages. In the first stage, the electroencephalogram (EEG) signal is decomposed into multi-resolution time-frequency subbands by discrete wavelet transform (DWT), and a lightweight vector quantization (VQ) encoder is designed. By combining the channel-level Transformer with the spatio-temporal convolution module, the spatio-temporal features across channels are captured. A compact representation is generated by using weighted K-means clustering and dynamic codebook quantization, solving the unified modeling problem of multi-source heterogeneous electroencephalogram. In the second stage, a three-dimensional parallel coding architecture (in the time, space and semantic dimensions) is constructed. Robust features are extracted by freezing the pre-trained model, brain region grouping interaction and spectral attention mechanism, respectively. A cross-modal alignment strategy is further designed. The semantic association between the electroencephalogram embedding and the clinical text is linked using contrastive learning, and the signal-semantic mapping is optimized by the mask reconstruction task, realizing end-to-end clinical description generation.
[0165] According to an embodiment of the present application, a method for classifying EEG signals, comprising: obtaining an EEG signal of a target, inputting the EEG signal of the target into a parallel feature encoder trained according to the training method of the preceding embodiment to extract parallel features of each EEG sub-signal thereof; fusing the parallel features of all EEG sub-signals of the target EEG signal to obtain an EEG embedding, inputting the EEG embedding into a pre-trained classifier to determine a classification result corresponding to the target. The pre-trained classifier can be trained using classification data sets of downstream tasks, the parameters of the parallel feature encoder are frozen during training, and only the parameters of the classifier are updated based on the classification cross-entropy loss function. The classifier can be implemented by using a linear layer and a Softmax function.
[0166] For example: the downstream task is electroencephalogram signal classification (such as sleep staging).
[0167] The task background is: sleep staging needs to distinguish the categories corresponding to "awake, N1, N2, N3, rapid eye movement (REM)" states, and the rhythm characteristics (such as δ, θ, α wave proportion) and spatio-temporal patterns (such as vertex spike distribution) of electroencephalogram signals of different categories are significantly different.
[0168] Of course, in addition to using the classifier, other ways of using the features extracted by the present application can also be used, such as directly determining whether it is abnormal based on the degree of difference between the features to be detected and the features related to the normal population.
[0169] For example: EEG signal anomaly detection (such as identification of epileptic spike waves) task
[0170] Task background: Epileptic spike waves are abnormal EEG events with sudden, high-frequency and high-amplitude, which need to be accurately located from long-term signals, and rely on the "transient spatiotemporal anomaly pattern" of the signal (such as a sudden spike in a certain brain channel, accompanied by synchronous anomalies in surrounding channels).
[0171] Scheme application: Characterization input: Use the encoded EEG embedding, which contains the "normal EEG spatiotemporal correlation learned by self-supervised learning" (such as the channel coordination pattern of background rhythm).
[0172] Downstream adaptation: Build an "abnormal scoring model" (such as calculating the distance between the EEG embedding of the test sample and the "normal sample characterization distribution"), if the EEG embedding of a certain signal deviates significantly from the normal distribution (such as the characterization vector deviating from the normal cluster center caused by the spike wave), it is determined to be abnormal.
[0173] Detection advantage: The pre-trained code has "understood the spatiotemporal semantics of normal EEG" (such as the posterior occipital dominance distribution of alpha waves), and self-supervised learning makes the model more sensitive to "normal waveform dependency relationships", and abnormalities such as spike waves that break these dependencies are easier to capture.
[0174] Embodiment 2
[0175] In the case of only solving the first problem mentioned in the background art, a better effect than the prior art can still be achieved. Therefore, according to one optional embodiment of the present application, the difference of this embodiment is that there is no parallel feature encoder and its training step. That is, this embodiment contains steps S1 and S2, but not S3.
[0176] According to one embodiment of the present application, a method for classifying EEG signals, the classification method comprises: obtaining the EEG signal of the target, inputting the EEG signal of the target into the wavelet quantization neural marker trained according to the training method of the foregoing embodiment to extract the latent representation of each EEG sub-signal and match the closest discrete vector from the neural codebook according to the latent representation; fuse the closest discrete vectors corresponding to all EEG sub-signals of the target EEG signal, input the pre-trained classifier, and determine the classification result corresponding to the target.
[0177] In addition, the inventors show through experiments that the method of the present application significantly outperforms mainstream baselines in the multi-modal semantic decoding task on EIT-1M and Chinese EEG datasets. For example, in the EIT-1M dataset, the accuracy of the method of the present application is improved by 12.9% compared with the single-modal enhanced LaBram method; and the F1 value is improved by 3.9% compared with the NeuroLM method. In the Chinese EEG dataset in the Chinese scenario, the ROUGE index of the method of the present application is improved by 9.6% compared with the LaBram, verifying the adaptability of the cross-modal alignment to the Chinese corpus. In addition, the method of the present application also shows superior performance on the Thought2Text and TUAB datasets, fully proving its effectiveness and robustness on different tasks and datasets.
[0178] In general, the scheme of the embodiments of the present application can at least achieve one of the following beneficial effects:
[0179] 1) The present application proposes an EEG unified representation framework based on wavelet neural quantization and dynamic semantic alignment, which effectively solves the key technical problems such as multi-source heterogeneous data unification, codebook bias of self-supervised representation learning and / or cross-modal feature alignment difficulty in EEG signal representation learning.
[0180] 2) The unified representation of multi-source EEG signals is realized by combining discrete wavelet transform and weighted codebook quantization, the time and space features of EEG signals are captured by using channel-level Transformer and spatial-temporal convolution, the time, space and semantic features are extracted by using the three-dimensional parallel coding architecture of frozen time model, brain region grouping and frequency band energy modeling, and the EEG-text semantic mutual information is maximized by using the double constraint alignment mechanism of contrastive learning and mask-based reconstruction.
[0181] It should be noted that although the above describes each step in a specific order, it does not mean that each step must be performed in the above specific order, and in fact, some of these steps can be performed concurrently or even in a changed order, as long as the desired function can be achieved.
[0182] The present application can be a system, a method and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for causing a processor to implement various aspects of the present application.
[0183] A computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves.
[0184] Embodiments of the application have been described above, with the understanding that these embodiments are exemplary only, and are not restrictive, in terms of the scope of the embodiments disclosed. Many modifications and variations of the described embodiments are possible, in light of the above teachings, without departing from the scope and spirit of the described embodiments. The choice of words in this document is intended to best explain the principles of the embodiments, practical application, or technical improvement in the art, or to enable others skilled in the art to utilize the embodiments disclosed herein.
Claims
1. A training method based on wavelet neural quantization, characterized in that, The method comprises: S1, obtaining training data comprising a plurality of EEG signal samples, each sample comprising a plurality of EEG sub-signals, each EEG sub-signal being a part of a single channel signal; S2, training a wavelet quantization neural marker and its neural codebook using the training data, comprising: using a vector quantization encoder of the wavelet quantization neural marker to perform discrete wavelet transform on the input sample to generate a feature block corresponding to each sub-signal of each sample, and extracting a latent representation of each feature block; obtaining the closest discrete vector matched from the plurality of discrete vectors in the neural codebook according to the latent representation, and using a neural decoder based on inverse wavelet transform to reconstruct the EEG sub-signal corresponding to each feature block according to the discrete vector matched by the feature block; training and updating the parameters of the neural marker and the discrete vectors of the neural codebook according to a first total loss function, wherein the first total loss function comprises a weighted sum of a codebook matching loss sub-function, a reconstruction loss sub-function of the EEG sub-signal and a time-frequency domain structure constraint loss sub-function.
2. The method of claim 1, wherein, The feature block corresponding to each sub-signal of each sample is generated in the following manner: performing discrete wavelet transform on the entire sample to obtain time-frequency sub-band sequences of multiple resolutions including approximation coefficients and detail coefficients; segmenting each frequency sub-band sequence to obtain a plurality of sub-blocks; splicing the sub-blocks related to each sub-signal in a predetermined order to obtain a feature block corresponding to the sub-signal, and the feature blocks corresponding to different sub-signals are of the same length.
3. The method of claim 2, wherein, The vector quantization encoder comprises: a channel-level Transformer module for independently mapping each input feature block to generate query, key and value vectors, capturing inter-channel dependencies according to the query, key and value vectors corresponding to each feature block by means of a cross-channel attention mechanism, obtaining attention output, and using a feedforward network to perform nonlinear transformation on the attention output to output a first feature matrix corresponding to each feature block; a time-space convolution module for sequentially performing one-dimensional time convolution to extract time features and spatial depth separable convolution to extract spatial features on each first feature matrix, and using gate weights generated by a gating mechanism to weight and fuse the time features and spatial features and filter invalid information to obtain spatio-temporal features; a position encoding module for adding position information to the spatio-temporal features by sinusoidal time encoding to obtain a second feature matrix containing spatio-temporal position information, and projecting the second feature matrix to obtain a latent representation of the corresponding feature block.
4. The method of claim 1, wherein, The first total loss function is: wherein, is a codebook matching loss sub-function for constraining the latent representation to have values as close as possible to the discrete vectors in the neural codebook, is a reconstruction loss sub-function for computing a point-wise amplitude error and a waveform similarity error weighted sum between the reconstructed EEG sub-signals and the EEG sub-signals of the sample, is a time-frequency domain structure constraint loss sub-function for computing a weighted sum of an approximation coefficient structure loss and a detail coefficient structure loss during the reconstruction process.
5. The method of claim 4, wherein, The method further comprises training a parallel feature encoder, comprising: using the wavelet quantization neural marker and its neural codebook trained in S2 to obtain the final matched discrete vector of each EEG sub-signal of each sample; performing random mask processing on the EEG sub-signals in each sample to obtain a mask-containing sample containing partially masked EEG sub-signals; The parallel feature encoder is used to extract time, space and semantic features of each EEG sub-signal from the mask sample respectively, and the stacked Transformer block is used to fuse the time, space and semantic features to obtain the parallel features of each EEG sub-signal. According to the error between the prediction discrete vector corresponding to each EEG sub-signal and the final matched discrete vector, the parameters of the parallel feature encoder are updated.
6. The method of claim 5, wherein, The parallel feature encoder comprises: a time encoder, which adopts a pre-trained time encoder, is used to extract the time features of each EEG sub-signal from the input mask sample; a space encoder, which is used to maintain the local space encoding parameters corresponding to each brain region group according to a pre-defined brain region grouping rule, capture the spatial correlation features between the channels within the corresponding brain region group from the mask sample by using the local space encoding parameters corresponding to each brain region group, obtain the local space features corresponding to each EEG sub-signal, and perform pooling operation on all local space features of the same sample after attention mechanism-based processing to obtain global space features, and aggregate the local space features and the global space features corresponding to each EEG sub-signal to obtain the space features of the EEG sub-signal; a semantic encoder, which is used to extract the semantic features of each EEG sub-signal from the mask sample; and a plurality of stacked Transformer blocks, which are used to fuse the time, space and semantic features of each EEG sub-signal according to the attention mechanism to obtain the parallel features of each EEG sub-signal.
7. The method according to claim 5 or 6, characterized in that, The training of the parallel feature encoder further comprises: obtaining the text embedding from the clinical diagnosis description text corresponding to each EEG signal sample by using a pre-set text encoder; fusing the parallel features of each EEG sub-signal extracted from the mask sample corresponding to each EEG signal sample to obtain the EEG embedding; updating the parameters of the parallel feature encoder according to the second total loss function, wherein the second total loss function comprises a vector reconstruction sub-loss function for calculating the error between the prediction discrete vector corresponding to each EEG sub-signal and the final matched discrete vector, and an alignment loss sub-function for calculating the deviation between the text embedding and the EEG embedding.
8. A method of classifying an EEG signal, characterized by, The classification method comprises: obtaining the EEG signal of the target, inputting the EEG signal of the target into the wavelet quantization neural marker trained according to any one of claims 1-4 to extract the latent representation of each EEG sub-signal thereof, and matching the closest discrete vector from the neural codebook according to the latent representation; fusing the closest discrete vectors corresponding to all EEG sub-signals of the target EEG signal, inputting the fused discrete vectors into the pre-trained classifier, and determining the classification result corresponding to the target.
9. A method of classifying an EEG signal, characterized by, The classification method comprises: obtaining the EEG signal of the target, inputting the EEG signal of the target into the parallel feature encoder trained according to any one of claims 5-7 to extract the parallel features of each EEG sub-signal thereof; Parallel features of all EEG sub-signals of the target EEG signal are fused to obtain an EEG embedding, and the EEG embedding is input into a pre-trained classifier to determine a classification result corresponding to the target.
10. An electronic device, comprising: Comprise: one or more processors; and a memory, wherein the memory is configured to store executable instructions; the one or more processors are configured to implement the steps of the method of any one of claims 1-9 by executing the executable instructions.
Citation Information
Cited By
Electroencephalogram signal processing method and device and computer equipment
CN121465609A