Explanatable zero-sample visual EEG decoding method
By using the EEG-to-image alignment framework of variational autoencoding and the multi-layer sparse encoding algorithm to unfold the EEG encoder in the EEG decoding method, the problems of low signal-to-noise ratio, insufficient generalization ability and poor interpretability in EEG decoding are solved, and visual EEG decoding with high accuracy, stability and interpretability are achieved.
Patent Information
- Application Number
- CN202510347998.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
The existing EEG decoding methods have problems such as low signal-to-noise ratio, complex background noise, insufficient generalization ability and poor interpretability, making it difficult to accurately decode visual information and promote it in practical applications.
Using a decoding framework (E2IVAE) based on variational autoencoding and a new electroencephalogram encoder (ISTANet) based on multi-layer sparse encoding algorithm, visual perception information is extracted through cross-modal alignment of EEG signals to stimulating images, and the accuracy and stability of decoding are improved through ISTANet, giving the decoding framework interpretability.
It realizes the extraction of visual perception information from EEG signals, improves the accuracy and stability of decoding, and gives the decoding framework interpretability, solving the problems of low signal-to-noise ratio, insufficient generalization ability and poor interpretability in the prior art.
Smart Images

Figure CN120196869A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of brain-computer interfaces and relates to an interpretable zero-shot visual EEG decoding method. Background Art
[0002] In the fields of brain-computer interfaces, cognitive science, and neuroengineering, EEG, as a non-invasive brain activity monitoring technology, has important value in studying brain functions and developing BCI systems due to its high temporal resolution and relatively low cost. However, the low signal-to-noise ratio (SNR) and complex background noise of EEG signals pose great challenges to accurately decoding visual information from EEG. In addition, existing neural network models often suffer from insufficient generalization ability and interpretability when processing EEG data, which limits their popularization in practical applications.
[0003] Traditional EEG decoding methods mainly rely on supervised learning and require a large amount of labeled data to train the model. This not only increases the workload of data preparation but also makes it difficult to adapt to new categories and tasks. In recent years, unsupervised learning and multimodal learning methods have made remarkable progress in fields such as computer vision, providing new ideas for solving the data annotation problem and model generalization problem in EEG decoding. By leveraging the shared information in multimodal data, self-supervised alignment and feature extraction of EEG signals can be achieved, thereby reducing the dependence on labeled data and improving the generalization ability of the model.
[0004] In terms of model interpretability, deep learning models are usually regarded as "black boxes", and their internal decision-making processes are difficult to understand and interpret. This is an important obstacle to the application of BCI systems because users need to have sufficient trust in the system's output to use it safely. Therefore, developing interpretable EEG decoding methods not only helps improve the credibility of the model but also provides deeper insights for neuroscience research. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an interpretable zero-shot visual EEG decoding method. By constructing a decoding framework (E2IVAE) for EEG-to-image alignment based on variational autoencoders, cross-modal alignment of EEG signals to stimulus images is achieved, forcing the extraction of visual perception information from EEG, so that zero-shot neural decoding can be realized through downstream cross-modal EEG decoding. At the same time, a new EEG encoder (ISTANet) based on algorithm unfolding is used to improve the accuracy and stability of decoding and endow the decoding framework with interpretability, and it is judged whether useful features are extracted by intuitively analyzing the reconstructed EEG features.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An interpretable zero-shot visual EEG decoding method, comprising the following steps:
[0008] S1: Preprocess the bimodal dataset of images-evoked EEGs to obtain a dataset with N independent and identically distributed (i.i.d.) samples where each sample has a stimulus image X (i) and an evoked EEG Y (i) Two modalities. Subsequent steps will be performed on a single multimodal sample, so the superscript is omitted.
[0009] S2: Build an E2IVAE framework, which includes an image codec, an EEG encoder, and a neural decoder for cross-modal training and application. By minimizing the Kullback-Leibler divergence (KL divergence) between the representation distribution generated by the EEG encoder and the true posterior distribution of the image variational autoencoder (VAE), EEG is aligned with the image in the latent space. This alignment method can not only capture the shared information between the two modalities, but also train the neural decoder cross-modally to achieve zero-shot neural decoding.
[0010] S3: Use the ISTANet encoder as the EEG encoder of this framework. As a new type of multi-channel EEG encoder, ISTANet is based on the multi-layer sparse coding algorithm and transforms the multi-layer sparse coding algorithm into an end-to-end structure. It can extract features from multi-channel noisy EEG data while maintaining the interpretability of traditional machine learning, significantly improving the robustness and accuracy of feature extraction.
[0011] S4: Conduct post hoc interpretability analysis on the features learned by ISTANet. To verify whether the model has learned meaningful features, the present invention provides a visualization method. By calculating the overall reconstructed features and multi-scale atomic features, the time-domain, frequency-domain, and spatial feature extraction effects of the model on EEG signals are intuitively displayed to judge whether the model effectively captures the key information conforming to the biological mechanism. This not only helps to improve the credibility of the results, but also provides deeper insights for neuroscience research.
[0012] For S2, constructing the E2IVAE includes the following detailed steps:
[0013] S201: The variational posterior estimator of an image VAE, i.e., its encoder, is taken as the image encoder, which can derive the approximate posterior distribution of the image. Correspondingly, the image decoder is taken from the decoder of this VAE. For the image modality X, the VAE defines the marginal log-likelihood logp θ (X) = log ∫ z p θ (X|z)p(z)dz. Since accurately estimating logp θ (X) is usually intractable, the variational autoencoder estimates it by maximizing the tractable evidence lower bound (ELBO):
[0014]
[0015] By training the image modality VAE on the stimulus image dataset of the visible classes, the image encoder and the image decoder p θ (X|z) can be obtained.
[0016] S202: Align the representation distribution output by the EEG encoder with the posterior distribution of the image, which is achieved by minimizing the KL divergence between the representation distribution output by the EEG encoder and the image posterior distribution. This self-supervised method enables the model to learn the shared information between the two modalities, rather than the traditional supervised learning with predefined labels. The optimization objective can be expressed as:
[0017]
[0018] Therefore, the overall objective function is
[0019]
[0020] S203: The weight allocation between the reconstruction term and the KL loss term in the variational autoencoder optimization objective has a significant impact on the final optimization result. Therefore, (3) a weight coefficient needs to be added, and the final objective function is
[0021]
[0022] S204: After training the encoders of the two modalities, the latent features output by them (sampled from the distributions output by the encoders) can be used for cross-modal training and testing of the neural decoder on new class data. The neural decoder can select any classification model, and this classification model will be trained on the image latent features of the new class dataset and tested on the EEG latent features, thus completing the zero-shot neural decoding of the new class EEG.
[0023] For S2, constructing ISTANet includes the following detailed steps:
[0024] S301: Under visual stimulation conditions, the collected EEG signals contain response characteristics to the stimulation, such as event-related potential (ERP), as well as irrelevant noise (such as spontaneous potential, environmental noise, acquisition device noise, etc.), which can be expressed as follows:
[0025] Y = y + ∈. (7)
[0026] Where is the noise signal collected by the EEG acquisition device from the scalp of the subject, which is a flattened one-dimensional signal, y is the feature signal we are interested in, and ∈ is the interference noise.
[0027] S302: For the collected noise signal Y, apply the sparse coding algorithm to learn an over-complete dictionary D to sparsely represent the feature y therein, that is
[0028] Y = Dα + ∈. (8)
[0029] Where Dα = y, And is the sparse coding (or sparse representation), usually N >> (r × c).
[0030] S303: Usually ∈ follows a normal distribution, while α follows a Laplace distribution, then we can obtain the classic sparse coding model Lasso:
[0031]
[0032] Where λ is the balance parameter. The first term on the right side of the equation constrains the reconstruction error loss, and the second term constrains the sparsity of the sparse coding.
[0033] S304: Preferably, apply multi-layer sparse coding instead of single-layer sparse coding to extract deep non-linear features. For an L-layer sparse coding, the feature y will go through L times of sparse coding, and the relationship of each layer of sparse coding can be expressed as follows:
[0034]
[0035] Where l rec is the error loss function between Y and y, and δ is the error constraint threshold; l spar is the penalty function that constrains the sparsity of α i and ξ i is the sparse constraint threshold. is the convolutional dictionary associated with the sparse coding.
[0036] S305: As can be seen from (9), the error loss constraint and each layer of sparse coding α iThe sparsity constraints will all be converted into a constraint term in the final objective function. Therefore, the estimation of the last-layer sparse coding α L is transformed into:
[0037]
[0038] where the joint dictionary D (i,L) represents the dictionary product D i ·D i+1 ·…·D L . In this way, the multi-layer sparse coding model (11) is obtained, which can gradually encode Y into α L .
[0039] S306: For a non-smooth function h(β) = f(β) + g(β), where f(β) is a convex smooth function with Lipschitz constant 1 / c, and g(β) is a continuous but non-smooth function. Then the gradient mapping of h(β) can be obtained:
[0040]
[0041] where prox cg (·) is the proximal operator of g(·). At the same time, (12) is also the generalized gradient of h(β).
[0042] S307: Let h i (α L ) = f(D (2,L) α L ) + g1(D (2,L) α L ) + g2(D (3,L) α L ) + … + g L-i (D (L-i+1,L) α L ), where
[0043] g i (·) = λ i ||·||1, then F(α L ) in (11) is transformed into:
[0044] F(α L ) = h1(α L ) + g L (α L ). (13)
[0045] S308: Use the Generalized Proximal Gradient Algorithm (GPGD) to optimize (13) as follows:
[0046]
[0047] S309: Let (14) can be expressed as:
[0048]
[0049] S310: From (12), after replacing the ordinary gradient with the gradient mapping, it can be further expressed as
[0050]
[0051] S311: Let A nested optimization step is obtained:
[0052]
[0053] S312: Since c i is a constant, let Then we can obtain and The nested optimization step between them is:
[0054]
[0055] where is the well-known soft-threshold activation function.
[0056] S313: After calculating the sparse coding α L of the last layer, the mean and standard deviation of the latent distribution can be obtained through linear transformation:
[0057]
[0058] where is the weight matrix.
[0059] S314: Determine the number of layers and the number of iterations of the multi-layer sparse coding, and construct an end-to-end neural network through (18) and (19). The input of this neural network is the flattened EEG signal.
[0060] For S4, the post hoc interpretability analysis of the features learned by ISTANet includes the following detailed steps:
[0061] S401: Use a one-dimensional convolutional neural network to implement transposed convolutional dictionary multiplication
[0062] S402: For a two-dimensional input EEG, in order to capture the inter-channel information, we unfold the EEG in the channel direction (column unfolding).
[0063] S403: Set the size of the atoms of the first-layer convolutional dictionary (i.e., the kernel of the convolutional dictionary) to an integer multiple of the number of EEG channels, and the stride is also an integer multiple of the number of channels to extract spatial features.
[0064] S404: The kernels of subsequent convolutional dictionaries then focus on the extraction of temporal features, and set the corresponding kernel size according to the Nyquist sampling theorem to extract temporal / frequency features.
[0065] S405: Calculate and observe the atoms of the joint dictionary D (1,i) to analyze the rationality of multi-scale atomic features.
[0066] S406: Calculate the overall reconstruction features through y = D (1,i) α i Visualize the overall reconstruction features to confirm whether the features learned by ISTANet are interpretable rather than noise or pseudo-features.
[0067] The beneficial effects of the present invention are as follows:
[0068] 1. The EEG-to-image alignment model framework (E2IVAE) based on variational autoencoder provided by the embodiments of the present invention can extract intrinsic shape information from EEG data rather than relying solely on abstract features. This method breaks through the limitations of traditional EEG decoding techniques and realizes the alignment of EEG to images in the latent space by minimizing the KL divergence between the representation distribution generated by the EEG encoder and the posterior distribution of the images. This alignment method can not only capture the shared information between the two modalities but also train the neural decoder cross-modally to achieve zero-shot neural decoding.
[0069] 2. The ISTA Net EEG encoder with an end-to-end training mode based on the multi-layer sparse coding algorithm provided by the embodiments of the present invention combines data-driven deep learning and model-driven traditional machine learning methods, and can more effectively capture key information in noisy EEG data. This innovative architecture of fusion significantly enhances the performance and stability of EEG decoding, improves the adaptability of the model in complex environments, and at the same time endows the E2IVAE decoding framework with interpretability.
[0070] 3. In the embodiments of the present invention, a novel EEG feature visualization method based on ISTANet is developed, which can simultaneously display the multi-scale features and overall reconstruction features of EEG signals. This method not only supports the intuitive display of the features extracted by the model from the time domain, frequency domain, and spatial dimensions, but also can effectively judge whether the model captures the key information that conforms to the biological mechanism. Through the visual analysis of the extracted features, it can be verified whether the model has learned meaningful features that conform to the EEG decoding mechanism, thereby improving the interpretability of the model and the credibility of the results, and providing a more in-depth theoretical support and practical reference for neuroscience research.
[0071] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0073] Figure 1 Schematic diagram of the E2IVAE decoding framework proposed by the present invention;
[0074] Figure 2 Schematic diagram of implementing matrix-vector multiplication with a one-dimensional convolutional neural network in the EEG encoder of ISTANet proposed by the present invention;
[0075] Figure 3 Two-dimensional diagram of multi-scale features and average PSD visually generated in the embodiments of the present invention; Figure 3 (a) D1 atomic features from subject 1; Figure 3 (b) Partial D (1,2) and D (1,3) atomic features from subject 1; Figure 3 (c) Channel average power spectral density of all atoms of D (1,2) and D (1,3) from subject 1;
[0076] Figure 4 Globally reconstructed features visually generated in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0078] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as limitations on the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0079] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as limitations on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0080] 1. Model construction:
[0081] Build the E2IVAE model framework as described in S2. Among them, except that the EEG encoder uses the ISTANet built as described in S3, the image codec is implemented using a fully connected neural network. Figure 1 Shows the schematic diagram of the E2IVAE framework and the architecture of the EEG encoder ISTANet. Among them, the flame indicates that the model weights will be updated, and the snowflake indicates that the model is frozen. Figure 2 Shows a schematic diagram of implementing matrix-vector multiplication using a one-dimensional convolutional neural network in ISTANet when the vector is α0, and shows the relationship between the convolutional matrix and the dictionary.
[0082] Preferably, a logistic regression model is used as the downstream neural decoder (classifier) for the following reasons:
[0083] (1) Linear Separability: The latent variables used for training are low-dimensional representations with inherent linear separability. Therefore, it is feasible to use a linear classifier such as logistic regression without data normalization because the latent space approximates an isotropic standard normal distribution;
[0084] (2) Cross-modal Training and Testing: The classifier will be trained and tested on different modal features. Although the self-supervised learning of previous models has aligned the data of each modality in the latent space, inevitable differences still exist. Therefore, choosing a highly non-linear classifier (such as a neural network or an SVM with a non-linear kernel) may lead to severe overfitting;
[0085] (3) Decision Probability: Logistic regression predicts the probability of the final decision, thus allowing the calculation of the top-5 classification accuracy to more comprehensively evaluate the model performance.
[0086] 2. Training, Validation, and Testing Processes:
[0087] During the training process, the training set is randomly drawn into mini-batches in each epoch. For each mini-batch, after the data is input into the model for forward propagation, batch-size optimization functions will be obtained, as shown in S203, and then the model parameters are updated through backpropagation and the gradient descent algorithm. After completing the training of all epochs, the best model is determined according to the results of the validation process.
[0088] During the testing process, we obtain an approximate posterior for the new class images through the updated image encoder, and then sample from this distribution multiple times to obtain an augmented training dataset for training a neural decoder (logistic regression classification model). Similarly, the new class EEG also obtains the representation distribution of the images through the trained EEG encoder, and the mean (the sampling point with the highest probability) of this distribution is extracted as the test set for cross-modal testing of this classifier. In this way, zero-shot neural decoding can be achieved on new classes without the need for neural data of new classes to train the neural decoder.
[0089] The validation process is roughly the same as the testing process, except that only the mean of the approximate posterior distribution of the images is taken for the training set to obtain a stable validation result. During the testing process, since the training dataset of the neural decoder is obtained through random sampling, the results may be unstable. Therefore, the sampling-training-testing process is repeated 10 times, and the mean is taken as the final stable result.
[0090] The validation process is carried out after each epoch ends, and the testing process is carried out after all epochs end.
[0091] 3. Datasets and Preprocessing:
[0092] The ThingsEEG2 (Gifford et al., 2022) electroencephalogram (EEG)-image bimodal dataset was used in this example. ThingsEEG2 is a new and comprehensive dataset based on the Rapid Serial Visual Presentation (RSVP) paradigm. It contains data from ten participants, which were collected through the time-efficient RSVP paradigm. The training set includes 1,654 concepts, with 10 images for each concept and 4 repetitions for each image; the test set includes 200 concepts, with 1 image for each concept and 80 repetitions for each image. The training and test images appear in a pseudo-random order, and target images are used to reduce eye blinks and other artifacts. Each image is displayed for 100 milliseconds, followed by a 100-millisecond blank screen, and the stimulus presentation frequency is 5 Hz. The original EEG data is filtered to [0.1, 100] Hz, with 63 channels and a sampling rate of 1,000 Hz.
[0093] During the preprocessing, we segmented the EEG data into trials from 0 to 600 milliseconds after the start of the stimulus. Baseline correction was performed using the mean value of the data 200 milliseconds before the stimulus. The data of all electrodes were retained, and the sampling rate was reduced to 100 Hz. Multivariate noise normalization (Guggenmos et al., 2018) was performed using the training data. To ensure the signal-to-noise ratio, we averaged all EEG repetitions for each image and compared the effects of the number of repetitions on the test set. Since the ThingsEEG2 dataset already provides 3,000-dimensional visual features extracted by different pre-trained models, we used the first 1,000 principal components of the CORnet-S features extracted by CORnet-S (Kubilius et al., 2019) instead of re-extracting them.
[0094] 4. Experimental parameters:
[0095] This experiment was implemented on an NVIDIA RTX A4000 GPU based on PyTorch. In the experiment, the training data contains 16,540 samples of 1,654 classes. We randomly selected 1,000 samples of 100 classes from the training data as the validation set. The total number of epochs is 400, the batch size is 256, and a validation process is performed on the validation set once per epoch. During the training process, the model with the lowest validation loss is saved as the best model. We calculated the results of the test set once after training; in the experiment, the size of the second-layer atomic window in ISTANet is 25 sampling points, which can capture features above 4 Hz. The model parameter settings of each module in E2IVAE are shown in Table 1.
[0096] Table 1
[0097]
[0098] Each module of the model is named independently. The meanings of the parameters in the parentheses after different modules are as follows: One-dimensional convolutional neural network: (kernel length × number of input channels × number of kernels, stride, padding), fully connected neural network: (input size, output size).
[0099] In the validation stage, the latent encoding used to train and test the classifier is the mean of its distribution. In the test stage, the latent variables for training the classifier are obtained by sampling from its distribution 10 times, while the test latent variables remain the mean of its distribution. This process is repeated 10 times to obtain a stable average result.
[0100] For ISTANet, the number of unfolding iterations (K) is set to 4, and the number of sparse coding layers (L) is 3. Logistic regression (C = 0.01, maximum number of iterations = 10) is used as the classifier.
[0101] To prevent the vanishing gradient, before the data is input into the model, the image feature values are multiplied by 50.0, and the EEG amplitudes are multiplied by 2.0. The Adam optimizer is used with a learning rate of 0.00005, and the other parameters are default values.
[0102] Experimental results:
[0103] In the zero-sample visual decoding task on the ThingsEEG2 dataset, compared with the current state-of-the-art methods, our method achieved the best decoding accuracy. The specific results are shown in Table 2.
[0104] Table 2
[0105]
[0106] We visualized the multi-layer features of ISTANet trained on the data of Subject 1, as Figure 3 shown.
[0107] The atoms of the first convolutional dictionary, as described in S3 above, can be represented as spatial filters. By plotting its topological map, the brain regions with the largest response to visual stimuli can be explored. The components of the atoms in the second and third convolutional dictionaries contain rich frequency information. We averaged the two-dimensional maps of the power spectral density (PSD) of the atoms in the second or third layer to view the overall atomic frequency domain characteristics.
[0108] D1, D (1,2) , D (1,3)Atoms have different lengths over the time window and can exhibit features at different scales. The D1 atom only involves one sampling point and can be regarded as a spatial filter. It incorporates the idea of blind source separation, projects the measured signal onto the signal sources (multiple), and then further extracts features from the signal sources. Figure 3 (a) shows the D1 atom features from Subject 1. It can be observed that the active brain regions are mainly located in the posterior part of the brain, covering the brain regions for visual information processing. (1,2) and (1,3) atoms represent the spatio-temporal components of different time windows. For example, Figure 3 (b) shows the data of some atoms from Subject 1's (1,2) and (1,3) ; while Figure 3 (c) shows the channel-averaged power spectral density of all (1,2) and (1,3) atoms from Subject 1. It can be found that the atomic activities are spatially concentrated in the occipital, parietal, and temporal lobes in the posterior part of the brain, which is related to the primary processing of visual signals. The frequencies involve the δ, θ, α, and β bands, and the most active aggregation is in the θ and α frequency bands, which is related to the RSVP presentation frequency and its high-speed primary visual processing activities. At the same time, we find that the deeper dictionaries have higher resolution in frequency, enabling these atoms to construct more discriminative features, demonstrating the superiority of the multi-layer dictionary.
[0109] In addition, we can also visualize the overall reconstruction features to explore whether the model has learned meaningful features. Figure 4 shows the overall reconstruction features of some subjects to explore the phenomenon of individual differences.
[0110] In these reconstruction features, we can clearly observe that our framework captures the ERP in the occipital lobe region, and at the same time, for the noise in other channels, E2IVAE can filter it out, including the low-frequency noise in the prefrontal lobe. The individual differences among subjects can be intuitively observed through the reconstruction features, including waveforms and activation regions, etc. These individual differences may come from physiological factors, the configuration of EEG acquisition devices, the environment, etc.
[0111] These visualization results fully demonstrate the ability of the proposed model in capturing spatio-temporal features and spectral features and its biological rationality, while revealing the existing phenomenon of individual differences.
[0112] The framework proposed by this method has achieved excellent decoding accuracy and interpretable effects in the zero-shot visual decoding task of the Things-EEG2 dataset, providing new ideas for scenarios such as cognitive science research, auxiliary device control, and virtual reality interaction.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An interpretable zero-shot visual EEG decoding method, characterized by: The method includes the following steps: S1: Preprocess the image-evoked EEG bimodal dataset to obtain a dataset with N independent and identically distributed (iid) samples Each sample has two modalities: stimulus image X and evoked EEG Y; S2: Build the E2IVAE framework, which includes an image codec, an EEG encoder, and a neural decoder for cross-modal training and application; align EEG to image in latent space by minimizing the Kullback-Leibler Divergence (KL Divergence) between the representation distribution generated by the EEG encoder and the true posterior distribution of the image variational autoencoder (VAE); S3: Use the ISTANet encoder as the EEG encoder of this framework, where ISTANet is based on the multi-layer sparse coding algorithm, converting the multi-layer sparse coding algorithm into an end-to-end structure; while maintaining the traditional machine learning interpretability, it extracts features from multi-channel noisy EEG data; S4: Perform post-hoc interpretability analysis on the features learned by ISTANet; provide a visualization method to verify whether the model has learned meaningful features. By calculating the overall reconstruction features and multi-scale atomic features, the model’s time domain, frequency domain, and spatial feature extraction effects on the EEG signal are intuitively displayed to determine whether the model effectively captures key information that conforms to the biological mechanism.
2. The interpretable zero-sample visual EEG decoding method according to claim 1, characterized in that: In S2, constructing E2IVAE includes the following steps: S201: The image encoder is taken from the variational posterior estimator of an image VAE, that is, its encoder, and the approximate posterior distribution of the image is derived; the image decoder is correspondingly taken from the decoder of the VAE; for the image modality X, the VAE defines the marginal log-likelihood logp through the latent variable z θ (X) = log∫ z p θ (X|z)p(z)dz, due to accurate estimation of log p θ (X) is usually intractable, and the variational autoencoder estimates it by maximizing the tractable evidence lower bound ELBO: The image encoder is obtained by training the image modality VAE on a stimulus image dataset of the visible class. and image decoder p θ (X|z); S202: EEG encoder The output representation distribution is aligned with the posterior distribution of the image, which is achieved by minimizing the KL divergence between the representation distribution output by the EEG encoder and the true posterior distribution of the image; this self-supervised method enables the model to learn the shared information between the two modalities, rather than traditional supervised learning with predefined labels. The optimization objective is expressed as: The overall objective function is: S203: The weight distribution between the reconstruction term and the KL loss term in the variational autoencoder optimization objective has a significant impact on the final optimization result. Therefore, (3) needs to add a weight coefficient, and the final objective function is: S204: After the encoders of the two modalities are trained, the latent features of their outputs are sampled from the distribution of the encoder outputs and used for cross-modal training and testing of the neural decoder on the new class of data; the neural decoder selects an arbitrary classification model, which is trained on the image latent features of the new class of data sets and tested on the EEG latent features to complete zero-sample neural decoding of the new class of EEG.
3. The interpretable zero-sample visual EEG decoding method according to claim 2, characterized in that: In S3, building ISTANet includes the following steps: S301: Under the condition of visual stimulation, the collected EEG signal contains the response characteristics to the stimulation, including event-related potential (ERP), and irrelevant noise, as shown below: Y=y+∈ (7) in is the noise signal collected from the subject's scalp by the EEG acquisition device, which is a flattened one-dimensional signal. y is the characteristic signal we are concerned about, and ∈ is the interference noise. S302: For the collected noise signal Y, a sparse coding algorithm is applied to learn an overcomplete dictionary D to sparsely represent the feature y therein, that is, Y=Dα+∈ (8) Where Dα=y, as well as is sparse coding, N>>(r×c); S303: ∈ follows the normal distribution, α follows the Laplace distribution, and the classic sparse coding model Lasso is obtained: Where λ is a balancing parameter; the first term on the right side of the equation constrains the reconstruction error loss, and the second term constrains the sparsity of sparse coding; S304: Apply multi-layer sparse coding instead of single-layer sparse coding to extract deep nonlinear features. For an L-layer sparse coding, feature y will undergo L sparse coding, and each layer of sparse coding The relationship is expressed as follows: Among them l rec is the error loss function between Y and y, δ is the error constraint threshold; l spar is the constraint α i The penalty function for sparsity, ξ i is the sparse constraint threshold; is the convolution dictionary associated with sparse coding; S305: Error loss constraint and sparse coding α for each layer i The sparsity constraints of will be converted into a constraint term in the final objective function, then the last layer of sparse coding α L The estimated conversion is: where λ i represents the balance parameter; obtain the multi-layer sparse coding model (11), and gradually encode Y into α L ; S306: Algorithm Unrolling converts the fixed iterative solution algorithm of multi-layer sparse coding into an end-to-end neural network, fully utilizing the data-driven characteristics of the neural network, automatically updating parameters and learning discriminative features; S307: Through the algorithm expansion method of the multi-layer sparse coding model, formula (11) is obtained and The nested optimization steps between are: in It is a well-known soft threshold activation function; k is related to the number of iterations; S308: Calculate the sparse coding α of the last layer L After that, the mean and variance of the potential distribution are obtained by linear transformation: in is the weight matrix; S309: Determine the number of layers and iterations of multi-layer sparse coding, and construct an end-to-end neural network through (12) and (13); the input of the neural network is the flattened EEG signal, which is ISTANet.
4. The interpretable zero-sample visual EEG decoding method according to claim 3, characterized in that: In S4, the post-hoc interpretability analysis of the features learned by ISTANet includes the following steps: S401: Implementing transposed convolution dictionary multiplication using a one-dimensional convolutional neural network S402: for a two-dimensional input EEG, in order to capture inter-channel information, the EEG is expanded from the channel direction; S403: setting the size of atoms of the first layer convolution dictionary to an integer multiple of the number of EEG channels and the stride to an integer multiple of the number of channels to extract spatial features; S404: The kernel of the subsequent convolution dictionary focuses on the extraction of time features, and the corresponding kernel size is set according to the Nyquist sampling theorem to extract time / frequency features; S405: through y=D (1,i) α i Calculate the overall reconstruction feature, where D (1,i) Represents the dictionary product D1×D2×…×D i ; By visualizing the overall reconstruction features, confirm whether the features learned by ISTANet are interpretable, rather than noise or pseudo features S406: Calculate and observe the joint dictionary D (1,i) atoms to analyze the plausibility of multi-scale atomic features.