A method for reconstructing human visual information based on electroencephalogram signals

Through the time-frequency domain encoder and IP-Adapter embedding alignment combined with the cascade diffusion model, the problem of large parameters and high computing resource consumption in visual stimulus reconstruction is solved, and efficient and robust visual information reconstruction is achieved, which is suitable for the visual reconstruction task of EEG signals.

CN119832551BActive Publication Date: 2025-07-18CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411902988.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-07-18
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The existing visual stimulus reconstruction methods have made certain progress on the coarse-grained scale, but they cannot accurately restore the details of the visual stimulus, resulting in the reconstruction of the image losing a large amount of important semantic information, and requires a large amount of computing resources. The fMRI equipment is high and the equipment requirements are strict. The EEG fine-tuning model is easy to overfit and difficult to apply in resource-constrained environments.

Method used

The EEG sequence features are extracted using a time-frequency domain encoder, the IP-Adapter model is used for embedding alignment, and visual reconstruction is carried out in combination with the cascaded diffusion model, which reduces the model parameters and improves classification accuracy. High-quality images are generated through the improved frequency domain encoder and cascaded diffusion model.

Benefits of technology

With fewer parameters, better classification accuracy and image reconstruction quality are achieved, and calculation costs are reduced, suitable for resource-constrained environments, and the efficiency and robustness of visual reconstruction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832551B_ABST
    Figure CN119832551B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for reconstructing human visual information based on electroencephalogram (EEG) signals, which is applied to the field of artificial intelligence. In view of the problem that although existing visual reconstruction tasks have made certain progress at a coarse-grained scale, these methods often fail to accurately restore the details of visual stimuli, resulting in a large amount of important semantic information being lost in the reconstructed images. The present invention adopts a new embedding alignment method to reduce the number of model parameters and enable the EEG signal features to better represent the fine-grained information of visual stimuli. At the same time, a better frequency-domain encoder is used to improve the classification accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a human visual information reconstruction technology. Background Art

[0002] The human brain is an extremely complex and mysterious biological structure. It consists of approximately 86 billion neurons, which are interconnected through trillions of synapses to form a highly complex network. This complex network structure endows humans with a series of advanced cognitive functions, including learning, memory, emotion, language, decision-making, and consciousness. The processing ability of the brain is extremely powerful, capable of processing and integrating a large amount of information from the senses in an extremely short time and generating refined responses. Brain decoding focuses on interpreting neural activity patterns to understand cognitive and sensory processes and attempts to reverse-engineer these processes. In-depth research on brain functions can not only improve the understanding of humans themselves in the present invention but also promote the progress of multiple scientific fields, especially in the fields of medicine and artificial intelligence.

[0003] The research on the human visual system is an important branch of brain decoding. Although the rapid development of deep learning in recent years has greatly improved the performance of computer vision systems, the human brain still has unparalleled efficiency and robustness in processing complex visual information. The encoding and decoding of visual information are two key links in the human brain's processing of visual data. The encoding process involves converting external visual stimuli, such as light and color, into neural signals in the brain, while decoding is to reconstruct the original visual image from these neural signals. Exploring how the human brain performs these complex visual tasks and attempting to simulate these processes are crucial for enhancing the intelligence level of machine vision systems. By constructing a model that mimics the visual processing ability of the human brain, the present invention can enhance the perceptual intelligence of machines, making them perform more efficiently and accurately in image recognition, scene understanding, and other artificial intelligence applications related to vision. In the field of visual decoding, researchers use neuroimaging techniques to measure the brain's neural responses to visual stimuli to reconstruct visual experiences. A common visual decoding method is to map brain activity to the latent space of a generative model, and then the generative model reconstructs the visual stimuli.

[0004] The visual reconstruction task has made certain progress at a coarse-grained scale, but these methods often cannot accurately restore the details of visual stimuli, resulting in a large amount of important semantic information being lost in the reconstructed images. This loss of information limits the understanding of how the brain processes complex visual scenes in the present invention and also affects the quality and practicality of the reconstructed images. In addition, these methods usually require a large amount of computing resources, which may be an insurmountable obstacle for some research teams.

[0005] The existing visual stimulus reconstruction methods include the following two categories:

[0006] A. Reconstructing Visual Stimuli from Brain Signals

[0007] A common method for decoding the brain's visual process is to reconstruct the visual stimuli received by the human brain. Most previous studies have been limited to classifying visual stimuli and have not truly reconstructed them. In the field of image generation, generating images using feature vectors has become a cutting-edge trend in research. Technologies such as variational autoencoders (VAEs) and generative adversarial networks (GANs) have demonstrated the ability to transform from abstract features to specific images. These methods can generate new images that are visually indistinguishable from real images by learning the underlying representation of the data. The development of these technologies provides a basis for visual stimulus reconstruction. In early visual reconstruction tasks, these two models were mostly used as generative models. However, the VAE has poor generation quality and is prone to overfitting when the dataset is small. Although the GAN has high generation quality, its training process is unstable and it is extremely difficult to find the global optimal solution. The performance of the generative model restricts the development of the task of decoding visual stimuli, and there are still challenges in achieving high-fidelity, lightweight models and comprehensive subject visual brain decoding.

[0008] B. Reconstructing Visual Stimuli Using Diffusion Models

[0009] The progress of image generation technology in recent years has greatly improved the quality of generated images. As an emerging image generation technology, diffusion models (Diffusion Models) have received extensive attention due to their unique advantages. Diffusion models are based on Markov chains and Bayesian principles. Starting from a simple noise distribution, they gradually introduce structural information and finally generate complex images that match the target distribution. In the context of conditional generation, diffusion models can accept additional conditional information, such as text descriptions, class labels, or other data features, to guide the generation process and thus produce images that are consistent with these conditions. The advantage of this method lies in the smoothness and controllability of its generation process, making the generated images not only visually realistic but also closely related to the given conditions in terms of content.

[0010] Applying the conditional generation principle of diffusion models to brain signals, especially EEG and fMRI features, opens up new possibilities for neuroscience research and brain-computer interface technology. EEG and fMRI, as key technologies for capturing brain activities, provide rich neurobiological data that reflect the dynamic changes of the brain under different cognitive tasks and emotional states. By taking these signals as conditional inputs into diffusion models, researchers can explore generating visual images that match specific neural activity patterns, which not only helps better understand how the brain processes and encodes information but also provides new ideas for developing personalized neurorehabilitation and cognitive training applications. (Chen et al., 2022) proposed MinD-vis, which demonstrated the feasibility of using diffusion models for visual signal reconstruction. Specifically, MinD-vis first converted fMRI data into sparse coding features with locality constraints by sparsely coding masked brain modeling to effectively learn brain features, and then used it as a conditional input to fine-tune a pre-trained latent diffusion model, finally completing the reconstruction of visual stimuli. (Liu et al., 2023) proposed rainCLIP, a task-agnostic brain decoding model based on fMRI that can reconstruct visual stimuli with high semantic fidelity, bridging the gap between brain activities, images, and texts. (Ozcelik et al., 2023) proposed a two-stage scene reconstruction framework named Brain Diffuser for reconstructing natural images based on fMRI signals. (Sun et al., 2023) proposed a contrast and diffusion framework to decode real images in fMRI recordings, thus providing insights into the relationship between diffusion models and the human brain visual system. (Han et al., 2024) proposed MindFormer, a model specifically designed to generate fMRI-conditioned feature vectors, which can be used for conditional generation of stable diffusion models. However, the cost of fMRI is relatively high, the equipment purchase and maintenance costs are expensive, and fMRI equipment requires special environmental conditions, such as a room with anti-magnetic field interference, which limits its availability in some research and clinical settings.

[0011] EEG is a method for measuring brain activities. It records the electrical signals generated by brain neurons by placing a series of electrodes on the scalp. These electrical signals reflect the electrical activities in the brain and can be used to study brain functions, diagnose nervous system diseases, and develop brain-computer interfaces, etc. Compared with fMRI, EEG has a lower cost and the acquisition equipment is more portable. Therefore, using EEG to reconstruct brain visual stimuli shows great potential. DreamDiffusion (Bai et al., 2023) uses CLIP to provide additional supervision, which enables the alignment of EEG, text, and image embedding spaces even with limited EEG-image pair data.

[0012] Although the above methods have achieved certain results, they are all based on fine-tuning large models to achieve visual stimulus reconstruction. The disadvantages of this method are obvious. First, it eliminates the original ability of the model to generate images using text, and this fine-tuning usually requires a large amount of computing resources. Second, the fine-tuned models are usually not reusable because the EEG prompting function cannot be directly transferred to other custom models derived from the same text-to-image base model. In addition, the new models are usually incompatible with existing structure control tools, which poses significant challenges to downstream applications. Finally, the pre-trained diffusion models have been maturely trained on a large number of datasets, and fine-tuning them with scarce EEG-image data is likely to lead to overfitting. Summary of the Invention

[0013] To solve the above problems, the present invention proposes a method for reconstructing human visual information according to EEG signals, which can achieve better classification accuracy and image reconstruction quality while having fewer model parameters.

[0014] The technical solution adopted by the present invention is: a method for reconstructing human visual information according to EEG signals, including:

[0015] S0. Obtain an image of the scene observed by the human eye at the same moment, and collect the corresponding EEG sequence of the human body at this moment;

[0016] S1. Use a time-frequency domain encoder to extract the time-domain features and frequency-domain features of the EEG sequence;

[0017] S2. Input the concatenation result of the time-domain features and frequency-domain features of the EEG sequence into a semantic classifier to obtain an EEG classification label;

[0018] S3. Use a pre-trained IP-Adapter model to generate an image feature embedding based on the image of the scene observed by the human eye obtained in step S0;

[0019] S4. Based on the image feature embedding obtained in step S3, perform alignment processing on the concatenation result of the time-domain features and frequency-domain features of the EEG sequence extracted in step S1;

[0020] S5. Use the concatenation result of the time-domain features and frequency-domain features of the aligned EEG sequence and the EEG classification label together as the conditions of a cascaded diffusion model for multi-level semantic visual reconstruction.

[0021] Advantages of the present invention: The method of the present invention achieves better classification accuracy and image reconstruction quality while having fewer model parameters. Specifically, the present invention adopts a new embedding alignment method to reduce the model parameters and enable the EEG signal features to better represent the fine-grained information of visual stimuli. At the same time, a better frequency-domain encoder is used to improve the classification accuracy of the model. The method of the present invention has the following advantages:

[0022] 1. By using IP-adapter embedding to replace CLIP embedding, the feature dimension is reduced, thus reducing the computational cost;

[0023] 2. By adopting an improved frequency-domain encoder, the performance of EEG signal classification is improved, thus enhancing the quality of coarse-grained visual reconstruction;

[0024] 3. To verify the method of the present invention, the present invention uses the ImageNetEEG dataset for qualitative and quantitative experiments, and the results show the superiority of the method of the present invention. Description of the Drawings

[0025] Figure 1 is a flowchart of the present invention;

[0026] Figure 2 is a qualitative comparison of the reconstruction results between the method of the present invention and the prior art;

[0027] Figure 3 is a quality comparison of the reconstruction results between the method of the present invention and the prior art. Detailed Embodiments

[0028] To facilitate the understanding of the technical content of the present invention by those skilled in the art, the content of the present invention will be further explained below with reference to the accompanying drawings.

[0029] The present invention proposes a new research method, which directly aligns the brain features with the generation condition embedding of the diffusion model. By directly aligning the EEG time-frequency domain features with the coarse-grained and fine-grained semantic features of visual stimuli, the reconstruction of visual stimuli is completed, greatly simplifying the scalability and efficiency of the visual reconstruction task.

[0030] Such as Figure 1As shown in the figure, the present invention consists of three parts: an EEG time-frequency domain encoder, an alignment network, and a cascaded diffusion model. The time-domain encoder extracts time-domain features by reconstructing masked signals, and the frequency encoder uses Conv-LSTM to obtain frequency-domain features through the Fast Fourier Transform (FFT). The time-frequency domain features are concatenated and then input into the alignment network. The alignment network is globally fine-tuned under the supervision of IP-Adapter image embeddings, which enables the fine-grained semantic alignment of EEG features with the IP-Adapter space. Finally, the aligned EEG fine-grained embeddings and the coarse-grained embeddings of the classification results are used as the conditions for the cascaded diffusion model for multi-level semantic visual reconstruction.

[0031] The implementation processes of the three parts are described in detail as follows:

[0032] 1. EEG Time-Frequency Domain Encoder

[0033] Time-domain encoder: Masked pre-training is a self-supervised learning method that trains the model by predicting the masked parts of the input, enabling it to learn rich context representations.

[0034] The present invention uses the same time-domain encoder structure as DreamDiffusion. Specifically, given an input EEG sequence x with a shape of (B, C, L), where B is the batch size, C is the number of signal channels, and L is the signal sequence length, it is divided into N signal blocks in the time domain, and each signal block is mapped into a D-dimensional space to obtain Z ∈ R N×D features. The mapping is completed using one-dimensional convolution, and the operation is as follows:

[0035] x′ = proj(x)

[0036] Both the convolution kernel and the stride size of the one-dimensional convolution are equal to the signal block size K. The result is a new tensor x′ with a shape of (B, D, P), where

[0037] Then, x′ is transposed to place the features in the last dimension to obtain:

[0038] Z = transpose(x′)

[0039] Subsequently, a part of the signal blocks is randomly masked. For a given input sequence z, the algorithm steps for randomly masking the sequence according to the mask ratio mask_ratio are as follows:

[0040]

[0041] Finally, the MAE (Masked Autoencoder) architecture is used to predict the masked signal blocks based on the context clues of the signal blocks. Subsequently, these embeddings are fed into multiple VIT (Vision Transformer) blocks to effectively capture the complex dependencies of the EEG signals and enhance the ability of feature representation. It utilizes the self-attention mechanism to obtain the interaction information between different parts of the input data, and its core calculation can be expressed as:

[0042]

[0043] where Q = ZW q 、K = ZW k 、V = ZW V are the query, key, and value matrices from the EEG signal block features. Z is the feature of the EEG signal block, and W q ,W k ,W v are learnable weight matrices. d is the scalar for dimensional normalization in the softmax function.

[0044] By reconstructing the masked signals, the pre-trained time-domain encoder can gain in-depth understanding of the EEG features of different populations and various brain activities, thus completing the extraction of EEG features.

[0045] The loss function of the reconstructed signal is as follows:

[0046]

[0047] where: B represents the batch size of the input samples, N represents the number of signal blocks, p ij represents the reconstructed value of the j-th signal block of the i-th sample, that is, the output result of the above encoder; t ij represents the actual value of the j-th signal block of the i-th sample, and m ij represents the mask value of the j-th signal block of the i-th sample, and m ij = 0 indicates that this signal block is not masked, and m ij = 1 indicates that this signal block is masked. The samples here are the EEG signal sequences mentioned above.

[0048] Frequency Domain Encoder: In the prior art, LSTM is used as the extraction network for frequency domain features. However, a major drawback of LSTM is that it does not consider the spatial dependence of frequency domain signals. Therefore, in the present invention, Conv-LSTM is used as the extraction network for frequency domain features. Conv-LSTM overcomes the limitations of traditional RNN and CNN by combining CNN and LSTM. Conv-LSTM introduces convolutional operations in space and recurrent operations in time, while retaining the memory units and gating mechanisms in LSTM, enabling it to capture features in both time and space and having better generalization ability. Given the input tensor X t and the previous state (H t-1 , C t-1 ), the core calculation of the ConvLSTM cell can be described as:

[0049] combined = concat(X t , H t-1 )

[0050] combined_conv = Convld(combined)

[0051] (I t , F t , O t , G t ) = split(combined_conv, 4)

[0052]

[0053]

[0054] where σ represents the sigmoid activation function, which is used to map real numbers to the range between 0 and 1, that is, I t , F t , O t , G t represent real numbers, represents mapping I t , F t , O t , G t to values between 0 and 1 respectively; tanh represents the hyperbolic tangent activation function, Convld represents the one-dimensional convolutional operation, and split represents splitting the convolutional output into four parts. During the training process, in the present invention, referring to the process of BrainVis, the visual stimulus labels of EEG are used as supervision, and CE (CrossEntropy, cross-entropy loss) is used as the loss function for training. Regarding the description of the visual stimulus labels of EEG, for example, for the EEG signal induced by seeing a dog image, the label of this segment of the label is dog.

[0055] Semantic Classifier: After the time-domain encoder and frequency-domain encoder are trained, the present invention concatenates the time-domain features and frequency-domain features, and then uses the labels of visual stimuli as supervision to train the semantic classifier. The classifier is a multi-layer perceptron (MP) stacked by two linear layers. The trained semantic classifier can obtain the coarse-grained semantic information of the class labels, and the loss function used in the training process is the CE cross-entropy loss.

[0056] 2. Alignment Network

[0057] Different from previous work that uses CLIP embeddings as alignment guidance, the work of the present invention uses the pre-trained IP-Adapter model to generate image feature embeddings as alignment guidance. The core idea of IP-Adapter revolves around its innovative decoupled cross-attention mechanism. Specifically, instead of using a single cross-attention layer to process text and image features simultaneously, IP-Adapter introduces a cross-attention layer dedicated to image features. Given the image feature f i , the output of the new cross-attention Z ip is calculated as follows:

[0058]

[0059] where Q = ZW q , K' = f i W' k , V' = f i W' v are the query, key, and value matrices from the image features; f i represents the image features. W' k and W' v are the corresponding weight matrices. d is a scalar used for dimensional normalization in the softmax function.

[0060] Those skilled in the art should note that the pre-trained IP-Adapter model is a known prior art, and the present invention does not elaborate on its specific training process. One can refer to: Ye, Hu, et al. "Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models." arXiv preprint arXiv:2308.06721 (2023).

[0061] In the present invention, the pre-trained IP-Adapter model takes the actual images seen by the human body of the collected EEG signals as input to generate image feature embeddings.

[0062] This separation enables the model to focus on learning more detailed specific image features, thereby enhancing its ability to capture the unique features of visual data, and solves the problem that DreamDiffusion cannot reconstruct fine-grained semantics when using CLIP image feature embeddings to fine-tune large pre-trained models. This training ensures that the neural representation closely matches the visual feature embeddings of IP-Adapter, thus contributing to the reconstruction of accurate and reliable images from EEG signals. In addition, the dimension of the IP-Adapter image feature embeddings used in this work is only 4 * 768, which is more compact than the 77 * 768 dimension of the CLIP text feature embeddings used by BrainVis, greatly reducing the number of parameters of the alignment network and alleviating the requirements for the device during the training process.

[0063] The present invention follows the training process of BrainVis and uses a network composed of fully connected layers with residual connections to achieve the alignment of embeddings, as Figure 1 shown. Specifically, this network includes two weight layers, where the weight layer is a fully connected layer, and the two fully connected layers are connected through a residual connection. Different from this, the present invention does not need to align the class label embeddings and the additionally generated text description embeddings simultaneously, but only needs to align the image feature embeddings. Maximize the cosine similarity between the EEG feature embeddings f eeg aligned and the image feature embeddings f img generated by the pre-trained IP-Adapter model. The loss function is defined as follows:

[0064]

[0065] Figure 1 The rectified linear unit function in

[0066] 3. Cascade diffusion model

[0067] It may be difficult to generate the desired image using an exact text prompt because it may not always capture the content to be expressed. Therefore, adding an image beside the text prompt helps the model better understand what it should generate and can produce more accurate results. On the contrary, although the work of the present invention uses IP-Adapter image embeddings as the guidance for alignment, due to the impossibility of the alignment process being completely accurate, the reconstruction process will inevitably cause the loss and noise addition of semantic information. To solve this problem, the present invention uses a cascade diffusion model for the image reconstruction task. Specifically:

[0068] The present invention first performs reverse diffusion in the latent space conditioned on the aligned EEG time-frequency domain embeddings.

[0069] The embedding is updated at each step of the iteration and injected into the current image state to affect the denoising direction and result. Finally, after multiple iterations of denoising, the latent representation generated by the model is decoded into an image that is both clear and conforms to the semantic information contained in the EEG time-frequency domain embedding.

[0070] However, due to semantic noise and information loss in the embedding, the quality of the reconstructed image is not high. Therefore, the present invention uses explicit EEG classification labels to perform secondary reconstruction on the reconstruction results of the EEG time-frequency domain embedding to improve the latent image that has not been fully denoised in the reverse diffusion process. This can reduce semantic noise and obtain higher-quality results.

[0071] The following combines specific data to illustrate the technical effects of the present invention:

[0072] Dataset: The dataset used in the present invention contains EEG recordings of six subjects (five males and one female) while viewing object images. In the experiment, 40 different and easily distinguishable categories from ImageNet were used. Easily distinguishable means that each category has obvious differences from other categories (such as German shepherd - parachute), rather than different subcategories of the same object (German shepherd - Siberian husky). Each category contains 50 images, for a total of 2000 images as visual stimuli. In each experiment, different images of one category were shown one by one, each shown for 0.5 seconds, and there was a 10-second black screen time after each category was shown to reset the visual pathway. The entire experimental process lasted for 1400 seconds, that is, about 23 minutes and 20 seconds. After the experiment, 536 records were discarded due to being too short or having too large variations, and finally 11,466 128-channel EEG sequences were obtained. To prevent interference from previously shown images, the present invention intercepted the data from 40 - 460 milliseconds of each sequence for the experiment and divided the data into a training set, a test set, and a validation set according to a ratio of 8:1:1.

[0073] Quantitative values: The present invention refers to the parameter settings of previous works (Bai et al., 2023 and Fu et al., 2024). For EEG signals, the present invention uses the signals filtered at 5 - 95 HZ, and the data length is uniformly set to 512 through interpolation. The masking rate of the time-domain encoder is 75%, and 500 rounds of training are performed. The frequency-domain encoder uses a 3-layer conv-LSTM for 1000 rounds of training, then the time-frequency domain features are aligned for 500 rounds, the linear classifier is trained for 50 rounds, and finally the stable diffusion model version 1.5 is used for image generation. Those skilled in the art should note that the version 1.5 here is only the version number of the stable diffusion model in this embodiment. In actual applications, the stable diffusion model obtained from the final training can be used for image generation.

[0074] Evaluation metrics

[0075] The present invention uses a variety of metrics for evaluation to measure the generation quality of the model from various perspectives:

[0076] EEG Classification Accuracy (CA): It is used to evaluate the precision of EEG classification, which shows the degree of correlation between EEG features and the coarse-grained sub-semantics of visual stimuli.

[0077] 50-way Top-1 Classification Accuracy (GA): It is used to evaluate the semantic accuracy of the generated images. In the present invention, the generated images and the actual images are fed into a classifier pre-trained on the ImageNet1K dataset, and then it is checked whether the top-1 predicted classes of the generated images and the actual images match among 50 predefined classes. If the classification labels of the generated images and the actual images are consistent, it can be considered that images conforming to the semantics of the actual images are generated.

[0078] Initial Score (IS): It is used to evaluate the diversity of the generated images by calculating the entropy of the class distribution of the generated images, and at the same time, the average entropy of the generated images is compared with that of the real images to evaluate the authenticity of the generated images.

[0079] Fréchet Inception Distance (FID): It is used to evaluate the perceptual difference between the generated images and the real images. One advantage of FID is that it does not require pairwise comparison of the generated images, which makes it more efficient than some pixel-based metrics when evaluating a large number of images. In addition, FID is also considered to be a more human-vision-perception-compliant evaluation metric because it focuses on the high-level semantic features of images rather than the low-level pixel features.

[0080] Results and Comparison with Other Methods

[0081] Most of the metrics used in previous image reconstruction work for evaluating generation quality only adopted GA or the initial score and could not comprehensively evaluate the quality of the reconstructed images. The work of BrainVis evaluated the reconstruction effect from two aspects of image reconstruction and EEG classification, solving this problem. The work of the present invention further follows each evaluation metric used by BrainVis and uses a new metric to evaluate the generation quality. The baselines for image reconstruction include DreamDiffusion, Brain2Image, NeuroVision, DCLS-GAN, ESG-ADA, and MB2C; the baselines for EEG classification include EEGCN, Brain2Image, S-EEGNet, and KDSTFT.

[0082] Quantitative Results

[0083] The present invention uses a trained model to perform four image reconstructions for each EEG sequence of each subject to conduct reconstruction quality inspection. For the image reconstruction quality, the results in Table 1 show that the BrainAdapter of the present invention achieves the best performance in terms of two metrics, GA and FID, being 2.4% and 4.1% higher than the existing best methods respectively. This indicates that the overall image reconstruction quality of the present invention is the best and can better capture the semantic information contained in visual data. The results in Table 2 show that the BrainAdapter of the present invention is 10.15%, 4.45%, 0.61% and 10.32% higher than the existing best methods in terms of Top1, Top3, Top5 CA and F1 score respectively. This indicates that the EEG encoding method proposed by the present invention has the best electroencephalogram representation ability. In addition, as shown in Table 3, the method of the present invention only uses 188M parameters, far lower than 297M parameters of reamDiffusion and 518M parameters of BrainVis. This indicates that the method of the present invention helps to reduce the energy consumption during model training and inference, which is particularly important for application scenarios that require deploying models in resource-constrained environments.

[0084] Table 1 Comparison of the effects of the method of the present invention and the prior art in terms of three metrics, GA, IS, and FID

[0085]

[0086] Table 2 Comparison of the effects of the method of the present invention and the prior art in terms of four metrics, Top1, Top3, Top5 CA and F1 score

[0087]

[0088] Table 3 Comparison of the number of model parameters

[0089]

[0090] Qualitative results

[0091] As Figure 2As shown, the image reconstruction results of the present invention significantly demonstrate the ability to finely capture GT, whether in terms of pose, color, or other details. The fine-grained reconstruction method of the present invention successfully captures the deep semantics of GT, which benefits from the advanced feature extraction ability of the model of the present invention. This ability makes the reconstructed image visually very similar to the original image, retaining key visual information and structural integrity. In addition, the reconstruction results of the present invention also perform excellently in terms of color accuracy and texture details, which are key factors in image realism. In the coarse-grained reconstruction stage, the present invention utilizes the powerful generation ability of the diffusion model to supplement and improve the results of the fine-grained reconstruction. This two-stage method not only improves the efficiency of reconstruction but also enhances the robustness of the results, ensuring high-quality output even when facing challenging image content.

[0092] In Figure 3 , the present invention selected the reconstructed images of three categories, namely horses, dogs, and jack-o'-lanterns, and compared them with DreamDiffusion based on diffusion. It is worth noting that the original model of DreamDiffusion used 120k data from MOAAB to train MSM, while in the comparison of the present invention, both the model of the present invention and DreamDifusion only used 20k data from ImageNetEEG to train MSM. According to Table 1, it can be seen that the performance of DreamDifusion drops significantly when the amount of MSM training data is insufficient, while the model of the present invention performs well under the same amount of data. According to Figure 3 it can be seen that the image quality of the present invention is higher, which highlights the effectiveness of the method of the present invention.

[0093] To prove that each part of the method of the present invention plays an active role in the image reconstruction task, the present invention conducted ablation experiments. In Table 4, the present invention conducted an ablation study through an electroencephalogram classification task to analyze the influence of the main components in the electroencephalogram encoder on the electroencephalogram representation. The results show that both time-domain features and frequency-domain features have a positive impact on performance, and among them, the extraction of frequency-domain features has a higher impact on electroencephalogram classification.

[0094] Table 4 Ablation analysis of electroencephalogram encoder components

[0095]

[0096] The present invention realizes a conversion method from EEG signals to visual images by combining time-frequency domain feature extraction, deep learning alignment network, and cascading diffusion model. This method can not only extract and align the deep features of EEG data, but also generate visual content with rich semantic information. It achieves better classification results and image reconstruction quality with a smaller number of parameters. Through this method, the present invention can understand more deeply how the brain processes visual information and provide new perspectives and tools for research in related fields.

[0097] Ordinary technicians in the art will realize that the embodiments described herein are to help readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. For technicians in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A method for reconstructing human visual information based on electroencephalogram signals, characterized in that, Including: S0. Obtain the image of the scene observed by the human eye at the same moment, and collect the corresponding EEG sequence of the human body at this moment; S1. Use a time-frequency domain encoder to extract the time-domain features and frequency-domain features of the EEG sequence. The time-frequency domain encoder adopts the same time-domain encoder structure as DreamDiffusion to extract the time-domain features of the EEG sequence; use Conv-LSTM as the frequency-domain encoder to extract the frequency-domain features of the EEG sequence; It also includes training the time-domain encoder and the frequency-domain encoder, concatenating the time-domain features and frequency-domain features output by the trained time-domain encoder and frequency-domain encoder, and then inputting the concatenated result into the semantic classifier; S2. Input the concatenated result of the time-domain features and frequency-domain features of the EEG sequence into the semantic classifier to obtain the EEG classification label; the semantic classifier is specifically a multi-layer perceptron stacked by several linear layers; S3. Use the pre-trained IP-Adapter model to generate image feature embeddings based on the image of the scene observed by the human eye obtained in step S0; S4. Based on the image feature embeddings obtained in step S3, perform alignment processing on the concatenated result of the time-domain features and frequency-domain features of the EEG sequence extracted in step S1; step S4 uses a network composed of fully connected layers with residual connections to align the time-domain features and frequency-domain features of the EEG sequence; S5. Take the concatenated result of the time-domain features and frequency-domain features of the aligned EEG sequence and the EEG classification label together as the conditions of the cascaded diffusion model for multi-level semantic visual reconstruction.

2. The method for reconstructing human visual information according to EEG signals as claimed in claim 1, wherein The network structure used for alignment in step S3 includes two fully connected layers, and the two fully connected layers are connected through residual connections.

3. A method for reconstructing human visual information based on electroencephalogram signals according to claim 2, characterized in that, It also includes training the network structure used for alignment, and the loss function used in the training process is: maximizing the cosine similarity between the alignment result of the time-domain features and frequency-domain features of the EEG sequence and the image feature embeddings in step S3.

Citation Information

Patent Citations

  • Simulation image reconstruction method and system based on artificial neural network

    CN115708687A

  • Method for carrying out image reconstruction by utilizing electroencephalogram signals and visual features

    CN116596046A