Visual scene reconstruction method based on electroencephalogram signal and combined with optimal transmission
By combining the diffusion model of EEG spatiotemporal frequency characteristics and optimal transmission theory, the problem of insufficient extraction of frequency domain features in the prior art is solved, and the image generation quality and robustness of visual scene reconstruction are improved.
Patent Information
- Application Number
- CN202510276075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The existing visual scene reconstruction technology based on EEG signals mainly focuses on spatiotemporal characteristics when extracting effective representation of EEG signals. The lack of extraction of frequency domain features leads to low quality of generated images, poor distribution matching effect, and limited robustness.
A visual scene reconstruction method combining EEG spatiotemporal frequency characteristics and optimal transmission theory is proposed. In the feature extraction stage, fast Fourier transform and LSTM network are introduced to extract frequency domain features and fuse them with spatiotemporal features. In the image generation stage, the quality and robustness of the generated samples are improved through feature alignment and optimal transmission processes using a diffusion model combined with optimal transmission theory.
By focusing on space-time and frequency domain characteristics, the accuracy of information acquisition and image generation of EEG signals is improved, the robustness and generation quality of the model are enhanced, and the sample space is covered is wider.
Smart Images

Figure CN120219765A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of electroencephalogram and computer technology, and particularly to a visual scene reconstruction method based on electroencephalogram signals and combining the optimal transport theory. Background Art
[0002] The visual scene reconstruction technology based on electroencephalogram signals is mainly applied to the fields of neuroscience and computer vision, providing a new way for aphasics to express their thoughts and emotions, and helping them improve their independence and self-care ability. Due to the impairment of language function, aphasics cannot conveniently express their ideas, while this technology can directly convert "thoughts" into visual images by analyzing electroencephalogram (EEG), realizing visual scene reconstruction. The combination of electroencephalogram signals and deep learning also provides an important way to study the relationship between neural activities and visual perception, and has broad application prospects.
[0003] Although good results have been achieved in electroencephalogram-based visual scene reconstruction based on diffusion models, there are still various problems hindering the quality of the generated images. First of all, some current methods mainly focus on spatio-temporal features when extracting effective representations of electroencephalogram signals, while lacking the extraction of features in the frequency domain of electroencephalogram signals. Therefore, it is necessary to establish an effective model that simultaneously focuses on spatio-temporal and frequency domain features, so that the extracted features can cover more information, better represent the relevant activities of the brain, and make sufficient preparations for downstream tasks such as generating pictures.
[0004] Previous studies have mainly been based on traditional diffusion models. Traditional latent diffusion models (LDMs) perform excellently in many aspects, but compared with diffusion models combined with the optimal transport theory, they have some disadvantages in the following aspects:
[0005] 1. Disadvantage of poor distribution matching: Traditional latent diffusion models may not be precise enough in distribution matching. Especially in the case of complex data distributions, the generated samples may not fully cover the diversity of real data.
[0006] Advantages of the optimal transport theory: The optimal transport theory provides an accurate measurement method, which can better match the data distribution and the generated distribution, thereby generating samples of higher quality and covering a wider sample space.
[0007] 2. Disadvantage of poor robustness: When traditional latent diffusion models process noise and complex backgrounds, the generated effects may be poor, and the robustness of the models is limited.
[0008] Advantages of Optimal Transport Theory: The optimal transport theory enhances the robustness of the model on different datasets and tasks, and can maintain high generation quality in the face of noise and complex backgrounds. Summary of the Invention
[0009] The present invention proposes an improved visual scene reconstruction method, which combines the EEG spatio-temporal frequency features with the diffusion model of the optimal transport theory, and solves the deficiencies of the prior art.
[0010] In the feature extraction stage, the visual scene reconstruction method introduces the frequency domain features of EEG. The frequency domain feature extraction method includes fast Fourier transform and LSTM network. Compared with the previous methods based on spatio-temporal features, this method simultaneously focuses on spatio-temporal and frequency domain features, can obtain EEG information more fully, and provides better spatio-temporal frequency collaborative representation for downstream tasks.
[0011] In the image generation stage, the visual scene reconstruction method uses a diffusion model combined with the optimal transport theory. There is a large difference between the data generated by general diffusion models and real data, while the new diffusion model combines the optimal transport theory to align the generated data with real data, thereby making the generated sample quality higher, covering a wider sample space, and the generated visualization images more accurate.
[0012] To achieve the above objectives, the present invention is realized through the following technical solutions: A visual scene reconstruction method based on electroencephalogram signals and combined with the optimal transport theory, specifically including:
[0013] The feature extraction method consists of two parts: the spatio-temporal feature extraction part and the frequency domain feature extraction part. Then the spatio-temporal features and frequency domain features are fused to obtain fused features.
[0014] In the spatio-temporal feature extraction part, a pre-trained Masked AutoEncoder (MAE) is used to obtain the spatio-temporal features of EEG.
[0015] In the frequency domain feature extraction part, the spatio-temporal is transformed to the frequency domain, and the designed Long Short-Term Memory (LSTM) network is used to extract the frequency domain features.
[0016] Feature alignment method: Align text, images, and EEG space. Use the clip (Contrastive Language-Image Pre-Training) image encoder to obtain the embedding of the real image. The fused features are transformed into embeddings of the same scale through a projection matrix, and then the parameters of the projection matrix are trained through the cosine similarity with the image embedding of clip, so that the aligned fused features are more conducive to generating images as conditions.
[0017] Diffusion process: This process uses the features output by the feature alignment stage as conditional features to guide the U-net network in the diffusion model to predict noise. In this stage, a U-Net network combined with the multi-head attention mechanism is used to predict the noise added in the forward stage to enhance the model's recovery ability and generate the hidden features of the real image.
[0018] Optimal transport process: Train a discriminator D, whose main task is to distinguish the data distribution differences between real images and images generated by the diffusion model in the hidden feature space. The discriminator D learns how to identify their distribution features in the high-dimensional feature space by comparing the fused features after EEG feature extraction and the hidden features of the images generated by the diffusion model. By inputting the fused features extracted in the feature extraction stage and the features generated by the diffusion model into the discriminator D to calculate the Kantorovich potential function, whose gradient can be used to optimize the solution of the optimal transport mapping, and use Brenier's theorem to calculate the mapping. According to Brenier's theorem, the gradient is added to the hidden features of the image generated by the diffusion model to obtain the hidden features of the final image.
[0019] Decoding process: Decode the hidden features of the final image into a real image through a pre-trained decoder.
[0020] The present invention provides a method for visual scene reconstruction based on EEG, which has the following beneficial effects compared with the prior art:
[0021] Some studies focus on extracting features from the spatio-temporal signals of EEG and ignore the frequency-domain features of EEG. The present invention fuses spatio-temporal features for visual scene reconstruction, pays attention to both the spatio-temporal features and frequency-domain features of EEG, and obtains spatio-temporal-frequency collaborative representations through feature extraction methods, thereby reducing the loss of key information in EEG signals and improving the accuracy of generated images.
[0022] Advantages of the diffusion model based on optimal transport theory compared with traditional diffusion models:
[0023] Improve interpretability: The optimal transport theory provides a clear way to measure and match data distributions. Through optimal transport, we can intuitively understand how the model maps the data distribution to the latent space distribution and the specific process of this mapping. This clear distribution matching helps to explain the operations of the model in the latent space and the mechanism of generating samples.
[0024] Higher robustness: The optimal transport theory can handle complex and high-dimensional data distributions, enhancing the robustness of the model on different datasets and tasks. Even in the face of noise and complex backgrounds, the model can maintain a high generation quality. Description of the Drawings
[0025] Figure 1It is a step flow chart.
[0026] Figure 2 It is the data flow block diagram of the present invention.
[0027] Figure 3 It is the optimal transmission process structure diagram.
[0028] Figure 4 It is the framework diagram of the visual scene reconstruction method based on EEG signals combined with optimal transmission. Specific implementation manners
[0029] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. The technical solutions in the embodiments of the present invention are described, and some details in the technical invention solutions are supplemented and explained in detail, and are completely conveyed to those skilled in the art.
[0030] S101, referring to Figure 1 and Figure 2 , about 120,000 EEG data samples from more than 400 subjects were collected on the MOABB platform, and the channel range was from 30 to 128. For the convenience of pre-training, we uniformly filled all the data to 128 channels by filling the missing channels with replicated values. During the pre-training stage, every 4 adjacent time steps were grouped into a token, and each token was converted into a 1024-dimensional embedding feature through a projection layer for downstream EEG mask signal modeling. The loss function calculates the MSE between the reconstructed EEG signal and the original signal. The loss is only calculated on the occluded patches. The reconstruction is performed on the entire set of 128 channels, rather than on a per-channel basis. The decoder is discarded after pre-training.
[0031] S201, uniformly divide the EEG signal into n time segments, project each segment into a d-dimensional space to generate a representation Apply a random masking strategy (masking ratio is r m ), divide z into visible and masked parts. Use a transformer block to extract the visible feature f v , extract the true feature f of the masked segment through a non-trainable model m . Introduce a new transformer with f v as the input to generate the predicted value f mp , calculate the mean square error (MSE) between the two features f m and f mp as the regression target L reg = MSE(f m , f mp)。Meanwhile, use the codebook-based semantic unit classification target to predict the codeword probability p m 。Finally, combine the regression and classification targets for joint optimization to enhance the EEG feature representation.
[0032] S202,
[0033] S203. Extract features in the frequency domain. First, perform a fast Fourier transform on the EEG signal to obtain the spatio-temporal signal of the EEG. To prevent overfitting, use a long short-term memory network (LSTM) with fewer parameters for frequency domain feature extraction.
[0034] S204. Fuse the obtained frequency domain feature embedding and spatio-temporal feature embedding. In terms of the feature fusion method, select the following concatenation formula for fusion. Concatenate the spatio-temporal embedding f t and the frequency domain embedding f f directly together to form a joint feature representation f c :
[0035] S205, f c =[f t ; f t
[0036] S301. Align the text, image, and EEG space:
[0037] S302. Obtain the feature embedding: First, use the pre-trained CLIP image encoder to encode the real image I to obtain the embedding feature vector E I (I) for alignment with the spatio-temporal-frequency features.
[0038] S303. Convert the EEG feature representation τ θ (y), where y is the output of the EEG encoder, into an embedding h(τ θ (y)) of the same scale as the CLIP image embedding through a projection matrix h.
[0039] S304. Calculate the similarity: Calculate the cosine similarity between the projected EEG embedding h(τ θ (y)) and the CLIP image embedding E I (I):
[0040] S305,
[0041] S306. Train the parameters of the projection matrix by minimizing the loss function L clip of the cosine similarity so that the aligned EEG embedding and the image embedding are closer in the same feature space:
[0042] S307,
[0043] S401, Obtain the hidden layer features of the image using the trained VQ encoder: Use the trained vector quantization encoder (VQ encoder) E(·) to convert the image x into the corresponding hidden layer feature representation z, and the formula is as follows:
[0044] S402, z = E(x)
[0045] S403, Introduce the conditional signal: In the UNet structure of the Stable Diffusion model, introduce the conditional signal of the EEG data through the cross-attention mechanism. Specifically, convert the output of the EEG feature extraction encoder into an embedding representation through a projection layer, and this representation can perform cross-attention with the intermediate value of the UNet.
[0046] S404, Cross-attention calculation: The cross-attention layer is implemented by the following formula:
[0047] S405,
[0048] S406, Among them, the query matrix Q comes from the intermediate value of the UNet, and the key matrix K and the value matrix V come from the embedding representation of the EEG.
[0049] S407, Train a U-net network combined with the multi-head attention mechanism. By introducing the multi-head attention mechanism, this network can effectively process the feature information of different regions during the image generation process, so as to more accurately predict the noise added during the image generation process. Through training the network, the network learns to recognize the noise patterns at each stage of image generation and predict their distribution in the image. This process helps to improve the quality of image generation, making the generated images more realistic and having higher detail fidelity. By continuously optimizing the network parameters during the training process, the U-net network can show strong noise suppression ability when dealing with complex image tasks.
[0050] S408,
[0051] S409, Use the EEG fusion feature after the projection matrix as conditional information to guide the generation of the hidden layer features of the real image: Use the EEG fusion feature after the projection matrix h(τ θ (y)) as conditional information and input it into the U-net network to guide the generation of the hidden layer features of the real image. In this way, the information in the EEG signal can be effectively integrated into the image generation process, making the generated images more in line with the content represented by the EEG signal.
[0052] S410, Reverse Diffusion: During the image generation process, sample latent features z from a simple distribution (such as the standard Gaussian distribution), T and then generate a new latent feature vector z′0 through the reverse diffusion process of the diffusion model. This process is commonly used in generative models. Especially in diffusion models, the reverse diffusion process is an important step for recovering data (such as images) from noise. Specifically, the main operation of this process is to gradually transform the sampled latent features (representing a simple form of the noise distribution) into new latent features through a reverse process. These latent features are further processed and finally generate the desired image or data content. Each stage of reverse diffusion can help remove noise, making the generated result closer and closer to the real data. The goal of this step is to precisely control the removal of noise during the generation process, thereby obtaining a clearer and more meaningful output result, which is widely applied in fields such as image generation and image restoration.
[0053] S411, Among them,
[0054] S501, As Figure 3 shown, in order to distinguish real data from a simple distribution (such as the standard Gaussian distribution) in the latent space, a discriminator D can be trained in the latent space. The role of this discriminator is to help the model learn how to effectively recover real latent features from noise by distinguishing latent features generated from the real data distribution and the simplified distribution (such as the Gaussian distribution).
[0055] S502, Calculate the Kantorovich potential. According to the optimal transport theory, by training the discriminator, we actually obtain a Kantorovich potential function The gradient of this function can help us find the optimal transport mapping.
[0056] S503, Calculate the mapping using Brenier's theorem: According to Brenier's theorem, we can calculate the optimal transport mapping through the gradient of the Kantorovich potential: In this case, is the output of the discriminator D. Therefore, the mapping T can be expressed as:
[0057] S504, Optimal transport mapping: Apply the optimal transport mapping T to the latent features z′0 generated by reverse diffusion to transform it into latent features z″0 that conform to the data distribution. z0″ = T(z0′) where the mapping T is calculated through the optimal transport theory.
[0058] S601, Decode and generate an image: Use the decoder of AE to decode the transformed latent feature z″0 into a new image sample
[0059] S602,
[0060] As Figure 4 is the framework diagram of this method, and these steps detail the specific process of generating high-quality images from EEG signals. By aligning the text, image, and EEG feature spaces, use the VQ encoder to extract the image hidden layer features, combine the U-net network with the multi-head attention mechanism to predict the noise, and use the projection matrix to fuse the EEG features to guide the generation of the target data distribution. Then, optimize the mapping by combining the optimal transport theory, and finally generate high-quality images through the decoder. This series of technical means ensures the accuracy and quality of the generated images.
Claims
1. A visual scene reconstruction method based on EEG signals combined with optimal transmission, characterized in that: The method at least comprises: (1) Feature extraction: Extract effective features of EEG through feature extraction methods; (2) Feature alignment: Align the EEG features with the features of the real image to guide the diffusion model to generate data; (3) Diffusion model generation: Use the extracted features as conditions to guide the generation of image hidden layer features; (4) Optimal transmission optimization: The discriminator learns the distribution difference between the real image and the generated image in the hidden feature space, and performs optimal transmission mapping on the latent features of the generated image. (5) Image decoding: Use the decoder to decode the optimized latent features into the final high-quality image.
2. The method for visual scene reconstruction based on EEG signals combined with optimal transmission according to claim 1, characterized in that: In the step (1), the feature extraction is configured to extract the spatiotemporal features of the EEG signal through a masked autoencoder; perform frequency domain transformation on the EEG signal, and extract frequency domain features using a fast Fourier transform and a long short-term memory network; and fuse the spatiotemporal features and frequency domain features to form a spatiotemporal-frequency collaborative feature representation.
3. The method for visual scene reconstruction based on EEG signals combined with optimal transmission according to claim 1, characterized in that: In the step (2), the feature alignment is configured to obtain an embedding representation of the real image using a contrastive learning model, and to convert the spatiotemporal-frequency collaborative features into an embedding representation of the same scale as the image embedding through a projection matrix; and to train the projection matrix parameters through cosine similarity, and the cosine similarity calculation formula is: Make the EEG signal embedding features closer to the real image embedding features in the aligned feature space.
4. The method for visual scene reconstruction based on EEG signals combined with optimal transmission according to claim 1, characterized in that: In the step (3), the diffusion model generation configuration is a diffusion model (U-Net) combined with a multi-head attention mechanism to guide the generation of image hidden layer features based on the spatiotemporal and frequency collaborative features; and the target potential features are gradually generated through the reverse diffusion process.
5. The method for visual scene reconstruction based on EEG signals combined with optimal transmission according to claim 1, characterized in that: In the step (4), the optimal transmission optimization configuration is to learn the distribution difference between the real image and the generated image in the hidden layer feature space through the discriminator, calculate the Kantorovich potential function, and use the Brenier theorem to perform optimal transmission mapping on the latent features of the generated image.
6. The method for visual scene reconstruction based on EEG signals combined with optimal transmission according to claim 1, characterized in that: In the step (5), the image decoding is configured to decode the optimized latent features into a final high-resolution image through a pre-trained decoder.