An EEG Visual Stimulation Hybrid Modal Decoding Method and System

Through a hybrid modal encoder and diffusion prior model, combined with a lightweight Transformer and a time-frequency domain dual-stream spatiotemporal convolutional network, multimodal information is extracted from EEG data, solving the problem of mode correlation neglect in EEG data decoding, and achieving better visual stimulation reconstruction effect.

CN119919510BActive Publication Date: 2025-06-24GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510405346.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-24
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the commonality between different modes and the real activity of learning brain regions in EEG data decoding, and ignores the correlation between modes and fine-grained expression, resulting in unsatisfactory reconstruction results.

Method used

Using a hybrid modal encoder, images, categories, and semantic modal information are extracted from EEG data through technical means such as global feature mapping network, hybrid mapping network, diffusion prior model, and aligned and contrasted learning in CLIP space. Combined with a lightweight Transformer network and a time-frequency domain dual-stream spatiotemporal convolutional network, the fusion and decoding of multimodal features are achieved.

Benefits of technology

The multimodal decoding capability of EEG data is improved, the model captures time series features is enhanced, the fine-grainedness and accuracy of image reconstruction is improved, the problem of neglecting modal correlation in single-modal training is solved, and the better visual stimulation reconstruction effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919510B_ABST
    Figure CN119919510B_ABST
Patent Text Reader

Abstract

The present invention discloses an EEG visual stimulation hybrid modality decoding method and system. The method includes the following steps: acquiring EEG data and constructing a three-state EEG dataset of image-category-semantics; the hybrid modality encoder learns the corresponding image, category, and semantic modality information from the EEG data; mapping the three-state EEG dataset to the corresponding CLIP space to obtain the target modality; performing contrastive learning on the modality information output by the hybrid modality encoder and the target modality, and performing image retrieval and classification tasks based on the image modality information; constructing a diffusion prior model, aligning the image, category, and semantic modality information output by the hybrid modality encoder with the target modality in the CLIP space, outputting the image, category, and semantic modality information processed by the diffusion prior, and inputting it into a pre-trained image generation model for generation tasks. The present invention can effectively decode multiple modalities of brain data and improve the effects of retrieval and reconstruction tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of EEG data decoding, and particularly to an EEG visual stimulation hybrid modality decoding method and system. Background Art

[0002] Early studies used long short-term memory networks and generative adversarial networks for temporal feature extraction of EEG data and natural image reconstruction. Long short-term memory networks are good at dealing with temporal dependencies, but may ignore the spatial correlations of multi-channel EEG data (such as the topological relationship between electrodes), and the feature representation ability is affected due to the scarcity of training samples. On the other hand, the adversarial training of generative adversarial networks needs to balance the generator and the discriminator, which is prone to mode collapse or training oscillation, and the high noise contained in EEG data will exacerbate this problem, resulting in generally unsatisfactory generated image effects.

[0003] Recent studies have adopted window slicing and random masking strategies to divide EEG data into different semantic units and focus on reconstructing the embeddings of individual semantic units. Although this reduces the difficulty of reconstruction, masking the EEG data with large noise and small data volume will increase the difficulty of extracting useful features from the EEG data, thus damaging the reconstruction effect. Some other studies pursue cross-modal alignment using the image modality to decode multi-modal information from the brain. For modal learning, currently, the image modality is used as a reference for alignment to achieve the decoding of meaningful visual stimuli. However, using only single-modal decoding of visual stimuli often easily ignores the correlations between different modalities when the actual brain is visually stimulated, and it is also difficult to capture the activity changes in real brain regions.

[0004] The current existing technologies have common limitations including:

[0005] (1) Although the Transformer model can fully model time series models, the Transformer model is prone to focusing on the global modeling relationship and thus ignoring the expression effect of EEG data at a fine-grained level.

[0006] (2) Training different modalities separately while ignoring that the model can learn the correlations between different modalities.

[0007] (3) Since different parts of the brain generate different modal electro-data features when receiving natural visual stimuli, training the model with only a single modality will make it difficult for the model to capture the activity changes in real brain regions.

[0008] Therefore, there is an urgent need for a hybrid modality decoding technology that can capture the commonalities between different modalities and learn the real activity of brain regions when receiving visual stimuli. Summary of the Invention

[0009] To overcome the defects and deficiencies of the prior art, the present invention provides an EEG visual stimulation hybrid modality decoding method and system, which can better understand the generation and distribution of different modalities in the human brain, effectively decode multiple modalities from brain data, and improve the retrieval and reconstruction task effects.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] The present invention provides an EEG visual stimulation hybrid modality decoding method, including the following steps:

[0012] Obtain EEG data and construct a three-state EEG dataset of image-category-semantics;

[0013] Construct a hybrid modality encoder, and the hybrid modality encoder learns the corresponding image, category, and semantic modality information from the EEG data;

[0014] Map the image-category-semantic data in the three-state EEG dataset of image-category-semantics to the corresponding CLIP space to obtain the target modality;

[0015] Compare and learn the image, category, and semantic modality information output by the hybrid modality encoder with the target modality, and perform image retrieval and classification tasks based on the image modality information;

[0016] Construct a diffusion prior model, align the image, category, and semantic modality information output by the hybrid modality encoder with the target modality in the CLIP space, and output the image, category, and semantic modality information processed by the diffusion prior;

[0017] Input the image, category, and semantic modality information processed by the diffusion prior into a pre-trained image generation model for generation tasks. The modality information processed by the diffusion prior is respectively used to provide the representation degree in the coarse-grained direction during image generation, the category during image generation, and the representation degree in the fine-grained direction during image generation, and complete visual stimulation reconstruction.

[0018] As a preferred technical solution, the hybrid modality encoder includes a global feature mapping network and a hybrid mapping network;

[0019] The global feature mapping network includes a lightweight Transformer network, a multi-time scale hybrid convolutional network, a convolutional network with multiple feature convolutional layers, and a time-frequency domain two-stream spatio-temporal convolutional network;

[0020] The lightweight Transformer network extracts the temporal features of the EEG data;

[0021] The multi-time-scale hybrid convolutional network enhances the details of the temporal features based on the local receptive field and outputs the global temporal features;

[0022] The convolutional network of the multi-feature convolutional layer extracts the spatial dimension features of the multi-modal characteristics of the EEG data, analyzes the cortical spatial topological information of the EEG data, and outputs the temporal features;

[0023] The time-frequency domain two-stream spatio-temporal convolutional network fuses the temporal and frequency domain information features of the EEG data and outputs the global features;

[0024] The hybrid mapping network decodes the global features to obtain the category modality and the semantic modality.

[0025] As a preferred technical solution, the lightweight Transformer network includes a downsampling layer, an attention layer, and a convolutional layer. The average pooling layers with different sizes and different strides are used to form the downsampling layers of different scales. The EEG data is decomposed into time pattern information of different scales through the downsampling layers of different scales, and the time pattern information of different scales is reshaped into the same shape through the MLP and added together to be integrated into new pattern information , and based on the Embedding module, the new pattern information is mapped to a variable , and the attention layer is used to learn the correlation between different tokens of the variable , and the convolutional layer outputs the temporal features extracted by the lightweight Transformer network .

[0026] As a preferred technical solution, the multi-time-scale hybrid convolutional network performs a convolutional operation on the temporal features output by the lightweight Transformer network, extracts the time series features of different time scales, reshapes them into a temporal feature vector with the same dimension through a flattening operation and a fully connected layer, and concatenates them into a vector group ;

[0027] For each vector of the vector group , a descriptor of the channel is generated through a convolutional operation and a sigmoid activation function, and the descriptors are formed into a new vector group , which is expressed as:

[0028] ;

[0029] The vectors in the vector group are concatenated to obtain a vector with dimensions ;

[0030] The vector is in Obtain the weight matrix through the softmax function in the dimension Based on the weight matrix and the corresponding input vector, obtain the global temporal features output by the multi-time-scale hybrid convolutional network It is expressed as:

[0031] ;

[0032] ;

[0033] Among them, Indicates that the softmax function operation is performed on the 0th dimension of the vector The represents the dot product operation, and the represents the extraction operation in sequence from the 0th dimension of to the 0 dimension of M -1 dimension.

[0034] As a preferred technical solution, the convolutional network of the multi-feature convolutional layer includes a multi-scale spatial convolutional group and a dynamic feature fusion layer. The multi-scale spatial convolutional group uses parallel depth convolutional layers with different receptive fields to extract the local micro-topological correlation and global macro-distribution pattern of the electrode array respectively. The dynamic feature fusion layer adaptively weights and fuses the multi-scale spatial features through learnable channel attention weights and outputs temporal features.

[0035] As a preferred technical solution, in the multi-scale spatial convolutional group, the feature information of different dimensions of the global temporal features is extracted through convolutional layers with multiple different kernel sizes, the feature information of different dimensions is added together to fuse the multi-scale information, the channel descriptor representation is obtained through global pooling, the correlation between channels is modeled based on two fully connected layers, and after dimensionality increase in the second fully connected layer, it is evenly divided into multiple parts. In the dynamic feature fusion layer, the respective weight representations are generated through the sigmoid function respectively, and the information of different scales is weighted and added together to output temporal features.

[0036] As a preferred technical solution, the time-frequency domain two-stream spatio-temporal convolutional network obtains the frequency domain features through fast Fourier transform and ResMLP network modeling. The temporal features and frequency domain features output by the convolutional network of the multi-feature convolutional layer obtain the temporal domain weight and frequency domain weight through the sigmoid function, are weighted to obtain the corresponding output features, and after splicing, matrix multiplication operations are performed to obtain the hybrid time-frequency domain features, and the global features are obtained based on the hybrid time-frequency domain features. Specifically, it is expressed as:

[0037] ;

[0038] ;

[0039] ;

[0040] ;

[0041] ;

[0042] ;

[0043] Among them, represents the time feature, represents the frequency domain feature, represents the linear operation, represents the sigmoid function, represents the average pooling operation, represents the feature concatenation operation, represents the hybrid time-frequency domain feature, represents the feature transpose, represents the convolution operation, represents the convolution kernel size, represents the dilation rate of the th convolution operation,

[0044] As a preferred technical solution, the hybrid mapping network includes a brain modality regression layer, a class embedder, a brain modality embedder, a semantic embedder, and a multi-modal decoder;

[0045] The global feature is passed through the brain modality regression layer and then input into the class embedder, the brain modality embedder, and the semantic embedder respectively to obtain the corresponding preliminary modality feature representations;

[0046] The respective preliminary modality feature representations are concatenated in the first dimension and input into the multi-modal decoder for decoding to obtain the class modality feature and the semantic modality feature.

[0047] As a preferred technical solution, the diffusion prior model is trained using the mean squared error loss.

[0048] The present invention also provides an EEG visual stimulation hybrid modality decoding system for implementing the above-mentioned EEG visual stimulation hybrid modality decoding method. The system includes: an EEG data acquisition module, a three-state EEG dataset construction module, a hybrid modality encoder construction module, a target modality construction module, a contrast learning module, an image retrieval and classification module, a diffusion prior model construction module, and an image generation module;

[0049] The EEG data acquisition module is used to acquire EEG data;

[0050] The three-state EEG dataset construction module is used to construct a three-state EEG dataset of image-class-semantic;

[0051] The hybrid modality encoder construction module is used to construct a hybrid modality encoder, which learns the corresponding image, category, and semantic modality information from EEG data;

[0052] The target modality construction module is used to map the image-category-semantic data in the three-state EEG dataset of image-category-semantic to the corresponding CLIP space to obtain the target modality;

[0053] The contrastive learning module is used to perform contrastive learning on the image, category, and semantic modality information output by the hybrid modality encoder and the target modality;

[0054] The image retrieval and classification module is used to perform image retrieval and classification tasks based on the image modality information;

[0055] The diffusion prior model construction module is used to construct a diffusion prior model, align the image, category, and semantic modality information output by the hybrid modality encoder with the target modality in the CLIP space, and output the image, category, and semantic modality information processed by the diffusion prior;

[0056] The image generation module is used to input the image, category, and semantic modality information processed by the diffusion prior into a pre-trained image generation model for generation tasks. The modality information processed by the diffusion prior is respectively used to provide the representation degree in the coarse-grained direction during image generation, the category during image generation, and the representation degree in the fine-grained direction during image generation, and complete the visual stimulus reconstruction.

[0057] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0058] (1) Aiming at the deficiencies of the existing EEG encoders in image reconstruction, based on the corresponding image, category, and semantic modality information learned from EEG data by the hybrid modality encoder, the present invention can better understand the generation and distribution of different modalities in the human brain, enabling the model to effectively decode multiple modalities from brain data.

[0059] (2) To solve the problems of high computational complexity and easy overfitting brought by the Transformer network, the present invention adopts a lightweight Transformer network, which suppresses the problem of easy overfitting existing in the ordinary Transformer network. While retaining the powerful ability of the Transformer network to extract time series features, it also avoids the information bottleneck caused by the simplicity of the linear mapping in the class MLP method, making it difficult to capture the dependencies of time series context relationships and complex time patterns.

[0060] (3) To address the problem that Transformer networks usually weaken temporal relationships due to overemphasis on mutation points, the EEG data is divided into multiple temporal patterns and they are connected through residual mixing to obtain enhanced time series data, enabling the Transformer network to comprehensively capture information of different temporal patterns and better extract time series features.

[0061] (4) In view of the fact that the internal connection between different brain regions in decoding and reconstructing visual stimuli was not considered in previous model designs, and most of the feature extraction of EEG data caused by visual stimuli still remained at the extraction of time domain features, ignoring the frequency domain information of EEG data, the present invention constructs a time-frequency domain two-stream spatio-temporal convolutional network to better model brain activities, further improving the performance of the model in retrieval and reconstruction tasks.

[0062] (5) In view of the fact that previous EEG encoders were all trained for a certain modality, ignoring the expression relationship between multiple modalities, the present invention can autonomously learn the weight relationship between multiple modalities to maximize the ability to extract effective information from EEG data, map the image modality, category modality, and semantic modality output by the hybrid modality encoder to the target CLIP space to align the distribution of the three EEG embedding information and the target modality in the embedding space, achieving a reliable image reconstruction task. The image modality information will consider the category modality information and semantic information during the training process, and can achieve the best results simultaneously in retrieval and generation tasks without the need for targeted training for a single task. Description of the Drawings

[0063] Figure 1 It is a schematic flow diagram of the EEG visual stimulus hybrid modality decoding method of the present invention;

[0064] Figure 2 It is a schematic overall architecture diagram of the hybrid modality encoder of the present invention;

[0065] Figure 3 It is a schematic overall architecture diagram of the lightweight Transformer network of the present invention;

[0066] Figure 4 It is a schematic network structure diagram of the multi-time scale hybrid convolutional network of the present invention;

[0067] Figure 5 It is a schematic network structure diagram of the convolutional network of the multi-feature convolutional layer of the present invention

[0068] Figure 6 It is a schematic network structure diagram of the time-frequency domain two-stream spatio-temporal convolutional network of the present invention;

[0069] Figure 7Schematic diagram of the network structure of the hybrid mapping network of the present invention;

[0070] Figure 8 Schematic diagram of the network structure of the embedder of the present invention;

[0071] Figure 9 Schematic diagram of the network structure of the multi-modal decoder of the present invention;

[0072] Figure 10 Schematic diagram of the overall implementation framework for visual stimulus reconstruction of the present invention. Detailed implementation manners

[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0074] Embodiment 1

[0075] As Figure 1 shown, this embodiment provides an EEG visual stimulus hybrid-modal decoding method, including the following steps:

[0076] S1: Construct a training set and a test set based on EEG data, and preprocess the EEG data;

[0077] In this embodiment, the original image-EEG dataset uses the THINGS-EEG dataset, which contains electroencephalogram data collected from 10 subjects under the RSVP paradigm;

[0078] Separate the training set EEG data and the test set EEG data. When collecting EEG data, when presenting images, the images are presented to the subjects in a pseudo-random order. Each image is displayed for 100 milliseconds, followed by a 100-millisecond blank screen, the purpose of which is to reduce the interference caused by other artifacts;

[0079] The training set includes 1654 different natural image categories. Each image category in the training set will contain 10 images of this category. During the process of the subjects viewing the training set images and collecting EEG data, each image will appear four times. The subjects have viewed a total of 1654×10×4 = 66160 images, and then 66160 EEG data collected corresponding to the viewed images are collected. For constructing the training set, each image and the EEG data collected by the subject under its visual stimulus are combined into a (target-training data) group of (image, EEG), and a total of 66160 groups of (image, EEG) are obtained;

[0080] The test set contains 200 other categories that do not appear in the training set. Each category in the test set is represented by an image, and each image appears 80 times when viewing the image categories in the test set. So the subjects viewed a total of 200×80 = 16000 test set images. When viewing the test set images, EEG data was collected, and a total of 16000 EEG data were collected. In the construction of the test set, the 16000 collected EEG data were subjected to average pooling to obtain 200 EEG data corresponding to viewing 200 test images, that is, 200 groups of (image, EEG);

[0081] The EEG data in the obtained original training set of 66160 groups (image, EEG) and the original test set of 200 groups (image, EEG) were both segmented into time periods from 0 to 1000 milliseconds after the start of the stimulus, and the average value of the 200 milliseconds before the stimulus was used for baseline correction to ensure that the collected EEG data did not change greatly due to environmental interference. The EEG data was downsampled to 250 Hz and multivariate noise normalization was applied to the EEG data to better enable the model to converge during the training process;

[0082] For each image in the training set and the test set, semantic description information corresponding to each image was generated through the BLIP2 image-to-text model, and the semantic description information was added to the original training set and the original test set to form a three-state EEG data training set and test set of image-category-semantics;

[0083] S2: Construct a hybrid modality encoder;

[0084] As Figure 2 shown, the EEG data is input into the hybrid modality encoder, and the hybrid modality encoder learns the corresponding three modality information of image, category, and semantics from the EEG data. The hybrid modality encoder includes a global feature mapping network and a hybrid mapping network;

[0085] The global feature mapping network includes a lightweight Transformer network, a multi-time scale hybrid convolutional network, a convolutional network with multi-feature convolutional layers, and a time-frequency domain two-stream spatio-temporal convolutional network;

[0086] As Figure 3 shown, the lightweight Transformer network includes: a downsampling layer, an attention layer, and a convolutional layer;

[0087] In this embodiment, average pooling layers with different sizes and different strides are used to form downsampling layers of different scales, and the EEG data is decomposed into time pattern information of different scales through the downsampling layers of different scales. The input EEG data is expressed as: , is the input batch dimension, is the number of channels composed of EEG data, is the timing information of EEG data, Figure 3 in represents the K th time pattern information after the EEG data passes through the K th downsampling layer. The time pattern information at different scales is reshaped into the same shape through MLP and added together to be integrated into new pattern information , enhancing the expressive ability of time series synthesis;

[0088] Use the Embedding module to map the new pattern information to the variable , and embed the position information for output. Use the attention layer to learn the correlation between different tokens among the variables , fuse the correlation between different variables and update the variable of . Subsequently, model the temporal features inside the variable through a convolutional layer to obtain efficient temporal representations. Finally, the Transformer network outputs features , specifically expressed as:

[0089] ;

[0090] ;

[0091] ;

[0092] The average pooling operation of the lightweight Transformer network realizes feature dimensionality reduction by taking the mean of the local receptive field, and has dual advantages at the level of temporal feature processing: First, the smoothing property of average pooling can effectively suppress the interference of high-frequency noise on temporal modeling; Second, the hierarchical time representation formed by multi-scale downsampling can provide complementary temporal context information for the Transformer network. The self-attention mechanism of the Transformer network realizes the modeling of global dependence relationships across time steps by dynamically allocating weights to multi-scale temporal features, thereby improving the model's ability to parse complex temporal patterns.

[0093] After completing the temporal feature modeling in this embodiment, a multi-time-scale hybrid convolutional network is introduced for multi-granularity feature refinement. Based on the local perception characteristics of convolutional operations, the multi-time-scale hybrid convolutional network enhances the details of the high-level features output by the Transformer through a controllable local receptive field, effectively compensating for the potential information loss during the pooling process, and finally realizing the collaborative optimization of global pattern perception and local feature preservation, which helps to enhance the model's learning of the intrinsic temporal features of EEG data;

[0094] such asFigure 4 As shown, in the multi-time-scale hybrid convolutional network, the output of the lightweight Transformer network is subjected to a convolution operation to extract time series features of different time scales and, through a flattening operation and a fully connected layer, the time series features are reshaped into a time series feature vector with the same dimension , and concatenated into a vector group . For the vector group , for each vector , , the correlation between channels is modeled through a 1×1 convolution operation, and a channel descriptor is generated through a sigmoid activation function , , and they are combined into a new vector group , specifically expressed as:

[0095] ;

[0096] After obtaining the new vector group , the vectors in the vector group are concatenated in the 0th dimension to obtain a vector with dimension , expressed as:

[0097] ;

[0098] At dimension, a weight matrix is obtained through a softmax function, and the generated weight matrix is multiplied element-wise with the corresponding input vector , , , added after passing through a global average pooling layer (GAP), and finally the output is used as the output of the multi-time-scale hybrid convolutional network, specifically expressed as:

[0099] ;

[0100] ;

[0101] Among them, represents that the softmax function operation is performed on the 0th dimension of the vector , represents the dot product operation, represents in sequence from 's 0 dimension to MExtract operation is performed in the -1 dimension, which is the final output result of the multi-time-scale hybrid convolutional network, and B represents the batch size;

[0102] In this embodiment, based on the convolutional network of the multi-feature convolutional layer, the spatial dimension features of the multi-modal characteristics of EEG data are extracted, and the cortical spatial topological information (specifically manifested as the spatial distribution correlation of multi-channel electrodes) hidden in the EEG data is analyzed;

[0103] The convolutional network of the multi-feature convolutional layer includes a multi-scale spatial convolutional group and a dynamic feature fusion layer. Among them, the multi-scale spatial convolutional group uses parallel depth convolutional layers with different receptive fields to capture the local micro-topological correlation and global macro-distribution pattern of the electrode array respectively. The depth convolutional kernel parameters are initialized according to the electrode coordinates of the international 10-20 system to conform to the neuroelectrophysiological characteristics; the dynamic feature fusion layer adaptively weights and fuses the multi-scale spatial features through learnable channel attention weights, and finally outputs a dimensionally regular cortical spatial representation tensor;

[0104] As Figure 5 shown, B represents the batch number of features T represents the temporal length of features C represents the number of channels of features The feature information of different dimensions contained in the global temporal features is extracted through three convolutional layers with different convolutional kernel sizes. Compared with directly inputting three identical unprocessed features into the network, the former represents different expressions of the same kind of information for processing, so it is meaningful to perform feature fusion or feature interaction on them;

[0105] After constructing the three inputs in the above way, they are added together to fuse the multi-scale information, and the channel descriptor representation is obtained through global pooling. This corresponds to the operation of dimension elevation. In the excitation stage, two fully connected layers are also used to model the correlation between channels. The dimension is elevated to 3C in the second layer of the fully connected layer, and then evenly divided into 3 parts, and the respective weight representations are generated through the sigmoid function. Finally, the information of the three scales is weighted and then added together to obtain the output time feature ;

[0106] As Figure 6 shown, this embodiment fuses two different information features of the time domain and frequency domain of EEG data based on the time-frequency domain two-stream spatio-temporal convolutional network. The EEG data is subjected to fast Fourier transform to obtain its expression form in the complex domain, and the frequency domain features are modeled through a ResMLP network , and the time features output by the convolutional network of the multi-feature convolutional layer and the frequency domain features are used to obtain the time domain weight and frequency domain weight through the sigmoid function, and the time features Multiply the time-weighted information to obtain the weighted output features , the frequency-domain features Multiply the frequency-domain weight information to obtain the output features , concatenate and , and perform a matrix multiplication operation after concatenation. The matrix multiplication can effectively mix two different-dimensional features to obtain the mixed time-frequency domain features . Concatenate through a convolutional kernel of the same scale but different dilation rates 1D convolutions to extract information of the same scale, and then through merge the results of all branches to obtain , and finally pass through output , expressed as:

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] ;

[0112] ;

[0113] wherein, represents transpose;

[0114] In this embodiment, the time-domain and frequency-domain subspace features of different scales are fused through matrix multiplication operations. Under the same network width, the matrix multiplication realizes the two-dimensional non-linear feature fusion through the cross-dimensional pairwise mapping multiplication operation, achieving the efficient mixed learning of the time-frequency domain with less memory overhead to obtain common features, and then solving the conflict between local features and global features through convolutional operations of different scales. Specifically, the size of the convolutional kernel represents the size of the features of the receptive field it can capture. The convolutional operation with a small convolutional kernel focuses more on the learning of local features of EEG data and is prone to ignoring the overall context expression relationship, while the convolutional operation with a large convolutional kernel can effectively capture the context expression relationship but is prone to ignoring the details. Therefore, convolutional networks with different sizes of convolutional kernels are constructed to capture mixed information of different scales. These convolutions can be dilated to increase their receptive fields without increasing the convolutional kernel size. As for this, the output of the time-frequency dual-domain feature mixed learning network is reset to the global feature , because it comprehensively considers the temporal features. Spatial features, frequency domain features and other important features are learned. It is worth noting that the global features are As the image modality information decoded from the EEG data by the encoder, this image modality information is used for the retrieval task and also as the input data for training the diffusion prior model.

[0115] like Figure 7 As shown, based on the hybrid mapping network in the global feature Decode more detailed category modes and semantic modes, global features After passing through a brain modality regression layer (brain regression ridge), it is divided into three paths and simultaneously input into the corresponding embedding device to obtain a preliminary modality feature representation, which is called fuzzy modality feature. Then it is concatenated in the first dimension and input into the multimodal decoder to further decode the fuzzy feature modality to obtain detailed category modality features and semantic modality features, which are used to train the subsequent diffusion prior model;

[0116] In this embodiment, a category embedder, a brain modality embedder, and a semantic embedder are constructed respectively, and each global feature is calculated for the three required modal information. The distance to the key feature information is calculated, and the clustering is performed in the vicinity of the specific key voxel according to the calculated distance. After aggregating these tokens as a group, the shape is B × G × K Key modal features of grouping , indicating that each global feature is divided into groups, each containing tokens, the modal features within the group may show similar characteristics. Based on this, the three required rough modal information can be obtained, namely brain modal information as the main modality, category modal information as the auxiliary modality, and semantic description modal information. Based on these three rough modal information, more accurate category modal information and semantic description modal information will be decoded in the multimodal decoder;

[0117] In this embodiment, the structure of each embedder is the same, such as Figure 8 As shown, taking the category embedder as an example, the global feature The local patterns of each group are aggregated to obtain the corresponding modal information features. The embedder consists of two main modules: the Token mapping network and the cross-attention layer. Dimensional operation to obtain , u is input into the Token mapping network for feature learning , , Represents the number of groups, represents the embedding dimension within the group, and subsequently the input is passed through a cross-attention layer module to decode the rough category modality information. Among them, is projected as values and keys , and a set of learnable tokens is randomly initialized as queries . The updated queries are used as output embeddings. Similarly, the same processing method can be used to process the other two rough modality information to obtain more refined modality embedding information;

[0118] In this embodiment, the three rough modality information obtained after cross-attention layer processing are concatenated in the dimension of to obtain a new global rough modality feature

[0119] As Figure 9 shown, represents the feature after concatenation of the three rough modality information in the group dimension. Through a cross-attention layer, the detailed category modality and semantic description modality are decoded. Subsequently, it is sent to a pooling layer for dimension integration to obtain and . Subsequently, it is passed through two MLP layers and then output to obtain the decoded detailed category features and semantic features, which are used as the input data for training the diffusion prior model;

[0120] S3: Construct a diffusion prior model to map the three modality information output by the hybrid modality encoder into the target CLIP space, so that the three modality information output by the hybrid modality encoder can be recognized by the pre-trained generative model;

[0121] In this embodiment, the diffusion prior model includes a simplified U-Net model. During training, the image modality, category modality, and semantic modality output by the hybrid modality encoder are separately trained one-to-one with the corresponding target modalities to align the modality information output by the hybrid modality encoder with the target information in the CLIP space, so that the modality information output by the model can be directly input into the pre-trained generative model for image generation. The diffusion prior model is trained using the classifier-free guidance method during the training process, effectively balancing the fidelity of the conditional data and the diversity of the generated output. The diffusion prior model is trained from scratch using the mean squared error loss, specifically expressed as:

[0122] ;

[0123] Among them, represents the mean squared error loss, represents the CLIP embedding of the perturbation after a given diffusion time step t, while represents the diffusion prior model network.

[0124] In this embodiment, after the hybrid multimodal features are processed by the diffusion prior, they can be directly used for image reconstruction just like the original CLIP embedding vector. Specifically, in order to ensure the reconstruction of high-quality visual stimuli, three types of modal information are adopted, namely, the image modality, the category modality, and the semantic modality output by the hybrid modal encoder. The image modality is used as the main modality for reconstructing the image. IP-Adapters and SDXL-Turbo are used to simultaneously utilize different modalities for reconstructing visual stimuli. In the generation stage, for the image embedding information of the main modality, the entire IP-Adapter is used to process the image embedding. For the category and semantic description modalities, following the process of how the brain understands natural things, that is, extracting the category information and then giving further detailed semantic descriptions, the modified IP-Adapters are used, namely, IP-Adapter-Style and IP-Adapter-Layout are used to handle the relationship between the two. The category embedding information acts on IP-Adapter-Style to guide the initial generation category direction of the image, and the semantic information embedding acts on IP-Adapter-Layout to fine-tune the generated image according to the semantic description information during generation.

[0125] Such as Figure 10As shown, the image modality, class modality, and semantic modality output by the hybrid modality encoder correspond to EEG2 image, EEG2 class, and EEG2 caption in the figure. The three modality information is compared and learned with the target modality. As the target modality for the contrastive learning of the hybrid modality encoder, the specific images, classes, and semantic descriptions in the three-state EEG dataset of image-class-semantic are obtained through the Open CLIP model to obtain the target modality, and the image modality (EEG2 Image in the figure) obtained by contrastive learning training is used for image retrieval and classification tasks. Subsequently, in the generation stage, the diffusion prior model maps the three EEG embedding information (EEG2 class, EEG2 image, and EEG2 caption in the figure) output by the hybrid modality encoder to the corresponding CLIP space to align the distribution of the three EEG embedding information and the target modality in the embedding space, and outputs the image, class, and semantic modality information processed by the diffusion prior. The image, class, and semantic modality information processed by the diffusion prior is input into the pre-trained image generation models (SDXL-Turbo and IP-Adapters) for generation tasks. In the generation tasks, the main embedding information of the image modality guides the entire generation direction, the class information is used to provide the class of image generation during the generation process to ensure that the generated image does not deviate from the true class, and the semantic description information is used to guide the representation degree in the fine-grained direction during image generation. Based on the above operations, visual stimuli can be decoded with higher accuracy.

[0126] In the process of training the image modality information in this embodiment, the class modality information and semantic information are considered, so that the output image modality information can fully consider the interoperability with other modalities, and only one encoder can be used to achieve high-quality decoding of the three modality information.

[0127] Embodiment 2

[0128] This embodiment provides an EEG visual stimulus hybrid modality decoding system for implementing the EEG visual stimulus hybrid modality decoding method of the above Embodiment 1. The system includes: an EEG data acquisition module, a three-state EEG dataset construction module, a hybrid modality encoder construction module, a target modality construction module, a contrastive learning module, an image retrieval and classification module, a diffusion prior model construction module, and an image generation module;

[0129] In this embodiment, the EEG data acquisition module is used to acquire EEG data;

[0130] In this embodiment, the three-state EEG dataset construction module is used to construct a three-state EEG dataset of image-class-semantic;

[0131] In this embodiment, the hybrid modality encoder construction module is used to construct a hybrid modality encoder, and the hybrid modality encoder learns the corresponding image, category, and semantic modality information from EEG data;

[0132] In this embodiment, the target modality construction module is used to map the image-category-semantic data in the image-category-semantic three-state EEG dataset to the corresponding CLIP space to obtain the target modality;

[0133] In this embodiment, the contrastive learning module is used to perform contrastive learning on the image, category, and semantic modality information output by the hybrid modality encoder and the target modality;

[0134] In this embodiment, the image retrieval and classification module is used to perform image retrieval and classification tasks based on the image modality information;

[0135] In this embodiment, the diffusion prior model construction module is used to construct a diffusion prior model, align the image, category, and semantic modality information output by the hybrid modality encoder with the target modality in the CLIP space, and output the image, category, and semantic modality information after diffusion prior processing;

[0136] In this embodiment, the image generation module is used to input the image, category, and semantic modality information after diffusion prior processing into a pre-trained image generation model for generation tasks. The modality information after diffusion prior processing is respectively used to provide the representation degree in the coarse-grained direction during image generation, the category during image generation, and the representation degree in the fine-grained direction during image generation, and complete the visual stimulus reconstruction.

[0137] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A mixed modality decoding method for EEG visual stimulation, characterized in that: The steps include: Obtain EEG data and construct a three-state EEG dataset of image, category and semantics; Construct a hybrid modality encoder that learns the corresponding image, category, and semantic modality information from EEG data; The hybrid modality encoder includes a global feature mapping network and a hybrid mapping network; The global feature mapping network includes a lightweight Transformer network, a multi-time scale hybrid convolutional network, a convolutional network with multiple feature convolutional layers, and a dual-stream spatiotemporal convolutional network in the time-frequency domain; The lightweight Transformer network extracts the temporal features of the EEG data; The multi-time-scale hybrid convolutional network enhances the details of the temporal features based on the local receptive field and outputs the global temporal features; The convolution network of the multi-feature convolution layer extracts spatial dimension features from the multimodal characteristics of the EEG data, analyzes the cortical spatial topology information of the EEG data, and outputs temporal features; The dual-stream spatiotemporal convolutional network in the time-frequency domain fuses the time-domain and frequency-domain information features of the EEG data, outputs global features, and uses the global features as image modality information decoded from the EEG data by the hybrid modality encoder; The hybrid mapping network decodes the global features to obtain the category modality and the semantic modality; Map the image-category-semantic data in the image-category-semantic three-state EEG dataset to the corresponding CLIP space to obtain the target modality; Compare the image, category, and semantic modality information output by the hybrid modality encoder with the target modality, and perform image retrieval and classification tasks based on the image modality information; Construct a diffusion prior model to align the image, category, and semantic modality information output by the mixed modality encoder with the target modality in the CLIP space, and output the image, category, and semantic modality information processed by the diffusion prior. The images, categories, and semantic modal information that have been processed with diffusion priors are input into the pre-trained image generation model to perform the generation task. The modal information that has been processed with diffusion priors is used to provide the degree of representation in the coarse-grained direction during image generation, the category during image generation, and the degree of representation in the fine-grained direction during image generation to complete the reconstruction of visual stimulation.

2. The EEG visual stimulation mixed modality decoding method according to claim 1, characterized in that: The lightweight Transformer network includes a downsampling layer, an attention layer, and a convolution layer. Downsampling layers of different scales are formed by using average pooling layers of different sizes and different step lengths. The EEG data is decomposed into time pattern information of different scales through downsampling layers of different scales. The time pattern information of different scales is reshaped into the same shape through MLP and added to integrate into new pattern information. , based on the Embedding module, the new mode information Mapping to variables , learning variables through attention layers The correlation between different tokens is obtained by outputting the time series features extracted by the lightweight Transformer network through the convolutional layer .

3. The EEG visual stimulation mixed modality decoding method according to claim 1, characterized in that: The multi-time-scale hybrid convolutional network performs convolution operations on the time series features output by the lightweight Transformer network, extracts time series features of different time scales, reshapes them into time series feature vectors with the same dimensions through flattening operations and fully connected layers, and connects them into vector groups. ; For vector group Each vector of the channel is generated through convolution operation and sigmoid activation function, and the descriptors are combined into a new vector group , expressed as: ; Group vectors The vectors in the concatenation are Vector of dimensions ; vector exist The weight matrix is ​​obtained by the softmax function in the dimension , based on the weight matrix and the corresponding input vector, the global temporal features of the multi-time scale hybrid convolutional network output are obtained , expressed as: ; ; in, Indicates that the softmax function operates on the vector The operation is performed on the 0th dimension of represents the dot product operation, Representatives in order from 0 dimension to M -1 dimension for extraction operations.

4. The EEG visual stimulation mixed modality decoding method according to claim 1, characterized in that: The convolution network of the multi-feature convolution layer includes a multi-scale spatial convolution group and a dynamic feature fusion layer. The multi-scale spatial convolution group uses parallel deep convolution layers with different receptive fields to extract local microtopological associations and global macroscopic distribution patterns of the electrode array respectively. The dynamic feature fusion layer adaptively weights and fuses multi-scale spatial features through learnable channel attention weights to output temporal features.

5. The EEG visual stimulation mixed modality decoding method according to claim 4, characterized in that: In the multi-scale spatial convolution group, feature information of different dimensions of global temporal features are extracted through multiple convolution layers with different convolution kernel sizes, and the feature information of different dimensions is added, multi-scale information is fused, and channel descriptor representation is obtained through global pooling. The correlation between channels is modeled based on two fully connected layers. After the second fully connected layer is upgraded in dimension, it is evenly divided into multiple parts. In the dynamic feature fusion layer, the sigmoid function is used to generate their own weighted representations, and the information of different scales is weightedly added to output the temporal features.

6. The EEG visual stimulation mixed modality decoding method according to claim 1, characterized in that: The dual-stream spatiotemporal convolutional network in the time-frequency domain obtains frequency domain features through fast Fourier transform and ResMLP network modeling, and the time features and frequency domain features output by the convolutional network of the multi-feature convolutional layer are obtained through the sigmoid function to obtain the time domain weights and frequency domain weights, and the corresponding output features are weighted. After splicing, matrix multiplication operation is performed to obtain mixed time-frequency domain features, and global features are obtained based on the mixed time-frequency domain features, which are specifically expressed as: ; ; ; ; ; ; in, Represents time characteristics, represents the frequency domain characteristics, represents a linear operation, represents the sigmoid function, represents the average pooling operation, represents the feature concatenation operation, represents the mixed time-frequency domain features, Representation characteristics The transpose of represents the convolution operation, represents the convolution kernel size, Representative The dilation rate of the convolution operation, Represents a max pooling operation.

7. The EEG visual stimulation mixed modality decoding method according to claim 1, characterized in that: The hybrid mapping network includes a brain modality regression layer, a category embedder, a brain modality embedder, a semantic embedder, and a multimodal decoder; After passing the global features through the brain modality regression layer, they are input into the category embedder, brain modality embedder, and semantic embedder respectively to obtain the corresponding preliminary modality feature representation; Each preliminary modal feature representation is concatenated in the first dimension and input into the multimodal decoder for decoding to obtain the category modal feature and semantic modal feature.

8. The EEG visual stimulation mixed modality decoding method according to claim 1, characterized in that: The diffusion prior model is trained using mean squared error loss.

9. An EEG visual stimulus mixed modality decoding system, characterized in that: Used to implement the EEG visual stimulus mixed modality decoding method according to any one of claims 1 to 8, the system comprises: an EEG data acquisition module, a three-state EEG data set construction module, a mixed modality encoder construction module, a target modality construction module, a contrast learning module, an image retrieval classification module, a diffusion prior model construction module, and an image generation module; The EEG data acquisition module is used to acquire EEG data; The three-state EEG data set construction module is used to construct an image-category-semantic three-state EEG data set; The hybrid modality encoder construction module is used to construct a hybrid modality encoder, and the hybrid modality encoder learns corresponding image, category, and semantic modality information from EEG data; The target modality construction module is used to map the image-category-semantic data in the image-category-semantic three-state EEG data set to the corresponding CLIP space to obtain the target modality; The contrastive learning module is used to compare the image, category, and semantic modality information output by the hybrid modality encoder with the target modality; The image retrieval and classification module is used to perform image retrieval and classification tasks based on image modality information; The diffusion prior model construction module is used to construct a diffusion prior model, align the image, category, and semantic modality information output by the hybrid modality encoder with the target modality in the CLIP space, and output the image, category, and semantic modality information processed by the diffusion prior; The image generation module is used to input the image, category, and semantic modal information that have been processed by diffusion prior into a pre-trained image generation model to perform a generation task. The modal information that has been processed by diffusion prior is used to provide the degree of representation in the coarse-grained direction during image generation, the category during image generation, and the degree of representation in the fine-grained direction during image generation, so as to complete the reconstruction of visual stimulation.