Eeg image decoding method based on source space and visual layer priori self-correction

CN122817876APending Publication Date: 2026-09-25HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611088019.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0008]本发明的目的在于针对现有脑电图像解码方法存在的跨受试者泛化能力不足、脑电特征与图像特征匹配不充分以及共享语义表示空间稳定性较差等问题,提供一种基于源空间与视觉层先验自校正的脑电图像解码方法,以提高零样本条件下脑电图像解码的准确性和模型鲁棒性

Benefits of technology

[0020]与现有技术相比,本发明具有如下有益效果:本发明通过引入源空间区域可靠性向量及其诱导的视觉层先验分布,使脑电信号样本中的视觉相关空间响应模式能够参与图像侧多粒度视觉目标的构建过程;同时,通过受试者偏置项调节的路由机制与视觉层先验驱动的目标自校正机制,使所述图像侧多粒度视觉目标能够在训练阶段吸收受试者差异,并在推理阶段保持对未知受试者的适用性,从而有效提高零样本及跨受试者条件下脑电图像解码的解码准确性、模型鲁棒性和共享语义表示空间稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817876A_ABST
    Figure CN122817876A_ABST
Patent Text Reader

Abstract

The application discloses an electroencephalogram image decoding method based on source space and visual layer priori self-correction. The method first acquires the electroencephalogram signal of a subject when the subject watches an image, extracts the spatial response information of the visual-related brain area, and forms a source space area reliability vector. Secondly, the image to be watched is input into a pre-trained visual encoder to extract multi-layer visual features, and based on the source space area reliability vector, an image-side multi-granularity visual target is constructed. Then, the preprocessed electroencephalogram signal is input into an electroencephalogram coding network to extract electroencephalogram features, and the electroencephalogram features are trained in cross-modal alignment with the image-side multi-granularity visual target through a shared coding module. Finally, the electroencephalogram features of the electroencephalogram sample to be decoded and the candidate image are input into the shared coding module, the similarity score is calculated, and the picture decoding is completed. The application effectively improves the decoding accuracy, model robustness and shared semantic representation space stability of the electroencephalogram image decoding under the conditions of zero sample and cross-subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of EEG signal processing and brain-computer interface technology, and relates to an EEG image decoding method based on source space and visual prior self-correction. It can be used for zero-sample EEG visual decoding, cross-subject EEG feature modeling, and cross-modal matching of EEG and image semantic representation. Background Technology

[0002] Visual information decoding is an important research direction in the fields of brain-computer interfaces and cognitive neuroscience, aiming to infer the visual content perceived by an individual from brain activity signals. This technology has significant application value in scenarios such as assisted communication, intelligent interaction, neurological function assessment, and cognitive state analysis. Compared with neuroimaging methods such as functional magnetic resonance imaging, electroencephalography (EEG) signals have advantages such as high temporal resolution, low acquisition cost, portable equipment, and ease of real-time application. Therefore, research on visual decoding based on EEG signals has attracted widespread attention.

[0003] Existing EEG visual decoding methods mainly fall into two categories: classification-based decoding and retrieval-based decoding. With the development of large-scale visual representation models and cross-modal representation learning, EEG image decoding methods based on pre-trained visual feature spaces have gradually become an important research direction. These methods typically map EEG features to an image semantic space and achieve candidate image selection or semantic recovery through similarity matching, thereby supporting zero-shot decoding under conditions where no category has been seen. This approach avoids the problems of fixed categories and limited generalization ability in traditional closed-set classification methods, demonstrating significant potential in large-scale open-vocabulary visual decoding tasks.

[0004] However, existing methods still have significant shortcomings in practical applications. First, EEG signals are characterized by high noise, low spatial resolution, and significant individual differences. Different subjects often exhibit large deviations in their EEG responses when perceiving the same visual stimulus, leading to insufficient robustness and generalization ability of the model across different subjects. Second, most existing EEG image decoding methods use a single, fixed image feature as the supervision target, or combine features from multiple visual layers using fixed rules, assuming all subjects correspond to the same visual semantic granularity. Such methods struggle to adequately adapt to the multi-level, multi-scale visual representation information contained in EEG signals, easily causing semantic granularity mismatch between EEG features and visual targets.

[0005] Furthermore, even when existing methods incorporate multiple intermediate-layer visual features or model subject differences, they mostly perform weighting or routing only at the global level. They lack a mechanism to extract stable spatial priors from the spatial response patterns of EEG signals and further map these priors onto visual hierarchical representations for target construction. In other words, existing methods lack a decoding method that can transform the visually relevant spatial response patterns of EEG signal samples into visual-level prior distributions and utilize these priors for dynamic screening and self-correction of multi-granularity visual targets.

[0006] Furthermore, existing cross-modal alignment methods typically perform end-to-end optimization directly with the final decoding goal in mind, rarely distinguishing between the two stages of shared semantic representation space construction and fine-grained instance discrimination. In situations with strong EEG signal noise and significant subject variability, directly performing single-stage alignment optimization can easily lead the model to prematurely emphasize sample-level discrimination capabilities before the shared space has been stably established, thus affecting cross-modal alignment stability and final decoding performance. This problem is particularly pronounced under zero-sample and cross-subject conditions.

[0007] Therefore, how to construct an EEG image decoding method that can induce visual layer priors based on the reliability of source spatial regions, and dynamically screen and self-correct multi-granularity visual targets based on these priors, in order to improve the accuracy of zero-sample EEG image decoding and cross-subject generalization ability, has become an urgent technical problem to be solved in this field. Summary of the Invention

[0008] The purpose of this invention is to address the problems of insufficient cross-subject generalization ability, inadequate matching of EEG features and image features, and poor stability of shared semantic representation space in existing EEG image decoding methods. This invention provides an EEG image decoding method based on source space and visual layer prior self-correction to improve the accuracy and robustness of EEG image decoding under zero-sample conditions. Specifically, addressing the difficulties of existing methods in fully utilizing the spatial response patterns of visually relevant brain regions in EEG signal samples, effectively mapping these spatial response patterns to visual layer prior distributions, and maintaining stable decoding performance during the inference phase without relying on subject identity information, this invention proposes an EEG image decoding mechanism that induces a visual layer prior distribution using a source space region reliability vector, and dynamically filters and self-corrects multi-granularity visual targets on the image side based on this visual layer prior distribution.

[0009] Step S1: Acquire the image stimuli viewed by the subject and the corresponding EEG signals, construct an EEG-image pairing sample set, and preprocess the EEG signals; based on the preprocessing, further extract the spatial response information of visually related brain regions in the EEG signal samples to form a source spatial region reliability vector for subsequent visual target construction.

[0010] Step S2: Input the corresponding image stimulus viewed by the subject into the pre-trained visual encoder to extract multi-layer visual features; generate a visual layer prior distribution based on the source space region reliability vector, and determine the routing weights of the multiple intermediate layer visual features in combination with the subject bias term; perform target self-correction on the routing weights based on the visual layer prior distribution, and dynamically filter and weightedly fuse the multiple intermediate layer visual features to construct an image-side multi-granularity visual target for matching with EEG features.

[0011] Step S3: Input the preprocessed EEG signal into the EEG encoder to extract EEG features, and map the EEG features and the image-side multi-granularity visual targets to a unified shared semantic representation space through a shared coding module to obtain the EEG-side shared embedding representation and the image-side shared embedding representation.

[0012] Step S4: Perform cross-modal alignment training on the shared embedding representations on the EEG side and the shared embedding representations on the image side in the shared semantic representation space to improve the matching ability between EEG signal samples and corresponding image stimuli; the cross-modal alignment training adopts a two-stage optimization method from coarse to fine.

[0013] Step S5: After cross-modal alignment training, the EEG signal sample to be decoded is input into the trained model to obtain the EEG-side shared embedding representation of the EEG signal sample to be decoded; for candidate images, image-side multi-granularity visual targets corresponding to each candidate image are constructed based on the trained model, and the corresponding image-side shared embedding representations are obtained by mapping through the shared coding module; the similarity score between the EEG-side shared embedding representation and each image-side shared embedding representation is calculated, and the image corresponding to the EEG signal sample to be decoded is output from high to low according to the similarity score, thus completing the EEG image decoding.

[0014] The further improvements to the technical solution of this invention are as follows: In step S1, after completing the EEG signal preprocessing, the spatial response of multiple visually related source regions is further calculated based on the preprocessed EEG signal, and the source spatial region reliability vector is constructed from the response intensity of each visually related source region; the source spatial region reliability vector is used to characterize the visually related spatial response pattern of the current EEG signal sample.

[0015] In step S2, a visual layer prior distribution is generated based on the source space region reliability vector, and the visual layer prior distribution is used to dynamically filter and weighted fuse multiple intermediate layer visual features. At the same time, the routing weights in the fusion process are self-corrected based on the visual layer prior distribution to construct image-side multi-granularity visual targets, thereby enhancing the adaptability of image-side multi-granularity visual targets to the multi-level visual representation characteristics contained in EEG signals.

[0016] In step S2, when constructing the image-side multi-granularity visual target, a subject bias term is introduced to adjust the routing weights of multiple intermediate layer visual features, so that the image-side multi-granularity visual target can adapt to the differences in EEG responses among different subjects; at the same time, the routing weights are self-corrected based on the prior distribution of the visual layer to improve the matching degree between the image-side multi-granularity visual target and the EEG signal sample.

[0017] In step S2, during the model training phase, the routing weights are jointly determined by the globally shared routing term and the subject bias term to absorb the visual granularity bias corresponding to different subjects. During the model inference phase, the subject bias term and its corresponding subject bias gating coefficient are discarded, and only the globally shared routing term and the visual layer prior distribution generated by the current EEG signal sample are retained to achieve EEG image decoding without the need for subject identity information.

[0018] In step S3, EEG features and image-side multi-granularity visual targets are mapped to a unified shared semantic representation space through a shared coding module, thereby reducing the representation differences between EEG modalities and image modalities and improving the consistency of the shared semantic representation space.

[0019] In step S4, the cross-modal alignment training adopts a two-stage optimization approach from coarse to fine. First, a stable shared semantic representation space is constructed, and then the instance-level matching ability between EEG signal samples and corresponding image stimuli is enhanced, thereby improving the stability of the model training process and the decoding performance in cross-subject scenarios.

[0020] Compared with existing technologies, the present invention has the following beneficial effects: By introducing a source spatial region reliability vector and its induced visual layer prior distribution, the present invention enables visually relevant spatial response patterns in EEG signal samples to participate in the construction process of multi-granular visual targets on the image side; at the same time, through the routing mechanism adjusted by the subject bias term and the target self-correction mechanism driven by the visual layer prior, the multi-granular visual targets on the image side can absorb subject differences during the training phase and maintain applicability to unknown subjects during the inference phase, thereby effectively improving the decoding accuracy, model robustness, and shared semantic representation space stability of EEG image decoding under zero-sample and cross-subject conditions. Attached Figure Description

[0021] Figure 1 This is a general framework diagram of the present invention. Detailed Implementation

[0022] The present invention will be further described in detail below with reference to the technical solution of the present invention: This embodiment is implemented under the premise of the technical solution of the present invention, and provides a detailed implementation plan and specific operation process.

[0023] Studies have shown that EEG signals contain multi-level semantic information related to visual stimuli, but significant individual differences exist among different subjects. Furthermore, existing EEG image decoding methods typically employ fixed-granularity visual targets, making it difficult to fully adapt to the multi-level visual representation characteristics in EEG signals. Therefore, an EEG image decoding method based on source space and visual layer prior self-correction can effectively improve EEG image decoding performance under zero-sample and cross-subject conditions, which is of great significance for brain-computer interface and cognitive brain function research. This invention proposes an EEG image decoding method based on source space and visual layer prior self-correction, mainly involving four parts: a source space region reliability vector construction module, an image-side multi-granularity visual target construction module, an EEG encoder, and a coarse-to-fine cross-modal alignment module. Among them, the image-side multi-granularity visual target construction module generates a prior distribution of the visual layer using the spatial response patterns of visually related brain regions in the EEG signal samples, and performs dynamic screening and self-correction of the image-side multi-granularity visual targets based on the prior distribution of the visual layer; at the same time, during the training phase, the subject bias term is used to absorb the visual granularity differences between different subjects, and during the inference phase, the subject bias term is discarded to maintain applicability to unknown subjects. The implementation of the present invention mainly includes the following processes: (1) EEG signal-image stimulus acquisition and preprocessing, and construction of source spatial region reliability vector; (2) generation of visual layer prior distribution and construction of image-side multi-granularity visual targets; (3) cross-modal alignment training and EEG image decoding output.

[0024] This embodiment provides a method for decoding electroencephalograms based on prior self-correction of source space and visual layer, such as... Figure 1 As shown. Let the EEG image pairing training sample set be: .in, Indicates the first n One EEG signal sample, This represents the image stimulus corresponding to the electroencephalogram (EEG) signal sample. This indicates the subject identifier corresponding to the electroencephalogram (EEG) signal sample. Represents the total number of samples. Indicates the number of brainwave channels. Indicates the number of time sampling points.

[0025] Step S1: Acquire EEG signals and image stimuli, and preprocess the EEG signals.

[0026] First, EEG signals collected when subjects viewed image stimuli, along with the corresponding image stimuli, were acquired to construct an EEG-image paired sample set. The EEG signals were then preprocessed, including filtering, segmentation, baseline correction, downsampling, and normalization, to reduce the impact of noise and artifacts on subsequent feature extraction, resulting in preprocessed EEG signals.

[0027] In one embodiment, the electroencephalogram (EEG) signal is formed after being captured by a time window. A three-dimensional EEG matrix is ​​constructed, where the rows correspond to EEG channels and the columns correspond to time sampling points. After preprocessing, this EEG matrix serves as the input to the EEG encoder. To further characterize the effectiveness of visually relevant spatial responses in the current EEG signal samples, after EEG preprocessing, the preprocessed EEG signal samples are used as the input. Calculate its source space response. Let... The source space projection matrix is ​​represented by the source space projection matrix, which can be obtained from the standard head model and electrode topological mapping relationship. Then, the EEG signal sample... The corresponding source space response matrix is: in, Indicates the first The source spatial response matrix of each EEG signal sample. Further, the visually related source regions are divided into... Each sub-region is denoted as: .

[0028] in Indicates the first A visually related source region, .set up Indicates the first The first EEG signal sample in the... Source space response at each source point Indicates the first The number of source points within the i-th visually relevant source region, then for the i-th The operator that aggregates visually related source regions uses average pooling and is defined as follows: .

[0029] Then the first The first EEG signal sample in the... Regional reliability on a visually relevant source region can be defined as... , .

[0030] in, Indicates the first The first EEG signal sample in the... Regional reliability on visually relevant source regions This is a stable term. The samples are composed of the regional reliability of each visually relevant source region. Source space region reliability vector: .in, Indicates the first The source space region reliability vector of an EEG signal sample.

[0031] Step S2: Extract image features and construct multi-granularity visual targets on the image side.

[0032] Images corresponding to EEG signal samples Input a pre-trained visual encoder and extract visual features from multiple intermediate layers. Let the frozen pre-trained visual encoder be... The selected set of intermediate layers is as follows: ,in Indicates the number of intermediate layers selected. Indicates the first A selected visual layer. For input image stimuli In the Extract corresponding visual features from each selected layer: .in, Indicates the image stimulus at the 1st Visual features on an intermediate layer.

[0033] Since the visual features of different intermediate layers are at different representation scales and semantic levels, this implementation method projects the visual features of each layer separately to facilitate subsequent unified processing. Let the first layer be... k The layer-specific projector corresponding to the layer is Then the projected candidate visual representation is: ,in, This represents the first [characteristic] after unification into the same feature space. Layer candidate visual representation.

[0034] In this embodiment, the multi-granularity visual target on the image side is not directly composed of intermediate layer visual features, but is obtained by fusing multiple intermediate layer visual features. To achieve adaptive fusion, the source spatial region reliability vector obtained in step S1 is first used as the basis. Construct the prior distribution of the visual layer. Let... This represents the mapping matrix from the region to the visual layer. This represents the hierarchical consistency constraint matrix. Let represent the bias vector. Then, the prior distribution of the visual layer for the nth EEG signal sample is: in, This represents the Hadamard product. Indicates the first The visual layer prior distribution of EEG signal samples. The hierarchical consistency constraint matrix. It is used to characterize the correspondence between the visually related source region hierarchy and the visual feature hierarchy, so that the early visually related source regions tend to correspond to the shallower visual layer, and the high-level visually related source regions tend to correspond to the deeper visual layer.

[0035] Furthermore, let the first The target visual layer index corresponding to each visually relevant source region is: , The attenuation coefficient is the hierarchical consistency constraint matrix. The element is defined as: in, Indicates the first The visually related source region for the first The hierarchical consistency constraint strength of each visual layer. Through the hierarchical consistency constraint matrix, the spatial response patterns of EEG signal samples can be mapped to the prior distribution of the visual layer in a hierarchical consistent manner.

[0036] To control the involvement of subject bias terms in routing calculation, a subject bias gating coefficient is further constructed. Let... Indicates a globally shared routing entry. This represents the prior distribution of the globally shared routes. First, the prior distribution of the visual layer is calculated. Information entropy: , Let represent the component of the prior distribution vector of the visual layer at the k-th visual layer, and calculate the prior distribution of the visual layer. Jensen-Shannon divergence relative to the global shared route prior distribution: .

[0037] Based on the information entropy and Jensen-Shannon divergence, the prior confidence deviation of the nth EEG signal sample is defined as: The more concentrated the visual layer prior of the current EEG signal samples, and the more significant the difference from the global shared routing prior distribution, the greater the prior confidence deviation.

[0038] Furthermore, set and For parameters, Let denot the Sigmoid function, then the subject bias gating coefficient for the nth EEG signal sample is: in, This represents the subject bias gating coefficient for the nth EEG signal sample. A larger prior confidence deviation indicates that the current EEG signal sample needs to utilize the subject bias term to absorb individual differences. Take a larger value; conversely, The value is relatively small.

[0039] To enable the prior constraint strength of the vision layer on route calculation to adaptively change according to the current sample, the prior constraint strength of sample n is further defined as follows: .

[0040] in, This represents the prior constraint strength of the nth EEG signal sample. and These represent the lower and upper bounds of the preset constraint strength, respectively. The greater the prior confidence deviation of the current EEG signal sample, the stronger the constraint of the visual layer prior distribution on the routing weight calculation process.

[0041] Based on the globally shared routing term, the subject bias term, the subject bias gating coefficient, and the prior distribution of the visual layer, the initial routing weight of the nth EEG signal sample is defined as follows: in, Indicates with the subject The corresponding subject bias, Indicates temperature parameter, Indicates the first The initial routing weight distribution of each EEG signal sample in the intermediate layers. Thus, the spatial response pattern of the current sample not only determines the prior distribution of the visual layer, but also the degree of participation of the subject's bias in the routing calculation and the strength of the constraint of the visual layer prior on the routing calculation.

[0042] Furthermore, to enhance the consistency between the multi-granularity visual targets in the image and the prior distribution of the visual layer of the current EEG signal samples, based on the prior distribution of the visual layer... Initial route weights Perform target self-calibration. Define the first... The self-correction coefficient for each EEG signal sample is: .in Indicates the first The self-correction coefficients of each EEG signal sample are calculated. The more concentrated the prior distribution of the visual layer, the larger the self-correction coefficient; the flatter the prior distribution of the visual layer, the smaller the self-correction coefficient. Based on these self-correction coefficients, corrected routing weights are constructed: in, Indicates the first The corrected routing weight distribution of each EEG signal sample. Through the aforementioned target self-correction mechanism, the final routing weights can maintain higher consistency with the visual layer prior induced by the source space region reliability vector while preserving the modeling of subject differences. Furthermore, to avoid the fusion routing weights being concentrated in a few visual layers, the distribution can be based on the visual layer prior distribution. From the above From the intermediate visual features, the top m candidate visual layers with the highest prior weights are selected, and weighted fusion is performed only within these candidate visual layers. Let the final normalized routing weights be... Then the multi-granularity visual targets on the image side are: in, Indicates the first The image side of the multi-granular visual target corresponding to the EEG signal sample.

[0043] During the inference phase, the subject bias gating coefficient is discarded, and only the globally shared routing term is used to construct multi-granularity visual targets on the image side; at the same time, the source space region reliability vector calculated from the current EEG signal sample itself is retained. and its induced visual layer prior distribution Furthermore, the prior distribution of the visual layer is used to perform target self-correction on the routing weights, thereby enhancing the matching between image-side multi-granular visual targets and EEG signal samples without relying on the subject's identity input. This forms an image-side multi-granular visual target generation mechanism that is subject-related during the training phase, subject-independent during the inference phase, and driven by source space and visual layer prior self-correction.

[0044] Step S3: Extract EEG features and map them to a shared semantic representation space.

[0045] The preprocessed EEG signal from step S1 is input into the EEG encoder to extract EEG features. Let the EEG encoder be... The EEG projection module is The EEG representation is as follows: .in, This represents the feature representation of an EEG signal sample after encoding and projection.

[0046] In step S2, a priori distribution of the visual layer is generated based on the reliability vector of the source space region, and a multi-granularity visual target on the image side is constructed through dynamic screening and target self-correction mechanism. To reduce the representational differences between EEG modalities and image modalities, this implementation method includes a shared encoding module. A unified mapping was performed on the EEG-side representation and the image-side multi-granularity visual targets to obtain the representation in the shared semantic representation space: in, This indicates shared embedding representations on the EEG side. This represents a shared embedding representation on the image side. By applying the same transformation to both modalities through a shared coding module, EEG features and image features can have a more consistent geometric structure in the shared semantic representation space.

[0047] Step S4: Perform cross-modal alignment training on the shared semantic representation space After obtaining the shared embedding representations on the EEG side and the shared embedding representations on the image side, cross-modal alignment training is performed on both. The cross-modal alignment training adopts a two-stage optimization approach from coarse to fine.

[0048] In the coarse alignment stage, a bidirectional contrastive retrieval loss is used to constrain the matching relationship between EEG signal samples and their corresponding image stimuli. Assume a batch contains... M For EEG-image paired samples, the bidirectional contrastive retrieval loss is: in, Represents the cosine similarity function. This indicates the temperature parameter being retrieved. The bidirectional contrastive retrieval loss simultaneously constrains the similarity relationship between paired samples in both the EEG-to-image and image-to-EEG directions.

[0049] To further reduce the distribution differences between EEG modalities and image modalities in the shared semantic representation space, a multi-kernel maximum mean difference loss is introduced during the coarse alignment stage: in, This represents the multinucleus radial basis function.

[0050] Furthermore, in order to ensure that the visual layer prior distribution induced by the source space region reliability vector in step S2 is... It can continuously constrain the construction results of multi-granularity visual targets on the image side during training, and introduce consistency constraint loss in the coarse alignment stage. Let... Indicates the first The corrected routing weight distribution of EEG signal samples is then defined as follows: ,in This represents the Kullback-Leibler divergence. The prior consistency constraint loss is used to measure the deviation between the corrected routing weight distribution and the prior distribution of the visual layer, thereby ensuring that the image-side multi-granularity visual targets maintain consistency with the spatial response patterns of EEG signal samples during training.

[0051] Based on the bidirectional comparison retrieval loss, multi-core maximum mean difference loss, and prior consistency constraint loss, the overall optimization objective for the coarse alignment stage is: in, This represents the weight coefficient of the corresponding loss term. Through training in this stage, a relatively stable shared semantic representation space structure can be established first, reducing cross-modal distribution bias and ensuring consistency between the multi-granularity visual targets on the image side and the prior visual layer distribution induced by the reliability vector of the source space region. In the fine alignment stage, the shared encoding module is frozen, and the model is further optimized using only the bidirectional contrastive retrieval loss. Its objective function is: .

[0052] Through the fine alignment stage, the instance-level discrimination ability between EEG signal samples and image stimuli can be further improved on the basis that the shared semantic representation space has been basically stabilized.

[0053] Step S5: Output the EEG image decoding results.

[0054] After model training is complete, the EEG signal samples to be decoded are input into the EEG encoder, EEG projection module, and shared encoding module to obtain the EEG-side shared embedding representation of the EEG signal samples to be decoded. Simultaneously, for each image in the candidate image set, multiple intermediate layer visual features are extracted using a pre-trained visual encoder and layer-specific projector, and combined with the source space region reliability vector calculated from the current EEG signal sample itself. Generate visual layer prior distribution During the inference phase, the subject bias term is discarded, and only the globally shared routing term is retained. The prior distribution of the visual layer is used to dynamically filter, weightedly fuse, and self-correct the visual features of multiple intermediate layers of the candidate image, thereby constructing an image-side multi-granularity visual target corresponding to the current EEG signal sample to be decoded.

[0055] Further, the image-side multi-granularity visual target is input into the shared encoding module to obtain the image-side shared embedding representation corresponding to the candidate image. Then, the similarity between the EEG-side shared embedding representation of the EEG signal sample to be decoded and the image-side shared embedding representation of each candidate image is calculated, and the images are sorted from high to low similarity. The images with the highest similarity are output as the EEG image decoding results.

[0056] Therefore, during the inference phase, this invention does not require inputting subject identity information. Instead, it relies solely on the source space region reliability vector of the current EEG signal sample to be decoded and its induced visual layer prior distribution to adaptively construct and self-correct multi-granularity visual targets on the image side. This enhances the matching between multi-granularity visual targets on the image side and EEG signal samples, and improves the decoding accuracy and model robustness of EEG image decoding under zero-sample and cross-subject conditions.

[0057] Description of a specific embodiment: In one specific embodiment, multiple intermediate layers of a pre-trained visual encoder can be selected as candidate visual layers. The visual features of each layer are mapped to a unified dimensional space using corresponding layer-specific projectors. On the EEG side, the EEG representation is obtained through the EEG encoder and EEG projection module, and the EEG representation and the multi-granularity visual targets on the image side are uniformly mapped to a shared semantic representation space through a shared encoding module. Unlike existing methods that directly construct image-side supervised targets based on fixed visual layers or fixed fusion rules, this embodiment first constructs a source space region reliability vector based on the source space response of the EEG signal sample itself, and generates a visual layer prior distribution based on the source space region reliability vector. Then, the candidate visual layers are dynamically screened using the visual layer prior distribution, and multi-granularity visual targets on the image side are constructed by combining the subject bias term and the target self-correction mechanism.

[0058] During the training phase, a coarse alignment stage, including bidirectional contrast retrieval loss, multi-kernel maximum mean difference loss, and prior consistency constraint loss, is first used to establish a shared semantic representation space. Then, a fine alignment stage is used to enhance instance-level matching capabilities. Among these, the subject bias term is used to absorb the visual granularity differences between different subjects, while the visual layer prior distribution is used to constrain the routing weights and target self-correction process in the image-side multi-granularity visual target construction process.

[0059] During the inference phase, subject bias terms are discarded, retaining only the globally shared routing terms. The reliability vector of the source spatial region generated by the current EEG signal sample and its induced prior visual layer distribution are used to dynamically filter candidate visual layers. Furthermore, multi-granularity visual targets on the image side undergo target self-correction, thereby completing EEG image decoding without relying on subject identity input. Thus, this embodiment can simultaneously utilize subject difference information from the training phase and visually relevant spatial response pattern information of the current EEG signal sample during the inference phase, improving decoding accuracy and model robustness under zero-sample and cross-subject conditions.

[0060] As shown in Table 1, in the zero-shot 200-class concept decoding task of the THINGS-EEG dataset, the method of this invention achieved a Top-1 accuracy of 91.3% and a Top-5 accuracy of 98.8% under the in-subject setting. Compared with the current state-of-the-art methods, the Top-1 accuracy of the method of this invention is improved by 8.7% under the in-subject setting, indicating that the method of this invention can more effectively utilize stable EEG response patterns within the same subject to achieve more accurate visual semantic decoding.

[0061] As shown in Table 2, in the inter-subject setting, the method of this invention achieved a Top-1 accuracy of 34.4% and a Top-5 accuracy of 64.8%. Compared with the current state-of-the-art methods, the Top-1 accuracy of the method of this invention is improved by 12.0% in the inter-subject setting, indicating that the method of this invention has stronger generalization ability in cross-subject scenarios and can maintain high image decoding accuracy even when there are significant differences in EEG among different subjects.

[0062] Table 1: Comparison of performance of in-subject settings and various optimal methods

[0063] Table 2: Comparison of Subject-to-Subject Settings and Performance of Various Optimal Methods

[0064] The embodiments described above are merely preferred examples of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for decoding electroencephalograms based on prior self-correction of source space and visual layer, characterized in that, Includes the following steps: Step S1: Obtain EEG signal data when the subject views the image, and preprocess the EEG signals; extract spatial response information of visually related brain regions to form a source spatial region reliability vector; Step S2: Input the viewed image into the pre-trained visual encoder, extract multi-layer visual features, generate the visual layer prior distribution based on the source space region reliability vector, and construct the image-side multi-granularity visual target. Step S3: The preprocessed EEG signal is input into the EEG coding network to extract EEG features. The EEG features and multi-granularity visual targets on the image side are mapped to a unified shared semantic representation space through the shared coding module. Step S4: Perform cross-modal alignment training on EEG features and image features in the shared semantic representation space; Step S5: After cross-modal alignment training, the EEG features of the EEG sample to be decoded and the candidate image are processed through a shared encoding module, and a similarity score is calculated. The image corresponding to the subject's EEG segment is output based on the highest score to complete the decoding.

2. The EEG image decoding method based on source space and visual layer prior self-correction according to claim 1, characterized in that, The specific implementation of step S1 is as follows: Training sample set is ,in, Indicates the first n One EEG sample, This represents the image stimulus corresponding to the EEG sample. This indicates the subject identifier corresponding to the EEG sample. N Represents the total number of samples; based on preprocessed EEG samples. Calculate the source space response matrix by mapping the source space projection matrix. The visually relevant source region is divided into R sub-regions, denoted as . ,in Let r-th visually relevant source region be denoted as r-th; let... Indicates the first The first EEG signal sample in the... Source space response at each source point This represents the number of source points within the r-th visually relevant source region. Average pooling is used to aggregate the r-th visually relevant source region, defined as... The regional reliability of sample n in the r-th visually relevant source region can be defined as follows: ,in, This indicates the regional reliability of sample n in the r-th visually relevant source region. For the stable term; the source space region reliability vector of sample n is composed of the region reliability on each visually relevant source region. .

3. The EEG image decoding method based on source space and visual layer prior self-correction according to claim 2, characterized in that, Step S2 is specifically implemented as follows: Let the frozen pre-trained visual encoder be... The selected set of intermediate layers is: ,in K Indicates the number of intermediate layers selected. Indicates the first k A selected visual layer, for the input image In the k Extract corresponding visual features from selected layers. Indicates the image at the 1st k Visual features on an intermediate layer; Projecting the visual features of each layer separately, let the first layer be... k The layer-specific projector corresponding to the layer is Then the projected candidate visual representation is ; set up This represents the mapping matrix from the region to the visual layer. This represents the hierarchical consistency constraint matrix. Let represent the bias vector, then the prior distribution of the visual layer for sample n is: in, Represents the Hadamard product; Let q represent a globally shared routing entry. Represents the prior distribution of the globally shared routes; first, the prior distribution of the visual layer is calculated. Information entropy And calculate the prior distribution of the visual layer. Jensen-Shannon divergence relative to the global shared route prior distribution Based on information entropy and Jensen-Shannon divergence, the prior confidence deviation of sample n is defined as... ; set up and For parameters, Let represent the Sigmoid function, then the bias gating coefficient of the sample is... The prior constraint strength of sample n is defined as follows: ;in and These represent the lower and upper bounds of the preset constraint strength, respectively; Based on the globally shared routing term, the subject bias term, the subject bias gating coefficient, and the visual layer prior distribution, the initial routing weight for sample n is defined as follows: in, Indicates with the subject The corresponding bias term, Indicates temperature parameter, This represents the initial fusion weight distribution of sample n in each intermediate layer; based on the prior distribution of the visual layer. Initial route weights Perform target self-calibration: Define the self-calibration coefficient for sample n as... Based on the self-correction coefficient, the corrected routing weights are constructed. Based on the prior distribution of the visual layer From the K intermediate visual features, the candidate visual layers with the top m prior weights are selected, and weighted fusion is performed only within these candidate visual layers. Let the final normalized routing weight be... Then the multi-granularity visual target on the image side is .

4. The EEG image decoding method based on source space and visual layer prior self-correction according to claim 3, characterized in that, Step S2 further includes, during the inference phase, discarding the subject bias gating coefficient and using only the globally shared routing term to construct the image-side visual target; simultaneously, retaining the source space region reliability vector calculated from the current EEG sample itself. and visual layer prior distribution Furthermore, the prior distribution of the visual layer is used to perform target self-correction on the routing weights.

5. The EEG image decoding method based on source space and visual layer prior self-correction according to claim 4, characterized in that, The specific implementation of step S3 is as follows: The preprocessed brain signals are input into the EEG coding network to extract EEG features. Configure shared encoding module By performing unified mapping on the EEG-side representation and the image-side visual target respectively, the representation in the shared semantic space is obtained: in, This indicates shared embedding representations on the EEG side. This indicates that the image side shares the embedding representation.

6. The EEG image decoding method based on source space and visual layer prior self-correction according to claim 5, characterized in that, Step S4 is specifically implemented as follows: The cross-modal alignment training employs a two-stage optimization approach, from coarse to fine. In the coarse alignment stage, a bidirectional contrastive decoding loss is used to constrain the matching relationship between EEG samples and their corresponding image samples. Assuming a batch contains a total of... M For EEG and image samples, the decoding loss is: in, Represents the cosine similarity function; Introduce multi-core maximum mean difference loss during the coarse alignment stage: in, Represents the kernel function of a multi-kernel radial basis function; In the coarse alignment stage, a consistency constraint loss is introduced, let... Let n represent the corrected routing weight distribution. Then, the prior consistency constraint loss is defined as: ,in Indicates the Kullback-Leibler divergence; The overall optimization objective for the coarse alignment stage is obtained by weighted summation of bidirectional comparison decoding loss, multi-core maximum mean difference loss, and prior consistency constraint loss. During the fine alignment phase, the shared encoder is frozen, and the model is further optimized using only the retrieval loss.