An electroencephalogram image reconstruction method and system based on flow matching and manifold alignment

CN122473292BActive Publication Date: 2026-09-29GUANGDONG ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY LAB (GUANGZHOU)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610952946.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-29
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种基于流匹配与流形对齐的脑电图像重构方法及系统,用于解决现有脑电图像重构技术特征表征能力不足,以及生成过程稳定性差且推理效率低的问题

Benefits of technology

[0018]本发明与现有技术相比,其有益效果在于:本发明通过动态图拓扑建模结合正交化正则约束的方式,改善了脑电特征提取过程中的信息损失问题。动态图拓扑建模能够根据输入信号实时推导样本特异性的邻接矩阵,捕捉不同脑区之间的瞬态功能连接关系,相比传统静态序列建模方法更符合神经活动的生理特性。正交化正则约束通过最小化特征维度间的互相关性,强制特征分量保持相对独立,避免了图神经网络聚合过程中常见的特征趋同现象。实验结果显示,加入该约束后,模型在零样本检索任务中的准确率有所提升,验证了其对特征表征能力的增强作用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473292B_ABST
    Figure CN122473292B_ABST
Patent Text Reader

Abstract

The application discloses a kind of electroencephalogram image reconstruction method and system based on flow matching and manifold alignment, belong to biomedical signal processing and brain-computer interface technical field.Method includes: the original electroencephalogram signal is preprocessed operation, obtains standardized electroencephalogram data;The standardized electroencephalogram data is dynamically graph topological feature modeling processing, obtains initial electroencephalogram embedding representation;The initial electroencephalogram embedding representation is orthogonal regular constraint processing, obtains de-correlated electroencephalogram feature;The de-correlated electroencephalogram feature is prototype-guided semantic manifold rectification processing, obtains semantic alignment electroencephalogram feature;The semantic alignment electroencephalogram feature is condition flow matching generation processing, obtains corresponding visual feature;The visual feature is image decoding processing, generates and outputs and the reconstruction image corresponding to original electroencephalogram signal.The application can be used for non-invasive brain-computer interface electroencephalogram decoding and visual image reconstruction, applicable to neural repair, mental state assessment scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical signal processing and brain-computer interface technology, specifically involving cross-modal semantic modeling of electroencephalogram (EEG) signals and visual image reconstruction technology. Background Technology

[0002] Electroencephalography (EEG), as an important means of recording physiological signals of brain neural activity, occupies an important position in the field of non-invasive brain-computer interfaces. In recent years, decoding EEG signals into visual images has become a cutting-edge research direction in this field, with significant application value in neural repair, mental state assessment, and human-computer interaction. Existing reconstruction techniques typically involve two main steps: first, extracting EEG features through an encoder and mapping them to the feature space of a pre-trained visual model; then, using a generative model to reconstruct the image based on these features.

[0003] Existing technologies have several shortcomings in reconstructing complex natural images. Traditional methods often treat EEG signals as independent channel sequences or static time windows, failing to explicitly model the rapidly changing functional connections between neurons in the cerebral cortex. Some techniques introduce graph neural networks to capture the spatial relationships between channels, but due to the strong correlation and low signal-to-noise ratio between EEG channels, graph neural networks are prone to oversmoothing when aggregating graph node information, leading to convergence of node features. This phenomenon manifests in the latent space as high-dimensional features having discriminative power only in a very few dimensions, with a large amount of effective information being redundantly covered, limiting the representational capabilities of the EEG encoder.

[0004] Due to the low information density and strong non-stationarity of EEG signals, a significant semantic gap exists between the EEG feature space and the visual semantic space. Existing cross-modal alignment methods mostly rely on end-to-end contrastive learning. However, in the absence of explicit semantic anchors, unconstrained EEG embeddings are often loosely distributed in the latent space, failing to form tight semantic clusters. This semantic ambiguity makes the model susceptible to noise interference during image retrieval or reconstruction, leading to erroneous semantic associations. Consequently, the decoded features deviate from the true support set of the visual manifold, resulting in semantic distortion during reconstruction.

[0005] Current mainstream visual reconstruction schemes typically employ latent diffusion models as generators. The essence of diffusion models is image generation through random backsampling. Under weak constraints such as EEG signals, the lack of a deterministic guiding path easily leads to random semantic shifts in the generated trajectory. This results in the same EEG segment potentially generating images with completely different semantics, lacking consistency in reconstruction. Furthermore, diffusion models usually require multiple iterative sampling steps, making each reconstruction time-consuming and difficult to meet the low latency and rapid response requirements of intraoperative neural monitoring or real-time assistance systems. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for reconstructing electroencephalogram (EEG) images based on flow matching and manifold alignment, in order to solve the problems of insufficient feature representation capabilities, poor stability of the generation process, and low inference efficiency in existing EEG image reconstruction technologies.

[0007] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for reconstructing electroencephalogram (EEG) images based on flow matching and manifold alignment, comprising the following steps: The raw EEG signals are preprocessed to obtain standardized EEG data; The standardized EEG data are subjected to dynamic graph topology feature modeling to obtain an initial EEG embedding representation; The initial EEG embedding representation is subjected to orthogonalization and regularization constraints to obtain decorrelated EEG features; The decorrelated EEG features are subjected to prototype-guided semantic manifold correction processing to obtain semantically aligned EEG features; The semantically aligned EEG features are subjected to conditional flow matching to generate corresponding visual features. The visual features are processed by image decoding to generate and output a reconstructed image corresponding to the original EEG signal.

[0008] In one possible implementation, the step of preprocessing the raw EEG signal to obtain standardized EEG data includes: The raw EEG signal was processed by bandpass filtering and notch filtering to remove power frequency interference and noise. The filtered EEG signal was subjected to time window truncation to extract signal segments related to visual stimulus response; The extracted EEG signals were format-converted and standardized to obtain standardized EEG data in a unified tensor format.

[0009] In one possible implementation, the step of performing dynamic graph topological feature modeling on the standardized EEG data to obtain an initial EEG embedding representation includes: The standardized EEG data were subjected to time-dimensional statistical aggregation processing to obtain the global response features of each channel; The inter-channel correlation matrix is ​​obtained by performing an inner product operation on the global response features of each channel. The correlation matrix between the channels is normalized to obtain a sample-specific dynamic adjacency matrix. Graph convolution operations are performed on the dynamic adjacency matrix and standardized EEG data to obtain EEG features that fuse topological information. Set a learnable residual adjustment factor; Based on the residual adjustment coefficient, the standardized EEG data and the EEG features of the fused topological information are weighted and summed to obtain the initial EEG embedding representation.

[0010] In one possible implementation, the step of performing orthogonalization and regularization constraints on the initial EEG embedding representation to obtain decorrelated EEG features includes: During the training phase, the initial EEG embedding representations within a batch are processed by autocorrelation matrix calculation to obtain the batch autocorrelation matrix. The first loss component is obtained by calculating the sum of squares of the differences between the diagonal elements of the batch autocorrelation matrix and 1. The second loss component is obtained by calculating the sum of squares of the off-diagonal elements of the batch autocorrelation matrix. The first loss component and the weighted second loss component are summed to obtain the Barlow Twins loss function value; By minimizing the Barlow Twins loss function value, the features of each dimension of the initial EEG embedding representation are forced to decorrelate, resulting in decorrelated EEG features.

[0011] One possible implementation also includes: Extract visual features from all images in the training set; The visual features were clustered using the K-means clustering algorithm to obtain K cluster centers; The K cluster centers are used as visual semantic prototypes and stored in a learnable memory to construct a visual prototype library.

[0012] In one possible implementation, the step of performing prototype-guided semantic manifold correction processing on the decorrelated EEG features to obtain semantically aligned EEG features includes: The similarity between the decorrelation EEG features and each visual semantic prototype in the visual prototype library is calculated to obtain the attention weight distribution; Based on the attention weight distribution, the visual semantic prototypes in the visual prototype library are weighted and summed to obtain anchor features. By using learnable gating parameters, fusion weights are assigned to the decorrelation EEG features and anchor features respectively; Based on the fusion weights, the decorrelation EEG features and anchor features are weighted and summed to obtain semantically aligned EEG features.

[0013] In one possible implementation, the training step of performing conditional flow matching generation processing on the semantically aligned EEG features includes: The semantically aligned EEG features are used as source distribution features, and the corresponding real image features are used as target distribution features. A linear interpolation path is constructed between the source distribution features and the target distribution features to obtain intermediate features that evolve continuously over time. Using the semantically aligned EEG features as conditional variables, a neural velocity estimator with a conditional U-Net structure is constructed. The mean square error between the velocity vector predicted by the neural velocity estimator and the actual semantic change velocity is calculated to obtain the conditional flow matching loss function value; By minimizing the value of the conditional flow matching loss function, the parameters of the neural velocity estimator are optimized, resulting in a trained neural velocity estimator.

[0014] In one possible implementation, the step of performing conditional flow matching to generate the corresponding visual features from the semantically aligned EEG features includes: During the reasoning phase, the semantically aligned EEG features are used as the initial state; Discretize the time interval into a preset number of time steps; At each time step, the trained neural velocity estimator is used to predict the rate of change of the current feature state; Based on the change rate and time step, perform Euler iterative update processing to obtain the feature state at the next moment; By iterating through all time steps, the ordinary differential equations are solved to obtain the visual features of the target.

[0015] In one possible implementation, the step of performing image decoding processing on the visual features to generate and output a reconstructed image corresponding to the original EEG signal includes: A pre-trained image decoding model is constructed by using a frozen diffusion model backbone network in conjunction with an image adapter; The pre-trained image decoding model is used to perform pixel-level reconstruction of the visual features; Generate and output a reconstructed image corresponding to the original EEG signal.

[0016] One possible implementation also includes: During the training phase, joint optimization processing is performed simultaneously, including dynamic graph topology feature modeling, orthogonalization regularization constraints, semantic manifold correction, and conditional flow matching generation. By using a joint loss function, the parameters of the EEG encoder, semantic correction module, and generation module are optimized simultaneously.

[0017] Secondly, the present invention provides an EEG image reconstruction system based on flow matching and manifold alignment, comprising: The signal preprocessing module is used to preprocess the raw EEG signals to obtain standardized EEG data. The dynamic graph topology modeling module is used to perform dynamic graph topology feature modeling processing on the standardized EEG data to obtain an initial EEG embedding representation. The orthogonalization regularization constraint module is used to perform orthogonalization regularization constraint processing on the initial EEG embedding representation to obtain decorrelated EEG features; A prototype-guided manifold correction module is used to perform prototype-guided semantic manifold correction processing on the decorrelation EEG features to obtain semantically aligned EEG features. The conditional flow matching generation module is used to perform conditional flow matching generation processing on the semantically aligned EEG features to obtain the corresponding visual features. The image decoding module is used to perform image decoding processing on the visual features, generate and output a reconstructed image corresponding to the original EEG signal.

[0018] Compared with existing technologies, the advantages of this invention are as follows: This invention improves the information loss problem in the EEG feature extraction process by combining dynamic graph topology modeling with orthogonalization regularization constraints. Dynamic graph topology modeling can derive sample-specific adjacency matrices in real time based on input signals, capturing transient functional connectivity relationships between different brain regions, which is more consistent with the physiological characteristics of neural activity compared to traditional static sequence modeling methods. Orthogonalization regularization constraints, by minimizing the cross-correlation between feature dimensions, force feature components to remain relatively independent, avoiding the feature convergence phenomenon commonly seen in graph neural network aggregation. Experimental results show that after adding this constraint, the model's accuracy in zero-shot retrieval tasks is improved, verifying its enhancement effect on feature representation capabilities.

[0019] This invention employs a prototype-guided semantic manifold correction mechanism to improve the accuracy of cross-modal semantic mapping. This mechanism uses stable semantic prototypes obtained from clustering large amounts of visual data as anchors, calculates the correlation between EEG features and each prototype through an attention mechanism, and corrects the original EEG embeddings through gating fusion. This approach can cluster loosely distributed EEG features towards corresponding visual semantic centers, forming a more discriminative feature distribution. Compared to methods that simply rely on end-to-end contrastive learning, this effectively reduces semantic ambiguity. Visual analysis shows that the corrected EEG embeddings form well-defined semantic clusters in the latent space, improving the accuracy of cross-modal matching.

[0020] This invention constructs an image generation paradigm based on conditional flow matching, improving the stability and inference efficiency of the generation process. This method models cross-modal generation as an optimal transmission problem from EEG distribution to visual distribution, replacing the random sampling process of traditional diffusion models by learning deterministic linear transmission paths, thus eliminating semantic drift that easily occurs under weak constraints. Simultaneously, by employing ordinary differential equations for solving, the number of iterations required for the inference process is significantly reduced. Experimental data shows that, while maintaining generation quality, the inference speed of this method is significantly improved compared to traditional diffusion models, better meeting the needs of real-time application scenarios. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating the EEG image reconstruction method based on flow matching and manifold alignment according to an embodiment of the present invention. Figure 2 The diagrams and flowcharts illustrate the principles of dynamic graph topological feature modeling, orthogonalization regularization constraint processing, and prototype-guided semantic manifold correction processing in embodiments of the present invention. Figure 2 (a) is a schematic diagram of the overall architecture of the processing procedure. Figure 2 (b) is a schematic diagram of the collaborative process of the processing; Figure 3 This is a schematic diagram and flowchart illustrating the principle of dynamic graph topology feature modeling processing according to an embodiment of the present invention. Figure 3 (a) A structural schematic diagram for dynamic graph topology modeling. Figure 3 (b) Processing flowchart for modeling topological features of dynamic graphs; Figure 4 The diagram and flowchart illustrate the principle of prototype-guided semantic manifold correction processing in this embodiment of the invention. Figure 4 (a) is a schematic diagram of the semantic manifold correction architecture. Figure 4 (b) is a flowchart of semantic manifold correction processing; Figure 5 The present invention provides a schematic diagram and flowchart of the conditional flow matching generation process according to an embodiment of the present invention. Figure 5 (a) is the architecture diagram generated for conditional flow matching. Figure 5 (b) Processing flowchart generated for conditional flow matching; Figure 6 This is a schematic diagram of the EEG image reconstruction effect output by the image decoding module in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0024] Example: It should be noted that the terms "comprising" and "having" and any variations thereof in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products, or devices.

[0025] See Figure 1 This embodiment of an EEG image reconstruction method based on flow matching and manifold alignment includes the following steps: Step 101: Perform preprocessing on the raw EEG signals to obtain standardized EEG data.

[0026] Specifically, the raw EEG signal can be the electrical signal of brain neural activity recorded by a multi-channel EEG acquisition device; the standardized EEG data can be EEG tensor data after filtering, truncation, and format standardization. For example, the raw EEG signal is acquired using a 63-channel international 1020 system; the standardized EEG data format is sample number × channel number × time length.

[0027] The step of preprocessing the raw EEG signals to obtain standardized EEG data includes: The raw EEG signal was processed by bandpass filtering and notch filtering to remove power frequency interference and noise. The filtered EEG signal was subjected to time window truncation to extract signal segments related to visual stimulus response; The extracted EEG signals were format-converted and standardized to obtain standardized EEG data in a unified tensor format.

[0028] Specifically, bandpass filtering can be a filtering operation that preserves EEG signals within a specific frequency range; notch filtering can be a filtering operation that removes power line interference at a specific frequency; time window truncation can be an operation that extracts EEG signals for a specific time period after visual stimulation; format conversion can be an operation that converts EEG data into a tensor format supported by a deep learning framework; and normalization can be an operation that normalizes the amplitude of EEG signals to a specific range. For example, the frequency range of bandpass filtering is 0.1Hz to 70Hz; the frequency of notch filtering is 50Hz; the time window truncation range is 0 seconds to 1 second after visual stimulation; the unified tensor format is sample number × 63 × 250; and the normalization process uses the Z-score normalization method.

[0029] Step 102: Perform dynamic graph topology feature modeling on the standardized EEG data to obtain an initial EEG embedding representation.

[0030] Specifically, dynamic graph topological feature modeling can be a feature extraction process involving constructing a sample-specific dynamic adjacency matrix and performing graph convolution; the initial EEG embedding representation can be a high-dimensional feature vector after topological feature extraction. For example, dynamic graph topological feature modeling uses dynamic graph convolution operators.

[0031] The step of performing dynamic graph topology feature modeling on the standardized EEG data to obtain an initial EEG embedding representation includes: The standardized EEG data were subjected to time-dimensional statistical aggregation processing to obtain the global response features of each channel; The inter-channel correlation matrix is ​​obtained by performing an inner product operation on the global response features of each channel. The correlation matrix between the channels is normalized to obtain a sample-specific dynamic adjacency matrix. Graph convolution operations are performed on the dynamic adjacency matrix and standardized EEG data to obtain EEG features that fuse topological information. Set a learnable residual adjustment factor; Based on the residual adjustment coefficient, the standardized EEG data and the EEG features of the fused topological information are weighted and summed to obtain the initial EEG embedding representation.

[0032] Specifically, time-dimensional statistical convergence processing can be an operation of performing global statistics on the time series of each channel; global response features can be the overall neural activity features of each channel within a time window; the inter-channel correlation matrix can be a matrix reflecting the functional connectivity strength between different EEG channels; normalization processing can be an operation of mapping matrix elements to the range of 0 to 1; and the dynamic adjacency matrix can be an adjacency matrix representing the instantaneous functional connectivity relationships of the brain regions in the current sample. In the formula, The sample-specific dynamic adjacency matrix is ​​represented by Softmax(), which represents the normalization exponential function. Let represent the global response feature matrix of each channel, and d represent the feature dimension; graph convolution operation can be a feature extraction operation based on graph structure for information propagation. In the formula, EEG features representing fused topological information Represents a standardized EEG data matrix. The graph convolution weight matrix represents the EEG features that fuse topological information, which can be EEG features that fuse spatial connectivity; the residual adjustment coefficient represents the EEG features that fuse topological information. It can be a learnable parameter that controls the fusion ratio of original features and topological features, and the initial EEG embedding representation. In the formula, This represents the initial EEG embedding representation. For example, global average pooling is used for time-dimension statistical aggregation; the inter-channel correlation matrix has a dimension of 63×63; the Softmax function is used for normalization; the first-order graph convolution operator is used for graph convolution; and the initial value of the residual adjustment coefficient is set to 0.5.

[0033] Step 103: Perform orthogonalization regularization constraint processing on the initial EEG embedding representation to obtain decorrelated EEG features.

[0034] Specifically, orthogonalization regularization constraint processing can be a forced decorrelation process for feature dimensions; decorrelation EEG features can be highly discriminative EEG features with independent dimensions. For example, the orthogonalization regularization constraint uses the BarlowTwins loss function.

[0035] The step of performing orthogonalization and regularization constraint processing on the initial EEG embedding representation to obtain decorrelated EEG features includes: During the training phase, the initial EEG embedding representations within a batch are processed by autocorrelation matrix calculation to obtain the batch autocorrelation matrix. The first loss component is obtained by calculating the sum of squares of the differences between the diagonal elements of the batch autocorrelation matrix and 1. The second loss component is obtained by calculating the sum of squares of the off-diagonal elements of the batch autocorrelation matrix. The first loss component and the weighted second loss component are summed to obtain the Barlow Twins loss function value; By minimizing the Barlow Twins loss function value, the features of each dimension of the initial EEG embedding representation are forced to decorrelate, resulting in decorrelated EEG features.

[0036] Specifically, the batch autocorrelation matrix can be the cross-correlation matrix of EEG embedding features within the batch; the first loss component can be a loss term measuring feature variance; the second loss component can be a loss term measuring the correlation between feature dimensions; the Barlow Twins loss function can be a loss function used to force orthogonality of feature dimensions; and the decorrelated EEG features can be EEG features whose dimensions are independent of each other. For example, In the formula, Let represent the Barlow Twins orthogonalized regularization loss, and represent the batch autocorrelation matrix. Represents the batch autocorrelation matrix. This represents the off-diagonal loss weighting coefficient. This represents the feature dimension index; the batch size is set to 32; the weight coefficient of the second loss component is set to 0.001; the weighting ratio of the Barlow Twins loss function to the supervised loss function is 1:10.

[0037] Step 104: Perform prototype-guided semantic manifold correction processing on the decorrelated EEG features to obtain semantically aligned EEG features.

[0038] Specifically, prototype-guided semantic manifold correction can be a cross-modal alignment process that uses visual prototypes to correct the distribution of EEG features; semantically aligned EEG features can be EEG features anchored to the visual semantic manifold. For example, semantic manifold correction can use a visual prototype library generated by K-means clustering.

[0039] This also includes: Extract visual features from all images in the training set; The visual features were clustered using the K-means clustering algorithm to obtain K cluster centers; The K cluster centers are used as visual semantic prototypes and stored in a learnable memory to construct a visual prototype library.

[0040] Specifically, visual features can be high-dimensional semantic features of images extracted by a pre-trained visual model; the K-means clustering algorithm can be an unsupervised clustering algorithm; visual semantic prototypes can be feature centers representing different visual semantic categories; the learnable memory can be an updatable parameter matrix storing visual semantic prototypes; and the visual prototype library can be a feature set containing multiple visual semantic prototypes. For example, visual features are extracted using the CLIP model, with a feature dimension of 1024; K is set to 1000; the learnable memory has a dimension of 1000×1024; and the training set contains 16540 natural images.

[0041] Further, the step of performing prototype-guided semantic manifold correction processing on the decorrelated EEG features to obtain semantically aligned EEG features includes: The similarity between the decorrelation EEG features and each visual semantic prototype in the visual prototype library is calculated to obtain the attention weight distribution; Based on the attention weight distribution, the visual semantic prototypes in the visual prototype library are weighted and summed to obtain anchor features. By using learnable gating parameters, fusion weights are assigned to the decorrelation EEG features and anchor features respectively; Based on the fusion weights, the decorrelation EEG features and anchor features are weighted and summed to obtain semantically aligned EEG features.

[0042] Specifically, similarity can be a numerical value measuring the degree of similarity between two feature vectors; attention weight distribution can be the contribution weight of each visual semantic prototype to the current EEG feature; anchor feature can be the visual semantic combination feature most relevant to the current EEG feature; learnable gating parameter can be a learnable parameter controlling the fusion ratio of the original feature and anchor feature; fusion weight can be the weight coefficient of the original feature and anchor feature in the fusion process. For example, In the formula, This indicates semantically aligned EEG features. This indicates the removal of relevant EEG features. Represents visual prototype anchor point features. The prototype fusion weight coefficient is represented by the similarity calculated using cosine similarity; the attention weight distribution is obtained by normalization using the Softmax function; the initial value of the learnable gating parameter is set to 0.3; and the fusion weights are 0.7 and 0.3 respectively.

[0043] Step 105: Perform conditional flow matching generation processing on the semantically aligned EEG features to obtain the corresponding visual features.

[0044] Specifically, the conditional flow matching generation process can be a cross-modal feature generation process based on a deterministic optimal transmission path; visual features can be image embedding features corresponding to EEG semantics. For example, conditional flow matching uses a neural velocity estimator based on a conditional UNet architecture.

[0045] The training step of performing conditional flow matching generation processing on the semantically aligned EEG features includes: The semantically aligned EEG features are used as source distribution features, and the corresponding real image features are used as target distribution features. A linear interpolation path is constructed between the source distribution features and the target distribution features to obtain intermediate features that evolve continuously over time. Using the semantically aligned EEG features as conditional variables, a neural velocity estimator with a conditional U-Net structure is constructed. The mean square error between the velocity vector predicted by the neural velocity estimator and the actual semantic change velocity is calculated to obtain the conditional flow matching loss function value; By minimizing the value of the conditional flow matching loss function, the parameters of the neural velocity estimator are optimized, resulting in a trained neural velocity estimator.

[0046] Specifically, the source distribution features can be the semantically aligned EEG feature distribution; the target distribution features can be the visual feature distribution corresponding to the real image; the linear interpolation path can be a continuous linear path connecting the source and target distributions; the intermediate features can be features at any time point on the linear interpolation path; the neural velocity estimator can be a neural network predicting the velocity of feature changes; the velocity vector can be the change in a feature per unit time; the true semantic change velocity can be the constant velocity of change from the source to the target distribution; and the conditional flow matching loss function can be a loss function that measures the difference between the predicted velocity and the true velocity. For example, intermediate features... In the formula, Indicates time Corresponding intermediate features, Represents source distribution features (semantic-aligned EEG features). Represents the target distribution characteristics (real image features). The time step parameter, with a value range of [0,1]; conditional flow matching loss. In the formula, This represents the conditional flow matching loss. This represents the mathematical expectation operation. This represents a neural velocity estimator; the source distribution feature dimension is 1024; the target distribution feature is extracted using the CLIP model; the time range of the linear interpolation path is 0 to 1; the neural velocity estimator adopts an 8-layer conditional U-Net structure; the mean squared error loss is used to measure the difference in velocity vectors.

[0047] Further, the step of performing conditional flow matching to generate the corresponding visual features from the semantically aligned EEG features includes: During the reasoning phase, the semantically aligned EEG features are used as the initial state; Discretize the time interval into a preset number of time steps; At each time step, the trained neural velocity estimator is used to predict the rate of change of the current feature state; Based on the change rate and time step, perform Euler iterative update processing to obtain the feature state at the next moment; By iterating through all time steps, the ordinary differential equations are solved to obtain the visual features of the target.

[0048] Specifically, the inference phase can be the stage where the prediction task is executed after the model training is completed; the initial state can be the starting feature state of the flow matching generation process; time interval discretization can be the operation of dividing continuous time into multiple equal time steps; the preset number of steps can be the pre-set total number of iterative inference steps; the Euler iterative update processing can be a numerical calculation method based on solving ordinary differential equations using the first-order Euler method; solving ordinary differential equations can be the process of solving differential equations describing the continuous evolution of features; and the target visual features can be the final feature vector mapped to the visual semantic space. For example, target visual features... In the formula, This represents the feature state of the (k+1)th iteration; This represents the characteristic state of the k-th iteration. Indicates the discrete time step. This represents the time parameter corresponding to the k-th iteration; the preset number of steps is set to 20 steps; the time step size is 0.05; the Euler iteration update formula is that the feature at the next moment equals the current feature plus the rate of change multiplied by the time step size; the target visual feature dimension is 1024.

[0049] Step 106: Perform image decoding processing on the visual features to generate and output a reconstructed image corresponding to the original EEG signal.

[0050] Image decoding can be the process of converting visual features into pixel-level images; image reconstruction can be a visual stimulus image restored based on electroencephalogram (EEG) signals. For example, image decoding can use a frozen SDXL backbone network in conjunction with an IP adapter.

[0051] The step of performing image decoding processing on the visual features to generate and output a reconstructed image corresponding to the original EEG signal includes: A pre-trained image decoding model is constructed by using a frozen diffusion model backbone network in conjunction with an image adapter; The pre-trained image decoding model is used to perform pixel-level reconstruction of the visual features; Generate and output a reconstructed image corresponding to the original EEG signal.

[0052] Specifically, the frozen diffusion model backbone network can be a pre-trained diffusion model with fixed parameters that does not participate in training; the image adapter can be a lightweight module used to inject external features into the diffusion model; the pre-trained image decoding model can be a pre-trained model that converts visual features into pixel images; pixel-level reconstruction processing can be the process of mapping high-dimensional visual features into pixel-space images; and the reconstructed image can be a visual stimulus image obtained based on EEG signals. For example, the diffusion model backbone network uses the SDXL model; the image adapter uses the IP-Adapter; all parameters of the pre-trained image decoding model are frozen; and the reconstructed image resolution is 512×512 pixels.

[0053] Based on the same inventive concept, this embodiment also provides an EEG image reconstruction system based on flow matching and manifold alignment, comprising: The signal preprocessing module is used to preprocess the raw EEG signals to obtain standardized EEG data. The dynamic graph topology modeling module is used to perform dynamic graph topology feature modeling processing on the standardized EEG data to obtain an initial EEG embedding representation. The orthogonalization regularization constraint module is used to perform orthogonalization regularization constraint processing on the initial EEG embedding representation to obtain decorrelated EEG features; A prototype-guided manifold correction module is used to perform prototype-guided semantic manifold correction processing on the decorrelation EEG features to obtain semantically aligned EEG features. The conditional flow matching generation module is used to perform conditional flow matching generation processing on the semantically aligned EEG features to obtain the corresponding visual features. The image decoding module is used to perform image decoding processing on the visual features, generate and output a reconstructed image corresponding to the original EEG signal.

[0054] In some embodiments, it also includes: During the training phase, joint optimization processing is performed simultaneously, including dynamic graph topology feature modeling, orthogonalization regularization constraints, semantic manifold correction, and conditional flow matching generation. By using a joint loss function, the parameters of the dynamic graph topology modeling module, the orthogonalization regularization constraint module, the prototype-guided manifold correction module, and the conditional flow matching generation module are optimized synchronously.

[0055] Since this device corresponds to the EEG image reconstruction method based on flow matching and manifold alignment in this embodiment of the invention, and the principle of this device in solving the problem is similar to that of this method, the implementation of this device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0056] Specifically, joint optimization can be a training process that simultaneously optimizes the parameters of multiple modules; the joint loss function can be a total loss function composed of a weighted combination of multiple sub-loss functions. For example, the joint loss function consists of Barlow Twins loss, conditional flow matching loss, and supervised classification loss; the weight of Barlow Twins loss is 0.1; the weight of conditional flow matching loss is 1.0; and the weight of supervised classification loss is 0.5.

[0057] This implementation method utilizes the NeuroFlow framework to perform visual EEG decoding and image reconstruction. The system hardware environment consists of an NVIDIA RTX 4090 GPU, an Intel CPU, and 48GB of RAM. The software environment runs on an Ubuntu 20.04 operating system and the PyTorch deep learning framework. The data source is the publicly available THING-EEG visual EEG pairing dataset, which contains 16,540 images with a resolution of 540×540 and 63 EEG signal channels. Data was collected from 10 subjects, with each subject participating in 10 data collection operations.

[0058] During data acquisition and preprocessing, EEG signal data from multiple subjects were obtained from the publicly available THING-EEG dataset. The international 10-20 standard lead system was used to collect EEG signals generated when subjects viewed visual stimuli. Raw data was stored per subject, with each dataset containing preprocessed EEG signals, timestamps, and channel names. When loading data, subjects participating in training or testing were selected according to the experimental settings. Exclusion of specific subjects was also supported to construct cross-subject verification scenarios. Subsequently, the EEG signal data from the corresponding files was read and converted to floating-point tensor format. Samples were organized by category: 10 samples per category during the training phase and 1 sample per category during the testing phase, with each sample assigned a corresponding category label. Temporally, EEG signals were processed by time window truncation according to a preset time window. Signal segments within a specified time range were retained by filtering through timestamp intervals, extracting EEG feature segments related to visual stimulus responses, and removing irrelevant interfering information. EEG signals were processed using bandpass and notch filtering, with bandpass filtering from 0.1Hz to 70Hz and a 50Hz notch filter applied to remove power frequency interference. EEG data from different categories and subjects were concatenated and rearranged into a standardized tensor form. Labels were simultaneously expanded and aligned to suit subsequent model training requirements. During training, labels were repeatedly expanded to match the number of augmented samples. Category labels were remapped as needed to ensure category index continuity. After these processes, standardized EEG data with uniform structure, time alignment, and consistent labels were obtained, providing standardized input for subsequent topological feature extraction and cross-modal alignment.

[0059] In the dynamic graph topology feature modeling process, the preprocessed standardized EEG data is represented as a multi-channel time-series signal composed of the number of channels and the time length. Statistical convergence is performed on the time dimension to calculate the global response features of each channel within the current time window, obtaining channel-level feature vectors. Based on the channel-level feature vectors, an inter-channel correlation matrix is ​​constructed through inner product operations. The correlation matrix is ​​normalized to obtain a sample-specific dynamic adjacency matrix, which is used to characterize the instantaneous functional connectivity between brain regions. Graph convolution is performed between the dynamic adjacency matrix and the standardized EEG data to complete topological information propagation and obtain an EEG feature representation with fused topological information. A residual connection mechanism is introduced during feature fusion, and a learnable residual adjustment coefficient is set. Based on the learnable residual adjustment coefficient, the standardized EEG data and the EEG features with fused topological information are weighted and summed to avoid the oversmoothing problem caused by graph convolution. During feature learning, orthogonalization regularization constraints are applied, introducing Barlow Twins-based orthogonal constraints. During the training phase, the initial EEG embedding representation within a batch is processed by calculating the autocorrelation matrix to obtain the batch autocorrelation matrix. The sum of squares of the differences between the diagonal elements and 1 in the batch autocorrelation matrix is ​​calculated to obtain the first loss component. The sum of squares of the off-diagonal elements in the batch autocorrelation matrix is ​​calculated to obtain the second loss component. The first loss component and the weighted second loss component are summed to obtain the Barlow Twins loss function value. By minimizing the Barlow Twins loss function value, the correlation between the dimensions of the output features is suppressed, ensuring that the features maintain a decoupled distribution in the high-dimensional space, avoiding dimensionality collapse and improving feature expressive power, ultimately yielding the initial EEG embedding representation. Figure 2 The diagrams and flowcharts illustrate the principles of dynamic graph topological feature modeling, orthogonalization regularization constraint processing, and prototype-guided semantic manifold correction processing in embodiments of the present invention. Figure 2 (a) is a schematic diagram of the overall architecture of the processing process, showing the module composition and feature transmission path of the EEG encoder. After the input of multi-channel EEG time-series data, the topological feature extraction, orthogonalization constraint and semantic manifold correction are completed in sequence, and the corresponding EEG feature vector is output. Figure 2 (b) is a schematic diagram of the collaborative process, showing the complete processing link from the original EEG signal input to the feature transformation layer by layer and finally the generation of semantically aligned features. Figure 3 This is a schematic diagram and flowchart illustrating the principle of dynamic graph topology feature modeling processing according to an embodiment of the present invention. Figure 3 (a) is a structural principle diagram for dynamic graph topology modeling, showing the correspondence between EEG electrode channels, inter-channel correlation matrices and graph convolution operations, and presenting the extraction structure of spatial topological information. Figure 3(b) A flowchart of the processing for modeling dynamic graph topology features, showing the complete processing steps of obtaining fused topology features after the original EEG signal undergoes time consensus calculation, dynamic adjacency matrix construction, graph convolution operation and residual fusion.

[0060] In the prototype-guided semantic manifold correction process, a visual prototype library is constructed using image features from the training phase via K-means clustering. Visual features from all images in the training set are extracted, and K-means clustering is applied to these features to obtain K cluster centers. These K cluster centers are used as visual semantic prototypes and stored in a learnable memory to construct the visual prototype library. Each visual semantic prototype represents a stable visual semantic pattern. During feature alignment, the similarity between the decorrelated EEG features and each visual semantic prototype in the visual prototype library is calculated. An attention weight distribution is obtained through normalization. Based on this attention weight distribution, combinations of visual semantic prototypes related to the current EEG feature are retrieved from the visual prototype library. Anchor features are obtained by weighted summation of the visual semantic prototypes in the library according to the attention weight distribution. By using learnable gating parameters, fusion weights are assigned to the decorrelation EEG features and anchor features respectively. Based on the fusion weights, the decorrelation EEG features and anchor features are weighted and summed to complete the prototype-guided semantic manifold correction processing of the EEG features, so that the EEG features converge to the corresponding visual semantic center, improving the clustering compactness and discriminativeness of the features in the latent space. After this step, semantically aligned EEG features are obtained, providing stable input for subsequent cross-modal generation. Figure 4 The diagram and flowchart illustrate the principle of prototype-guided semantic manifold correction processing in this embodiment of the invention. Figure 4 (a) is a schematic diagram of the semantic manifold correction architecture, showing the construction process of visual features being generated into visual semantic prototypes by K-means clustering algorithm and stored in the learnable memory bank during the training phase, as well as the matching process of EEG features and prototypes to calculate similarity and generate attention weights during the inference phase. Figure 4 (b) is a flowchart of semantic manifold correction, showing the complete processing steps of outputting semantically aligned EEG features after adaptive gating fusion of initial EEG features and anchor features.

[0061] In the conditional flow matching generation process, the semantically aligned EEG features are used as source distribution features, and the corresponding real image features are used as target distribution features. A linear interpolation path is constructed between the source and target distribution features, forming a continuous-time evolution process. This process corresponds to a deterministic vector field, which represents the direction of change from EEG semantics to visual semantics. During the training phase, a neural velocity estimator with a conditional U-Net structure is constructed. The semantically aligned EEG features are used as conditional variables to model intermediate state features, predict the corresponding semantic change velocity, and calculate the mean square error between the velocity vector predicted by the neural velocity estimator and the real semantic change velocity to obtain the conditional flow matching loss function value. The optimal transmission path is learned by minimizing the conditional flow matching loss function value. This process abandons the random sampling method relied upon by traditional diffusion models, improves the stability of the generated path, and reduces computational complexity. After the model training is completed, the inference phase begins. The semantically aligned EEG features are used as the initial state, and the time interval is discretized into a preset number of time steps. At each time step, the trained neural velocity estimator is used to predict the change direction of the current state, and Euler iteration is performed to update the data. The feature representation of the next time step is obtained by multiplying the current feature with the step size and the predicted velocity. After a finite number of iterations, the ordinary differential equation is solved, the EEG features are mapped to the target location in the visual semantic space and the target visual features are obtained. The target visual features are then converted into the final image result through a pre-trained image decoding model. This generation process is completed by relying on a deterministic path and has fewer iterations than the traditional diffusion model. It reduces inference time while ensuring the semantic consistency and stability of the generated results. Figure 5 This is a schematic diagram and flowchart of the conditional flow matching generation module according to an embodiment of the present invention. Figure 5 (a) is a schematic diagram of the architecture generated for conditional flow matching, showing the process of constructing a linear interpolation path between the brain power distribution features and the visual target distribution features during the training phase, learning the feature change velocity field through the conditional U-shaped network structure, and the process of obtaining visual features from the initial brain power features through iterative solution during the inference phase. Figure 5 (b) is the processing flowchart generated for conditional flow matching, showing the complete inference steps of inputting EEG feature vectors into the vector field estimation unit, updating the feature state through multiple Euler iterations, and outputting the target image feature vectors after traversing all time steps.

[0062] The image decoding module performs image decoding processing on the visual features. It uses a frozen diffusion model backbone network in conjunction with an image adapter to construct a pre-trained image decoding model. The pre-trained image decoding model is used to perform pixel-level restoration processing on the visual features, generating and outputting a reconstructed image corresponding to the original EEG signal. Figure 6 This is a schematic diagram of the EEG image reconstruction effect output by the image decoding module in an embodiment of the present invention. Figure 6 This demonstrates the matching effect between visual stimuli and the corresponding generated reconstructed images.

[0063] During the training phase, dynamic graph topological feature modeling, orthogonalization regularization constraints, prototype-guided semantic manifold correction, and conditional flow matching generation are jointly optimized. The parameters of the EEG encoder, prototype-guided manifold correction module, and conditional flow matching generation module are simultaneously optimized through a joint loss function.

[0064] During the technical effectiveness verification process, this implementation method was validated on the THING-EEG dataset, which includes 10 subjects and 1854 visual concepts. In the zero-shot retrieval task, this implementation method achieved a Top-1 accuracy of 32.4%, demonstrating its generalization ability in cross-modal semantic alignment scenarios. In the image reconstruction task, the structural similarity index of the generated results reached 0.414, and the semantic consistency index reached 0.841. In terms of inference efficiency, image generation can be completed in 20 iterations, with an overall time of 7.7 seconds. Compared with the traditional generation method based on a 50-step diffusion process, this method reduces computational overhead while ensuring generation quality, demonstrating real-time performance and practicality.

[0065] In expanded application scenarios, this implementation can be used to analyze the consistency between subjects' EEG signals and visual stimulus responses in clinical settings, assisting in the assessment of visual function-related diseases. In practical applications, it can be deployed in mobile brain-computer interface devices with a lightweight model structure to complete real-time image retrieval and generation operations based on EEG signals, supporting human-computer interaction-related application needs.

[0066] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0067] The above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made based on the essence of the content of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for reconstructing electroencephalogram (EEG) images based on flow matching and manifold alignment, characterized in that, Includes the following steps: The raw EEG signals are preprocessed to obtain standardized EEG data; The standardized EEG data are subjected to dynamic graph topology feature modeling to obtain an initial EEG embedding representation; The initial EEG embedding representation is subjected to orthogonalization and regularization constraints to obtain decorrelated EEG features; The decorrelated EEG features are subjected to prototype-guided semantic manifold correction processing to obtain semantically aligned EEG features; The semantically aligned EEG features are subjected to conditional flow matching to generate corresponding visual features. The visual features are processed by image decoding to generate and output a reconstructed image corresponding to the original EEG signal; The step of performing orthogonalization and regularization constraint processing on the initial EEG embedding representation to obtain decorrelated EEG features includes: During the training phase, the initial EEG embedding representations within a batch are processed by autocorrelation matrix calculation to obtain the batch autocorrelation matrix. The first loss component is obtained by calculating the sum of squares of the differences between the diagonal elements of the batch autocorrelation matrix and 1. The second loss component is obtained by calculating the sum of squares of the off-diagonal elements of the batch autocorrelation matrix. The first loss component and the weighted second loss component are summed to obtain the Barlow Twins loss function value; By minimizing the Barlow Twins loss function value, the features of each dimension of the initial EEG embedding representation are forced to decorrelate, resulting in decorrelated EEG features.

2. The method according to claim 1, characterized in that, The step of preprocessing the raw EEG signals to obtain standardized EEG data includes: The raw EEG signal was processed by bandpass filtering and notch filtering to remove power frequency interference and out-of-band noise. The filtered EEG signal was subjected to time window truncation to extract signal segments related to visual stimulus response; The extracted EEG signals were format-converted and standardized to obtain standardized EEG data in a unified tensor format.

3. The method according to claim 1, characterized in that, The step of performing dynamic graph topology feature modeling on the standardized EEG data to obtain an initial EEG embedding representation includes: The standardized EEG data were subjected to time-dimensional statistical aggregation processing to obtain the global response features of each channel; The inter-channel correlation matrix is ​​obtained by performing an inner product operation on the global response features of each channel. The correlation matrix between the channels is normalized to obtain a sample-specific dynamic adjacency matrix. Graph convolution operations are performed on the dynamic adjacency matrix and standardized EEG data to obtain EEG features that fuse topological information. Set a learnable residual adjustment factor; Based on the residual adjustment coefficient, the standardized EEG data and the EEG features of the fused topological information are weighted and summed to obtain the initial EEG embedding representation.

4. The method according to claim 1, characterized in that, Also includes: Extract visual features from all images in the training set; The visual features were clustered using the K-means clustering algorithm to obtain K cluster centers; The K cluster centers are used as visual semantic prototypes and stored in a learnable memory to construct a visual prototype library.

5. The method according to claim 4, characterized in that, The step of performing prototype-guided semantic manifold correction processing on the decorrelated EEG features to obtain semantically aligned EEG features includes: The similarity between the decorrelation EEG features and each visual semantic prototype in the visual prototype library is calculated to obtain the attention weight distribution; Based on the attention weight distribution, the visual semantic prototypes in the visual prototype library are weighted and summed to obtain anchor features. By using learnable gating parameters, fusion weights are assigned to the decorrelation EEG features and anchor features respectively; Based on the fusion weights, the decorrelation EEG features and anchor features are weighted and summed to obtain semantically aligned EEG features.

6. The method according to claim 1, characterized in that, The training step of generating conditional flow matching for the semantically aligned EEG features includes: The semantically aligned EEG features are used as source distribution features, and the corresponding real image features are used as target distribution features. A linear interpolation path is constructed between the source distribution features and the target distribution features to obtain intermediate features that evolve continuously over time. Using the semantically aligned EEG features as conditional variables, a neural velocity estimator with a conditional U-Net structure is constructed. The mean square error between the velocity vector predicted by the neural velocity estimator and the actual semantic change velocity is calculated to obtain the conditional flow matching loss function value; By minimizing the value of the conditional flow matching loss function, the parameters of the neural velocity estimator are optimized, resulting in a trained neural velocity estimator.

7. The method according to claim 6, characterized in that, The step of performing conditional flow matching to generate corresponding visual features on the semantically aligned EEG features includes: During the reasoning phase, the semantically aligned EEG features are used as the initial state; Discretize the time interval into a preset number of time steps; At each time step, the trained neural velocity estimator is used to predict the rate of change of the current feature state; Based on the change rate and time step, perform Euler iterative update processing to obtain the feature state at the next moment; By iterating through all time steps, the ordinary differential equations are solved to obtain the visual features of the target.

8. The method according to claim 1, characterized in that, The step of performing image decoding processing on the visual features to generate and output a reconstructed image corresponding to the original EEG signal includes: A pre-trained image decoding model is constructed by using a frozen diffusion model backbone network in conjunction with an image adapter; The pre-trained image decoding model is used to perform pixel-level reconstruction of the visual features; Generate and output a reconstructed image corresponding to the original EEG signal.

9. A brainwave image reconstruction system based on flow matching and manifold alignment, characterized in that, include: The signal preprocessing module is used to preprocess the raw EEG signals to obtain standardized EEG data. The dynamic graph topology modeling module is used to perform dynamic graph topology feature modeling processing on the standardized EEG data to obtain an initial EEG embedding representation. The orthogonalization regularization constraint module is used to perform orthogonalization regularization constraint processing on the initial EEG embedding representation to obtain decorrelated EEG features; A prototype-guided manifold correction module is used to perform prototype-guided semantic manifold correction processing on the decorrelation EEG features to obtain semantically aligned EEG features. The conditional flow matching generation module is used to perform conditional flow matching generation processing on the semantically aligned EEG features to obtain the corresponding visual features. An image decoding module is used to perform image decoding processing on the visual features, generate and output a reconstructed image corresponding to the original EEG signal; Specifically, in the orthogonalization and regularization constraint module, the orthogonalization and regularization constraint processing of the initial EEG embedding representation to obtain decorrelated EEG features includes: During the training phase, the initial EEG embedding representations within a batch are processed by autocorrelation matrix calculation to obtain the batch autocorrelation matrix. The first loss component is obtained by calculating the sum of squares of the differences between the diagonal elements of the batch autocorrelation matrix and 1. The second loss component is obtained by calculating the sum of squares of the off-diagonal elements of the batch autocorrelation matrix. The first loss component and the weighted second loss component are summed to obtain the Barlow Twins loss function value; By minimizing the Barlow Twins loss function value, the features of each dimension of the initial EEG embedding representation are forced to decorrelate, resulting in decorrelated EEG features.

Citation Information

Patent Citations

  • Electroencephalogram image reconstruction method based on frequency steering and bidirectional diffusion

    CN122066822A