EEG function connection prediction method based on fMRI depth cross-modal representation learning

By constructing a deep cross-modal representation learning model, the problem of EEG functional connectivity map reconstruction was solved, achieving high-fidelity mapping from fMRI data to EEG functional connectivity and improving the application capabilities of cross-modal neuroimaging.

CN121622004APending Publication Date: 2026-03-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately reconstruct EEG functional connectivity maps at the frequency domain level, and methods for inferring EEG features based on fMRI are still immature, resulting in low efficiency in the synchronous acquisition and information fusion of cross-modal neuroimaging data.

Method used

A deep cross-modal representation learning model based on fMRI is constructed. Through a time projection module, a cross-attention fusion encoder, and a connection synthesis module, a high-fidelity mapping from fMRI data to EEG functional connectivity is achieved, including data preprocessing, feature extraction, and model training.

Benefits of technology

It significantly improves the numerical accuracy and network topology fidelity of cross-modal reconstruction, realizes high spatiotemporal resolution prediction of brain functional connectivity, and expands the application boundaries of cross-modal neuroimaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121622004A_ABST
    Figure CN121622004A_ABST
Patent Text Reader

Abstract

The invention provides an EEG function connection prediction method based on fMRI deep cross-modal representation learning, and relates to the crossing field of neuroiconography and artificial intelligence. According to the method, a sample pair is constructed through time alignment of EEG and a blood oxygen level dependent signal, an end-to-end deep neural network architecture composed of a time projection unit, a cross-attention fusion encoder and a connection synthesis head is designed, and direct mapping from a BOLD signal to EEG function connection is achieved. The model adopts a composite loss function, and gives consideration to the consistency of the prediction connection matrix with the real EEG function connection in numerical precision and topological mode, thereby achieving the high-fidelity reconstruction of the EEG function connection map at the frequency domain level. According to the invention, the topological structure stability of the brain function network can be effectively maintained. Under the conditions of incomplete EEG data, serious noise interference or complete loss, the frequency domain function connection characteristics can still be stably recovered, and a reliable analysis substitution path is provided for brain function network research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of neuroimaging and artificial intelligence, specifically to a method for predicting EEG functional connectivity based on deep cross-modal representation learning using functional magnetic resonance imaging (fMRI). Background Technology

[0002] Functional magnetic resonance imaging (fMRI) and electroencephalography (EEG), as two widely used non-invasive detection methods in the field of neuroimaging, each possess unique advantages and inherent limitations. EEG records electrical potential fluctuations on the scalp, accurately depicting the brain's dynamic response processes with millisecond-level temporal resolution; however, its spatial resolution is limited due to the conductivity characteristics of the skull and tissues, making it difficult to accurately locate neural activity in deep or localized fine brain regions. In contrast, fMRI indirectly reflects neural activity based on blood oxygen level-dependent (BOLD) signals, achieving whole-brain functional area localization with millimeter-level spatial resolution, but its temporal resolution is typically limited to the second level, making it difficult to track rapid dynamic processes of neural information processing. Currently, simultaneous acquisition of EEG and fMRI is widely used to combine their spatiotemporal advantages to obtain high-spatiotemporal-precision brain functional imaging. However, in practical applications, simultaneous acquisition of multimodal data faces challenges such as high equipment costs, long experimental cycles, and limited subject cooperation, resulting in many existing neuroimaging databases containing only single-modality data (fMRI or EEG only), wasting resources and limiting the effective integration and utilization of multimodal information.

[0003] To overcome the aforementioned bottlenecks, researchers have recently focused on cross-modal neural information prediction and reconstruction, aiming to infer neural activity characteristics of one modality based on data from another, thereby achieving information complementarity and modality fusion. Some studies have attempted to predict fMRI activity from EEG signals, preliminarily verifying a certain mapping relationship between electrophysiological activity and hemodynamic response. However, inferring EEG features, especially frequency-specific functional connectivity patterns, from fMRI is still in its early stages, and the research methods are not yet mature. The main technical bottlenecks include: first, existing models struggle to accurately reconstruct EEG functional connectivity maps in the frequency domain; second, the predicted brain networks show significant differences in topology compared to actual neural activity. Functional connectivity, as a core indicator in EEG research, can reveal functional synergies and information transmission mechanisms between different brain regions. Therefore, reconstructing frequency-specific EEG functional connectivity from fMRI data could not only provide effective alternative information when EEG data is unavailable but also expand the analytical capabilities of fMRI at the frequency-specific brain network level, enhancing the overall understanding of brain spatiotemporal patterns. Therefore, this paper proposes a deep cross-modal representation learning method for predicting EEG band-specific functional connectivity based on fMRI data. By constructing an end-to-end deep learning model, an effective mapping from fMRI features to EEG functional connectivity is achieved, reliably recovering functional connectivity patterns even when EEG data is missing or of poor quality. This method has significant theoretical value and broad application prospects for overcoming the limitations of single-modal analysis, improving the spatiotemporal resolution of brain function research, and expanding the application boundaries of cross-modal neuroimaging. Summary of the Invention

[0004] To address the core bottleneck of simultaneous acquisition and effective fusion of multimodal data under the current demand for high spatiotemporal resolution brain functional imaging, this invention proposes an EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning.

[0005] This invention proposes an EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning, comprising the following steps:

[0006] Step S1: Preprocess the fMRI data;

[0007] Remove the data from the first 10 time points and perform time correction and head movement correction to ensure timing consistency;

[0008] Subsequently, the images were spatially normalized and registered to the MNI standard space, and spatial smoothing was performed using a 4mm full-width half-height Gaussian kernel to improve the signal-to-noise ratio. Finally, the time series was bandpass filtered by 0.01–0.1Hz to remove low-frequency drift and high-frequency noise, and fMRI signals of 200 brain regions of interest were extracted based on the Schaefer standard brain region template; fMRI refers to functional magnetic resonance imaging.

[0009] Step S2: Preprocess the header table EEG data;

[0010] EEG data were acquired synchronously using a 64-channel MR-compatible amplifier at a sampling rate of 2000Hz. First, reference electrodes and channels with quality below the threshold were removed, along with gradient and cardiac artifacts. The cleaned EEG signal was downsampled to 250Hz and bandpass filtered from 0.5 to 40Hz to remove low-frequency drift and high-frequency noise. Subsequently, the EEG data was divided into non-overlapping signal segments according to a 2-second time window, providing a time window for subsequent functional connectivity calculations. EEG represents an electroencephalogram.

[0011] Step S3: Divide the frequency band-specific EEG signal;

[0012] The preprocessed EEG signal was divided into frequency bands of 1-4 Hz delta, 4-8 Hz theta, 8-13 Hz alpha, 13-30 Hz beta and 30-45 Hz gamma according to its frequency distribution, thus obtaining frequency-specific EEG signals.

[0013] Step S4: Calculate the functional connectivity matrix;

[0014] For each data segment, the Pearson correlation values ​​between the EEG signals of all electrode pairs are calculated, forming a connectivity matrix. ;

[0015] Step S5: Perform EEG-BOLD time alignment;

[0016] Based on the EEG frequency-specific functional connectivity obtained in steps S1-S3, pair them with the corresponding BOLD signal segments. Match each functional connectivity segment with an fMRI segment delayed by 10 seconds to form sample pairs of EEG functional connectivity and BOLD signal segments, and construct a dataset; BOLD represents blood oxygen level dependence.

[0017] Step S6: Construct a deep cross-modal representation learning model;

[0018] A predictive model for EEG functional connectivity based on fMRI deep cross-modal representation learning was constructed. The predictive model consists of three parts: a temporal projection module, a cross-attention fusion encoder module, and a connectivity synthesis module. The temporal projection module is responsible for temporal compression and feature extraction of fMRI data. The cross-attention fusion encoder module optimizes the features through a multi-head self-attention mechanism and a local feature extraction module to further improve the model's expressive power. The connectivity synthesis module transforms the optimized feature mapping into the final predicted EEG functional connectivity strength.

[0019] Step S7: Model training and testing;

[0020] Regarding dataset partitioning, the EEG-fMRI dataset obtained in step S5 was used for training and testing via leave-one-out cross-validation. In each round of cross-validation, the data of one subject was selected as the test set, and the data of the remaining subjects were used as the training set. Supervised learning was employed during training, using the Adam optimizer to train the model. The optimization process used adaptive moment estimation to adjust the learning rate of each parameter and updated the neural network weights through backpropagation. A composite loss function was adopted for the loss function design. Composite loss function Optimize using the following formula:

[0021] ;

[0022] in, The mean squared error between the predicted and actual values. A negative Pearson correlation coefficient The adjustment coefficients are used to control the weights of each part in the loss function;

[0023] Step S8: Evaluation of prediction results;

[0024] The predicted EEG functional connectivity is quantitatively evaluated using three evaluation metrics:

[0025] (1) Root mean square error: The root mean square error between the predicted EEG functional connectivity and the actual calculated EEG functional connectivity is calculated to measure the difference between the predicted value and the actual value.

[0026] (2) Cosine similarity: To further evaluate the similarity between the predicted results and the actual results, cosine similarity is used to measure the directional similarity between the two sets of functional connectivity matrices.

[0027] (3) Pearson similarity, in order to more comprehensively evaluate the linear relationship between the prediction results and the true functional connectivity matrix.

[0028] Furthermore, the steps for constructing the EEG frequency-specific functional connection network in step S4 are as follows:

[0029] For the For each segment of data, calculate the Pearson correlation value between the EEG signals of all electrode pairs. This forms a 60×60 functional connection matrix. The calculation formula is as follows:

[0030] ;

[0031] in, For the first Channel in time EEG signal, The value is the average value of the channel within the segment, where T is the segment length; the absolute value of the functional connectivity matrix is ​​taken. Characterizes the connection strength.

[0032] Furthermore, the EEG-BOLD time alignment step in step S5 is as follows:

[0033] Step S51: Based on the results of steps S1-S2, obtain the preprocessed and segmented EEG signal segments into fixed time windows, as well as the corresponding BOLD signal segments;

[0034] Step S52: Set the hemodynamic delay time The center time point of each EEG signal segment Compared with delayed fMRI time points correspond;

[0035] Step S53: Extract the BOLD signal segment that matches the EEG time window within the corresponding delayed fMRI window to ensure that the two types of data are aligned in the time dimension;

[0036] Step S54: Combine the matched EEG functional connection matrix with the corresponding BOLD signal segment to form a pair of samples and store them in the dataset;

[0037] Step S55: Repeat the above operation until all EEG–BOLD sample pairs are constructed, and finally the complete training and testing dataset is obtained.

[0038] Furthermore, the construction steps of the EEG functional connectivity model based on fMRI deep cross-modal representation learning in step S6 are as follows:

[0039] Step S61: The time projection module compresses the fMRI input signal in the time dimension;

[0040] For input tensor ,in For batch size, To determine the number of brain regions of interest extracted in step S1, adaptive pooling is first performed on the time series of each ROI, compressing several time points into a single feature representation, and then... The convolutional layer expands the number of channels to obtain a feature representation suitable for subsequent convolution and attention processing; finally, the non-linear activation function ReLU is applied to the result to enhance the feature representation capability of the model.

[0041] Step S62: The cross-attention fusion encoder module is composed of multiple identical encoding blocks stacked together. Each encoding block includes a context-aware local-global fusion module and a feature aggregation module. The context-aware local-global fusion module consists of local convolutional branches and global self-attention branches.

[0042] 1) In the local convolution branch, a combination of pointwise convolution and depthwise convolution is first used to extract local features within the ROI, and a gating mechanism is used to enhance the effective feature representation. ,in, Pointwise convolution is a cross-channel fusion. convolution, For depthwise convolution is Depth convolution of local spatial patterns, It is a zero-parameter gating mechanism, in which and These represent that the feature tensor Z is equally divided into two parts along the channel dimension. This represents the element-wise multiplication operation, used to enhance effective feature representation and suppress redundant features;

[0043] 2) In the global self-attention branch, based on the query-key-value multi-head attention mechanism, the dependencies between all ROIs are modeled to obtain global interaction information. ;

[0044] ;

[0045] in, It is a query, key, and value matrix. The number of heads receiving multi-head attention. It is the feature dimension of each head. It is a projection layer;

[0046] 3) Finally, the results of the local convolutional branch and the global attention branch are fused according to the learnable weights. , It is a learnable parameter used to balance local and global branches;

[0047] 4) Next, the feature fusion module uses a pointwise fully connected layer with residual connections to re-fuse the features and obtain the fused features. ;

[0048] ;

[0049] in, For layer normalization, and It is a cross-channel integration Pointwise convolution;

[0050] After multiple layers are stacked, a fused representation with local and global multi-scale features is obtained;

[0051] Step S63: Connect the synthesis module to fuse the above representation Flattened into feature vectors The input is fed into a multilayer perceptron structure, undergoes linear transformation, nonlinear activation, and a sigmoid function, and the output dimension is equal to the number of EEG functional connection edges. The output result is in vector form, which is further mapped to the upper triangular part of the EEG functional connection matrix, and then the complete matrix structure is obtained through symmetry filling.

[0052] Furthermore, the training process of the deep learning model for fMRI–EEG functional connectivity based on deep cross-modal representation learning in step S7 is as follows:

[0053] Step S71: Divide the dataset into training and testing sets, set the number of iterations, and initialize the model hyperparameters;

[0054] Step S72: Organize the sample pairs obtained in step S2 into... , For the first One BOLDROI-time window tensor; To connect the upper triangular vectorization results for the corresponding EEG function, the number of channels is... ,but The numerical value is taken as the absolute correlation strength;

[0055] Step S73: Input the processed data into the neural network to obtain the functional connectivity of the fitted EEG signal, with the loss function being... Where MSE is the standard mean square error. It is a negative Pearson correlation coefficient;

[0056] Step S74: An early stopping strategy is introduced during the training process. Training is stopped when the loss function does not decrease within consecutive iterations of round a.

[0057] Furthermore, the process of evaluating the prediction results in step S8 is as follows:

[0058] Step S81: Calculate the root mean square error between the predicted EEG functional connectivity and the actual functional connectivity to measure the difference between the predicted and actual values.

[0059] ;

[0060] in, The first element of the real functional connection matrix represents the... One element, For the prediction function connection matrix, the first There are n elements, where Q is the total number of elements in the matrix;

[0061] Step S82: Use cosine similarity to measure the directional similarity between the two sets of functional connectivity matrices;

[0062] ;

[0063] Step S83: Calculate the Pearson correlation coefficient to evaluate the linear relationship between the prediction results and the true functional connectivity matrix;

[0064] ;

[0065] in, and These are the mean values ​​of the connection matrices for the true and predicted functions, respectively.

[0066] The advantages of this invention are:

[0067] The deep cross-modal representation learning framework proposed in this invention systematically solves the key bottlenecks in current fMRI-based EEG functional connectivity inference research by introducing a highly structure-adaptive temporal projection module, a cross-attention fusion encoder, and a connection synthesis head. This significantly improves the numerical accuracy and network topology fidelity of cross-modal reconstruction, demonstrating several outstanding advantages. First, the temporal projection module design breaks through the limitations of traditional fMRI temporal modeling, enabling temporal compression and feature expansion of fMRI signals. Second, the cross-attention fusion encoder, as the core architecture of the model, captures local patterns of ROIs through local convolutional branches, models long-range dependencies through global multi-head attention branches, and combines a channel recalibration mechanism to achieve the fusion of local and global features, significantly enhancing the decoding capability of implicit electrophysiological information in fMR spatial patterns. Finally, the connection synthesis head can map multi-scale fused features into an EEG functional connectivity matrix, ensuring that the prediction results are close to the real EEG in numerical accuracy and consistent in network topology, ultimately achieving high-fidelity cross-modal functional connectivity reconstruction. This invention not only represents a significant breakthrough in methodology, but also demonstrates broad scope and profound scientific value in brain science research and neural engineering applications. Attached Figure Description

[0068] Figure 1 This is a schematic diagram of an EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning provided by the present invention.

[0069] Figure 2 A schematic diagram of the time projection module in the neural network structure is provided for this invention.

[0070] Figure 3 This invention provides a schematic diagram of a context-aware local-global fusion module in a neural network structure.

[0071] Figure 4 A schematic diagram of the feature aggregation module in the neural network structure is provided for this invention.

[0072] Figure 5 A schematic diagram of the connection synthesis module in the neural network structure is provided for this invention.

[0073] Figure 6 This is a performance comparison chart between our method and the baseline method at different frequency bands.

[0074] Figure 7 This is a visualization comparing the top 10% of the strongest connections in the actual functional connections and predicted connections of this method in different frequency bands. Detailed Implementation

[0075] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0076] In a first aspect, according to embodiments of the present invention, a method for predicting EEG functional connectivity based on fMRI deep cross-modal representation learning is proposed. (See also...) Figure 1 This includes the following steps:

[0077] Step S1: Preprocessing fMRI data

[0078] DPABI (v5.1) software was used to preprocess the resting-state fMRI data. First, the first 10 time points were removed, and time correction and head motion correction were performed to ensure temporal consistency. Then, the images were spatially normalized and registered to the MNI standard space (resolution 3×3×3 mm³), and spatial smoothing was performed using a 4 mm full-width half-height (FWHM) Gaussian kernel to improve the signal-to-noise ratio. Finally, the time series were bandpass filtered at 0.01–0.1 Hz to remove low-frequency drift and high-frequency noise, and fMRI signals from 200 brain regions of interest were extracted based on the Schaefer standard brain region template.

[0079] Step S2: Preprocess the header table EEG data

[0080] EEG data were acquired synchronously using a 64-channel MR-compatible amplifier (Neuroscan) at a sampling rate of 2000 Hz. First, reference electrodes (FCz, AFz) and channels with poor quality (such as EOG and ECG) were removed, and gradient and cardiac artifacts were removed using Curry 7 software. The cleaned EEG signals were downsampled to 250 Hz and bandpass filtered from 0.5 to 40 Hz to remove low-frequency drift and high-frequency noise. Subsequently, the EEG data were divided into non-overlapping signal segments according to a 2-second time window (corresponding to one time repetition (TR) on fMRI), providing a time window for subsequent functional connectivity calculations.

[0081] Step S3: Divide the frequency band-specific EEG signal

[0082] Bandpass filtering was used to divide the preprocessed EEG signals into the delta (1-4 Hz), theta (4-8 Hz), alpha (8-13 Hz), beta (13-30 Hz) and gamma (30-45 Hz) frequency bands according to their frequency distribution, thus obtaining frequency-specific EEG signals.

[0083] Step S4: Calculate the functional connectivity matrix

[0084] For each EEG data segment of the band of interest, the Pearson correlation coefficient was used to evaluate the EEG signals at two electrodes located at different positions. and The synchronization strength. Specifically, for the first... For each segment of data, the Pearson correlation values ​​between the EEG signals of all electrode pairs are calculated to form a connectivity matrix. .

[0085] Step S5: Perform EEG-BOLD time alignment

[0086] Based on steps S1-S3, EEG frequency-specific functional connectivity segments are paired with their corresponding BOLD signal segments. Considering hemodynamic delay, each functional connectivity segment is matched with an fMRI segment delayed by 10 seconds to form sample pairs of (EEG functional connectivity, BOLD signal segments), thus constructing a dataset.

[0087] Step S6: Construct a deep cross-modal representation learning model

[0088] A novel EEG functional connectivity prediction model based on fMRI deep cross-modal representation learning is constructed. This model bridges the spatiotemporal resolution gap between fMRI and EEG signals by utilizing multi-stage feature alignment and cross-modal representation learning, thereby achieving efficient mapping from fMRI to EEG functional connectivity. The core of the model consists of three parts: a temporal projection module, a cross-attention fusion encoder module, and a connectivity synthesis module. The temporal projection module is primarily responsible for temporal compression and feature extraction of fMRI data; the cross-attention fusion encoder module optimizes features through a multi-head self-attention mechanism and a local feature extraction module, further improving the model's expressive power; and the connectivity synthesis module transforms the optimized feature mapping into the final predicted EEG functional connectivity strength.

[0089] Step S7: Model Training and Testing

[0090] For dataset partitioning, the EEG-fMRI dataset obtained in step S5 was used for training and testing via leave-one-out cross-validation. In each round of cross-validation, the data of one subject was selected as the test set, and the data of the remaining subjects were used as the training set. Supervised learning was employed during training, using the Adam optimizer. The optimization process used adaptive moment estimation to adjust the learning rate of each parameter and updated the neural network weights through backpropagation. To ensure model stability and prevent overfitting, an early stopping strategy was introduced: training was stopped and the current optimal model parameters were saved when the loss function did not decrease significantly within several consecutive iterations. A composite loss function was used, combining mean squared error (MSE) and Pearson correlation coefficient loss. Specifically, the composite loss function was optimized using the following formula:

[0091] ;

[0092] in, The mean squared error between the predicted and actual values. A negative Pearson correlation coefficient This is an adjustment coefficient used to control the weights of each part of the loss function.

[0093] Step S8: Evaluation of Prediction Results

[0094] The predicted EEG functional connectivity was quantitatively evaluated using three metrics: (1) Root mean square error (RMSE), which calculated the root mean square error between the predicted and actual EEG functional connectivity to measure the difference between the predicted and actual values; (2) Cosine similarity, which was used to measure the directional similarity between the two sets of functional connectivity matrices to further evaluate the similarity between the predicted and actual results; and (3) Pearson similarity, which was used to more comprehensively evaluate the linear relationship between the predicted and actual functional connectivity matrices. These metrics help to comprehensively evaluate the accuracy and stability of the cross-modal prediction method, thereby verifying the performance and applicability of the model.

[0095] Furthermore, the steps for constructing the EEG frequency-specific functional connection network in step S4 are as follows:

[0096] Pearson correlation coefficient was used to assess EEG signals from electrodes at two different locations. and The synchronization strength. Specifically, for the first... For each segment of data, calculate the Pearson correlation value between the EEG signals of all electrode pairs. This forms a 60×60 functional connection matrix. The calculation formula is as follows:

[0097] ;

[0098] in, For the first Channel in time EEG signal, This is the average value of the channel within the segment, with a segment length of T = 500 (corresponding to 2 seconds × 250 Hz). Take the absolute value of the functional connectivity matrix. Characterizes the connection strength.

[0099] Furthermore, the EEG-BOLD time alignment step in step S5 is as follows:

[0100] Step S51: Based on the results of steps S1-S2, obtain the preprocessed and segmented EEG signal segments into fixed time windows, as well as the corresponding BOLD signal segments.

[0101] Step S52: Set the hemodynamic delay time (e.g., 10 seconds), the center time point of each EEG signal segment , and delayed fMRI time points correspond.

[0102] Step S53: Extract the BOLD signal fragment that matches the EEG time window within the corresponding delayed fMRI window to ensure that the two types of data are aligned in the time dimension.

[0103] Step S54: Combine the matched EEG functional connection matrix with the corresponding BOLD signal segment to form a pair of samples and store them in the dataset.

[0104] Step S55: Repeat the above operation until all EEG–BOLD sample pairs are constructed, and finally the complete training and testing dataset is obtained.

[0105] Furthermore, the EEG functional connectivity model based on fMRI deep cross-modal representation learning in step S6 consists of a temporal projection module and a cross-attention fusion encoder module (stacked). The module is composed of individual blocks connected in series with the connection and synthesis module.

[0106] Further, please refer to Figure 2 The time projection module compresses the fMRI input signal along the time dimension. Specifically, it compresses the input tensor... (in For batch size, This refers to the number of regions of interest (ROIs) extracted in step S1. (The hemodynamic delay time set in step S52) is first used to adaptively pool the time series of each ROI, compressing several time points into a single feature representation, and then... The convolutional layers expand the number of channels to obtain feature representations suitable for subsequent convolution and attention processing. Finally, a non-linear activation function, ReLU, is applied to the result to enhance the model's feature representation capabilities.

[0107] Further, please refer to Figure 3 and Figure 4 The cross-attention fusion encoder module consists of multiple identical coding blocks stacked together, each coding block including a context-aware local-global fusion module. Figure 3 ) and feature aggregation module ( Figure 4 The context-aware local-global fusion module consists of local convolutional branches and global self-attention branches.

[0108] 1) In the local convolution branch, a combination of pointwise convolution and depthwise convolution is first used to extract local features within the ROI, and a gating mechanism is used to enhance the effective feature representation. Pointwise convolution ( (This is a cross-channel fusion) Convolution, depthwise convolution ( ) is a Depth convolution of local spatial patterns, It is a zero-parameter gating mechanism, in which and These represent that the feature tensor Z is equally divided into two parts along the channel dimension. This represents the element-wise multiplication operation, used to enhance effective feature representation and suppress redundant features.

[0109] 2) In the global self-attention branch, based on the query-key-value multi-head attention mechanism, the dependencies between all ROIs are modeled to obtain global interaction information:

[0110] ;

[0111] in, It is a query-key-value matrix. The number of heads receiving multi-head attention. For characteristic number, It is the feature dimension of each head. It is the projection layer.

[0112] 3) Finally, the results of the local convolutional branch and the global attention branch are fused according to the learnable weights. , It is a learnable parameter used to balance local and global branches, and its value is between 0 and 1.

[0113] 4) Next, through the feature aggregation module, please refer to [link / reference]. Figure 4 The feature fusion module uses a pointwise fully connected layer with residual connections to re-fuse features, ensuring training stability.

[0114] ;

[0115] in, For layer normalization, and It is a cross-channel integration Pointwise convolution, It is a zero-parameter gating mechanism, in which and These represent that the feature tensor Z is equally divided into two parts along the channel dimension. This represents the element-wise multiplication operation, used to enhance effective feature representation and suppress redundant features.

[0116] After multiple layers are stacked, a fused representation with local and global multi-scale features is obtained.

[0117] Further, please refer to Figure 5 The connection synthesis module will fuse the above representation. Flattened into feature vectors The input is fed into a multilayer perceptron structure, undergoes linear transformation, nonlinear activation, and a sigmoid function, and the output dimension equals the number of EEG functional connection edges. The output result is in vector form, which can be further mapped to the upper triangular part of the EEG functional connection matrix, and then padded with symmetry to obtain the complete matrix structure.

[0118] Furthermore, the training process of the deep learning model for fMRI–EEG functional connectivity based on deep cross-modal representation learning in step S7 is as follows:

[0119] Step S71: Divide the dataset into training and testing sets, set the number of iterations, and initialize the model hyperparameters.

[0120] Step S72: Organize the sample pairs obtained in step S2 into... , For the first A BOLD ROI-time window tensor T represents the number of ROIs, and T represents the TR number of fMRI, which is the hemodynamic delay time set in step S22. To correspond to the upper triangular vectorization result of the EEG functional connection (FC), the number of channels is ,but The numerical value is taken as the absolute correlation strength.

[0121] Step S73: Input the processed data into the neural network to obtain the functional connectivity of the fitted EEG signal, with the loss function being... Where MSE is the standard mean square error. This is the negative Pearson correlation coefficient. The loss function is a weighted combination of the mean squared error and the negative Pearson correlation coefficient, ensuring that the predicted results are numerically close to the actual EEG function while maintaining overall structural consistency with the real network. The weighting coefficients can be experimentally set to achieve a balance between the two types of indicators.

[0122] Step S74: Introduce an early stopping strategy during training. Stop training when the loss function does not decrease within several consecutive iterations to avoid overfitting and save the current optimal model parameters.

[0123] Furthermore, the process of evaluating the prediction results in step S8 is as follows:

[0124] Step S81: Calculate the root mean square error between the predicted EEG functional connectivity and the actual functional connectivity to measure the difference between the predicted and actual values.

[0125] ;

[0126] in, The first element of the real functional connection matrix represents the... One element, For the prediction function connection matrix, the first One element, The total number of matrix elements.

[0127] Step S82: To further evaluate the similarity between the predicted results and the actual results, cosine similarity is used to measure the directional similarity between the two sets of functional connectivity matrices.

[0128] ;

[0129] in, and The first and second connections of the real and predictive function matrices are respectively Each element is evaluated for similarity in the vector space.

[0130] Step S83: Use the Pearson correlation coefficient to evaluate the linear relationship between the prediction results and the true functional connectivity matrix.

[0131] ;

[0132] in, and These are the mean values ​​of the connection matrices for the true and predicted functions, respectively.

[0133] Further, please refer to Figure 5 The connection synthesis module will fuse the above representation. Flattened into feature vectors The input is fed into a multilayer perceptron structure, undergoes linear transformation, nonlinear activation, and a sigmoid function, and the output dimension equals the number of EEG functional connection edges. The output is in vector form, which can be further mapped to the upper triangular part of the EEG functional connection matrix, and then padded with symmetry to obtain the complete matrix structure.

[0134] Please see Figure 6 To compare the performance of the proposed method and the baseline method in different frequency bands, three indices were compared: mean Pearson correlation coefficient (a), cosine similarity (b), and root mean square error (c). The two baseline methods are a linear model (ridge regression with ℓ2 regularization) used to establish a reference benchmark for linear learnability and a general nonlinear model (shallow MLP with comparable number of parameters) to test whether simple nonlinearity is sufficient.

[0135] Please see Figure 7 This is a comparison diagram of the top 10% of edges predicted and actual functional connections in different frequency bands using the method proposed in this invention. Row (a) shows the actual connections, and row (b) shows the predicted connections.

[0136] Therefore, this invention proposes an EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning. By extracting BOLD signals from fMRI, a constructed neural network model is used to learn the cross-modal spatiotemporal information between the functional connectivity of BOLD signals and EEG signals in a supervised learning manner, and the functional connectivity of the corresponding BOLD segment in the EEG signal is directly estimated.

Claims

1. An EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning, comprising the following steps: Step S1: Preprocessing fMRI data; Remove the first 10 time point data, and perform time correction and head motion correction to ensure time consistency; Then, normalize the image in space, register it to the MNI standard space, and perform spatial smoothing processing with a 4mm full-width half-height Gaussian kernel to improve the signal-to-noise ratio; finally, perform band-pass filtering of 0.01-0.1 Hz on the time series to remove low-frequency drift and high-frequency noise, and extract fMRI signals of 200 brain regions of interest based on the Schaefer standard brain region template; fMRI represents functional magnetic resonance imaging; Step S2: Preprocessing of head EEG data; EEG data is synchronously collected by a 64-channel MR compatible amplifier with a sampling rate of 2000 Hz; first, remove the reference electrode and channels with quality below the threshold, and remove gradient artifacts and cardiac artifacts; downsample the cleaned electroencephalogram signal to 250 Hz, and use 0.5-40 Hz band-pass filtering to remove low-frequency drift and high-frequency noise; then, divide the EEG data into non-overlapping signal segments according to a 2-second time window to provide a time window for subsequent functional connectivity calculation; EEG represents electroencephalogram; Step S3: Divide the frequency-specific EEG signal; The preprocessed electroencephalogram signal is divided into 1-4 Hz delta, 4-8 Hz theta, 8-13 Hz alpha, 13-30 Hz beta, and 30-45 Hz gamma bands according to its frequency distribution using band-pass filtering, obtaining frequency-specific EEG signals; Step S4: Calculate the functional connectivity matrix; For each segment of data, the Pearson correlation value between all pairs of electrodes is calculated, resulting in a connectivity matrix ; Step S5: Perform EEG-BOLD time alignment; According to the EEG frequency-specific functional connectivity obtained in steps S1-S3 and the corresponding BOLD signal segment, each functional connectivity segment is paired with a 10-second delayed fMRI segment to form a sample pair of EEG functional connectivity and BOLD signal segment, and a dataset is constructed; BOLD represents blood oxygen level dependence; Step S6: Construct a deep cross-modal representation learning model; An EEG functional connectivity prediction model based on fMRI deep cross-modal representation learning is constructed, which consists of a time projection module, a cross-attention fusion encoder module, and a connection synthesis module; the time projection module is responsible for time compression and feature extraction of fMRI data; the cross-attention fusion encoder module optimizes the features through a multi-head self-attention mechanism and a local feature extraction module, further improving the expression ability of the model; the connection synthesis module maps the optimized features to the final predicted EEG functional connectivity strength; Step S7: Model training and testing; In data set division, the EEG-fMRI data set obtained in step S5 is trained and tested by leave-one-out cross-validation; in each round of cross-validation, the data of one subject is selected as the test set, and the data of the remaining subjects is selected as the training set; The training process adopts supervised learning, the model is trained by an Adam optimizer, the optimization process uses adaptive moment estimation to adjust the learning rate of each parameter, and the weights of the neural network are updated through back propagation; in the design of the loss function, a composite loss function is adopted Optimized by the following formula: ; wherein, is the mean squared error between predicted and true values, is the negative Pearson correlation coefficient, is a regulation coefficient that controls the weight of each part in the loss function; Step S8: Prediction result evaluation; For the predicted EEG functional connectivity, three evaluation indicators are used for quantitative evaluation: (1) Root mean square error, the root mean square error between the predicted EEG functional connectivity and the real calculated EEG functional connectivity is calculated to measure the difference between the predicted value and the actual value; (2) Cosine similarity, to further evaluate the similarity between the predicted results and the true results, cosine similarity is used to measure the directional similarity of the two groups of functional connectivity matrices; (3) Pearson similarity, in order to more comprehensively evaluate the linear relationship between the predicted results and the true functional connectivity matrix.

2. The EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning according to claim 1, wherein the step S4 of constructing the EEG frequency-specific functional connectivity network is as follows: For the first segment data, the Pearson correlation value between all electrode pairs of the electroencephalogram signal is calculated to form a 60x60 functional connection matrix , the calculation formula of which is: ; wherein For the first channel in time EEG signal, is the average value of this channel within the segment, T is the segment length; taking the absolute value of the functional connectivity matrix characterizes the connection strength.

3. The EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning according to claim 1, wherein the step S5 of EEG-BOLD time alignment is as follows: Step S51: According to the results of steps S1-S2, obtain the preprocessed and cut into fixed time window EEG signal segments, and the corresponding BOLD signal segments; Step S52: Set hemodynamic delay time The center time point of each EEG signal segment Corresponds to the delayed fMRI time point Corresponds to the delayed fMRI time point Step S53: Extract the BOLD signal segments matching the EEG time window within the corresponding delay fMRI window, ensuring that the two types of data are aligned in the time dimension; Step S54: Combine the matched EEG functional connectivity matrix and the corresponding BOLD signal segment into a sample, and store it in the data set; Step S55: Repeat the above operation until the construction of all EEG-BOLD sample pairs is completed, and finally obtain the complete training and test data set.

4. The method of claim 1, wherein the method is based on fMRI deep cross- modality representation learning for EEG functional connectivity prediction. The step S6 of constructing the EEG functional connectivity model based on fMRI deep cross-modal representation learning is as follows: Step S61: The time projection module is to compress the fMRI input signal in the time dimension; for the input tensor wherein is the batch size, is the number of ROIs extracted in step S1, the time series of each ROI is first adaptively pooled to compress several time points into a single feature representation, and then the convolutional layer expands the number of channels to obtain a feature representation suitable for subsequent convolution and attention processing; finally, a nonlinear activation function ReLU is applied to the result to enhance the feature expression ability of the model; Step S62: The cross-attention fusion encoder module is stacked by multiple identical encoding blocks, each encoding block includes a context-aware local-global fusion module and a feature aggregation module; the context-aware local-global fusion module is composed of a local convolution branch and a global self-attention branch; 1) In the local convolution branch, a combination of pointwise convolution and depthwise convolution is first used to extract local features within the ROI, and a gating mechanism is used to enhance the effective feature representation. ,in, Pointwise convolution is a cross-channel fusion. convolution, For depthwise convolution is Depth convolution of local spatial patterns, It is a zero-parameter gating mechanism, in which and These represent that the feature tensor Z is equally divided into two parts along the channel dimension. This represents the element-wise multiplication operation, used to enhance effective feature representation and suppress redundant features; 2) In the global self-attention branch, the query-key-value-based multi-head attention mechanism is used to model the dependency between all ROIs to obtain global interaction information ; ; wherein, is a query, key, value matrix, is the number of heads of multi-head attention, is the feature dimension of each head, is a projection layer; 3) finally, the results of the local convolution branch and the global attention branch are fused with learnable weights , is a learnable parameter for balancing the local and global branches; 4) Then through the feature fusion module, the feature fusion module re-fuses the features using the point-wise fully connected layer with residual connection to obtain the fused features ; ; wherein, is layer normalization, and is cross-channel fusion point-wise convolution; After multi-layer stacking, a fusion representation with local and global multi-scale features is obtained; Step S63: the connection synthesis module fuses the above representations Flattened into a feature vector , input into a multi-layer perception structure, linear transformation, nonlinear activation and sigmoid function, output dimension equal to the number of EEG functional connection edges; The output result is in vector form, which is further mapped into the upper triangular part of the EEG functional connection matrix, and the complete matrix structure is obtained through symmetry filling.

5. The method of claim 1, wherein the method is based on fMRI deep cross- modality representation learning for EEG functional connectivity prediction. The training process of the deep learning model of fMRI-EEG functional connectivity based on deep cross-modal representation learning in step S7 is as follows: Step S71: Divide the data set into training set and test set, set the number of iterations and initialize the model hyperparameters; Step S72: arrange the samples obtained in step S2 into , is the first BOLD ROI-time window tensor; is the first BOLD ROI-time window tensor; is the upper triangular vectorization result of the corresponding EEG functional connectivity, and the number of channels is then the numerical value is the absolute correlation strength; Step S73: input the processed data into the neural network to obtain the functional connectivity of the fitted EEG signal, and the loss function is where MSE is the standard mean square error, is the negative Pearson correlation coefficient; Step S74: The training process introduces an early stopping strategy, which stops training when the loss function does not decrease for a consecutive a number of iterations.

6. The EEG functional connectivity prediction method based on fMRI deep cross-modal representation learning according to claim 1, wherein the step S8 of predicting the results is as follows: Step S81: Calculate the root mean square error between the predicted EEG functional connectivity and the real functional connectivity to measure the difference between the predicted value and the actual value; ; wherein, The first element of the real functional connection matrix represents the... One element, For the prediction function connection matrix, the first There are n elements, where Q is the total number of elements in the matrix; Step S82: Cosine similarity is used to measure the directional similarity of the two sets of functional connectivity matrices; ; Step S83: Calculate the Pearson correlation coefficient to evaluate the linear relationship between the predicted results and the true functional connectivity matrix; ; wherein, and are the mean values of the true and predicted functional connectivity matrices, respectively.

Citation Information

Cited By

  • Unmanned aerial vehicle control method and system based on EEG and fNIRS multi-modal feature fusion

    CN122020129A