A brain decoding system based on functional magnetic resonance imaging and latent diffusion models

By combining spherical convolution and mask autoencoder to extract fMRI features and using potential diffusion models to generate images, the problem of reconstructing visual images from fMRI data in the prior art is solved, and efficient and accurate image reconstruction and brain decoding are achieved.

CN119273783BActive Publication Date: 2025-06-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411338463.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-06-13
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately reconstruct external visual images from functional magnetic resonance imaging (fMRI) data, especially when processing low signal-to-noise ratio and high dimensional fMRI signals.

Method used

A brain decoding system based on fMRI images and potential diffusion models is adopted, which combines spherical convolution and masked autoencoder to extract neural activity features from fMRI data and uses the latent diffusion model to generate images of high semantic information.

Benefits of technology

By reducing the dimension of fMRI data, accurately capturing the distance between brain regions and local global activation patterns, mitigating the impact of noise, generating images with more accurate semantic information, and improving the accuracy of brain decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119273783B_ABST
    Figure CN119273783B_ABST
Patent Text Reader

Abstract

The present invention discloses a brain decoding system based on functional magnetic resonance imaging and latent diffusion model, comprising: a functional magnetic resonance imaging feature extraction module, an image generation module and an image output module; the functional magnetic resonance imaging feature extraction module combines spherical convolution and masked autoencoder to extract neural activity features related to visual stimuli from functional magnetic resonance imaging (fMRI) data, the image generation module uses these neural activity features as conditions to generate images with high semantic information, and the output module processes and outputs the generated images. The present invention solves the major challenges of reconstructing visual stimuli from fMRI data, including its low signal-to-noise ratio and high-dimensional problems. By introducing a feature learning algorithm based on surface fMRI data, the dimension of fMRI data is effectively reduced, the distance between brain regions is accurately captured, and local and global brain activation patterns are learned; in addition, the design of the noise-resistant latent diffusion model (NR-LDM) further improves the quality of the reconstructed images by reducing the influence of fMRI noise, and can generate images with more accurate semantic information, which helps to improve the accuracy of brain decoding based on fMRI data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of neuroscience and computer vision technologies, and particularly to a brain decoding system based on functional magnetic resonance imaging and latent diffusion models. Background Art

[0002] With the rapid development of neuroscience and artificial intelligence technologies, researchers have gradually started to explore how to decode and reconstruct brain activities, especially to predict and reproduce the external visual images perceived by humans through functional magnetic resonance imaging (fMRI) data. The fMRI technology can capture the changes in blood oxygen level in the brain when processing visual information, reflecting the activities of specific brain regions. However, the fMRI signals have characteristics such as high dimensionality, low signal-to-noise ratio, and complex time series, making it a highly challenging task to extract the neural activity features directly related to visual stimuli from them and convert them into images.

[0003] Traditional brain activity decoding methods usually rely on linear models or simple neural networks. Although these methods can decode fMRI signals to a certain extent, the generated images usually have low quality and lack high-level semantic information. To solve these problems, in recent years, researchers have introduced more complex deep learning models, such as generative adversarial networks (GANs) and variational autoencoders (VAEs), to enhance the decoding ability and improve the quality of the generated images. However, when dealing with complex brain activity signals, these models still have problems such as insufficient generalization ability, unstable training, and difficulty in guaranteeing the quality of the generated images.

[0004] In this context, diffusion models, as an emerging generative model, have demonstrated their great potential in high-quality image generation. Diffusion models generate clear images in reverse by gradually adding perturbations to the noise in the latent space, and can capture complex data distributions and generate high-fidelity images. However, how to effectively apply diffusion models to the field of brain activity decoding, especially in the process of generating high-semantic information images from fMRI data, remains an unsolved problem. Summary of the Invention

[0005] Object of the Invention: The present invention provides a brain decoding system based on functional magnetic resonance imaging and latent diffusion models, which can efficiently and accurately reconstruct external visual images from functional magnetic resonance imaging (fMRI) data.

[0006] Technical solution: A brain decoding system based on functional magnetic resonance imaging and latent diffusion model according to the present invention includes: a functional magnetic resonance imaging feature extraction module, an image generation module, and an image output module; the functional magnetic resonance imaging feature extraction module combines spherical convolution and masked autoencoder to extract neural activity features related to visual stimuli from functional magnetic resonance imaging (fMRI) data, the image generation module uses these neural activity features as conditions to generate images with high semantic information, and the output module processes and outputs the generated images.

[0007] Furthermore, the functional magnetic resonance imaging feature extraction module includes a spherical convolution part, a masked autoencoder part, and a spherical deconvolution part; the spherical convolution part includes a spherical convolution layer and a spherical pooling layer, the masked autoencoder part includes a Vision Transformer (ViT) encoder and a ViT decoder, and the spherical deconvolution part includes a series of spherical deconvolution layers; the fMRI data first passes through the spherical convolution part, the spherical convolution layer effectively captures the local features of the data by smoothly moving the DirectNeighbor (DiNe) convolution kernel on the sphere while maintaining the integrity of the spatial information, the spherical pooling layer aggregates small regions in the feature map into a single value using average pooling, reducing the computational amount and extracting the main features, the feature map after spherical convolution will pass through the spherical pooling layer, and through the spherical convolution part, the dimension of the feature map gradually increases, indicating that higher-level features are gradually extracted; the obtained feature map is sent to the masked autoencoder part, the masked autoencoder part is implemented based on ViT, through the ViT encoder, an fMRI representation is obtained, and then this representation is sent to the ViT decoder to reconstruct the feature map, this process forces the model to learn how to reconstruct the lost information, thereby capturing the hidden structure of the data; the reconstructed feature map is processed by the spherical deconvolution part, and these feature maps are gradually restored to the dimension and spatial resolution of the original data.

[0008] Furthermore, the intermediate features extracted by the ViT encoder in the masked autoencoder are used for subsequent generation tasks, and the mean squared error is used as the reconstruction loss (L recon ) to measure the similarity between the input data x and the reconstructed data x':

[0009]

[0010] where S is the number of samples, and s represents the s-th sample.

[0011] Furthermore, the image generation module is based on the latent diffusion model (LDM), and the LDM operates in the image latent space, including a forward diffusion process and a reverse denoising process.

[0012] Furthermore, the forward diffusion process generates a series of latent variables y by iteratively adding Gaussian noise to the image 0 , y 1 ......, y T , while the reverse process aims to denoise by learning a denoising function ∈ θ to restore y T to y 0 ; After receiving the fMRI representation extracted by the feature extraction module, the image generation module first obtains the main fMRI representation through pseudo - principal component analysis (pseudo - PCA), and then uses the main fMRI representation as a control condition to control the LDM to generate pictures from noise;

[0013] The pseudo - PCA part uses linear projection (W k ) to approximate PCA to extract the main neurons in the network, identify and retain the key features (principal components) that capture most of the data, while discarding the less important components usually associated with noise, thereby effectively removing noise. For the fMRI representation e fMRI obtained in the feature extraction module, the main fMRI representation z fMRI is obtained through pseudo - PCA. The specific formula is as follows:

[0014] z fMRI = e fMRI ×W k .

[0015] Furthermore, an alignment loss (L Align ) is introduced to ensure that pseudo - PCA learns the main semantic concepts while minimizing the impact of noise; L Align quantifies the difference between the embedding of the generated image and the main fMRI embedding, thereby improving the effectiveness of the fMRI features for visual stimulus reconstruction; For the representation e gen of the generated image, it becomes the same dimension as z θ through the projection layer ρ fMRI , and the mean square error is calculated as the loss. The specific formula is as follows:

[0016]

[0017] Furthermore, to provide more powerful guidance for the reconstruction task, cross - attention conditioning and time - step conditioning are also integrated into the LDM, and the mean square error between the real noise ∈ and the predicted noise ∈ θ (y t , t, z fMRI , σ θ (z fMRI )) is calculated as the loss. The specific formula is as follows:

[0018]

[0019] where y 0 represents the latent variable of the given image, z is the conditional control variable, and t represents the time step. is the corresponding mathematical expectation, and z fMRI represents the main fMRI representation, and σ θ (z fMRI ) represents the time step embedding from the dimensional projector σ θ , ∈ represents the real noise, and ∈ θ is a set of denoising functions, which are implemented as U-Net.

[0020] Furthermore, a perceptual loss (L PSM ) is added. L PSM calculates the similarity between the representation (e gen ) of the generated image and the representation (e orig ) of the real image, guiding the model to prioritize high-level semantic information. The specific formula is as follows:

[0021]

[0022] Furthermore, the overall objective function is as follows:

[0023] L total = L DC-LDM + λ 1 * L Align + λ 2 * L PSM

[0024] where λ 1 , λ 2 are used to balance the relative importance of the alignment loss and the perceptual loss.

[0025] Furthermore, the image output module performs a series of processes on the generated image to ensure the best visualization effect of the data, facilitating subsequent analysis and interpretation.

[0026] Advantageous effects: Compared with the prior art, the present invention has the following remarkable advantages: The present invention solves the major challenges of reconstructing visual stimuli from fMRI data, including its low signal-to-noise ratio and high dimensionality problems. By introducing a feature learning algorithm based on surface fMRI data, the dimensionality of fMRI data is effectively reduced, the distances between brain regions are accurately captured, and local and global brain activation patterns are learned. In addition, the design of the denoising latent diffusion model (NR-LDM) further improves the quality of the reconstructed image by reducing the influence of fMRI noise, and can generate images with more accurate semantic information, which helps to improve the accuracy of brain decoding based on fMRI data. Description of the Drawings

[0027] Figure 1 This is a schematic diagram of the brain decoding system of the present invention.

[0028] Figure 2 This is a schematic diagram of the functional magnetic resonance imaging feature extraction module of the present invention.

[0029] Figure 3 This is a schematic diagram of the image generation module of the present invention.

[0030] Figure 4 This is a comparison graph of the generated images of all comparison methods. Detailed implementation manners

[0031] As Figure 1 shown, a brain decoding system based on functional magnetic resonance imaging and latent diffusion model includes: a functional magnetic resonance imaging feature extraction module, an image generation module, and an image output module; the functional magnetic resonance imaging feature extraction module combines spherical convolution and masked autoencoder to extract neural activity features related to visual stimuli from functional magnetic resonance imaging (fMRI) data, the image generation module uses these neural activity features as conditions to generate images with high semantic information, and the output module processes and outputs the generated images.

[0032] As Figure 2 shown, the functional magnetic resonance imaging feature extraction module includes a spherical convolution part, a masked autoencoder part, and a spherical deconvolution part; the spherical convolution part includes a spherical convolution layer and a spherical pooling layer, the masked autoencoder part includes a ViT encoder and a ViT decoder, and the spherical deconvolution part includes a series of spherical deconvolution layers; the fMRI data first passes through the spherical convolution part, the spherical convolution layer effectively captures the local features of the data by smoothly moving the DiNe convolution kernel on the sphere while maintaining the integrity of the spatial information, the spherical pooling layer uses average pooling to aggregate small regions in the feature map into a single value, reducing the computational amount and extracting the main features, the feature map after spherical convolution will pass through the spherical pooling layer, and through the spherical convolution part, the dimension of the feature map gradually increases, indicating that higher-level features are gradually extracted; the obtained feature map is sent to the masked autoencoder part, the masked autoencoder part is implemented based on ViT, through the ViT encoder, the fMRI representation is obtained, and then this representation is sent to the ViT decoder to reconstruct the feature map, this process forces the model to learn how to reconstruct the lost information, thereby capturing the hidden structure of the data; the reconstructed feature map is processed by the spherical deconvolution part, and these feature maps are gradually restored to the dimension and spatial resolution of the original data.

[0033] The intermediate features extracted by the ViT encoder in the masked autoencoder are used for subsequent generation tasks, and the mean square error is used as the reconstruction loss (L recon) to measure the similarity between the input data x and the reconstructed data x':

[0034]

[0035] where S is the number of samples, and s represents the s-th sample.

[0036] As Figure 3 shown, the image generation module is based on the latent diffusion model LDM, which operates in the image latent space and includes a forward diffusion process and a reverse denoising process.

[0037] The forward diffusion process generates a series of latent variables y 0 , y 1 ......, y T by iteratively adding Gaussian noise to the image, while the reverse process aims to denoise by learning a denoising function ∈ θ to restore y T to y 0 ; after receiving the fMRI representation extracted by the feature extraction module, the image generation module first obtains the main fMRI representation via pseudo-PCA, and then uses the main fMRI representation as a control condition to control LDM to generate pictures from noise;

[0038] The pseudo-PCA part uses linear projection (W k ) to approximate PCA to extract the main neurons in the network, identify and retain the key features (principal components) that capture most of the data, and discard the less important components that are usually related to noise, thereby effectively removing noise. For the fMRI representation e fMRI obtained in the feature extraction module, the main fMRI representation z fMRI is obtained via pseudo-PCA, and the specific formula is as follows:

[0039] z fMRI = e fMRI × W k .

[0040] An alignment loss (L Align ) is introduced to ensure that pseudo-PCA learns the main semantic concepts while minimizing the influence of noise; L Align quantifies the difference between the embedding of the generated image and the main fMRI embedding, thereby improving the effectiveness of the fMRI features for visual stimulus reconstruction; for the representation e gen of the generated image, it becomes the same dimension as z fMRI via the projection layer, and the mean square error is calculated as the loss. The specific formula is as follows:

[0041]

[0042] To provide more powerful guidance for the reconstruction task, cross-attention conditioning and time-step conditioning are also integrated into LDM, and the mean squared error between the real noise ∈ and the predicted noise ∈ θ (y t ,t,z fMRI ,σ θ (z fMRI )) is used as the loss, and the specific formula is as follows:

[0043]

[0044] where y 0 represents the given image latent variable, z is the conditional control variable, t represents the time step, is the corresponding mathematical expectation, z fMRI represents the main fMRI characterization, σ θ (z fMRI ) represents the time-step embedding from the dimensional projector σ θ , ∈ represents the real noise, ∈ θ is a set of denoising functions, which are implemented as U-Net.

[0045] The perceptual loss (L PSM ) is added. L PSM calculates the similarity between the characterization (e gen ) of the generated image and the characterization (e orig ) of the real image, guiding the model to prioritize high-level semantic information. The specific formula is as follows:

[0046]

[0047] The overall objective function is as follows:

[0048] L total = L DC-LDM + λ 1 * L Align + λ 2 * L PSM

[0049] where λ 1 , λ 2 are used to balance the relative importance of the alignment loss and the perceptual loss.

[0050] The image output module performs a series of processes on the generated image to ensure the best visualization effect of the data for subsequent analysis and interpretation.

[0051] Numerous experiments were conducted on three public datasets for comprehensive evaluation. The datasets include NSD, BOLD5000, and GOD. The NSD dataset collected fMRI images of 8 subjects while they viewed over 70,000 high-resolution natural scene images. The BOLD5000 dataset includes fMRI imaging data from 4 subjects, recording their brain activities while viewing images from the COCO, ImageNet, and Scene Recognition databases, with a total of 5,254 images. The GOD dataset contains 1,250 different images from 200 different categories. The N-way-Top-K was used as an indicator, which is a classification metric that can effectively measure the semantic fidelity of the generated images, and the larger the value, the better. As shown in Table 1 and Figure 4 As can be seen from the experimental results, compared with the existing methods, the present invention has better performance on all datasets and can generate images with higher semantic information and more credible

[0052] Table 1 Comparative experiments of the system of the present invention and existing brain decoding methods on three datasets

[0053]

[0054]

[0055] At the same time, ablation experiments were conducted on the effectiveness of the system to verify each part of the image generation module. The experimental results are shown in Table 2. In Table 2, LDM was used as the baseline model, and perceptual loss, pseudo-PCA, and alignment loss were added to reduce the noise impact in the learned fMRI feature representation. After adding the perceptual loss, the semantic fidelity of the generated images was improved. After further adding pseudo-PCA to extract the main fMRI representations, the quality was further improved. Finally, the alignment loss was incorporated, effectively learning the main semantic concepts. As can be seen from the experimental results, each part can effectively improve the model performance.

[0056] Table 2 Verification of the effectiveness of each part of the system image generation module

[0057]

Claims

1. A brain decoding system based on functional magnetic resonance imaging and latent diffusion model, characterized in that: include: Functional magnetic resonance imaging feature extraction module, image generation module and image output module; the functional magnetic resonance imaging feature extraction module combines spherical convolution and mask autoencoder to extract neural activity features related to visual stimulation from functional magnetic resonance imaging (fMRI) data, the image generation module uses these neural activity features as conditions to generate images with high semantic information, and the output module processes and outputs the generated images; The functional magnetic resonance imaging feature extraction module includes a spherical convolution part, a masked autoencoder part and a spherical deconvolution part; the spherical convolution part includes a spherical convolution layer and a spherical pooling layer, the masked autoencoder part includes a ViT encoder and a ViT decoder, and the spherical deconvolution part includes a series of spherical deconvolution layers; the fMRI data first passes through the spherical convolution part, the spherical convolution layer smoothly moves the DiNe convolution kernel on the sphere, and the spherical pooling layer uses average pooling to aggregate small areas in the feature map into a single value. The feature map after spherical convolution will pass through the spherical pooling layer. Through the spherical convolution part, the dimension of the feature map gradually increases, indicating that higher-level features are gradually extracted; the obtained feature map is sent to the masked autoencoder part, which is implemented based on ViT. Through the ViT encoder, the fMRI representation is obtained, and then the representation is sent to the ViT decoder to reconstruct the feature map; The reconstructed feature maps are processed by the spherical deconvolution part, and these feature maps are gradually restored to the dimension and spatial resolution of the original data.

2. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 1, characterized in that: The intermediate features extracted by the ViT encoder in the mask autoencoder are used for subsequent generation tasks, and the mean square error is used as the reconstruction loss L recon To measure the similarity between the input data x and the reconstructed data x': Where S is the number of samples and s represents the sth sample.

3. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 1, characterized in that: The image generation module is based on the latent diffusion model (LDM), which operates on the image latent space and includes a forward diffusion process and a reverse denoising process.

4. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 3, characterized in that: The forward diffusion process generates a series of latent variables y0, y1, ..., y by iteratively adding Gaussian noise to the image. T , while the reverse process aims to learn a denoising function ∈ θ To remove noise, y T After receiving the fMRI representation extracted by the feature extraction module, the image generation module first obtains the main fMRI representation through pseudo principal component analysis pseudo PCA, and then uses the main fMRI representation as a control condition to control LDM to generate images from noise; The pseudo PCA part uses linear projection W k To approximate PCA to extract the main neurons in the network, identify and retain the key features that capture most of the data, while discarding less important components that are usually related to noise. For the fMRI representation obtained in the feature extraction module, fMRI , the main fMRI representation z was obtained through pseudo PCA fMRI , the specific formula is as follows: with fMRI =e fMRI ×W k 。 5. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 4, characterized in that The alignment loss L Align , L Align We quantify the difference between the embeddings of generated images and the primary fMRI embeddings; for the representation of the generated images, we gen , through the projection layer ρ θ becomes and z fMRI The same dimension, calculate the mean square error as the loss, the specific formula is as follows:

6. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 5, characterized in that: In order to provide more powerful guidance for the reconstruction task, cross-attention conditioning and time-step conditioning are also integrated into LDM to calculate the true noise ∈ and the predicted noise ∈ θ (y t ,t,z fMRI ,σ θ (z fMRI )) is taken as the loss, and the specific formula is as follows: Where y0 represents the latent variable of a given image, z is the conditional control variable, t represents the time step, and E y0,z,t is the corresponding mathematical expectation, z fMRI represents the main fMRI representation, σ θ (z fMRI ) represents the dimension projector σ θ The time step embedding of ∈ represents the real noise, ∈ θ is a set of denoising functions, which are implemented as U-Net.

7. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 6, characterized in that: Added perceptual loss L PSM , L PSM Computationally generated image representations gen and real image representation orig The similarity between them guides the model to give priority to high-level semantic information. The specific formula is as follows:

8. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 7, characterized in that: The overall objective function is as follows: L total =L DC-LDM +λ1*L Align +λ2*L PSM Among them, λ1 and λ2 are used to balance the relative importance of alignment loss and perceptual loss.

9. The brain decoding system based on functional magnetic resonance imaging and latent diffusion model as claimed in claim 1, characterized in that: The image output module performs a series of processing on the generated images to ensure the best visualization of the data and facilitate subsequent analysis and interpretation.

Citation Information

Patent Citations

  • Method and system for reconstructing visual stimulation image based on human brain fMRI

    CN118135052A