Machine learning model for reconstructing video or audio data based on neuroimaging data
A machine learning model using a masked autoencoder and diffusion model with spatiotemporal attention and multimodal learning effectively reconstructs high-quality video and audio from neuroimaging data, addressing the limitations of existing methods.
Patent Information
- Application Number
- PCT/SG2024/050328
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-20
AI Technical Summary
Existing methods struggle to accurately reconstruct dynamic video or audio data from neuroimaging data, particularly due to the limitations of non-invasive tools like fMRI, which capture limited information susceptible to noise and have temporal resolution issues, and the challenge of decoding high-quality audio from non-invasive brain recordings.
A machine learning model using a masked autoencoder and diffusion model trained with unsupervised learning and masked data modeling, combined with spatiotemporal attention and multimodal contrastive learning, to generate neuroimaging data embeddings for reconstructing video or audio data.
The model achieves high-quality video reconstruction with accurate semantics and frame rates, outperforming previous state-of-the-art methods by 45% in semantic classification and pixel-level metrics, and effective audio reconstruction from neuroimaging data.
Smart Images

Figure SG2024050328_20112025_PF_FP_ABST
Abstract
Description
MACHINE LEARNING MODEL FOR RECONSTRUCTING VIDEO OR AUDIO DATA BASED ON NEUROIMAGING DATATECHNICAL FIELD
[0001] The present invention generally relates to a machine learning model for reconstructing video or audio data based on ncuroimaging data of a subject, and more particularly, a method and a system for training the machine learning model and a method and a system for using the machine learning model trained for reconstructing video or audio data based on neuroimaging data of a subject.BACKGROUND
[0002] Life unfolds like a film reel, each moment seamlessly transitioning into the next, forming a “perpetual theater” of experiences. This dynamic narrative forms our perception, explored through the naturalistic paradigm, painting the brain as a moviegoer engrossed in the relentless film of experience. Understanding the information hidden within our complex brain activities is a big puzzle in cognitive neuroscience. The task of recreating human vision from brain recordings (neuroimaging data), especially using non-invasive tools like functional Magnetic Resonance Imaging (fMRI), is an exciting but difficult task. Non-invasive methods, while less intrusive, capture limited information, susceptible to various interferences like noise. Furthermore, the acquisition of ncuroimaging data is a complex, costly process. Despite these complexities, progress has been made, notably in learning valuable fMRI features with limited fMRI-annotation pairs. Deep learning and representation learning have achieved significant results in visual class detections and static image reconstruction, advancing our understanding of the vibrant, ever-changing spectacle of human perception.
[0003] Unlike still images, human vision is a continuous, diverse flow of scenes, motions, and objects. To recover dynamic visual experience, the challenge lies in the nature of fMRI, which measures blood oxygenation level dependent (BOLD) signals and captures snapshots of brain activity every few seconds. Each fMRI scan essentially represents an “average” of brain activity during the snapshot. In contrast, a typical video has about 30 frames per second (FPS). If an fMRI frame takes 2 seconds, during that time, 60 video frames - potentially containing various objects, motions, and scene changes - are presented as visual stimuli. Thus, decoding fMRI and recovering videos at an FPS much higher than the fMRI’s temporal resolution is a complex task.
[0004] Hemodynamic response (HR) refers to the lags between neuronal events and activation in BOLD signals. When a visual stimulus is presented, the recorded BOLD signal will have certain delays with respect to the stimulus event. Moreover, the HR varies across subjects and brain regions. Thus, the common practice that shifts the fMRI by a fixed number in time to compensate for the HR would be sub-optimal
[0005] Audio, the symphony of frequencies, harmonizes our understanding of the world, giving voice to the silent echo of the universe and unlocking the doors to the realms of human experiences and emotions. Neuroscientists have continuously explored the field of auditory processing, working to uncover the complexities within the brain that manage our cognitive functions and abilities. Notably, significant strides have been made to decode high-quality human speech from neural activities in speech-related cortical areas using invasive electrocorticography (ECoG) arrays. These advances hold meaningful implications, especially for individuals with severe speech impairments. In contrast, robust audio reconstruction based on non-invasive brain recordings still warrants deeper investigations.
[0006] A need therefore exists to provide a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject that is effective with accurate semantics. It is against this background that the present invention has been developed.SUMMARY
[0007] According to a first aspect of the present invention, there is provided a method of training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, the method comprising: training a neuroimaging data encoder based on neuroimaging data from a neuroimaging training dataset for generating neuro imaging data embeddings; and training a diffusion model based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model, the diffusion model being trained to reconstruct video or audio data based on neuroimaging data embeddings of neuroimaging data of a subject obtained in response to a visual or audio stimulus, wherein the neuroimaging data encoder comprises a masked autoencoder, and the above-mentioned training the neuroimaging data encoder comprises training an encoder of the masked autoencoder based on the neuroimaging training dataset using unsupcrviscd learning with masked data modeling for generating the ncuroimaging data embeddings, the unsupervised learning with masked data modeling comprising generatingneuroimaging data embeddings from neuroimaging data from the neuroimaging training dataset, masking a portion of the neuroimaging data embeddings into masked neuroimaging data embeddings and training the masked autoencoder to recover the masked neuroimaging data embeddings.
[0008] According to a second aspect of the present invention, there is provided a system for training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, the system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to: train a neuroimaging data encoder based on neuroimaging data from a neuroimaging training dataset for generating neuroimaging data embeddings; and train a diffusion model based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model, the diffusion model being trained to reconstruct video or audio data based on neuroimaging data embeddings of neuroimaging data of a subject obtained in response to a visual or audio stimulus, wherein the neuroimaging data encoder comprises a masked autocncodcr, and the above-mentioned train the neuroimaging data encoder comprises training an encoder of the masked autocncodcr based on the neuroimaging training dataset using unsupcrviscd learning with masked data modeling for generating neuroimaging data embeddings, the unsupervised learning with masked data modeling comprising generating neuroimaging data embeddings from neuroimaging data from the neuroimaging training dataset, masking a portion of the neuroimaging data embeddings into masked neuroimaging data embeddings and training the masked autoencoder to recover the masked neuroimaging data embeddings.
[0009] According to a third aspect of the present invention, there is provided a computer program product, embodied in one or more non-transitory computer-readable storage mediums, comprising instructions executable by at least one processor to perform the method of training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject according to the above-mentioned first aspect of the present invention.
[0010] According to a fourth aspect of the present invention, there is provided a method of using the machine learning model trained according to the above-mentioned first aspect of the present invention for reconstructing video or audio data based on neuroimaging data of a subject, tire method comprising:generating, by the neuroimaging data encoder of the machine learning model, neuroimaging data embeddings based on neuroimaging data of a subject obtained in response to a visual or audio stimulus; and reconstructing video or audio data, by the diffusion model of the machine learning model, based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model.
[0011] According to a fifth aspect of the present invention, there is provided a system for using the machine learning model trained according to the above-mentioned first aspect of the present invention for reconstructing video or audio data based on neuroimaging data of a subject, the system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to: generate, by the neuroimaging data encoder of the machine learning model, neuroimaging data embeddings based on neuroimaging data of a subject obtained in response to a visual or audio stimulus; and reconstruct video or audio data, by the diffusion model of the machine learning model, based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Embodiments of the present invention will be better understood and readily apparent to one of ordinary skill in the art from the following written description, by way of example only, and in conjunction with the drawings, in which:FIG. 1 depicts a schematic flow diagram of a method of training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, according to various embodiments of the present invention;FIG. 2 depicts a schematic block diagram of a system for training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, according to various embodiments of the present invention;FIG. 3 A depicts a schematic block diagram of a first exemplary computer system which may be used to realize or implement the system for training a machine learning model for reconstructing video or audio data, according to various embodiments of the present invention;FIG. 3B depicts a schematic block diagram of a second exemplary computer system which may be used to realize or implement the system for training a machine learning model for reconstructing video or audio data, according to various embodiments of the present invention;FIG. 4A shows an overview of brain decoding and video reconstruction;FIG. 4B shows an example visual stimuli and the reconstructed video, according to various example embodiments of the present invention;FIG. 5 depicts an overview of an example machine learning model for reconstructing video (referred to as Mind-Video), according to various example embodiments of the present invention;FIG. 6 illustrates the phenomenon due to hemodynamic response, causing a discrepancy between the fMRI data and the corresponding stimulus;FIG. 7 shows visual comparisons of video frames of reconstructed videos obtained using Mind-Video and those obtained using existing models;FIG. 8 shows a few examples of reconstructed frames using Mind-Video, including an example involving a scene transition, according to various example embodiments of the present invention;FIG. 9 shows structural similarity index (SSTM) comparisons of Mind-Video and existing models;FIG. 10 shows a table (Table 1) presenting results of the ablation study on window sizes, multimodal contrastive learning and adversarial guidance (AG) of Mind-Video, according to various example embodiments of the present invention;FIG. 11 shows attention visualization of the attention maps for different transformer layers and learning stages with bar charts and brain flat maps, according to various example embodiments of the present invention;FIG. 12 shows a table (Table 2) presenting the hyperparameters used in a large-scale pre-training of Mind-Video, according to various example embodiments of the present invention;FIG. 13 shows reconstruction samples of different subjects obtained using Mind-Video, according to various example embodiments of the present invention;FIG. 14 shows reconstruction samples obtained using Mind-Video for ablation studies, according to various example embodiments of the present invention;FIG. 15 shows reconstruction samples obtained using Mind-Video for certain fail cases;FIG. 16 depicts an ovendew of an example machine learning model for reconstructing audio (referred to as MinD-Audio), according to various example embodiments of the present invention;FIG. 17 shows a tabic (Tabic 3) presenting the evaluation results on fMRI-guided speech reconstruction for the NeuroSynth model and the AudioLDM model, according to various example embodiments of the present invention;FIG. 18 shows a table (Table 4) presenting the evaluation results on fMRI-guided music reconstruction for the NcuroSynth model and the AudioLDM model, according to various example embodiments of the present invention;FIG. 19A shows a comparison of generated mel-spectrograms between the NeuroSynth model and the AudioLDM model on the fMRI-to-Speech dataset, according to various example embodiments of the present invention;FIG. 19B shows a comparison of generated mel-spectrograms between the NeuroSynth model and the AudioLDM model on the fMRI-to-Music dataset, according to various example embodiments of the present invention;FTGs. 20A to 20C depict plots comparing between the NeuroSynth model and the AudioLDM model across different epochs using SI data of fMRl-to-Spccch dataset with respect three metrics FD (Frechet Distance), FAD (Frechet Audio Distance) and KL (Kullback- Lciblcr), according to various example embodiments of the present invention;FIG. 21 shows a table (Table 5) presenting the comparison results between different fMRI encoding models using SI data of the fMRI-to-Speech dataset, according to various example embodiments of the present invention;FIG. 22 shows a table (Table 6) presenting the comparison results between different fine-tuned components in the AudioLDM model using SI data of the fMRI-to-Speech dataset, according to various example embodiments of the present invention;FIGs. 23 A and 23B show the analysis of self-attention from SI in the NeuroSynth model and the AudioLDM model, according to various example embodiments of the present invention;FIG. 24A shows a table (Table 7) presenting the performance of the NcuroSynth model on the fMRI-to-Speech dataset, according to various example embodiments of the present invention;FIG. 24B shows a table (Table 8) presenting the performance of the NeuroSynth model on the fMRI-to-Music dataset, according to various example embodiments of the present invention;FIG. 25 shows the analysis of cross-attention maps from SI in the NcuroSynth model, according to various example embodiments of the present invention; andFIG. 26 depicts a schematic flow diagram of an example method of training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, according to various example embodiments of the present invention.DETAILED DESCRIPTION
[0013] Various embodiments of the present invention provide a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, and more particularly, a method and a system for training the machine learning model and a method and a system for using the machine learning model trained for reconstructing video or audio data based on neuro imaging data of a subject.
[0014] As explained in the background, there exist various technical challenges or complexities in effectively reconstructing video or audio data (or reconstructing perceived video or audio stimuli) based on neuroimaging data of a subject with accurate semantics. To seek to overcome, or at least ameliorate, these technical challenges, various embodiments of the present invention provide a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject that is effective with accurate semantics.
[0015] FIG. 1 depicts a schematic flow diagram of a method 100 of training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject using at least one processor, according to various embodiments of the present invention. The method comprising: training (at 106) a neuroimaging data encoder based on neuroimaging data from a neuroimaging training dataset for generating neuroimaging data embeddings; and training (at 108) a diffusion model based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model. In this regard, the diffusion model is trained to reconstruct video or audio data based on neuroimaging data embeddings of neuroimaging data of a subject obtained in response to a visual or audio stimulus. The neuroimaging data encoder comprises a masked autoencoder. In this regard, the above- mentioned training (at 106) the neuroimaging data encoder comprises training an encoder of the masked autoencoder based on the neuroimaging training dataset using iinsupcrviscdlearning with masked data modeling for generating the neuroimaging data embeddings, the unsupervised learning with masked data modeling comprising generating neuroimaging data embeddings from neuroimaging data from the neuroimaging training dataset, masking a portion of the ncuroimaging data embeddings into masked ncuroimaging data embeddings and training the masked autoencoder to recover the masked neuroimaging data embeddings.
[0016] In various embodiments, the method 100 further comprising determining activated brain regions of the neuroimaging data from the neuroimaging training dataset. In this regard, the neuroimaging data encoder is trained based on the activated brain regions of the neuroimaging data from the neuroimaging training dataset.
[0017] In various embodiments, the method 100 is fortraining the machine learning model for reconstructing video data. In this regard, the above-mentioned training (at 106) the neuroimaging data encoder further comprises augmenting the encoder of the masked autoencoder with spatiotemporal attention heads for processing the neuroimaging data embeddings generated by the encoder of the masked autoencoder in a sliding window. In various embodiments, for each of a plurality of windows of neuroimaging data embeddings generated by the encoder of the masked autoencoder, the above-mentioned training (at 106) further comprises training the encoder of the masked autocncodcr based on a set of embeddings comprising the window of neuroimaging data embeddings, corresponding image embeddings and corresponding text embeddings projected into a shared latent space based on contrastive learning. For example, each type (modality type) of data (e.g., fMRI, image, and text) is processed through its respective encoder (e.g., the above-mentioned encoder of the masked autoencoder, an image encoder, and a text encoder, respectively) to generate the fMRI embeddings, the image embeddings and the text embeddings. In this regard, the fMRI data and the video data (based on which the image embeddings are generated) are collected in synchronized sessions where the subject is exposed to visual stimuli (video data) while their brain activity is simultaneously recorded (fMRI data). Therefore, the images inputted to the image encoder correspond to video frames that the subject is viewing during the fMRI scan for obtaining the fMRI data. Concurrently, the text (video caption) for each video frame is generated by a pre-trained image captioning model (e.g, the BLIP (Bootstrapping languageimage pre-training) model) as a caption that describes a content of the video frame at that particular moment in time. Accordingly, these encoders standardize their output embeddings to the same dimensional space, facilitating direct comparison. The training involves multimodal contrastive learning, which utilizes the set of embeddings of different modalities comprisingthe aligned fMRI (neuroimaging data) embeddings, alongside similarly processed (e g , standardizing the input shape) image embeddings and text embeddings corresponding to the fMRI embeddings Tn various embodiments, these embeddings are aligned using the CLIP (contrastive language-image pre-training) methodology, which facilitates a robust training regime by leveraging contrastive learning. For example, the cosine similarity between these embeddings across the different modalities is calculated to ensure semantic coherence, minimizing the distance between semantically similar embeddings across the fMRI, image, and text data, while maximizing the distance between dissimilar ones. In various embodiments, each window captures a sequence of fMRI data corresponding to a specific timeframe, for example, specifically aligning two fMRI frames (e.g., covering a total of four seconds) with two seconds of video data. This alignment is provided for synchronizing the temporal dynamics of brain activity with the corresponding video frames, enhancing the model's ability to accurately reconstruct dynamic visual content from static fMRI images.
[0018] In various embodiments, for training the machine learning model for reconstructing video data, the diffusion model comprises a U-Net comprising self-attention layers, crossattention layers and temporal -attention layers; and the diffusion model is further trained based on a target video dataset.
[0019] Tn various embodiments, the method 100 is for training the machine learning model for reconstructing audio data. In this regard, in a first approach, the above-mentioned training (at 108) the diffusion model comprises training the diffusion model and the neuroimaging data encoder based on estimating noise at each time step of a denoising process of the diffusion model. For example, this first approach leverages raw neuroimaging data directly in training the diffusion model, thereby avoiding potential biases and limitations introduced by pre-trained data. For example, the diffusion model may undergo an extended training period of 1,000 epochs to ensure that the diffusion model teams the integration and synthesis of the necessary information for accurate signal decoding. Additionally, a higher initial teaming rate may be employed that is systematically reduced according to a predefined schedule, helping the diffusion model to rapidly converge on effective patterns and relationships early in training. Moreover, the diffusion model may be trained using full precision (32-bit floating-point) calculations, rather than using mixed-precision training in fine-tuning.
[0020] Tn various embodiments, for training the machine teaming model for reconstructing audio data in the first approach, the diffusion model comprises a U-Nct compnsing a set of encoder blocks and a set of decoder blocks. Each encoder block comprises a downsample blockand each decoder block comprising an upsample block. Furthermore, each encoder block and each decoder block corresponding to a low-resolution block further comprises a cross-attention layer. Tn this regard, the above-mentioned training the diffusion model further comprises performing cross-attention conditioning with respect to the cross-attention layer of the above- mentioned each encoder block and the above-mentioned each decoder block corresponding to a low-resolution block based on the ncuroimaging data embeddings generated by the neuroimaging data encoder. This advantageously enables the diffusion model to compute crossattention in low-resolution feature maps. In various embodiments, for example, low-resolution feature maps represent feature maps having a dimension (heightxwidth) of 64 - 64. 32x32, 16x 16 or smaller.
[0021] In various embodiments, the method 100 is fortraining the machine learning model for reconstructing audio data. In this regard, in a second approach, the diffusion model is pretrained and is further trained by fine-tuning jointly with the neuroimaging data encoder using the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model. For example, this fine-tuning is executed by adjusting the self-attention heads of the diffusion model to better align with the neuroimaging data embeddings generated by the ncuroimaging data encoder. For example, in contrast to the method for training the machine learning model for reconstructing video data which may fine-tune cross-attention mechanisms to integrate videos, the diffusion model of this second approach for training the machine learning model for reconstructing audio data uses a concatenation of audio embeddings and fMRI embeddings through its self-attention mechanism. For example, the audio embeddings may be obtained from recordings (audio data) of auditory stimuli such as spoken stories, which the subject listened to during fMRI scans. This audio data is processed to extract features, typically through signal processing techniques like Mel-frequency cepstral coefficients, which may then be projected to match the dimensionality of the fMRI embeddings. The concatenated audio and fMRI embeddings, representing both audio data and fMRI data, are then inputted into a self-attention mechanism within the diffusion model.
[0022] FIG. 2 depicts a schematic block diagram of a system 200 for training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject according to various embodiments of the present invention, corresponding to the method 100 of training a machine learning model for reconstructing video or audio data as described hereinbefore according with reference to FIG. 1 according to various embodiments of the present invention. Hie system 200 comprises: at least one memory 202; and at least oneprocessor 204 communicatively coupled to the at least one memory 202 and configured to perform the method 100 of training a machine learning model for reconstructing video or audio data as described hereinbefore according to various embodiments of the present invention. Accordingly, the at least one processor 204 is configured to: train a ncuroimaging data encoder based on neuroimaging data from a neuroimaging training dataset for generating neuroimaging data embeddings; and train a diffusion model based on the ncuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model. In this regard, the diffusion model is trained to reconstruct video or audio data based on neuroimaging data embeddings of neuroimaging data of a subject obtained in response to a visual or audio stimulus. The neuroimaging data encoder comprises a masked autoencoder. In this regard, the above-mentioned train the neuroimaging data encoder comprises training an encoder of the masked autoencoder based on the neuroimaging training dataset using unsupervised learning with masked data modeling for generating neuroimaging data embeddings, the unsupervised learning with masked data modeling comprising generating neuroimaging data embeddings from neuroimaging data from the neuroimaging training dataset, masking a portion of the neuroimaging data embeddings into masked neuroimaging data embeddings and training the masked autocncodcr to recover the masked ncuroimaging data embeddings.
[0023] It will be appreciated by a person skilled in the art that the at least one processor 204 may be configured to perform various functions or operations through sct(s) of instructions (e g., software modules) executable by the at least one processor 204 to perform various functions or operations. Accordingly, as shown in FIG. 2, the system 200 may comprise: a neuroimaging data encoder training module (or a neuroimaging data encoder training circuit) 206 configured to train the neuroimaging data encoder based on neuroimaging data from the neuroimaging training dataset for generating neuroimaging data embeddings: and a diffusion model training module (or a diffusion model training circuit) 208 configured to train a diffusion model based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model.
[0024] It will be appreciated by a person skilled in the art that the above-mentioned modules are not necessarily separate modules, and two or more modules may be realized by or implemented as one functional module (e.g., a circuit or a software program) as desired or as appropriate without deviating from the scope of the present invention. For example, the ncuroimaging data encoder training module 206 and the diffusion model training module 208 may be realized (e.g., compiled together) as one executable software program (e g., softwareapplication or simply referred to as an “app”), which for example may be stored in the at least one memory' 202 and executable by the at least one processor 204 to perform the corresponding functions or operations as described herein according to various embodiments.
[0025] In various embodiments, the system 200 for training a machine learning model for reconstructing video or audio data corresponds to the method 100 for training a machine learning model for reconstructing video or audio data as described hereinbefore with reference to FIG. 1, therefore, various operations, functions or steps configured to be performed by the least one processor 204 may correspond to various operations, functions or steps of the method 100 described hereinbefore according to various embodiments, and thus need not be repeated with respect to the system 200 for clarity and conciseness. In other words, various embodiments described herein in context of methods (e.g., the method 100) are analogously valid for the corresponding systems or devices (e.g., the system 200), and vice versa. For example, in various embodiments, the at least one memory' 202 may have stored therein the neuroimaging data encoder training module 206 and the diffusion model training module 208, which respectively' correspond to various operations, functions or steps of the method 100 as described hereinbefore according to various embodiments, which are executable by the at least one processor 204 to perform the corresponding operations, functions or steps as described herein.
[0026] A computing sy stem, a controller, a microcontroller or any other system providing a processing capability may be provided according to various embodiments in the present invention. Such a system may be taken to include one or more processors and one or more computer-readable storage mediums. For example, the system 200 described hereinbefore may include at least one processor (or controller) 204 and at least one computer-readable storage medium (or memory) 202 which are for example used in various processing earned out therein as described herein. A memory or computer-readable storage medium used in various embodiments may be a volatile memory , for example a DRAM (Dynamic Random Access Memory) or a non-volatile memory, for example a PROM (Programmable Read Only Memory), an EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or a flash memory, e.g., a floating gate memory', a charge trapping memory', an MRAM (Magnetoresistive Random Access Memory) or a PCRAM (Phase Change Random Access Memory). Furthermore, it will be appreciated by a person skilled in the art that the system 200 may be implemented by a high-performance computer known in the art for perfonning training, especially when a large-scale training is performed.
[0027] In various embodiments, a “circuit” may be understood as any kind of a logic implementing entity, which may be special purpose circuitry or a processor executing software stored in a memory, firmware, or any combination thereof. Thus, in an embodiment, a “circuit” may be a hard-wired logic circuit or a programmable logic circuit such as a programmable processor, e.g., a microprocessor (e.g., a Complex Instruction Set Computer (CISC) processor or a Reduced Instruction Set Computer (RISC) processor). A “circuit” may also be a processor executing software, e.g., any kind of computer program, e.g., a computer program using a virtual machine code, e.g., Java. Any other kind of implementation of various functions or operations may also be understood as a “circuit” in accordance with various other embodiments. Similarly, a “module” may be a portion of a system according to various embodiments in the present invention and may encompass a “circuit” as above, or may be understood to be any kind of a logic-implementing entity therefrom.
[0028] Some portions of the present disclosure are explicitly or implicitly presented in terms of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and functional or symbolic representations are the means used by those skilled in the data processing arts to convey most effectively the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities, such as electrical, magnetic or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated.
[0029] The present specification also discloses a system (e.g., which may also be embodied as one or more devices or apparatuses), such as the system 200, for performing various operations, functions or steps of various methods described herein. Such a system may be specially constructed for the required purposes or may comprise a general purpose computer system selectively activated or reconfigured by a computer program stored in the computer system. In general, various algorithms that may be presented herein are not limited to being implemented or executed by any particular computer system. Alternatively, the construction of more specialized computer system to perform various operations, functions or steps of various methods described herein may be provided as desired or as appropriate without going beyond the scope of the present invention.
[0030] In addition, the present specification also at least implicitly discloses computer program(s) or software / functional module(s), in that it would be apparent to a person skilled inthe art that various operations, functions or steps of various methods described herein may be put into effect by computer code. The computer program(s) is not intended to be limited to any particular programming language and implementation thereof, and it will be appreciated by a person skilled in the art that a variety of programming languages and coding thereof may be used to implement the computer program(s). Moreover, the computer program(s) is not intended to be limited to any particular control flow as there arc a variety of programming languages which can use different control flows. It will be appreciated by a person skilled in the art that a computer program may be stored on any computer-readable storage medium (non- transitory computer-readable storage medium), such as but not limited to, a magnetic disk, an optical disk or a memory chip. For example, a computer program stored on a computer-readable storage medium may be loaded and executed on a computer system to implement various operations, functions or steps of various methods described herein according to various embodiments of the present invention.
[0031] Accordingly, in various embodiments, there is provided a computer program product, embodied in one or more computer-readable storage mediums (non-transitory computer-readable storage medium), comprising instructions (e.g., the neuroimaging data encoder training module 206 and the diffusion model training module 208) executable by one or more computer processors to perform the method 100 of training a machine learning model for reconstructing video or audio data as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention. Accordingly, various computer programs or software modules described herein may be stored in a computer program product receivable by a system therein, such as the system 200 as shown in FIG. 2, for execution by at least one processor 204 of the system 200 to perform various operations, functions or steps of various methods described herein according to various embodiments of the present invention.
[0032] It will be appreciated by a person skilled in the art that various modules described herein (e.g., the neuroimaging data encoder training module 206 and the diffusion model training module 208) may be software module(s) realized by computer program(s) or set(s) of instructions executable by a computer processor to perform various functions or operations. Various modules described herein (e.g., the neuroimaging data encoder training module 206 and the diffusion model training module 208) may also be implemented as hardware module(s) being functional hardware unit(s) designed to perform various functions or operations. More particularly, in the hardware sense, a module is a functional hardware unit designed for use with other components or modules. For example, a module may beimplemented using discrete electronic components, or it can form a portion of an entire electronic circuit such as an Application Specific Integrated Circuit (ASIC). Numerous other possibilities exist. It will also be appreciated by a person skilled in the art that a combination of hardware and software modules may be implemented. Furthermore, various operations, functions or steps of various methods described herein may be performed in parallel rather than sequentially as desired or as appropriate (c.g., as long as it docs not render the mcthod(s) inoperable or unsatisfactory for its intended purpose).
[0033] In various embodiments, the system 200 for training a machine learning model for reconstructing video or audio data may be realized by any computer system (e.g., desktop or portable computer system) including at least one processor and at least one memory, such as an example computer system 300 as schematically shown in FIG. 3A as an example only and without limitation. Various methods / steps or functional modules may be implemented as software, such as a computer program being executed within the computer system 300, and instructing the computer system 300 (in particular, one or more processors therein) to conduct various functions or operations as described herein according to various embodiments. For example, the computer system 300 may comprise a system unit 302, one or more input devices 304 such as a keyboard, a touchscreen and / or a mouse, and a plurality of output devices such as a display 308. The system unit 302 may be connected to a computer network 312 via a suitable transceiver device 314, to enable access to c.g., the Internet or other network systems such as Local Area Network (LAN) or Wide Area Network (WAN). The system unit 302 may include a processor 31 for executing various instructions, a Random Access Memory (RAM) 320 and a Read Only Memory (ROM) 322. The system unit 302 may further include a number of Input / Output (I / O) interfaces, for example I / O interface 324 to the display device 308 and I / O interface 326 to the one or more input devices 304. The components of the system unit 302 typically communicate via an interconnected bus 328 and in a manner known to a person skilled in the art.
[0034] As mentioned hereinbefore, it will be appreciated by a person skilled in the art that the system 200 may be implemented by a high-performance computer known in the art for performing training, especially when a large-scale training is performed. In this regard, FIG. 3B depicts a schematic drawing of an example computer system 350 configured for high- performance computing, by way of an example only and without limitation. As shown, the example computer system 350 comprises a computer cluster that includes multiple GPUs and CPUs (i.e., processors) housed within a single unit, optimized for advanced computing tasks.Specifically, the example computer system 350 may comprise an array of GPUs (Graphics Processing Units) dedicated to handling parallel processing tasks to facilitate the rapid execution of machine learning algorithms. Additionally, the example computer system 350 may include powerful multi-corc CPUs (Central Processing Units) for managing complex computations and coordinating overall system operations. The architecture may be strategically designed with segregated processing units for GPUs and CPUs on the motherboard to facilitate efficient data handling and processing. Each processing unit is equipped with its own dedicated memory modules, further enhancing the system's capability to perform multiple tasks simultaneously without bottlenecking. The GPUs are directly connected through high-speed channels, enabling swift data transfer and minimizing latency to facilitate processing large datasets typically used in machine learning and neural network training.
[0035] It will be appreciated by a person skilled in the art that the terminology used herein is for the purpose of describing various embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0036] Any reference to an element or a feature herein using a designation such as “first”, “second” and so forth does not limit the quantity or order of such elements or features, unless stated or the context requires otherwise. For example, such designations may be used herein as a convenient way of distinguishing between two or more elements or instances of an element. Thus, a reference to first and second elements does not necessarily mean that only two elements can be employed, or that the first element must precede the second element, unless stated or the context requires otherwise. In addition, a phrase referring to “at least one of’ a list of items refers to any single item therein or any combination of two or more items therein.
[0037] In order that the present invention may be readily understood and put into practical effect, various example embodiments of the present invention will be described hereinafter by way of examples only and not limitations. It will be appreciated by a person skilled in the art that the present invention may, however, be embodied in various different forms or configurations and should not be construed as limited to the example embodiments set forth hereinafter. Rather, these example embodiments are provided so that this disclosure will bethorough and complete, and will fully convey the scope of the present invention to those skilled in the art.Machine Learning Model for Reconstructing Video Data
[0038] Various example embodiments provide high-quality video reconstruction from brain activity. Reconstructing human vision from brain activities has been an appealing task that helps to understand our cognitive process. Even though recent research has seen great success in reconstructing static images from non-invasive brain recordings, work on recovering continuous visual experiences in the form of videos is limited. Accordingly, various example embodiments provide a machine learning model for reconstructing video (which may herein be referred to as Mind-Video) that leams spatiotemporal information from continuous neuroimaging data (e.g., fMRI data) of the cerebral cortex progressively through masked brain modeling, multimodal contrastive learning with spatiotemporal attention, and co-training with an augmented Stable Diffusion model that incorporates network temporal inflation. It will be demonstrated later below that high-quality videos of arbitrary frame rates can be reconstructed with Mind-Video using adversarial guidance. The recovered videos were evaluated with various semantic and pixel-level metrics. For example, Mind-Video achieved an average accuracy of 85% in semantic classification tasks and 0.19 in structural similarity index (SSTM), outperforming the previous state-of-the-art by 45%. It will also be shown that Mind-Video is biologically plausible and interpretable, reflecting established physiological processes.
[0039] As example illustrations, FIG. 4A shows an overview of brain decoding and video reconstruction and FIG. 4B shows an example visual stimuli and the reconstructed video by Mind- Video according to various example embodiments of the present invention. In this regard, various example embodiments provide a progressive learning approach to recover continuous visual experience from fMRI data. High-quality videos with accurate semantics (e.g., including motions) are reconstructed.
[0040] In various example embodiments, Mind-Video comprises a two-module pipeline (namely, a first module being a fMRI encoder and a second module being a diffusion model) designed to bridge the gap between image and video brain decoding. The model progressively leams from brain signals, gaining a deeper understanding of the semantic space through multiple stages in the first module. Initially, the fMRI encoder leverages large-scale unsupcrviscd learning with masked brain modeling to learn general visual fMRI features. The fMRI encoder then distills semantic-related features using the multimodality of an annotatedembedding dataset (e.g., sets of embeddings of different modalities ( {fMRI embeddings, image embeddings, text (video caption) embeddings} triplets) and is trained in the Contrastive Language-Image Pre-Training (CLIP) space with contrastive learning. Tn the second module, the learned fMRI encoder features arc fine-tuned through co-training with an augmented stable diffusion model, which is specifically tailored for video generation under fMRI guidance from the fMRI encoder.
[0041] Accordingly, various example embodiments introduce a flexible and adaptable brain decoding pipeline decoupled into two modules, namely, an fMRI encoder and an augmented stable diffusion model, trained separately and then fine-tuned together. For the fMRI encoder, a progressive learning scheme is designed where the fMRI encoder learns brain features through multiple stages, including multimodal contrastive learning with spatiotemporal attention for windowed fMRI data. The stable diffusion model may be augmented for scene-dynamic video generation with near-frame attention. Adversarial guidance for distinguishable fMRI conditioning is also provided. In experiments conducted, Mind-Video was able to recover high- quality videos with accurate semantics, e.g., motions and scene dynamics. The results were also evaluated with semantic and pixel metrics at video and frame levels. For example, an accuracy of 85% was achieved in semantic metrics and 0.19 in SS1M, outperforming the previous state- of-the-art approaches by 45%. Furthermore, the attention analysis revealed mapping to the visual cortex and higher cognitive networks suggesting that Mind-Video is biologically plausible and interpretable.
[0042] Conventional methods formulated the video reconstruction as multiple image reconstructions, leading to low frame rates and frame inconsistency. Nonetheless, it has been shown that low-level image features and classes can be decoded from fMRI collected with video stimulus. Using fMRI representations encoded with a linear layer as conditions, higher quality and frame rate videos with a conditional video GAN were generated. However, the results are limited by data scarcity, especially for GAN training, which generally requires a large amount of data. There has also been disclosed a method of video reconstruction from brain activity which relied on a separable autoencoder that enables self-supervised learning in fMRI. Even though better results were achieved, the generated videos were of low visual fidelity and semantic meanings.
[0043] FIG. 5 depicts an overview of an example Mind-Video 500 according to various example embodiments of the present invention. Aiming for a flexible design, Mind-Video 500 is decoupled into two modules, namely, an fMRI encoder 506 and a video generative model(diffusion model) 508. In various example embodiment, these two modules 506, 508 are trained separately and then fme-tuned together, which allows for easy adaption of new models if better architectures of either one are available. As a representation learning model, the fMRI encoder 506 (c.g., corresponding to the ncuroimaging data encoder described hereinbefore according to various embodiments) in the first module transforms the pre-processed fMRI into fMRI embeddings (or fMRI representations, which arc encoded representations that summarize the neural activity associated with the stimuli) (e.g., corresponding to the neuroimaging data embeddings described hereinbefore according to various embodiments), which are used as a condition on the video generative model 508 for video generations. For this purpose, the fMRI embeddings preferably have the following traits: 1) contain rich and compact information about the visual stimulus presented during the scan; and 2) close to the embedding domain that the generative model 508 is trained with. In various example embodiments, the video generative model 508 is designed to produce not only diverse, high-quality videos with high computational efficiency but also handles potential scene transitions, mirroring the dynamic visual stimuli experienced during scans.
[0044] Tn various example embodiments, as shown in FIG. 5, the fMRI encoder 506 is configured to progressively learn fMRI features through multiple stages, including SC-MBM (scene-dynamic masked brain modeling) pre-training of the SC-MBM encoder 512 and multimodal contrastive learning. A spatiotemporal attention 518 may be designed to process multiple fMRI frames in a sliding window. In various example embodiments, the augmented stable diffusion model 508 is trained with videos and then tuned with the fMRI encoder 506 using the sets of embeddings of different modalities ({fMRI embeddings, image embeddings, text (video caption) embeddings} triplets) generated by the fMRI encoder 506. The example Mind-Video 500 will now be described in further detail below according to various example embodiments of the present invention. flvlRI Pre-processing
[0045] The fMRI captures whole-brain activity with BOLD signals (voxels). Each voxel is assigned to a region of interest (ROI) for focused analysis. Various example embodiments concentrate on voxels activated during visual stimuli. There are two ways to define the ROIs: one uses a pre-defmed parcellation to obtain the visual cortex; and the other relies on statistical tests to identify activated voxels during stimuli. In various example embodiments, the large- scale pre-training may be based on a parcellation such as that described in Glasser et al., “Amulti-modal parcellation of human cerebral cortex”, Nature, vol. 536, no. 7615, pp. 171-178, 2016 (herein referred to as the Glasser reference), while statistical tests may be employed for the target dataset (e.g., the dataset containing fMRI-video pairs disclosed in Wen et al. , “Neural encoding and decoding with deep learning for dynamic natural vision”, Cerebral cortex, vol. 28, no. 12, pp. 4136-4160, 2018 (which may herein be referred to as the Wen dataset)). To determine activated regions, various example embodiments calculate intra-subjcct reproducibility of each voxel, correlating fMRI data across multiple viewings. The correlation coefficients may then be converted to z-scores and averaged. The statistical significance may be computed using a one-sample t-test (PO OL DOF=17, Bonferroni correction). For example, the top 50% of the most significant voxels may be selected after the statistical test. As shown in FIG. 11 A, most of the identified voxels are from the visual cortex. Accordingly, activated brain regions of the fMRI data from the fMRI training dataset are determined and the fMRI encoder 506 is trained based on the activated brain regions of the fMRI data from the fMRI training dataset. fMRI Encoder 506
[0046] Progressive learning is used as an efficient training scheme where general knowledge is learned first, and then more task-specific knowledge is distilled through finetuning. To generate meaningful embeddings specific for visual decoding, various example embodiments design a progressive learning pipeline, which leams fMRI features in multiple stages, starting from general features (by unsupcrviscd (e.g., the self-supervised) learning with masked data modeling (or more specifically, masked brain modeling)) to more specific and semantic-related features (by the multimodal contrastive learning). It will be shown that the progressive learning process is reflected biologically in the evolution of fMRI attention maps.
[0047] In various example embodiments, the fMRI encoder 506 comprises a MBM encoder 512 configured to perform a large-scale pre-training with masked brain modelling (MBM) to learn general features of the visual cortex. For the MBM encoder 512, an asymmetric visiontransformer-based autoencoder (compnsing an encoder (as the MBM encoder 512) and a decoder) may be trained on a fMRI dataset (e.g., the pre-training dataset presented in the Human Connectome Project (Van Essen etal., “The wu-minn human connectome project: anoverview”, Neuroimage, vol. 80, pp. 62-79, 2013, herein referred to as the Van Essen reference)) with the visual cortex (VI to V4) as defined in the Glasser reference. More specifically, fMRI data of the visual cortex is rearranged from 3D into ID space in tire order of visual processing hierarchy.which is then divided into patches of the same size. The patches may then be transfonned into fMRI embeddings (which may also be referred to as tokens), and a large portion (e.g., about 75%) of the fMRI embeddings is randomly masked in the fMRI encoder 506 during training. With the autocncodcr architecture, a simple decoder aims to recover the masked fMRI embeddings based on the unmasked fMRI embeddings generated by the MBM encoder 512. The main idea behind the masked brain modelling is that if the training objective can be achieved with high accuracy using a simple decoder, the fMRI embeddings generated by the MBM encoder 512 will be a rich and compact description of the original fMRI data. For example, it is interesting to note that the masked brain modelling approach leverages a training objective that, when met with high accuracy, indicates that the MBM encoder 512 has successfully captured the intricate details necessary for a rich representation of the fMRI data. Accordingly, the fMRI encoder 506 is trained based on fMRI data from an fMRI training dataset for generating fMRI embeddings. In particular, the fMRI encoder 506 comprises a masked autoencoder and the encoder of the masked autoencoder is trained based on the fMRI training dataset using unsupervised learning with masked data modeling for generating the fMRI embeddings. Tn this regard, the unsupervised learning with masked data modeling comprises generating fMRI embeddings from fMRI data from the fMRI training dataset, masking a portion of the fMRI embeddings into masked fMRI embeddings and training the masked autoencoder to recover the masked fMRI embeddings.
[0048] Spatiotemporal attention for windowed fMRI will now be described according to various example embodiments of the present invention. Forthe purpose of image reconstruction, a fMRI encoder may shift the fMRI data by 6 seconds, which may then be averaged every 9 seconds and processed individually. However, this process only considers the spatial information in its attention layers. In the video reconstruction, if one fMRI is directly mapped to the video frames presented (e.g., 6 frames), the video reconstruction task may be formulated as a one-to-one decoding task, where each set of fMRI data corresponds to 6 frames. Each {fMRI-frames} pair may be referred to as a fMRI frame window. However, this direct mapping is sub-optimal because of the time delay between brain activity and the associated BOLD signals in fMRI data due to the nature of hemodynamic response. Thus, when a visual stimulus (i ,e . , a video frame) is presented at time t, the fMRI data obtained at t may not contain complete information about this frame. Namely, a lag occurs between the presented visual stimulus and the underlying information recorded by the fMRI data and this phenomenon is depicted in FIG. 6. In particular, FIG. 6 illustrates that due to hemodynamic response, the BOLD signal lags afew seconds behind the visual stimulus, causing a discrepancy between the fMRI data and the corresponding stimulus.
[0049] Tire hemodynamic response function (HRF) may be used to model the relationship between neural activity and BOLD signals, such as described in Lindquist et al., “Modeling the hemodynamic response function in fMRI: Efficiency, bias and mis-modeling”. Neuroimage, vol. 45, no. 1, pp. S187-S198, 2009. In an LTI system, the signal y(t) may be represented as the convolution of a stimulus function s(t) and the hemodynamic response (HR) h(t), i.c., y (t) = (s * h) (t) . The h(t) may be modeled with a linear combination of some basis functions, which can be collated into a matrix form where Y represents the observed data,ft is a vector of regression coefficients, and e is a vector of unexplained error values. However, e varies significantly across individuals and sessions due to age, cognitive state, and specific visual stimuli, which influence the firing rate, onset latency, and neuronal activity duration. These variations impact the estimation of e and ultimately affect the accuracy of the fMRIbased analysis. There are two ways to address individual variations: using personalized HRF models or developing algorithms that adapt to each participant. In various example embodiments, the latter is selected due to its superior flexibility' and robustness.
[0050] Aiming to obtain sufficient information to decode each scan window and account for the hemodynamic response, various example embodiments provide a spatiotemporal attention layer to process multiple fMRI frames in a sliding window. Consider a sliding window defined as {xt, xt+1, •■■ , xt+lv-1}, where xt6 F / iXfJxfldenote fMRI embeddings (which may also be referred to as token embeddings of the fMRI) at t and n, p, b denote the batch size, patch size, and embedding dimension, respectively. Therefore, xtG ]R>nXM / xpxi) , where w is the window size. Note that the attention is given by attn=softmax
[0051] To calculate spatial attention, the network inflation technique maybe employed (e.g., as described in Wu et al., “Tune-a-video: One-shot tuning of image diffusion models for text- to-video generation’’, arXiv preprint arXiv:2212.11565, 2022 (herein referred to as the Wu reference)), where the first two dimensions of xtare merged to obtainG ]]^nH,><Pxi\ Then the query' and key may be calculated using Equation (1) as follows:(Equation 1)
[0052] Likewise, the first and the third dimension of xtmay be merged to calculate the temporal attention, obtaininge R"PxvvxiAgain, the query and key may be calculated using Equation (2) as follows:(Equation 2)
[0053] The spatial attention learns correlations among the fMRI tokens, describing the spatial correlations of the fMRI patches. Then the temporal attention leams the correlations of fMRI from the sliding window, including sufficient information to cover the lag due to the hemodynamic response, as illustrated under fMRI Attention 518 in FIG . 5. In various example embodiments, regarding the spatial attention mechanism, the MBM encoder 512 first generates fMRI embeddings from segmented fMRI data as described hereinbefore. These fMRI embeddings represent localized brain activity within each patch. Spatial attention is applied to these fMRI embeddings to model the interdependencies between different brain regions captured in the fMRI embeddings. For example, this may be achieved by computing attention scores that measure the importance of each fMRI embedding's features in relation to all other fMRI embeddings in the sequence. Therefore, the spatial attention mechanism weights these fMRI embeddings based on their spatial relevance to each other, enhancing the model's ability to capture the spatial organization of brain activity. The output may thus be a set of spatially attended fMRI embeddings, where each fMRI embedding is adjusted to reflect the contextual significance of surrounding brain regions, emphasizing features that are spatially pertinent. In various example embodiments, regarding the temporal attention mechanism, a sliding window approach is utilized to address the dynamic nature of brain activity over time. The temporal attention mechanism considers the sequence of fMRI embeddings (more specifically, the spatially attended fMRI embeddings obtained from the spatial attention mechanism) over a defined time frame, allowing the model to incorporate information about changes in brain activity. For example, the attention scores for the temporal attention are calculated to highlight how brain regions' activities evolve over time, accounting for the inherent delay introduced by the hemodynamic response. The output may thus be a set of temporally attended fMRI embeddings, which incorporate an understanding of how brain activity patterns develop and change over the course of the fMRI scan For example, these temporally attended fMRI embeddings arc better suited for tasks that require knowledge of temporal dynamics, such as predicting brain responses to ongoing / continuous stimuli.
[0054] Multimodal contrastive learning will now be described according to various example embodiments of the present invention. As described hereinbefore, the fMRI encoder 506 is pre-trained to learn general features of the visual cortex, and then it is augmented with temporal attention heads to process a sliding window of fMRI embeddings. In this step, the augmented fMRI encoder 506 is further trained with the sets of embeddings of different modalities ({fMRI embeddings, image embeddings, text (video caption) embeddings} triplets) to pull the fMRI embeddings closer to a shared CLIP (Contrastive Language -Image PreTraining) space containing rich semantic information. In this regard, the fMRI data and video data are collected in synchronized sessions where the subject is exposed to visual stimuli (video data) while their brain activity is simultaneously recorded (fMRI data). Therefore, each image embedding corresponds to the corresponding video frame viewed by the subject during the fMRI scan, and concurrently, the text for each video frame is generated by a pre-trained image captioning model (e.g., the BLIP model) as a caption that describes a content of the video frame at that particular moment in time, thereby ensuring that the fMRI, image, and text embeddings in a set are inherently aligned and belong to a cohesive dataset. Accordingly, for each of a plurality of windows of neuroimaging data embeddings generated by the encoder of the masked autocncodcr, the encoder of the masked autocncodcr is trained based on a set of embeddings (of different modalities) comprising the window of fMRI embeddings, corresponding image embeddings and corresponding text embeddings projected into a shared latent space based on contrastive learning.
[0055] CLIP is a pre-training technique that builds a shared latent space for images and natural languages by large-scale contrastive learning. The training aims to minimize the cosine distance of paired image and text latent while maximizing permutations of pairs within a batch. The shared latent space (CLIP space) contains rich semantic information on both images and texts. Accordingly, the CLIP space, with rich semantic content, is engineered to facilitate a common ground for images and natural languages through large-scale contrastive learning. The principal aim during this training phase is to minimize the cosine distances between paired image and text embeddings, ensunng that these modalities are not only aligned but also semantically coherent across the batch. Simultaneously, the training process endeavors to maximize the distance between non-pairing embeddings, thus refining the model's ability' to distinguish between unrelated content. The utility of this approach lies in its capacity to create a robust, shared latent space that encapsulates comprehensive semantic information pertinent to both images and texts. Accordingly, this latent space enables the fMRI embeddings,generated by the fMRI encoder 512, to be seamlessly integrated with corresponding image and text embeddings produced by the respective CLIP-based image and text encoders 524, 522. Each encoder, whether processing textual descriptions through natural language techniques or visual inputs through vision-transformer methodologies, projects its output embeddings into this shared latent space, ensuring that each set of embeddings of different modalities (fMRI embeddings, image embeddings, text (video caption) embeddings) is not only internally consistent but also harmoniously aligned with embeddings from other modalities. Additionally, the generative model 508 may be pre-trained with text conditioning. Thus, pulling the fMRI embeddings closer to the text-image shared space which facilitates their understandability by the generative model 508 during conditioning.
[0056] To prepare the training set for the multimodal contrastive learning model, as described hereinbefore, each set of embeddings of different modalities may be a {fMRI embeddings, image embeddings, text (video caption) embeddings} triplet. The video frames may be originally shot at a high frame rate of 30 frames per second (FPS). To streamline the processing and adapt the video data for effective integration with fMRI and textual data, these video clips may be first segmented into a predefined (e.g., 2-second) intervals. Following this segmentation, the videos may be downsampled from their original frame rate to 3 FPS. Each video frame may then be captioned using the BLIP model (such as described in Li et al.. “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation”, in International Conference on Machine Learning, PMLR, 2022, pp. 12888 - 12900). The BLIP-generated captions help to bridge the gap between the visual content represented by the video frames and the complex neural activity patterns depicted in the fMRI data. Generated captions in most scanning windows are similar, except when scene changes occur in the window. In that case, for example, two different captions may be concatenated with the conjunction “Then”. In various example embodiments, the training process in the multimodal contrastive learning involves the use of the sets of {fMRI embeddings, image embeddings, text (video caption) embeddings} triplets to teach the MBM encoder 512 how to align and interpret multimodal data effectively. As described hereinbefore, the image embedding corresponds to the corresponding video frame that the subject views during the fMRI scan. Concurrently, the text (from which the text embedding is generated by the text encoder 522) for each video frame is generated as a caption that describes a content of the video frame at that particular moment in time. By integrating fMRI embeddings with corresponding image embeddings and text (video caption) embeddings, the MBM encoder 512 learns tounderstand and predict tire underlying neural mechanisms that correspond to both observed actions and described scenarios. Accordingly, the training of the MBM encoder 512 within this multimodal framework involves aligning fMRI embeddings with corresponding video and text embeddings in a shared semantic space facilitated by the CLIP model. As shown in FIG. 5, the image and text embeddings are generated by the CLIP-based image and text encoders 524, 522, which project the image and textual embeddings into a latent space optimized for cross-modal comparability.
[0057] With the CLIP-based text encoder 522 and image encoder 524 being fixed, the CLIP loss may be calculated for fMRI -image and fMRI-text, respectively. Denote the pooled text embedding, image embedding, and fMRI embedding by emht. embt , embf G R’1 X / 1. The contrastive language-image-fMRI loss may be given by Equation (3) as follows:(Equation 3) where _£CLIP(a, b) = CrossEntropy(ea ■ bT, [0,1, ••• , n]). with e being a scaling factor. Extra care may be needed to reduce similar frames in a batch for better contrastive pairs. From Equation (3), it can be seen that the loss largely depends on the batch size n. Thus, according to various example embodiments, a large n with data augmentation on all modalities may be preferred.Video Generative Module 508
[0058] Diffusion models are emerging probabilistic generative models defined by a reversible Markov chain. As a variant, stable diffusion generates a compressed version of the data (data latent) instead of generating the data directly. As it works in the data latent space, the computational requirement is significantly reduced, and higher-quality images with more details can be generated in the latent space. In various example embodiments, the stable diffusion model (e.g., described in Rombach, “High-resolution image synthesis with latent diffusion models”, in Proc. CVPR’22, 2022, pp. 10684-10695 (herein referred to as the Rombach reference)) is used as the base generative model considering its excellence in generation quality, computational requirements, and weights availability. However, as the stable diffusion model is an image-generative model, according to various example embodiments, temporal constraints are applied in order for video generation.
[0059] Scene-dynamic sparse causal (SC) attention will now be described according to various example embodiments of the present invention. The Wu reference mentionedhereinbefore uses a network inflation technique with sparse temporal attention to adapt the stable diffusion to a video generative model. Specifically, the sparse temporal attention effectively conditions each frame on its previous frame and the first frame, which ensures frame consistency and also keeps the scene unchanged. However, the human vision includes possible scene changes, so the video generation should also be scene-dynamic. To address this technical problem, the temporal-attention constraint in the Wu reference is relaxed and each frame is conditioned on its previous two frames, ensuring the video smoothness while allowing scene dynamics. Using notations from the Wu reference, the SC attention may be calculated with the query, key, and value given by:(Equation 4) where zrdenotes the latent of the / -th frame during the generation.
[0060] Adversarial guidance using fMRI data will now be described according to various example embodiments of the present invention. Classifier-free guidance is used in the conditional sampling of diffusion models for its flexibility and generation diversity, where the noise update function may be given by:(Equation 5) where c is the condition, s- is the guidance scale and eg(■) is a score estimator implemented with UNet, and ztis the generated latent at time step t. According to various example embodiments, Equation (5) may be changed to an adversarial guidance version as follows:(Equation 6) where c is the negative guidance. In effect, generated contents can be controlled through “what to generate” (positive guidance) and “what not to generate” (negative guidance). When c is a null condition, the noise update function falls back to Equation (5), the regular classifier-free guidance. Accordingly, the diffusion model is trained based on fMRI embeddings generated by the fMRI encoder as conditions on the diffusion model. Tn various example embodiments, in order to generate diverse videos for different fMRI, guaranteeing the distinguishability of the inputs, all the fMRI data in the testing set may be averaged and the averaged fMRI data may be used as the negative condition. Specifically, for each fMRI input, the fMRI encoder 506 may be configured to generate an unpooled embedding x G IPi,xh. where I is the latent channelnumber. Denote the averaged fMRI data as x, the noise update function can be obtained as ee(zt, x, x) = e9(zt, x) + s(ee(zt, x) - eg(zt, x)) .
[0061] As described hereinbefore, in various example embodiments, Mind-Video 500 has a decoupled structure whereby the fMRI encoder 506 and the video generative model 508 are trained separately in the first phase. In this regard, the fMRI encoder 506 is trained on a large- scale dataset and then tuned in a target dataset with contrastive learning, and separately, the video generative model 508 is trained with videos from the target dataset using text conditioning. In the second phase, the fMRI encoder 506 and the video generative model 508 are tuned together with fMRI-video pairs (the above-mentioned training set for the multimodal contrastive learning model), where the fMRI encoder 506 and the attention heads of the generative model 508 are trained. For example, in contrast to the above-mentioned Wu reference, various example embodiments time the whole self-attention, cross-attention, and temporal-attention heads instead of only the query projectors, as a different modality is used for conditioning. The second phase is also the last stage of encoder progressive learning, after which the fMRI encoder 506 finally generates fMRI embeddings that contain rich semantic information and are easy to understand by the generative model 508.
[0062] Accordingly, from the combined effect of the two stages, namely, the fMRI encoder 506 and the stable diffusion model 508, Mind-Video 500 is able to reconstruct video data based on the fMRI data of a subject effectively with accurate semantics. As described hereinbefore, in the first stage, the fMRI encoder 506 plays a crucial role in learning representations (fMRI embeddings) from the brain. It effectively captures the complex spatiotemporal information embedded in the fMRI data, allowing the Mind-Video 500 to understand and interpret the underlying neural activities. Accordingly, as the fMRI encoder 506 is designed to leam intricate representations from brain activity, these representations go beyond simple categorical information to encompass more nuanced semantic details that could not be adequately captured by discrete class labels (e.g. image texture, depth, etc.). Due to the richness and diversity of human thought and perception, a model that can handle continuous semantics, rather than discrete ones, is advantageous. As an illustrative comparison, a fMRI-to-object classifier suffers from a crucial trade-off between classification complexity and solution space as follows:• Classification Complexity: classifying fMRI data into a large number of classes (e.g., 1000 classes) is non-trivial. It has been reported that reasonable performance can only be achieved in a smaller classification task (less than 50-way), due to the limited data per category and the complexity of the task.• Limited Solution Space: the solution space of discrete classes is significantly more restricted than that of continuous semantics. Thus, a classifier may not capture the complex, multi -faceted nature of brain activities .
[0063] Accordingly, the above trade-off between classification complexity and solution space illustrates why a fMRI-to-objcct classifier may not be suitable. In contrast, the fMRI encoder 506 is configured to learn continuous semantic representations from the brain, which better reflects the complexity and diversity of neural processes. This approach not only improves the quality of the generated videos but also provides more meaningful and interpretable insights into brain decoding.
[0064] The fMRI-to-video generation task performed by Mind-Video 500 involves mapping brain activity to dynamic videos. This task is considerably more complex compared to a fMRI-to-image generation task as it requires the model to capture both spatial information and temporal dynamics. This does not only involve predicting which brain regions are active, but also involves understanding how these activations change over time and how they relate to moving elements in a video. Adding to the complexity is the hemodynamic response inherent in fMRI data, which introduces a delay and blur in the timing of neural activity. Furthermore, the temporal resolution of fMRI is quite low, making it challenging to capture fast-paced changes in neural activity.
[0065] Tn the second stage, the stable diffusion model 508 steps in to generate videos. For example, one of the key advantages of the stable diffusion model 508 over other generative models, such as GANs, lies in its ability to produce higher-quality' videos. It leverages the representations learned by the fMRI encoder 506 and utilizes its unique diffusion process to generate videos that are not only of superior quality but also better align with the original neural activities. Furthermore, since the stable diffusion process is a probabilistic model, this means that the generation process involves a degree of randomness, leading to slight differences during each generation for video frames. Accordingly, in video generation, various example embodiments seek to ensure consistency across video frames.
[0066] Various example embodiments extend beyond brain decoding and reconstruction. In this regard, various example embodiments seek to understand the decoding process’s biological principles. To this end, various example embodiments visualize average attention maps from the first, middle and last layers of the fMRI encoder 506 across all testing samples. This approach allows us to observe the transition from capturing local relations in early layers to recognizing more global, abstract features in deeper layers. Additionally, attention maps arevisualized for different learning stages: large-scale pre-training, contrastive learning, and cotraining. By projecting attention back to brain surface maps, each brain region’s contributions and the learning progress through each stage can be observed.Experiments
[0067] Various experiments conducted according to various example embodiments of the present invention will now be described.
[0068] For the pre-training dataset, the Human Connectome Project (HCP) 1200 Subject Release presented in the Van Essen reference mentioned hereinbefore was used. In this regard, for the upstream pre-training dataset, resting-state and task-evoked fMRI data from the HCP were employed. Furthermore, 600,000 fMRI segments were obtained from a substantial amount of fMRI scan data.
[0069] For the paired fMRI-Video dataset, a publicly available benchmark fMRI-video dataset, namely, the Wen reference mentioned hereinbefore, was used, comprising fMRI and video clips. The fMRI were collected using a 3T MRI scanner at a repetition time (TR) of 2 seconds with three subjects. The training data included 18 segments of 8-minute video clips, totaling 2.4 video hours and yielding 4,320 paired training examples. The test data comprised 5 segments of 8-minute video clips, resulting in 40 minutes of test video and 1 ,200 test fMRIs. The video stimuli were diverse, covering animals, humans, and natural scenery, and featured varying lengths at a temporal resolution of 30 FPS.
[0070] With respect to the implementation details, the original videos were downsampled from 30 FPS to 3 FPS for efficient training and testing, leading to 6 frames per fMRI frame. In the implementation, a video of 2 seconds (6 frames) was reconstructed from one fMRI frame. However, thanks to the spatiotemporal attention head design that encodes multiple fMRI frames at once, Mind-Video is able to reconstruct longer videos from multiple fMRI frames if more GPU memory' is available.
[0071] A ViT-based fMRI encoder with a patch size of 16, a depth of 24, and an embedding dimension of 1024 was used. After pre-training with a mask ratio of 0.75, the fMRI encoder was augmented with a projection head that projects the token embedding into the dimension of 77 x 768. The Stable Diffusion VI -5 (as described in the Rombach reference mentioned hereinbefore) trained at the resolution of 512 x 512 was used. The augmented Stable Diffusion was tuned at the resolution of 256 X 256 with 3 FPS Note that the FPS and image were downsampled for efficient experiments and that Mind-Video can also work with full temporaland spatial resolution. All parameters in the MBM pre-training were the same as that described in Chen et al. , “Seeing beyond the brain: Masked modelling conditioned diffusion model for human vision decoding”, in Proceedings ofthe TEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023 (herein referred to as the Chen reference) and trained with eight RTX3090, while other stages were trained with one RTX3090. The inference was performed with 200 DD1M steps.
[0072] With respect to evaluation metrics, following prior studies, both frame-based and video-based metrics are utilized. Frame-based metrics evaluate each frame individually, providing a snapshot evaluation, whereas video-based metrics assess sequences of frames, encapsulating the dynamics across frames. Both are used for a comprehensive analysis. Unless stated otherwise, all test set videos are used for evaluating the three subjects.
[0073] With respect to frame-based metrics, the frame-based metrics utilized are divided into two classes, pixel-level metrics and semantics-level metrics. The structural similarity index measure (SSIM) is used as the pixel-level metric and the N-way top-K accuracy classification test is used as the semantics-level metric. Specifically, for each frame in a scan window, the SSTM and classification test accuracy are calculated with respect to the groundtruth frame. To perform the classification test, the classification results of the groundtruth (GT) and the predicted frame (PF) are compared using an TmageNet classifier. If the GT class is within the top-K probability of the PF classification results from N randomly picked classes, including the GT class, a successful trial is declared. In this regard, since videos cannot be well described with a single ImageNet class, the top-K classification results are used as the GT class, and a successful event is declared if the test succeeds with any of the GT class. The test was repeated for 100 times, and the average success rate is reported.
[0074] With respect to video-based metric, the video-based metric measures the video semantics using the classification test as well, except that a video classifier is used. The video classifier based on VideoMAE (described in Tong et al., “Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training”, arXiv preprint arXiv: 2203.12602, 2022) was trained on Kinetics-400 (Kay et al. , “The kinetics human action video dataset”, arXiv preprint arXiv: 1705.06950, 2017), an annotated video dataset with 400 classes, including motions, human interactions, etc.
[0075] Various experimental results will now be discussed. Tn particular, the experimental results obtained for the Mind-Video according to various example embodiments arc compared against three fMRI-video baselines: the model of the Wen reference mentioned hereinbefore,the model of Wang et al., “Reconstructing rapid natural vision with fMRI -conditional video generative adversarial network”, Cerebral Cortex, vol. 32, no. 20, pp. 4502-4511, 2022 (herein referred to as the Wang reference) and the model of Kupershmidt et al., “A penny for your (visual) thoughts: Self-supervised reconstruction of natural movies from brain activity”, arXiv preprint arXiv:2206.03544, 2022 (herein referred to as the Kupershmidt reference).
[0076] Visual comparisons arc shown in FIG. 7, and quantitative comparisons arc shown in FIG. 9, where publicly available data and samples are used for comparison. In particular, FIG. 9 shows the SSIM comparison, whereby the left graph shows a comparison of Subject 1 and test video 1 and the right graph shows a comparison of all subjects and all test videos. As shown in FIG. 7, Mind-Video generated high-quality videos with more semantically meaningful content and match with the groundtruth. Following the literature, the SSIM of Subject 1 was evaluated with the first testing video, achieving a score of 0.186, outperforming the state-of-the-art by 45%. When comparing with the Kupershmidt model, all test videos for different subjects were evaluated, and Mind-Video outperforms by 35% on average, as shown in FIG. 9. Using semantic level metrics, Mind-Video achieves a success rate of 0.849 and 0.2 in 2-way and 50-way top-1 accuracy classification tests, respectively, with the video classifier. The image classifier gives a success rate of 0.795 and 0.19, respectively, with the same tests, which significantly surpasses the chance level of these two tests (2-way: 0.5, 50-way: 0.02). Full results arc shown in Table 1 shown in FIG. 10. In particular, Table 1 shows the results of the ablation study on window sizes, multimodal contrastive learning and adversarial guidance (AG). Evaluations on different subjects are also shown. For the full model, win=2 and Sub 1. In FIG. 10, the number of reflects the degree of statistical significance (two-sample t-test) compared to the full model, whereby more ‘*’ reflects higher degree of statistical significant and no ‘*’ indicating not significant, where p < 0.0001 (***); p < 0.01 (**): p < 0.05 (*); p > 0.05. The results from Mind-Video is also compared with an image-fMRI model by the Chen reference mentioned hereinbefore. An image is produced for each fMRI, samples of which are shown in FIG. 7 with the “walking man” as the groundtruth. Even though the results and groundtruth are semantically matching, frame consistency and image quality are not satisfying. A lag due to the hemodynamic response is also observed. The first frame actually corresponds to the previous groundtruth.
[0077] For different subjects, after fMRI pre-processing, each subject varies in the size of ROIs, where Subject 1, 2 and 3 have 6016, 6224, and 3744 voxels, respectively. As Subject 3 has only half the voxels of the others, a larger batch size can be used during contrastive learning,which may lead to better results as shown in Table 1 in FIG. 10. Nonetheless, all subjects are consistent in both numeric and visual evaluations.
[0078] For recovered semantics, in the video reconstruction, the semantics are defined as the objects, animals, persons, and scenes in the videos, as well as the motions and scene dynamics, e g., people running, fast-moving scenes, close-up scenes, long-shot scenes, etc. It is shown that even though the fMRI data has a low temporal resolution, it contains enough information to recover the mentioned semantics using Mind-Video. FIG. 8 shows a few examples of reconstructed frames using Mind-Video, including an example involving a scene transition. Firstly, it can be seen that the basic objects, animals, persons, and scene types can be well / correctly recovered. More importantly, the motions, such as running, dancing, and singing, and the scene dynamics, such as the close-up of a person, the fast-motion scenes, and the long- shot scene of a city view, can also be reconstructed correctly. This result is also reflected in the numerical metrics, which consider both frame semantics and video semantics, including various categories of motions and scenes.
[0079] Regarding ablations, Mind-Video was tested using different window sizes, starting from a window size of 1 up to 3. Table 1 in FIG. 10 shows that when all other parameters are fixed, a window size of 2 gives the best performance in general, which is reasonable as the hemodynamic response usually will not be longer than two scan windows. Additionally, the effectiveness of multimodal contrastive learning was also tested. As shown in Tabic 1 in FIG. 10, without contrastive learning, the generation quality degrades significantly. When two modalities are used, either text-fMRI or image-fMRI, the performance is inferior to the full modalities used in contrastive learning. Actually, the reconstructed videos are visually worse than the full model. Thus, it shows that the full progressive learning pipeline is significantly advantageous for the fMRI encoder to leam usefill representations for this task. The reconstruction results without adversarial guidance were also assessed. As a result, both numeric and visual evaluations decrease substantially. In fact, the generated videos can be highly similar sometimes, indicating that the negative guidance is important in increasing the distinguishability of fMRI embeddings.
[0080] To interpret the experimental results, the attention analysis is summarized FIG. 11. The sum of the normalized attention w ithin Yeo 17 networks are presented in the bar charts. Voxel-wise attention is displayed on abrain flat map, where comprehensive structural attention throughout the whole region can be seen. The average attention across all testing samples and attention heads is computed, revealing three insights on how transfonners decode fMRI data.
[0081] Regarding tire dominance of visual cortex, the visual cortex emerges as the most influential region. This region, encompassing both the central (VisCent) and peripheral visual (VisPeri) fields, consistently attracts the highest attention across different layers and training stages (shown in region B of FIG. 11). In all cases, the visual cortex is always the top predictor, which aligns with prior research, emphasizing the vital role of the visual cortex in processing visual spatiotemporal information. However, the visual cortex is not the sole determinant of vision. Higher cognitive networks, such as the dorsal attention network (DorsAttn) involved in voluntary visuospatial attention control, and the default mode network (Default) associated with thoughts and recollections, also contribute to visual perceptions process as shown in FIG. 11.
[0082] Regarding layer-dependent hierarchy, the layers of the fMRI encoder of Mind- Video function in a hierarchical manner, as shown in portions C, D and E of FIG. 11. In the early layers of the netw ork (portions A and C of FIG. 11), a focus on the structural information of the input data can be observed, marked by a clear segmentation of different brain regions by attention values, aligning with the visual processing hierarchy. As the network dives into deeper layers (portions D and E of FIG. 11), the learned information becomes more dispersed. The distinction between regions diminishes, indicating a shift toward learning more holistic and abstract visual features in deeper layers.
[0083] Regarding learning semantics progressively, to illustrate the learning progress of the fMRI encoder of Mind- Video, the first-layer attention after all learning stages was analyzed, as show n in portion B of FIG. 11 : before contrastive learning, after contrastive learning, and after co-traimng with the video generation model. An increase in attention in higher cognitive networks and a decrease in the visual cortex can be observed as learning progresses. This indicates that the fMRI encoder assimilates more semantic information as it evolves through each learning stage, improving the learning of cognitive-related features in the early layers.
[0084] For better understanding, further implementation details and further experiments performed will now be described according to various example embodiments of the present invention.
[0085] Regarding large-scale pre-training, the large-scale pre-training uses the same setup as the masked brain modeling (MBM) described in the above-mentioned Chen reference. A ViT-large-based model with a 1 -dimensional patchifier was trained with hyperparameters showm in Table 2 in FIG. 12. In particular, FIG. 12 shows the hyperparameters used in the large- scale pre-training. The training took around 3 days using 8RTX3090 GPUs. The training is performed on the 600,000 fMRI from HCP. Same as the literature, after the large-scale pre-training, the autoencoder is tuned with fMRI data from the target dataset, namely, that of the above-mentioned Wen reference, using MBM as well. The tuning was performed using a small learning rate and epochs.
[0086] Regarding multimodal contrastive learning, the pre-trained fMRI encoder was taken from the previous step and augment it with temporal attention heads to accommodate multiple fMRI frames. Then contrastive learning was performed with fMRl-imagc-tcxt triplets. The image was a randomly-picked frame from an fMRI scan window. As mentioned hereinbefore, there are two important factors in the contrastive: batch size and data augmentations. Therefore, data augmentations were applied for all modalities. Random sparsification was used for fMRI, where 20% of voxels are randomly set to zeros each time. The random crop was applied to videos with a probability of 0.5. To augment the frame captions, synonym augmentation and random word swapping were applied. Due to a small dataset size (about 4000), a dropout rate of 0.6 was used to avoid overfitting. For Subject 1 and 2, the training was performed with a batch size of 20, while a batch size of 32 was used due to fewer fMRI voxels with Subject 3. Training for all subjects was performed for 12,000 steps with a learning rate of 2 x 10-s. The training took around 10 hours using one RTX3090.
[0087] Regarding the training of the augmented stable diffusion model, the stable diffusion model was augmented with temporal attention heads for video generation. The augmented stable diffusion model was trained with videos from the target datasets. Hie videos were downsampled from 30 FPS to 3 FPS at a resolution of 256 x 256 due to limited GPU memory’, even though Mind-Video can work with the full time resolution. This step is important or advantageous for two reasons: 1) the augmented temporal heads are untrained; 2) the stable diffusion is pre-trained at a resolution of 512 x 512 , therefore, it is adapted to a lower resolution.
[0088] During the training, the self-attention heads (“attnl”), cross-attention heads (“attn2”), and temporal attention heads (“attn_temp”) were updated. Accordingly, the diffusion model compnscs a U-Nct comprising self-attention layers, cross-attention layers and temporalattention layers, and the diffusion model is further trained based on a target video dataset. The training was performed with text conditioning for 800 steps. A learning rate 2 x 10-5and a batch size of 14 were used. The training took around 2 hours using one RTX3090. Visual results show that videos of high quality can be generated with text conditioning after this step.
[0089] Regarding co-training of the fMRI encoder and the augmented stable diffusion model, the fMRI encoder produced embeddings of dimensions 77 x 768, which were used tocondition the augmented stable diffusion model during co-training. Hie whole fMRI encoder was updated, and only the attention heads of the stable diffusion model was updated (same as the last step). The training was performed with a batch size of 9 and a learning rate of 3 x 10 “5for 15,000 steps. The training took around 16 hours using one RTX3090.
[0090] Regarding inference, all samples were generated with 200 diffusion steps using fMRI adversarial guidance. Tire fMRI adversarial guidance used an average fMRI data as the negative guidance with a guidance scale of 12.5.
[0091] For the analysis of visual results, all three subjects in the dataset disclosed in the above-mentioned Wen reference were tested. Around 6000 voxels were identified as RO1 for Subject 1 and 2, while around 3000 voxels were identified for Subject 3. Thus, a larger batch size can be used when training with Subject 3, which may be the reason for its better numeric evaluation results. Nonetheless, all three subjects show consistent generation results and some samples are shown in FIG. 13.
[0092] Visual results of the ablation studies are shown in FIG 14. The Full model was trained with the full pipeline and inference with adversarial guidance according to various example embodiments of the present invention. In contrastive learning ablation, incomplete modality' was tested, namely, image-fMRT and text-fMRT, respectively. Similar to the numeric evaluations, using incomplete contrastive gave an unsatisfactory visual result compared to using all three modality (fMRI-image -text triplets). However, incomplete modality still outperformed inferencing without adversarial guidance significantly, w'hich generated visually meaningless results.
[0093] Some fail cases are shown in FIG. 15. It is observed that even though some fail cases generated different animals and objects compared to the groundtruth, other semantics like the motions, color, and scene dynamics can still be correctly reconstructed. For example, even though the airplane and flying bird are not reconstructed, similar fast-motion scenes are recovered in FIG. 15.
[0094] Accordingly, various example embodiments provide Mind-Video 500, which reconstructs high-quality videos with arbitrary frame rates from fMRI. Starting from a large- scale pre-training to multimodal contrastive learning with augmented spatiotemporal attention, the fMRI encoder 506 leams fMRI features progressively. Then, an augmented stable diffusion model 508 is fine-tuned for video generations, which is co-trained together with the fMRI encoder 506. In this regard, it has been shown that w ith fMRI adversarial guidance, Mind-Video 500 recovers videos with accurate semantics, motions, and scene dynamics compared with thegroundtruth, establishing a new state-of-the-art in this domain. It has also been shown through attention maps that the trained model decodes fMRI data with reliable biological principles. Accordingly, Mind-Video 500 has practical applications from neuroscience to brain -computer interfaces.Machine Learning Model for Reconstructing Audio Data
[0095] Various example embodiments relate to audio decoding from functional magnetic resonance imaging (fMRI) using conditional diffusion model.
[0096] Several attempts have been made for audio reconstruction from brain activity. However, it is still challenging for current generative models to navigate the complex relationship between specific brain regions and audio input for successful reconstruction and interpretation. To address this, various example embodiments provide a method of training a machine learning model for reconstructing audio data based on neuroimaging data of a subject, which may be herein referred to as Mind-Audio. Mind-Audio is a framework designed to not only recover perceived audio signals using fMRI but also enable detailed interpretation of brain regions responsible for auditory processing. To achieve this, in various example embodiments, Mind-Audio comprises self-supervised learning to effectively encode the complex fMRI data and a diffusion model due to their controllability with the cross-attention layers For the diffusion model, various example embodiments provide two strategics for diffusion model training and compared their performance in reconstruction and interpretation. In a first strategy, a first diffusion model (which is referred to herein as the NeuroSynth model) is trained from scratch using fMRI-audio data, in which a hierarchical interpretation method according to various example embodiments may be employed to characterize the fundamental associations between the cerebral cortex and audio generation. In various example embodiments, the NeuroSynth model utilizes a diffusion model comprising a customized U-Net architecture, conditioned on fMRI embeddings extracted by a trainable fMRI encoder. In a second strategy, a state-of-the-art AudioLDM is fine-tuned. Empirical results demonstrate that Mind-Audio effectively reproduces audio signals with 93.72% in 50-way classification accuracy in fMRI- to-Speech generation and 98.18% in fMRI-to-Music generation. From experiments conducted, fine-tuning the AudioLDM (i.e., the second strategy) slightly outperformed NeuroSynth (i.e., the first strategy) in terms of reconstruction performance but the former lacks detailed model interpretability. In contrast, training NcuroSynth from scratch coupled with the hierarchical interpretation method allows the specific role of brain regions for decoding individual audiostimuli to be determined Leveraging cutting-edge self-supervised learning and generative models, various example embodiments advantageously provide a method or roadmap for fMRIbased audio reconstruction and nuanced delineation of neural substrates underlying auditory processing. In this regard, various example embodiments enable robust audio reconstruction based on non-invasive brain recordings, focusing on blood-oxygen-level-dependent (BOLD) signals captured through fMRI.
[0097] Brain decoding is crucial for advancing human-machine communication, standing at the intersection of artificial intelligence and neuroscience. Recent innovations in this field leverage generative models, specifically for tasks related to text-to-image generation. The essence of this method is the application of generative models, conditioned on fMRI data gathered while subjects experience various stimuli such as images, videos, and language, aiming to accurately reconstruct these perceived stimuli. This approach has demonstrated remarkable success in numerous studies. Building on these advancements, researchers are exploring the possibilities of fMRI-to-audio reconstruction. For instance, Denk et al., “Brain2music: Reconstructing music from human brain activity”, arXiv prepnnt arXiv:2307.1 1078, 2023 (herein referred to as the Denk reference) refined the pre-trained MusicLM using fMRI-music paired data and applied autoregressive audio and w2v-Bcrt language modeling techniques. Similarly, Park etal., “Sound reconstruction from human brain activity via a generative model with brain-like auditory features”, arXiv prepnnt arXiv:2306: 11629, 2023 (herein referred to as the Park reference) have implemented a codebook encoder and decoder to reconstruct audio signals, conditioned on fMRI data, further expanding the application scope of brain decoding technologies.
[0098] While recent innovations have highlighted considerable capabilities in audio reconstruction, there still exists a noticeable gap in interpreting the activated brain regions. Typically, existing methodologies implement a linear model to encode brain activation and map feature importance back to the brain voxels. In such instances, brain encoding models can only provide insights into how each brain region contributes to the overall generation, lacking specificity and depth in interpretation. In contrast, by incorporating cross-attention layers in diffusion models according to various example embodiments of the present invention, more precise and nuanced decoding is enabled, allowing the recognition of how each brain patch contributes to the generation of particular audio stimuli through the denoising process. This process not only considers brain signals from multiple brain regions but also the relationships between them. Accordingly, various example embodiments investigate the diffusion model-based audio decoding, conditioned on fMRI data. The superior controllability afforded by diffusion models facilitates a more comprehensive interpretation of the relationships between different brain regions and audio reconstruction.
[0099] Fine-tuning currently dominates the implementation of diffusion models in brain decoding tasks in the era of large generative models. There have been several attempts to finetune the cross-attention layer in latent diffusion models (LDM). In parallel, some researchers mapped fMRI voxels directly to the language embeddings of CLIP to facilitate the use of LDMs without necessitating further adjustments. The fine-tuning of state-of-the-art audio generation models typically produces results comparable or potentially superior to those of models trained from scratch but with fewer trainable parameters and fewer training steps. However, a crucial yet unexplored domain in this area is the training of a diffusion model from scratch, specifically for brain decoding. This raises a compelling interest in examining how the performance of the train-from-scratch model compares with the fine-tuning of large generative models. Moreover, inherent biases in pre-trained models, derived from the characteristics and distribution of the training data, can potentially distort interpretative reliability. If the models are trained on data that is not wholly representative or contains such inherent biases, these inaccuracies may be reflected in the model’s interpretation reliability. Thus, various example embodiments arc centered around not merely achieving precise audio generation but also trustworthy interpretation.
[0100] Beyond the scope of generative models, building an effective fMRI learner also stands as another important component in brain decoding. Recent studies primarily use ridge regression to convert fMRI data into a suitable feature space that matches the embedding of the signals intended for reconstruction. However, due to the intricate nature of fMRI data, linear mappings may not capture all the nuances, leading to less optimal fMRI embeddings. Given this, Mind-Audio opts for the implementation of self-supervised learning approaches, aiming to better capture the underlying patterns within fMRI data.
[0101] Accordingly, to address the existing research gaps in fMRI-based audio decoding, various example embodiments introduce the Mind-Audio framework, seeking to set a standard in fMRI-conditioned audio reconstruction. The Mind-Audio framework is not merely a mechanism for reconstructing audio through fMRI conditioning but a comprehensive solution, particularly engineered to decode and interpret auditory processing at both overall and individual stimuli levels. It bridges the research gaps by providing a refined analytical perspective, allowing a detailed understanding of the correlation between cerebral activity andauditory stimuli. The Mind-Audio framework incorporates a masked autoencoder to effectively capture the underlying complexities of fMRI data representations (fMRI embeddings). In various example embodiments, by employing a three-phase training approach, refined, individualized fMRI embeddings arc extracted, paving the way for the meticulous reconstruction of perceived audio signals. The Mind-Audio framework encapsulates two distinct models, the selection of which is adaptive to different scenarios. For example, the NeuroSynth model, trained from scratch, provides enhanced hierarchical interpretability, mapping intricate associations between cerebral activities and audio generation. On the other hand, for example, fine-tuning a pre-trained diffusion model (e.g., the efficient AudioLDM model described in Liu et al., “AudioLDM: Text-to-audio generation with latent diffusion models”, in International Conference on Machine Learning, PMLR, 2023 (herein referred to as the Liu reference) offers expedited training and generally superior performance, suited for applications prioritizing speed and efficacy.
[0102] Accordingly, various example embodiments pioneer the use of diffusion models for fMRI-conditioned audio decoding, being the first to train a diffusion model from scratch or fine-tune a diffusion model (AudioLDM) using fMRT-to-Audio data. In addition, various example embodiments introduce a hierarchical cerebral interpretation method, enabling a multigranularity analysis of the brain regions associated with audio reconstruction, which provides insights into both global and detailed interactions between brain regions and audio stimuli. Furthermore, various example embodiments provide a self-supervised learning-based fMRI representation learning approach, utilizing a masked autoencoder to capture the underlying intricacies of complex fMRI data accurately.
[0103] Various methods have been employed to generate audio conditioned on external input. Some relevant examples are provided in the context of the text-to-audio task, in which spectrogram-conditioned audio generation has been intensively studied. Among different conditions on audio generation tasks, text-to-audio generation has gained a lot of attention recently. AudioGen (disclosed in Kreuk et al., “Textually guided audio generation”, in Tire Eleventh International Conference on Learning Representations, 2023) proposed an autoregressive audio generation model for textually guided audio generation. Moreover, their work collected 10 datasets and proposed audio mixing for driving the model to internally leam to separate multiple sources. DiffSound (disclosed in Yang et al., “Diffsound: Discrete diffusion model fortext-to-sound generation”, IEEE / ACM Transactions on Audio, Speech and Language Processing, vol. 31, pp. 1720-1733, 2023) proposed a framework that consists of a text encoder,a vector quantized variational autoencoder (VQ-VAE), a token-decoder, and a vocoder. Based on the discrete diffusion model, they developed a non-autoregressive token-decoder to transfer the text features to a mel-spectrogram. Because of the success of latent diffusion models in image generation, attempts based on this idea for audio generation have also been explored recently. For instance, the above-mentioned AudioLDM was proposed to improve the quality of synthesized audio. It exploited the joint embedding of text obtained by pre-trained CLAP models as the condition in sampling.
[0104] In the specific domain of music generation, there are some attempts that have been made recently. MusicLM (disclosed in Agostinelli el al., “MusicLM: Generating music from text”, arXiv preprint arXiv:2301: 11325, 2023) proposed a hierarchical sequence-to-sequence conditional music generation by exploiting separate decoder-only Transformers. Moreover, a cascaded diffusion model is employed for high-quality music generation. This was facilitated by using a cascader model conditioned on text embedding and low-fidelity audio obtained by the generator model.
[0105] Decoding natural modalities from the brain signals is pivotal for explonng human brain functionality. With the advancements in the generative model, several attempts have been made to generate natural-modality signals based on fMRI data.
[0106] For image generation, based on the pre-trained latent diffusion model on text-to- imagc generation, researchers developed different approaches for injecting the human brain information into the pre-trained LDM model. Chen et al., (“Seeing beyond the brain: Conditional diffusion model with sparse marked modelling for vision decoding”, in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 22710-22720) employed a two-stages strategy, which was to train a vision-transformer based MAE first on a large fMRI dataset and then fine-timed it on the individual-level fMRI data. Takagi et al. (High-resolution image reconstruction with latent diffusion models from human brain activity,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 14453-14463) exploited a Ridge regression to map the voxels of ROI to the conditional embedding for LDM. Involving fMRI embedding the joint embedding of image and text (Scotti et al., “Reconstructing tire mind’s eye: finri-to-image with contrastive learning and diffusion priors”, arXiv preprint arXiv:2305.18274, 2023), i.e., CLIP embedding (Radford et al., “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meilaand T. Zhang, Eds., vol. 139. PMLR, 18-24 Jul 2021, pp. 8748-8763”, was also investigated for better fine-tuning the LDM.
[0107] For text generation, Tang et al. (“Semantic reconstruction of continuous language from non-invasivc brain recordings”, Nature Neuroscience, vol. 26, no. 5, pp. 858-866, May 2023) introduced a linear encoding model to learn the mapping from text features to fMRI responses. A beam search decoder was employed to output the candidate sequence with the highest likelihood.
[0108] Although there are already image and text generations from fMRI data, generating / reconstructing high-quality audio is still a research gap. Some attempts to work on fMRI-to-audio have been made. The Denk reference mentioned hereinbefore developed Brain2Music, which is also based on the linear mapping of the fMRI voxels to predict the music embedding as the input of MusicLM for reconstruction. The Park reference mentioned hereinbefore employed a codebook encoder and decoder for audio reconstruction, in which the fMRI data were embedded by Ridge regression.
[0109] For better understanding, the basic concepts of diffusion models will now be described.
[0110] Diffusion models, also referred to as the score-based generative models, arc probabilistic models that involve two processes: i) a forward process to transform the data distribution into standard Gaussian distribution by using a predefined noise scheduler and ii) a reverse process to gradually denoise the Gaussian-distributedvariable, which corresponds to learning the reverse process of a fixed Markov Chain of length N. Accordingly, this reverse process may be referred to as denoising diffusion probabilistic model (DDPM). These two steps are described in more detail below.
[0111] In the forward process, at each time step n £ [1, ... , / V], by denoting the sample of mel-spectrogram at step n as mn, the transition probability is:(Equation 7) where an'■= epresent the noise level and the accumulated noiselevel at each step, respectively. In the reverse process, starting from the Gaussian noise distribution p(z0), a denoising process conditioned on the fMRI embedding E1gradually generates mel-spectrogram prior m0by:(Equation 8)
[0112] Hie mean function and variance are parameterized as:where eg(mn, n, E^~) is the predicted noise in generation.
[0113] AudioLDM performs the forward and reverse processes in the latent space. A mcl- spectrogram-based VAE was used to compress the audio data into this latent space. Furthermore, by exploiting the joint embedding of language and audio by contrastive language-audio pretraining (CLAP), the condition was performed by CLAP latent audio embedding. Formally, denoting the audio embedding compressed from the melspectrogram m by VAE as z0and the CLAP audio embedding as Em, Equations (7)-(9) arc rewritten as follows to express the training of AudioLDM:(Equation 13)
[0114] An example Mind -Audio framework 1600 will now be described with reference to FIG. 16 according to various example embodiments of the present invention. In FIG. 16, FC 1612 represents the fully connected layers. In the Mind-Audio framework 1600, activated fMRI signals arc encoded into embeddings to guide the denoising process of both diffusion models 1608a, 1608b. In this regard, activated brain regions of the fMRI data from the fMRI training dataset are determined and the fMRI encoder 1606 is trained based on the activated brain regions of the fMRI data from the fMRI training dataset. In FIG. 16, £ and T> represent the encoder and the decoder of the masked autoencoder (MAE), respectively. The setting for the number of RcsBlocks followed the UNct in the Rombach reference. In experiments performed, the fMRI encoding model 1606 was trained utilizing the fMRI data on the HCP dataset, as well as the fMRl-to-Spccch or the fMRl-to-Music datasets. Subsequently, the NcuroSynth model1608a and the AudioLDM model 1608b were trained on the paired fMRI-audio data from the fMRI-to-Speech or the fMRI-to-Music datasets, respectively.
[0115] Various example embodiments seek to construct a model capable of reconstructing audio clips that subjects perceived with the guidance of corresponding fMRI response. In this regard, various example embodiments seek to resolve two challenges for the fMRI-to-audio reconstruction task. Specifically, the first one is how to effectively encode the fMRI signal (fMRI data) into distinctive representations to serve as a condition for the generative model. After that, a generative model, conditioned on fMRI embeddings, aims to accurately reproduce audio while ensuring semantic correctness.
[0116] Fonnally, considering a monophonic audio sequence x 6 IRrpaired with fMRI signals / £ RF. the following two components of the Mind- Audio 1600 were exploited to process the paired signals.
[0117] A fMRI encoding model, which compresses f into a latent embedding E? in a distinctive representational space IF . This space is conceptualized to encode the intrinsic structures in the fMRI data corresponding to the perceived audio stimuli, thereby paving the way for subsequent reconstruction of human audio signals.
[0118] A diffusion model, which is conditioned on E? and generates the audio signals during inference. Various example embodiments explore and compare between training a proposed NeuroSynth model from scratch and fine-tuning state-of-the-art AudioLDM.JMRI Encoder 1606
[0119] Given the intricate nature of the information contained within fMRI data, it is crucial to extract pivotal features from the data. Such extraction is instrumental in enhancing the effectiveness of the conditioning process. In this regard, various example embodiments employ self-supervised learning to encode fMRI signals (fMRI data) using the masked autoencoder (MAE). The masked autoencoder comprises an encoder configured to extract the embeddings for masked data patches and a decoder to reconstruct the original patches from these embeddings. After that, the encoder can generate the representations (fMRI embeddings) for unmasked input data. Furthermore, vision transformer (ViT) may be employed as the backbone network Accordingly, the fMRI encoder 1606 is trained based on fMRI data from an fMRI training dataset for generating fMRI embeddings. In particular, the fMRI encoder 1606 comprises a masked autoencoder and the encoder of the masked autoencoder is trained based on the fMRI training dataset using unsupervised learning with masked data modeling forgenerating the fMRI embeddings. In this regard, the unsupervised learning with masked data modeling compnses generating fMRI embeddings from fMRI data from the fMRI training dataset, masking a portion of the fMRI embeddings into masked fMRI embeddings and training the masked autocncodcr to recover the masked fMRI embeddings.
[0120] According to various example embodiments, a 1 -dimensional (ID) masked autocncodcr is implemented for fMRI encoding. The 3 -dimensional (3D) voxels of fMRI signals, represented as f3De US‘x / xA:. are flattened into a ID sequence f e Rd. with d = i x j x k. To align the inputs of the MAE, the vectorized voxels are divided into ID patches.
[0121] To obtain more effective fMRI representations, according to various example embodiments, a three-phase training approach, involving pre-training followed by fine-tuning, is implemented. Tn the pre-training phase, the masked autoencoder is trained on the Human Conncctomc Project’s data, which includes fMRI data with the auditory cortex defined by the Glasser reference mentioned hereinbefore. In particular, fMRI data corresponding to the auditory cortex is reformulated from 3D to ID space, following the sequence of the audio processing hierarchy, and subsequently segmented into uniformly sized patches. These patches are then converted into fMRI embeddings (which may also be referred to as tokens), with a substantial proportion (e.g., approximately 75%) being randomly masked in the encoder throughout the training phase. Employing the architecture of an unsymmetric masked autoencoder (He el al., “Masked autoencoders are scalable vision learners”, in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), lune 2022, pp. 16000-16009, herein referred to as the He reference), a lightweight decoder is designed to restore the masked fMRI embeddings by leveraging the unmasked embeddings produced by the encoder. This is beneficial for building a fundamental understanding of fMRI signals for the masked autoencoder model. Subsequently, the pre-trained masked autoencoder model is fine-tuned on all the fMRI data in the fMRI-audio paired dataset (e.g., in the manner as described in the above-mentioned He reference). Tire latent representations generated from the unmasked fMRI patches are then employed as a conditional input to guide the downstream audio generation process. In the third phase, the fMRI encoder 1606 is further adjusted with the end-to-end training of the diffusion models.Conditional Diffusion Models 1608a, 1608b
[0122] In fMRI-to-audio generation, according to various example embodiments, the objective is to reconstruct the audio input given encoded fMRI embedding Eff In variousexample embodiments, two diffusion models with different training strategies are provided for reconstruction, namely, training a NeuroSynth model 1608a from scratch or fine-tuning AudioLDM 1608b
[0123] The training of the NcuroSynth model 1608a from scratch will now be described according to various example embodiments of the present invention. At first, a NeuroSynth model 1608a based on DDPM described hereinbefore is provided. During training, the NeuroSynth model 1608a, the fMRI encoder 1606 and the fully connected (FC) layers 1612 are trained by estimating the noises at step n. The reweighted training objective was employed:(Equation 14) where e e J(0, I) denotes injected noise and ed(■) represents the function to predict the noise. Accordingly, the NeuroSynth model 1608a and the fMRI encoder 1606 are trained based on estimating noise at each time step of the denoising process of the NeuroSynth model 1608a.
[0124] A modified UNet architecture (e.g., from the Rombach reference mentioned hereinbefore) is adapted and modified for ee(-) . An upsample block (i.e., RU) and a downsample block (i.e., DR) are incorporated to the decoder and the encoder, respectively, to suit the dimensions of the input mel-spectrogram before and after the DCR and RCU blocks as shown in FIG. 16. Furthermore, the cross-attention layers were shifted to the low-resolution blocks. The architecture of the modified UNet is shown in FIG. 16. As shown, the NeuroSynth model 1608a comprises a U-Net comprising a set of encoder blocks and a set of decoder blocks, each encoder block comprising a downsample block and each decoder block comprising an upsample block. In this regard, each encoder block and each decoder block corresponding to a low-resolution block further comprising a cross-attention layer. Accordingly, the NeuroSynth model 1608a is trained by performing cross-attention conditioning with respect to the crossattention layer of the above-mentioned each encoder block and the above-mentioned each decoder block corresponding to a low-resolution block based on the fMRI embeddings generated by the fMRI encoder 1606. This modification enables the model to compute crossattention in low-resolution feature maps, enhancing computational efficiency during training while preserving the conditional generation capability of the model. In various example embodiments, for example, low-resolution feature maps represent feature maps having a dimension (height* width) of 64*64, 32*32, 16*16 or smaller. In the implementations of the FC layer 1612 after the fMRI encoder 1606, the fMRI embeddings are transformed from the number of patches to the dimensions of 77 for conditioning the NeuroSynth model 1608a.Additionally, cross-attention facilitates correspondence between cortical regions and the generated audio signals. Accordingly, the training a diffusion model (NeuroSynth model) 1608a from scratch is advantageously capable of fMRI-conditioning through cross-attention, which also enables the hierarchical interpretation for fMRl-to-audio reconstruction based on crossattention.
[0125] The fine-tuning AudioLDM 1608b will now be described according to various example embodiments of the present invention: As discussed in hereinbefore, AudioLDM 1608b comprises VAE (encoder and decoder), UNet and neural vocoder. During fine-tuning, the UNet, fMRI encoder and projection head are jointly optimized. In contrast to the training of the original AudioLDM, which used the audio embedding output from CLAP as the condition, the fMRI embedding output from the fMRI encoding model was directly used as the condition for pretraining. The objective function of this LDM may be described by:(Equation 15)
[0126] It is worth noting that AudioLDM 1608b utilized a concatenation of fMRI embedding with VAE audio embedding as the conditioning to the diffusion model. To achieve a fair comparison with NeuroSynth 1608a, this conditioning approach was also implemented in NeuroSynth 1608a. Within this context, the fMRI embeddings were dimensionally transformed to 1 utilizing the FC layer. In various example embodiments, variants of the NeuroSynth models, incorporating cross-attention conditioning and concatenation conditioning, arc designated herein as NeuroSynth CA and NeuroSynth Co, respectively.
[0127] Understanding the activated brain regions is another significant objective in brain encoding and decoding. A hierarchical brain interpretation method involving globally activated brain regions and detailed voxels instructing the reconstruction is provided according to various example embodiments of the present invention. The basic idea is based on the cross-attention mechanism. Specifically, a sample-level detailed interpretation is performed first. For each paired fMRI-audio sample, the cross-attention maps in NeuroSynth 1608a fully trained on fMRI-audio data are visualized to connect the reconstruction process and the specific conditioning patches. All 64 x 64 attention maps across all heads are averaged for interpretation. The mapped 77 patches are used for calculating cross-attention maps. The inverse computation of the FC layer 1612 may be applied to further obtain the attention maps corresponding to the original fMRI patches. Moreover, the attention maps associated with fMRI patches are further interpolated to 128x 128. The Pearson correlation coefficients between the interpolated mapsand the reconstructed mel-spectrogram may be used to assign importance to different brain regions. Consequently, the detailed contribution of specific fMRI patches for the audio reconstruction at the sample level can be described.
[0128] Following that, the global guidance of the fMRI patches to audio decoding may be analyzed by averaging the correlation coefficients (importance) of the specific brain regions across all samples. Accordingly, the general contribution of each region through the denoising process can be better understood.Experiments
[0129] Various experiments conducted on the Mind- Audio model 1600 according to various example embodiments of the present invention will now be described.
[0130] Regarding the fMRI-to-Speech dataset, publicly available pre-processed fMRI data from three individuals who participated in a listening experiment was employed. The Blood- Oxygen Level Dependent (BOLD) signals from the brain were acquired using gradient-echo EPI on a 3T Siemens TIM Trio scanner, housed at the UC Berkeley Brain Imaging Center, and operated with a 32-channel volume coil. The acquisition parameters included a repetition rate (TR) of 2 seconds, time to echo (TE) of 31 ms, flip angle of 70 degrees, voxel size of 2.24x2.24x4.1 mm, with slice thickness at 3.5 mm and an 18% slice gap, and a matrix size of 100x 100, comprising 30 axial slices. All procedural and compensatory aspects of the experiment were sanctioned by the UC Berkeley Committee for the Protection of Human Subjects. For this experiment, subjects were exposed to 86 stories, each ranging from 10 to 15 minutes, and were instructed to listen with their eyes closed. Among these 86 stories, 4 stories were chosen as the test set and the rest were chosen as the training set. Each story was delivered during a single fMRI scan session. To maintain alignment with the fMRI TR, the audio stimuli were segmented to match each TR with a 2-second length.
[0131] Regarding the fMRI-to-Music dataset, an additional fMRI-to-Music dataset named the music genre neuroimaging dataset from the Denk reference mentioned hereinbefore was employed. The preprocessing pipeline outlined in the Denk reference was also followed. During the scanning sessions, five participants were directed to concentrate on a fixation cross at the center and were exposed to music via MRI-compatible earphones. Identical musical stimuli were administered to each participant. Scanning was conducted using a 3.0T MR! scanner (TIM Tno; Siemens), equipped with a 32-channcl head coil, capturing 68 axial slices employing a T2*-weighted MB-EPI sequence Hie parameters were TR=1, 500ms, TE=30ms, FA=62°,FOV= I 92 192mm2, voxel sizc=2 '2 -'2nim 2 and a multi-band factor of 4, resulting in the acquisition of 410 volumes per run. The dataset encompasses a diverse range of musical stimuli across ten genres. Each genre is represented by 60 pieces, each 30 seconds long and sampled at 22.050 kHz, totaling 600 pieces. Training was performed on the 540 pieces and the rest 60 pieces was used for testing. From each of these, a clip of 15 seconds was extracted and processed for further analysis . To maintain alignment with the fMRI TR, the music stimuli were segmented to match the length of each TR, specifically 1.5 seconds.
[0132] Regarding model setting, the implementation details of the example Mind-Audio framework 1600 are described as follows. There were 1 RD / RU block and 3 RDC / RUC blocks in the NeuroSynth UNet backbone. The channel dimensions corresponding to these four blocks are [128, 256, 512, 512] (following the block orders in the UNet encoder). For transformer layers, the number of heads was eight. In the forward process, the number of time steps was set to A = 1000. A linear noise schedule from= 0.0015 to p;, = 0.02 was used. Dcnoising Diffusion Implicit Model (DDIM) sampler with 200 sampling steps was employed in sampling. Regarding AudioLDM 1608b for fine-tuning, the pre-trained AudioLDM-S on AudioCaps dataset (disclosed in Kim et al., “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 119-132) was employed in the experiment. The UNet was optimized in AudioLDM 1608b during fine -tuning.
[0133] The backbone of the fMRI encoding model was similar to ViT-Large (disclosed in Dosovitskiy et al., “An image is worth 16x16 words: Transfonners for image recognition at scale,” in International Conference on Learning Representations, 2021) with a ID patch embedder. The model used a patch size of 16, embedding dimension of 1024, encoder depth of 24, and mask ratio of 0.75.
[0134] The parameter settings of AudioLDM 1608b (disclosed in the Liu reference) was followed in the experiment. Regarding the transformation from audio waveform to mcl- spectrogram in NeuroSynth 1608a, all the audios in this work were resampled to 16,000 Hz. The window size, number of fast Fourier transform (FFT) points, and hop size were set to 1024, 1024, and 250, respectively. 128 mel filterbands were used to calculate the mel-spectrograms.
[0135] Regarding the evaluation metnes, two categories of evaluation metrics were employed to evaluate the performance of the Mind-Audio framework 1600. The first category'was the N-way accuracy classification test, a semantics-level metric employed to measure the semantic accuracy of the reconstructed samples. The classification test is fundamentally a comparison between the ground truth (GT) and the generated audio, executed using a PANNs classifier. Out of N randomly selected classes, if the GT class has the highest classification probability, a trial was deemed successful. This test was reiterated 100 times, with the subsequent average success rate reported. The second category consisted of three metrics to assess the quality of the reconstructed audio signals, which were frechet distance (FD), kullback-leibler (KL) divergence and frechet audio distance (FAD). The audio classifier PANN (disclosed in Kong et al., “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2880-2894, 2020) was exploited as the backbone model for computing the metrics of two categories. Furthermore, FAD was built upon VGGish (Hershey et al., “Cnn architectures for large-scale audio classification,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 131-135).
[0136] The evaluation results on fMRI-guided speech reconstruction and music reconstruction are shown in Tables 3 and 4 in FIGs. 17 and 18, respectively. Regarding different N-way classification accuracy, the NcuroSynth 1608a showed better accuracy on the fMRl-to- Music dataset, while AudioLDM 1608b presented a slightly better perfonnance on the fMRl- to- Speech dataset. The high accuracy demonstrated that the proposed framework could effectively recover the semantic information involved in the audio signals. Regarding the metrics that evaluated the quality of the reconstructed samples, it was observed that AudioLDM 1608b showed generally better performance in terms of FD, FAD and KL. Overall, both models exhibited strong performance, effectively showcasing the effectiveness of the framework and setting a benchmark for the fMRI-to-audio reconstruction task. It was concluded that AudioLDM 1608b with prior knowledge can be beneficial to achieve generally superior performance.
[0137] The reconstruction examples of two datasets are presented in FIGs. 19A and 19B. In particular, FIG. 19A shows a comparison of generated mel-spectrograms between the NeuroSynth model 1608a and the AudioLDM model 1608b on the fMRI-to-Speech dataset. FIG. 19B shows a comparison of generated mel-spectrograms between the NeuroSynth model 1608a and the AudioLDM model 1608b on the fMRl-to-Music dataset. The results of different subjects, i.c., S1-S3 in FIG. 19A and S1-S5 in FIG. 19B arc presented as well. Qualitatively, it was observed that the proposed framework could effectively produce similar mel-spectrogramsas the stimuli across all subjects. Tire audios corresponding to the mel-spectrograms in FIGs. 19A and 19B.
[0138] To better understand the capability of both models, the metrics evaluating reconstruction quality in different training epochs arc presented in FIGs. 20A to 20C. In particular, FIGs. 20A to 20C depict plots comparing between the NciiroSyntli model 1608a and the AudioLDM model 1608b across different epochs using SI data of fMRl-to-Spccch dataset with respect three metrics FD (Frechet Distance), FAD (Frechet Audio Distance) and KL (Kullback-Leibler). Accordingly, three metrics (FD, FAD, KL) were evaluated during the training process. It was observed that AudioLDM 1608b could already achieve outstanding performance at 200 epochs while more than 500 epochs were used to achieve the best in Neuro Synth 1608a training.
[0139] Regarding conditioning model, as discussed hereinbefore, linear mapping has been widely used for encoding the fMRI data into joint spaces of audio and language or image and language. Therefore, the self-supervised fMRI encoding model 1606 was compared with linear mapping to gain a deeper understanding. Slightly different from previous usage, the masked autoencoder was replaced with a linear layer which was trained together with the diffusion models. This more effectively ensured that the linear mapping could straightforwardly learn the encoding of fMRI data for the reconstruction. Tire comparison results are presented in Table 5 in FIG. 21. In particular, FIG. 21 shows a table (Table 5) presenting the comparison results between different fMRI encoding models using SI data of the fMRI-to-Speech dataset, according to various example embodiments of the present invention. The term ‘Params’ denote the number of parameters in fMRI encoding. Linear fMRI encoding model was applied in both NeuroS nth 1608a and AudioLDM 1608b. Given that the semantic accuracy of the trained generative models has been demonstrated above, the focus here primarily lies on the similarity metrics. It was observed that the diffusion models paired with MAE-based fMRI encoding models generally yielded superior results, underscoring the efficacy of the proposed fMRI encoding model. The utilization of prior knowledge from a large-scale unlabeled fMRI dataset, combined with individualized fine-tuning, effectively enhanced the guidance for more accurate reconstruction.
[0140] Regarding fine-tuning components in AudioLDM 1608b, as the whole UNet was trained during fine-tuning AudioLDM 1608b, it was worth investigating the performance of fine-tuning different components. Tabic 6 in FIG. 22 showed the performance comparison across different components for fine-tuning. In particular, FIG. 22 shows a table (Table 6)presenting the comparison results between different fine-tuned components in tire AudioLDM model 1608b using Si data of the fMRI-to-Speech dataset, according to various example embodiments of the present invention. The term ‘Params' denotes the number of parameters in the fine-tuned components of AudioLDM 1608b and ‘Attn’ refers to self-attention layers. It was observed that fine-tuning the whole UNet demonstrated better performance. Nonetheless, fine-tuning self-attention layers also presented comparable results and was superior to the NeuroSynth model 1608a. Thus, for the objective of expediting the training process, exclusively fme-tuning the self-attention layers emerges as a practical compromise.
[0141] Regarding attention mapping and neural correlation insights, through our interpretation analysis, several key findings were uncovered. As demonstrated in FIGs. 23 A and 23B via self-attention map analysis in the masked autoencoder. In particular, FIGs. 23 A and 23B show the analysis of self-attention from S 1 in both NeuroSynth 1608a and AudioLDM 1608b. This exploited the first self-attention layer in ViT of the fMRI encoding model 1606. The self-attention values were averaged across all testing data and all heads. It was found that within both NeuroSynth 1608a and AudioLDM 1608b, the auditory network and the language / speech association areas were significantly influential in fMRI-audio generation. This is consistent with the existing neuroscience literature on auditory’ processing and language comprehension. Upon receiving auditory’ information, our brain processes it through the primary auditory network, acting as the initial processing center responsible for the basic interpretation of auditory signals. Following this, the refined auditory’ data are directed to the superior temporal gyrus and Broca’s area for more advanced processing, facilitating understanding of complex auditory stimuli such as speech and music.
[0142] The hierarchical interpretation analysis is then performed. Given that the NeuroSynth_CA served as the underlying model for this analysis, its audio reconstruction performance is detailed for reference in Tables 7 and 8 in FIGs. 24A and 24B. Furthermore, the visualization of the hierarchical interpretation is presented in FIG. 25. In particular, FIG. 25 shows the analysis of cross-attention maps from SI in NeuroSynth. Portion A of FIG. 25 highlights a specific example of audio stimuli. Portion B of Fig. 25 displays the averaged correlation between brain patches and the cross-attention map across all test stimuli. Portion C illustrates the increase in detail within the attention maps as the denoising process progresses. Several aspects of the attention maps and their correlations to stimuli can be observed. From the detailed sample-level view, a given example of audio stimuli was focused, as highlighted in portion A of FIG. 2 . It was observed that there was a pronounced correlation between thegenerated me 1 -spectrogram and the attention map. This correlation suggests that when the model is exposed to these mel-spectrograms, the related fMRI patches are responsible for generating different structure information from outline to details and rendering comparable cross-attention maps. Such observations extend our comprehension of the interactive dynamics between neural responses and generated mel-spectrogram within the model.
[0143] To gam a comprehensive understanding of the brain patches’ global contribution, the correlations across all test stimuli were averaged. As shown in portion B of FIG. 25, it can be observed that the averaged cross-attention map showed a more widespread feature importance compared to the averaged self-attention map. In this phase, the self-attention layers of the masked autoencoder directly interact with the fMRI tokens. This allows for the monitoring of the underlying relationships between each voxel. Concurrently, the crossattention layers in the UNet are responsible for evaluating how these brain patches contribute to audio generation. The resulting self-attention and cross-attention maps revealed some overlapping voxels, mainly located in the auditory cortex and the auditory association area. This overlap suggests that these regions play a crucial role not only in the brain’s auditory processing but also in the audio reconstruction process. However, there are noticeable differences between the self-attention and cross-attention maps as well. For instance, some voxels, notably in the prefrontal cortex regions, which are significant in the cross-attention map, exhibit low values in the self-attention map. This discrepancy implies that while these voxels may not typically be central to auditory processing, they are critical for audio reconstruction. This process elucidated how different focal points on the structural and spectral elements of audio recordings are characterized by individual brain patches.
[0144] Lastly, a progression of the granularity’ of the attention maps was noted as the denoising process unfolded. Analogous to the process in text-to-image generation, the preliminary’ denoising steps of the fMRI-to-audio reconstruction primarily focus on discerning the outline or foundational structure of the reconstructed mel-spectrograms. In contrast, the subsequent stages are dedicated to refinement, capturing intricate details. This process is showcased in portion C of FIG. 25.
[0145] Accordingly, various example embodiments introduce the Mind- Audio model 1600 as a unique approach to not only proficiently recover perceived speech and music signals utilizing fMRI but also to provide a comprehensive interpretation of the brain regions, which is crucial for auditory’ processing. This research paves the way’ for advancements in non-invasivcaudio decoding and represents a step forward to unravel the complex relationship between auditory stimuli and human brain activity'.
[0146] Accordingly, various example embodiments provide a method of training a machine learning model for reconstructing video or audio data based on ncuroimaging data of a subject. In this regard, various example embodiments provide a pioneering framework that utilizes a self-supervised learning approach combined with diffusion models to reconstruct perceived visual and / or auditory' information from fMRI data, thus enabling multimodal brain decoding. For example, the high-dimensional brain activity patterns are translated into high-fidelity representations across multiple senses (visual and audio), thereby enabling a deeper understanding of how the brain processes various stimuli. This method has substantial applications in the field of cognitive neuroscience and can significantly enhance brain-computer interface systems. The empirical evidence demonstrates the model's capability to reproduce visual and audio signals, including videos, speeches and music, with a high degree of fidelity, marking a significant advancement over existing methodologies.
[0147] In various example embodiments, self-supervised representation learning (masked brain modelling) is employed to allow for the efficient and effective extraction of valuable features from complex fMRI data without needing large amounts of unlabcllcd data. This approach significantly reduces the need for manual unlabelled and thus cuts down on time and resources while enhancing the model's ability' to understand and predict from the data. It also promotes the ability' to leam directly from the environment, increasing the model's versatility and adaptability. In various example embodiments, stable diffusion-based video / audio generation is employed which allows for the generation of high-quality, detailed videos and audio directly' from fMRI embeddings, supporting higher levels of detail and fidelity' in generated content. This helps to bridge the gap betw een brain activity and perceivable output, opening the door to more accurate and detailed visualizations of cognitive processes. The stability of the diffusion-based approach ensures consistency in the generated output, increasing the reliability' and usability' of the model's results. In various example embodiments, fMRI for perception decoding is employed. In this regard, fMRI represents a powerful tool in the decoding of perceptual information. With its high spatial resolution, fMRI provides the finegrained details that make it an ideal modality for perception decoding. The use of fMRI in this context allows the mapping of neural activity with great precision to distinct perceptual phenomena. Each voxel in an fMRI scan can be associated with a specific location in the brain, thereby allowing a detailed understanding of tire spatial distribution of neural activation. Thisgranular perspective opens a window into how different brain regions respond to various stimuli, enhancing our understanding of brain function. In various example embodiments, hierarchical interpretation method is employed, which enables the understanding of the interactions within fMRI patches, as well as between these patches and audio spectrums or videos, offers a detailed understanding of both global and fine-scale auditory and visual processing in the brain, facilitating the development of targeted interventions and improvements in brain-computer interface technologies.
[0148] For better understanding, FIG. 26 depicts a schematic flow diagram of an example method 600 of training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, corresponding to the Mind-Video and the Mind-Audio as described hereinbefore according to various example embodiments of the present invention. As shown in FIG. 26, the method 2600 comprises a pre-training phase, a Mind-Video model path and a Mind-Audio path.
[0149] The pre-training phase may include: data preparation from multiple public datasets; data quality control (for ensuring that only high-quality, usable data is processed); core preprocessing (involving normalization, alignment, and possibly augmentation to prepare the data for feature extraction); and generating slicc-wisc fMRI Data (extract and organize fMRI data into usable slices for model input).
[0150] Forthc Mind-Video model path, the model (Mind-Video) training may include: data acquisition (record fMRI data while the subject is watching videos); feature extraction (extract visual context and neuroimage data from the videos and corresponding neuroimaging data); model pre-training (utilize self-supervised learning to pre-train the model based on the extracted features, improving initial feature representation); MBM encoder tuning (fine-tune the MBM encoder to align and enhance the fMRI and video embeddings); CLIP fine-tuning (integrate and fme-tune using CLIP for enhancing cross-modal compatibility between fMRI and video data); and video diffusion model training (train the video diffusion model to generate video outputs from fMRI data using the tuned embeddings). Model (Mind-Video) testing may also be performed, including: generate videos from fMRI Data; use the trained model to generate video outputs; evaluate generated videos (assess the quality and accuracy ofthe reconstructed videos); interpret important brain patterns (analyze and report how brain data correlates with video generation).
[0151] Forthc Mind-Audio model path, the model (Mind-Audio) training may include: data acquisition (record fMRI data while tire subject listen to audio); feature extraction (extractauditory features and neuroimage data from the audio and corresponding neuroimaging data); model pre-training (apply self-supervised learning to pre-train the model on these extracted features); MBM fine-tuning (adjust the MBM encoder to better align fMRI and audio embeddings); and audio diffusion model training (tram the audio diffusion model using the fine-tuned embeddings to generate audio from fMRI data). Model (Mind-Audio) testing may also be performed, including: generate audio from fMRI Data (use the trained model to produce audio outputs); evaluate generated audio (test the audio for fidelity and accuracy); interpret important brain patterns on audio generation (analyze how brain data influences audio output characteristics).
[0152] The trained machine learning model for reconstructing video or audio data based on neuroimaging data of a subject according to various example embodiments has a wide variety of practical applications. One practical application is human-machine communication. In this regard, by accurately decoding and reconstructing visual information from fMRI data, this trained machine learning model lays the groundwork for more sophisticated brain-computer interfaces, thus enhancing the capacity for human-machine communication. Another practical application is in enhancing neurofeedback strategies. In this regard, the trained machine learning model provides more detailed insights into brain processes, which can aid in the development of more precise neurofeedback strategies to enhance human perceptual and cognitive performance, with applications in healthcare, education, and professional training environments. Another practical application is in cognitive neuroscience research. In this regard, the trained machine learning model can become a valuable tool in cognitive neuroscience research, helping to deepen understanding of complex brain activities and advancing the field. The commercial benefits include new treatment methodologies, advanced research capabilities, and facilitate new product development. For example, in consumer electronics, the trained machine learning model may be applied in the development of nextgeneration smart devices that can adapt to a user's preferences and responses by interpreting their neural signals related to, for example, auditory experiences. For example, in mental health therapies, the trained machine learning model may be applied to help better understand the human brain's processing, which can lead to more effective therapies for a variety of mental health disorders.
[0153] While embodiments of the invention have been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in fonn and detail may be made therein without departing from the scope ofthe invention as defined by the appended claims. The scope of the invention is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced.
Claims
CLAIMS1 . A method of training a machine learning model for reconstructing video or audio data based on ncuroimaging data of a subject, the method comprising: training a neuroimaging data encoder based on neuroimaging data from a neuroimaging training dataset for generating ncuroimaging data embeddings; and training a diffusion model based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model, the diffusion model being trained to reconstruct video or audio data based on neuroimaging data embeddings of neuroimaging data of a subject obtained in response to a visual or audio stimulus, wherein the neuroimaging data encoder comprises a masked autoencoder, and said training the neuroimaging data encoder comprises training an encoder of the masked autoencoder based on the neuroimaging training dataset using unsupervised learning with masked data modeling for generating the neuroimaging data embeddings, the unsupervised learning with masked data modeling comprising generating neuroimaging data embeddings from neuroimaging data from the neuroimaging training dataset, masking a portion of the ncuroimaging data embeddings into masked ncuroimaging data embeddings and training the masked autoencoder to recover the masked neuroimaging data embeddings.
2. The method according to claim 1, further comprising determining activated brain regions of the neuroimaging data from the neuroimaging training dataset, wherein the neuroimaging data encoder is trained based on the activated brain regions of the neuroimaging data from the neuroimaging training dataset.
3. The method according to claim 1 or 2, wherein the method is for training the machine learning model for reconstructing video data, and said training the neuroimaging data encoder further comprises augmenting the encoder of the masked autoencoder with spatiotemporal attention heads for processing the neuroimaging data embeddings generated by the encoder of the masked autoencoder in a sliding window.
4. The method according to claim 3, wherein for each of a plurality of windows of ncuroimaging data embeddings generated by the encoder of the masked autocncodcr, said training the neuroimaging data encoder further comprises training tire encoder of tire maskedautoencoder based on a set of embeddings comprising tire window of neuroimaging data embeddings, corresponding image embeddings and corresponding text embeddings projected into a shared latent space based on contrastive learning.
5. The method according to claim 4, wherein the diffusion model compnscs a U-Nct comprising self-attention layers, cross-attention layers and temporal-attention layers; and the diffusion model is further trained based on a target video dataset.
6. The method according to claim 1 or 2, wherein the method is for training the machine learning model for reconstructing audio data, and said training the diffusion model comprises training the diffusion model and the neuroimaging data encoder based on estimating noise at each time step of a denoising process of the diffusion model.
7. Tire method according to claim 6, wherein the diffusion model comprises a U-Nct comprising a set of encoder blocks and a set of decoder blocks, each encoder block comprising a downsample block and each decoder block comprising an upsample block, and each encoder block and each decoder block corresponding to a low-resolution block further comprising a cross-attention layer, said training the diffusion model further comprises performing cross-attention conditioning with respect to the cross-attention layer of said each encoder block and said each decoder block corresponding to a low-resolution block based on the neuroimaging data embeddings generated by the neuroimaging data encoder.
8. The method according to claim 1 or 2, wherein the method is for training the machine learning model for reconstructing audio data, and the diffusion model is pre-trained and is further trained bv fine-tuning jointly with the neuroimaging data encoder using the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model.
9. A system for training a machine learning model for reconstructing video or audio data based on neuroimaging data of a subject, the system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to: train a ncuroimaging data encoder based on ncuroimaging data from a ncuroimaging training dataset for generating neuroimaging data embeddings; and train a diffusion model based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model, the diffusion model being trained to reconstruct video or audio data based on neuroimaging data embeddings of neuroimaging data of a subject obtained in response to a visual or audio stimulus, wherein the neuroimaging data encoder comprises a masked autoencoder, and said train the neuroimaging data encoder comprises training an encoder of the masked autoencoder based on the neuroimaging training dataset using unsupervised learning with masked data modeling for generating the neuroimaging data embeddings, the unsupervised learning with masked data modeling comprising generating neuroimaging data embeddings from ncuroimaging data from the ncuroimaging training dataset, masking a portion of the neuroimaging data embeddings into of masked neuroimaging data embeddings and training the masked autocncodcr to recover the masked ncuroimaging data embeddings.
10. The system according to claim 9, wherein the at least one processor is further configured to determine activated brain regions of the neuroimaging data from the neuroimaging training dataset, wherein the neuroimaging data encoder is trained based on the activated brain regions of the neuroimaging data from the neuroimaging training dataset.
11. The system according to claim 9 or 10, wherein the system is for training the machine learning model for reconstructing video data, and said train the neuroimaging data encoder further comprises augmenting the encoder of the masked autoencoder with spatiotemporal attention heads for processing the neuroimaging data embeddings generated by the encoder of the masked autoencoder in a sliding window.
12. The system according to claim 11, wherein for each of a plurality of windows of neuroimaging data embeddings generated by the encoder of the masked autoencoder, said trainthe neuroimaging data encoder further comprises training tire encoder of the masked autoencoder based on a set of embeddings comprising the window of neuroimaging data embeddings, corresponding image embeddings and corresponding text embeddings projected into a shared latent space based on contrastive learning.
13. The system according to claim 12, wherein the diffusion model comprises a U-Net comprising self-attention layers, cross-attention layers and temporal-attention layers; and the diffusion model is further trained based on a target video dataset.
14. The system according to claim 9 or 10, wherein the system is for training the machine learning model for reconstructing audio data, and said train the diffusion model comprises training the diffusion model and the neuroimaging data encoder based on estimating noise at each time step of a denoising process of the diffusion model.
15. The system according to claim 14, wherein the diffusion model comprises a U-Net comprising a set of encoder blocks and a set of decoder blocks, each encoder block comprising a downsample block and each decoder block comprising an upsample block, and each encoder block and each decoder block corresponding to a low-resolution block further comprising a cross-attention layer, said training the diffusion model further comprises performing cross-attention conditioning with respect to the cross-attention layer of said each encoder block and said each decoder block corresponding to a low-resolution block based on the neuroimaging data embeddings generated by the neuroimaging data encoder.
16. The system according to claim 9 or 10, wherein the system is for training the machine learning model for reconstructing audio data, and the diffusion model is pre-trained and is further trained by fine-tuning jointly with the neuroimaging data encoder using the neuroimaging data embeddings generated by the ncuroimaging data encoder as conditions on the diffusion model.
17. A computer program product, embodied in one or more non-transitory computer- readable storage mediums, comprising instructions executable by at least one processor to perform the method of training a machine learning model for reconstructing video or audio data based on neuro imaging data of a subject according to any one of claims 1 to 8.
18. A method of using the machine learning model trained according to any one of claims 1 to 8 for reconstructing video or audio data based on neuroimaging data of a subject, the method comprising: generating, by the neuroimaging data encoder of the machine learning model, neuroimaging data embeddings based on neuroimaging data of a subject obtained in response to a visual or audio stimulus; and reconstructing video or audio data, by the diffusion model of the machine learning model, based on the neuroimaging data embeddings generated by the neuroimaging data encoder as conditions on the diffusion model.
Citation Information
Patent Citations
Image generation method and device based on electroencephalogram, computer equipment and storage medium
CN117472181A
Generating mask information
US20240096064A1
Cited By
Fetal heart rate signal data enhancement method based on conditional diffusion model
CN121313132A
FMRI video nerve decoding method based on visual perception and semantic consistency
CN121814971A