Methods, systems, and computer readable media for prospective motion estimation and correction in functional magnetic resonance imaging (FMRI) using deep neural networks

The dual-stream deep neural network addresses the slow response times of current prospective motion correction methods by directly estimating motion parameters from aliased SMS images, achieving accurate and real-time motion correction in fMRI scans.

WO2026015733A1PCT designated stage Publication Date: 2026-01-15THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/037153
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-07-10
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Current methods for prospective motion correction in functional magnetic resonance imaging (fMRI) are inadequate due to slow response times, typically requiring 300 to 500 milliseconds for image reconstruction and co-registration, making them ineffective for rapid head movements during scans.

Method used

A dual-stream deep neural network (DSDN) is employed to directly estimate motion parameters from aliased simultaneous multi-slice (SMS) images, bypassing traditional reconstruction and co-registration processes, enabling real-time motion correction by adjusting magnetic coil fields of view.

Benefits of technology

The DSDN achieves motion estimation accuracy below 0.2 voxel/degree, facilitating real-time motion correction and enhancing image quality and statistical significance in ultrahigh-resolution fMRI.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025037153_15012026_PF_FP_ABST
    Figure US2025037153_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for prospective motion estimation and correction in functional fMRI using deep neural networks includes inputting aliased SMS fMRI images into a first branch of a CNN trained to generate a feature vector from the aliased SMS fMRI images and receiving, as output, a feature vector for the aliased SMS fMRI images. The method further includes inputting a reference image into a second branch of the CNN trained to generate a feature vector from the reference image, and receiving, as output, a feature vector for the reference image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]METHODS, SYSTEMS, AND COMPUTER READABLE MEDIA FOR PROSPECTIVE MOTION ESTIMATION AND CORRECTION IN FUNCTIONAL MAGNETIC RESONANCE IMAGING (fMRI) USING DEEP NEURAL NETWORKS GOVERNMENT INTEREST This invention was supported by grant number AG073297 awarded by the National Institutes of Health. The government has certain rights in the invention. PRIORITY CLAIM This application claims the priority benefit of U.S. Provisional Patent Application Serial No.63 / 669,680 filed July 10, 2024, the disclosure of which is incorporated herein by reference in its entirety. TECHNICAL FIELD The subject matter described herein relates to motion estimation and correction for fMRI. More particularly, the subject matter described herein relates to methods, systems, and computer readable media for prospective motion estimation and correction in fMRI using deep neural networks. BACKGROUND fMRI images can be adversely affected by motion artifacts when the subject moves during an fMRI scan. One way to address this issue is to correct the fMRI images after the scan. This type of correction is referred to as retrospective motion correction. While retrospective motion correction can correct for slow movements during an fMRI scan, retrospective motion correction cannot effectively correct for rapid movements, such as head movements, that occur during a scan. In light of the issues with retrospective motion correction, prospective motion correction has been investigated. Prospective motion correction involves correction of images to remove motion artifacts during a scan. However, existing methods for prospective motion correction of fMRI images require image reconstruction and co-registration and may not be fast enough to perform prospective motion correction in real time. Accordingly, in light of these and other difficulties, there exists a need for improved methods, systems, and computer readable media for prospective motion correction for fMRI images. SUMMARY A method for prospective motion estimation and correction in functional magnetic resonance imaging (fMRI) using deep neural networks includes inputting aliased simultaneous multi-slice (SMS) fMRI images into a first branch of a convolutional neural network (CNN) trained to generate a feature vector from the SMS fMRI images. The method further includes inputting a reference image into a second branch of the CNN trained to generate a feature vector from the reference image. The method further includes inputting the feature vectors into a motion estimator trained to predict motion parameters of the SMS fMRI images relative to the reference image and receiving, as output, values for the predicted motion parameters of the aliased SMS fMRI images. The method further includes outputting the values for the predicted motion parameters to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to perform the fMRI scan. According to another aspect of the subject matter described herein, the first branch of the CNN includes a plurality of CNN blocks, each including a 2D convolution layer, a batch normalization layer, a rectilinear unit (ReLU) activation layer, and a max pooling layer. According to another aspect of the subject matter described herein, the second branch of the CNN includes a plurality of CNN blocks, each including a 3D convolution layer, a batch normalization layer, a rectilinear unit (ReLU) activation layer, and a max pooling layer. According to another aspect of the subject matter described herein, the motion estimator is configured to concatenate the feature vectors and reduce dimensionality of the concatenated feature vectors to produce the values of the predicted motion parameters. According to another aspect of the subject matter described herein, the values for the predicted motion parameters include predicted values of shift and rotation parameters of the aliased SMS fMRI images relative to the reference image. According to another aspect of the subject matter described herein, the method for prospective motion estimation and correction in fMRI includes adjusting the fields of view of the magnetic coils by shifting and rotating the fields of view of the magnetic coils using the predicted values of the shift and rotation parameters. According to another aspect of the subject matter described herein, the reference image comprises a T2-weighted MR image. According to another aspect of the subject matter described herein, the reference image comprises a T1-weighted MR image. According to another aspect of the subject matter described herein, the method for prospective motion estimation and correction in fMRI includes inputting aliased single slice fMRI images into a third branch of the convolutional neural network trained to generate a feature vector from the aliased single slice fMRI images. According to another aspect of the subject matter described herein, inputting the feature vectors into the motion estimator comprises inputting the feature vectors from the aliased single slice fMRI images into the motion estimator and the motion estimator is configured to predict motion parameters of the aliased single slice images relative to the reference image and receiving, as output, values for the predicted motion parameters of the SMS fMRI images. According to another aspect of the subject matter described herein, a system for prospective motion estimation and correction in fMRI using neural networks is provided. The system includes a computing platform including at least one processor and a memory. The system further includes a trained multi-branch feature extractor CNN implemented by the at least one processor and including a first branch for receiving, as input, aliased SMS fMRI images and generating, as output, a feature vector from the aliased SMS fMRI images and a second branch for receiving, as input, a reference image and generating, as output, a feature vector from the reference image. The system further includes a trained motion estimator for receiving the feature vectors as input and generating, as output, predicted values for motion parameters of the SMS fMRI images relative to the reference image, and providing the values of the predicted motion parameters of the SMS fMRI images to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to perform the fMRI scan. According to another aspect of the subject matter described herein, the first branch of the multi-branch feature extractor CNN includes a plurality of CNN blocks, each including a 2D convolution layer, a batch normalization layer, and a rectilinear unit (ReLU) activation layer, and a max pooling layer. The system of claim 11 wherein the second branch of the multi-branch feature extractor CNN includes a plurality of CNN blocks, each including a 3D convolution layer, a batch normalization layer, and a rectilinear unit (ReLU) activation layer, and a max pooling layer. According to another aspect of the subject matter described herein, the trained motion estimator is configured to concatenate the feature vectors and reduce dimensionality of the concatenated feature vectors to produce the values of the predicted motion parameters. According to another aspect of the subject matter described herein, the values for the predicted motion parameters include predicted values of shift and rotation parameters of the aliased SMS fMRI images relative to the reference image. According to another aspect of the subject matter described herein, the fMRI image acquisition controller is configured to use the values of the predicted motion parameters to adjust the fields of view of the coils by shifting and rotating the fields of view of the coils using the predicted values of the shift and rotation parameters. According to another aspect of the subject matter described herein, the trained feature extractor CNN comprises a third branch for receiving aliased single slice fMRI images and generating a feature vector from the aliased single slice fMRI images. According to another aspect of the subject matter described herein, a non-transitory computer readable medium having stored thereon computer- executable instructions that when executed by a processor of a computer control the computer to perform steps is provided. The steps include inputting aliased SMS fMRI images into a first branch of a CNN trained to generate a feature vector from the aliased SMS fMRI images and receiving, as output, a feature vector for the aliased SMS fMRI images. The steps further include inputting a reference image into a second branch of the CNN trained to generate a feature vector from the reference image and receiving, as output, a feature vector for the reference image. The steps further include inputting the feature vectors into a motion estimator trained to predict motion parameters of the SMS fMRI images relative to the reference image and receiving, as output, values for the predicted motion parameters of the aliased SMS fMRI images. The steps further include outputting the values for the predicted motion parameters to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to perform the fMRI scan. The subject matter described herein can be implemented in software in combination with hardware and / or firmware. For example, the subject matter described herein can be implemented in software executed by a processor. In one exemplary implementation, the subject matter described herein can be implemented using a non-transitory computer readable medium having stored thereon computer executable instructions that when executed by the processor of a computer control the computer to perform steps. Exemplary computer readable media suitable for implementing the subject matter described herein include non-transitory computer-readable media, such as disk memory devices, chip memory devices, programmable logic devices, and application specific integrated circuits. In addition, a computer readable medium that implements the subject matter described herein may be located on a single device or computing platform or may be distributed across multiple devices or computing platforms. BRIEF DESCRIPTION OF THE DRAWINGS Exemplary implementations of the subject matter described herein will now be explained with reference to the accompanying drawings, of which: Figure 1 is a block diagram illustrating one example of an architecture for a dual-stream deep network (DSDN) for prospective motion estimation of multi-slice images in ultrahigh resolution fMRI, with (1) a 2D convolutional neural network (CNN) for feature extraction of aliased multi-slice images, (2) a 3D CNN for reference image feature learning, and (3) motion estimation of six motion parameters; Figure 2 is a block diagram illustrating an exemplary process for synthesizing motion-affected aliased multi-slice images. The motion parameters are used as ground truth to train and test the AI model illustrated in Figure 1; Figures 3A and 3B are graphs illustrating training loss in terms of root mean squared error (RMSE) changes. In Figure 3A, the y axis ranges from 0 to 2 and in Figure 3B, the y axis ranges from 0 to .5; Figure 4 is a graph comparing results of the DSDN and its three variants (with different architectures) in terms of RMSE; Figure 5 includes graphs illustrating a visualization of motion tracking on a test dataset with motion parameters estimated by the DSDN; Figure 6 is a block diagram illustrating another example architecture for a three-branch deep neural network for prospective motion estimation and correction of fMRI images where the reference image is a T1-weighted MR image and the neural network includes 3 branches – one for learning features of the T1-weighted reference images, one for learning features of aliased simultaneous multi-slice images, and one for learning features of single-slice fMRI images; Figures 7A and 7B are graphs of training error for the architecture illustrated in Figure 6; Figure 8 is a block diagram illustrating a computing platform including a deep neural network trained to predict motion parameters and perform prospective motion correction of fMRI images; and Figure 9 is a flow chart illustrating an exemplary process for prospective motion estimation and correction of fMRI images. DETAILED DESCRIPTION Ultrahigh-resolution functional MRI (fMRI) remains unparalleled to date in spatiotemporal resolution and rich information for examining neuronal activity and functional connectivity in the brain; but as spatial resolution improves, its sensitivity to head motion usually increases. Traditional retrospective motion correction methods can correct slow inter-volume motion but cannot handle rapid intra-volume head movement in ultrahigh-resolution fMRI that may disrupt spin history and cause slice misalignment. Prospective motion correction (PMC) offers a promising solution to maintain a fixed relationship between subject and imaging geometry, thus reducing the level of residual motion during image acquisition. However, current PMC methods typically require 300∼500 ms for image reconstruction and co-registration, which is insufficient to handle fast head movement. To address this issue, we propose a dual-stream deep network (DSDN) that performs motion estimation directly from aliased simultaneous multi-slice (SMS) images, avoiding time- consuming image reconstruction and co-registration procedures. Specifically, the DSDN is designed to estimate six head motion parameters in terms of translation and rotation. It consists of two parallel branches, one of which handles the aliased SMS image stream, and the other inputs a 3D reference volume. Experimental results on a total of 12,400 synthetic aliased SMS images suggest that our method can achieve motion estimation accuracy below 0.2 voxel / degree, which is deemed adequate for reliable head motion tracking. Additionally, with the capability of fast forward computing capabilities of DSDN, we can estimate head motion in milliseconds and then update the field of view (FoV) with each radio frequency shot, thus facilitating real-time motion correction. This approach enables the direct estimation of motion parameters, bypassing traditional reconstruction and co-registration processes which can streamline the fMRI motion correction workflow. Experimental validation on synthetic data simulating real-world head motion during fMRI acquisition demonstrates that our model achieves impressive results, with a marked reduction in Root Mean Squared Error (RMSE) indicative of precise motion parameter estimation. This study paves the way for real-time, robust perspective motion estimation and correction in fast functional MRI (with multi- channel aliased images), promising to enhance image quality and the statistical significance of diagnostic findings. Introduction Ultrahigh-resolution functional MRI (e.g., with less than 1 mm resolution) plays an essential role in brain research by providing unparalleled insights into the neural basis of cognition, behavior, and neurological disorders. However, the increased spatial resolution comes with sensitivity to subject motion, which may severely degrade the image quality and undermine the statistical significance. Traditional approaches often involve post- processing steps that attempt to adjust for motion artifacts after the data acquisition phase. However, these techniques are inadequate in dealing with quick head movements. Rapid movements can disrupt the magnetic spin history and cause misalignment between slices, leading to significant data loss and inaccuracies. Recognizing these challenges, the field has seen a growing interest in prospective motion correction (PMC) techniques. PMC aims to adjust the scanning parameters in real time. This approach holds the promise of significantly reducing motion artifacts by adjusting for head movements. Despite its potential, current PMC methods are hindered by their relatively slow response times, typically requiring 300 to 500 milliseconds for image reconstruction and co-registration. This delay renders these methods ineffective for quick head movements. To tackle these limitations, our study introduces an approach that utilizes artificial intelligence (AI) to enhance the speed and accuracy of perspective motion estimation in ultrahigh-resolution fMRI. Specifically, we design a dual-stream deep network (DSDN) that replaces the traditional steps of image reconstruction and co-registration. It innovatively performs motion estimation directly from aliased simultaneous multi-slice (SMS) images. One branch of our network processes the aliased SMS images, while the other works with a 3D reference volume, facilitating rapid motion estimation that can keep pace with every radio frequency shot. Experiments on a total of 12,400 synthetic aliased SMS images indicate that this approach can achieve motion estimation accuracy below 0.2 voxel / degree, which is deemed reliable for head motion tracking. By enabling real-time motion correction, our method holds the potential to significantly enhance the quality and statistical significance of fMRI data. We conduct experiments to validate the effectiveness of our method through rigorous testing on synthetic datasets, which are designed to emulate realistic head movements. The results show a significant reduction in root mean squared error (RMSE) and confirm the model’s ability to accurately estimate subject head motions. These promising outcomes validate our model as a robust solution for real-time motion correction, potentially transforming the landscape of fast brain MR imaging. Methodology Problem Formulation In functional magnetic resonance imaging, patient motion is a prevalent source of artifacts that degrade the quality of acquired images. The challenge addressed herein is the formulation of a prospective motion estimation approach, which, rather than correcting the motion after image acquisition, aims to estimate motion parameters six coordinates in terms of rotation and shift) during the MRI scan to adjust for patient movements. Let ^ denote the reference 3D MR image obtained in the absence of motion. A sequence of moving images where ^ indicates the sequence index, is acquired with potential motion-induced aliasing, noted as^^^. The motionpresent during the scan is represented by a set of parameters ^^ ൌ with ^^^and ^^^specifying the rotation and shift vectors, respectively, for each image in the sequence. The objective is to estimate the six motion parameters with high precision. This is formalized as an optimization problem where the estimated parameters ^^^minimize the discrepancy between the acquired moving image and the reference image. The optimization can be expressed as: ^^^ ൌ argmin  ^^^, ^^; ^^,^ where ^ is a loss function that quantifies the difference between the reference image ^ and the moving image ^^under the motion parameters ^.In this work, we develop a deep neural network ^ to estimate ^^^ as: where ^ represents the learnable parameters of the network. Throughaccurate motion parameter estimation, we aim to enhance the quality and speed of fMRI scanning, thereby improving the accuracy of subsequent MRI analysis. Overall Framework Figure 1 illustrates the architecture of our dual-stream deep network (DSDN) for perspective motion estimation. Its architecture is designed to estimate the six motion parameters (3 for rotation and 3 for shift) of aliased multi-slice fMRIs. To achieve that, DSDN involves two parallel branches for MRI feature learning, specifically tailored for multi-channel (coils) aliased images and a 3D reference image (i.e., T2-weighted MRI). Features learned by the two branches are then concatenated and fused to estimate motion parameters. (1) Multi-Channel Aliased fMRI Processing. The first branch of DSDN is dedicated to processing multi-channel aliased fMRIs. It consists of seven sequential 2D convolution blocks, each comprising a convolutional layer (Conv), batch normalization (BN), Rectified Linear Unit (ReLU) activation, and max pooling. Formally, each block can be represented as: Block^^^ ൌ MaxPool ^ReLU^^ି^^ଶ^ ൬BN ^Convଶ^൫Input൯^^^, where Input^^ି^^denotes the input to the ^-th block. The inputs to the first block are the multi-channel aliased images with 32 channels (coils). After the seventh convolution block, a global average pooling (GAP) operation is applied, resulting in a 128-dimensional feature vector. (2) 3D Reference MRI Processing. The second branch parallels the first in structure but is adapted for 3D T2-weighted MRI which is used as the reference (without motion). It includes seven 3D convolution blocks, each with a similar composition of Conv, BN, ReLU, and max pooling, utilizing 3D operations. Analogous to the first branch, the blocks are defined as: Block^^^^^ି^^ଷ^ ൌ MaxPool൯^^^. The input to this branch is the 3D reference MRI. Following the final 3D Conv block, a GAP is employed, yielding another 128-dimensional feature vector. (3) Motion Estimation. The extracted features from both branches are concatenated to form a 256-dimensional vector. This combined feature vector is then fed into two fully connected (FC) layers for feature abstraction and motion estimation. The FC layers transform the feature vector from adimensionality of 128 ൈ 2 to 128, and finally to 6-dimensional outputcorresponding to the motion parameters (i.e., three coordinates for shift and three for rotation) as: Motion Parameters ൌ FC^൫FC^ଶ଼^Concatenated Features^൯Attention Mechanism To enhance the model’s focus on relevant features within the images, an attention mechanism is incorporated after the first convolutional block in each branch. The attention layers generate spatial attention maps which guide the network to emphasize salient regions during feature extraction. The attention mechanism is formalized as: Attention Map^^^ ൌ Attention൫OutputBlock^^^൯ where OutputBlock^^^is the output from the ^-th convolutional block, which, after applying the attention layer, produces the attention map Attention Map^^^. This architecture aims to leverage both local and global features from the alias images and the 3D volume, enhancing the network’s capacity to estimate motion parameters accurately. Objective Function The training of the DSDN is formulated as a regression problem where the Mean Squared Error (MSE) loss function is employed to quantify the discrepancy between the estimated motion parameters and the ground truth. Mathematically, the MSE loss is formulated as: where ^ is the number of training samples, ^^^ is the vector of estimatedmotion parameters for the ^ -th sample, and ^^ is the corresponding 6-dimensional vector of ground-truth parameters. For training, note that, the input consists of a 2D multi-channel aliased MRI (with motion) and the 3D reference MRI. Inference for Unseen Aliased fMRIs with Motion Upon completion of the training phase, the DSDN is utilized to perform inference on unseen MRIs exhibiting motion. For each inference operation, it takes as input a pair of images: a multi-channel aliased MRI with motion and the reference MRI without motion. Based on these inputs, the network outputs the estimated motion parameters. Estimated Motion Parameters ൌ Network^MRI-Motion,MRI-Reference^Implementation The proposed DSDN is implemented using PyTorch. The Adam algorithm is used as the optimizer with a learning rate of 0.0001 and a batch size of 8. (1) During training, all the training data are shuffled before feeding to the network. The network undergoes training for 145 epochs, which is found to be adequate for convergence. This process is executed on a workstation equipped with an Intel i-7 CPU and a GPU with 12 GB of memory. (2) During inference, the test time for processing each image pair (with the same hardware configuration used for model training) is approximately 0.5 milliseconds, demonstrating the model’s potential for real-time application scenarios. Experiment Materials To facilitate training and performance validation of our AI model for perspective motion estimation, we synthesize motion-affected aliased multi- slice images using our in-house MATLAB code, as illustrated in Figure 2. This process begins with a fully-sampled 3D reference image. With the Berkeley Advanced Reconstruction Toolbox (BART), we generate the sensitivity profile from a multi-channel 3D reference image. From this reference, a series of aliased simultaneous multi-slice (SMS) images with induced motions were generated. These motions incorporate both shift and rotation. In our simulations, the translation / shift and rotations range from -10 to 10 voxels / degrees. In this way, we have a total of 9,920 (320 volumes, 31 groups of aliased SMS images for each volume) aliased images for training, and 2,480 (80 volumes, 31 groups of aliased SMS images for each volume) aliased images for inference / test. The experimental dataset was initiated with a standard MRI image serving as the reference. From this reference, a series of MRI images with induced motions were generated. These motions, incorporating both shifts and rotations, were applied to simulate realistic patient movement scenarios. For each simulated MRI with motion, the corresponding motion parameters were recorded, serving as the ground truth for the network’s training and evaluation. Experimental Settings The primary goal is to assess the network’s ability to accurately predict the motion parameters, i.e., shifts (in the x-, y-, and z-axis) and rotations (in the x-, y-, and z-axis), that were artificially introduced in the simulated aliased multi-slice fMRI data. During training and testing, the input is comprised of aliased MRIs, which can result from sub-sampling in k-space, against the volume reference MRI. For performance evaluation, we use the Root Mean Square Error (RMSE) and Euclidean distance as the metrics. A smaller value of RMSE or ED indicates better performance. Competing Methods Given the scarcity of AI research on perspective motion estimation for aliased multi-slice fMRIs, we develop several variants of the proposed DSDN and compare their performance on the same test dataset. These methods include: 1) DSDN-2000 trained with 2,480 randomly selected alias SMS images, 2) DSDN-600 trained with 620 randomly selected alias SMS images, and 3) DSDN-5CB with five 2D convolution blocks and five 3D convolutional blocks for feature extraction. For a fair comparison, all these methods have the same training settings (e.g., batch size, optimizer, training epochs). Convergence Analysis We employ the MSE as the loss function for training the deep learning network. The overall training session spanned over three days, encompassing 145 epochs. The variations in the loss throughout the training period are illustrated in Figures 3A and 3B. Figures 3A and 3B demonstrate a gradual convergence of the network, with the final loss value being reduced to less than 0.1. This steady decrease in loss indicates the network’s improving accuracy in predicting the target variables, highlighting the effectiveness of our training strategy. Table 1: Euclidean distance results of perspective motion estimation (in terms of meanേstandard deviation) achieved by four different methods on the test set with 2,480 SMS images. ^ Prediction Result After training, we apply the DSDN (as well as the other competing methods) to predict motion parameters of test data (2,480 aliased images). We determined the Euclidean distance between each sample’s predicted motion parameters and their corresponding ground truth values within the test dataset. Subsequently, we calculated the average and standard deviation of these distances, as reported in Table 1. From the result achieved by the proposed AI model, the accuracy of motion estimation, both in translations and rotations, remains under 0.2 voxel / degree. Given that such accuracy is sufficient for stable motion tracking, the preliminary results suggest that the AI-driven perspective motion estimation approach is promising. Comparative Study We compare the perspective motion estimation performance (in terms of RMSE) of the proposed method and the other competing methods. The comparative results are shown in Figure 4. From the results, we have the following observations. 1) The proposed method trained with 9,920 images achieves the best performance (the smallest RMSE in most cases), while the methods trained with fewer data have inferior performance. Notably, when the training set only has 620 SMS images, the performance is significantly worse, and the motion estimation fails.2) Our model, which consists of 7 convolution blocks, obtains better performance than the model (i.e., DSDN-5CB) with 5 convolution blocks. The results indicate that maintaining enough training data (i.e., SMS images) and a suitable network architecture are essential factors in achieving promising performance for the perspective motion estimation task. Motion Tracking Analysis To enhance the understanding of the motion estimation capabilities of the proposed DSDN model, we construct a scatter plot accompanied by a regression line for all parameters (i.e., three shift parameters and three rotation parameters). In each plot, it positions the ground truth values of a motion parameter on the x-axis against the predicted values of that same parameter on the y-axis. A regression line, as depicted in each of the graphs in Figure 5, is then drawn to represent the accuracy of the model’s estimations, offering a visual measure of the relationship between the predicted and actual motion parameters. Based on the result, it is evident that the trained deep network is good at identifying the motion patterns of previously unseen samples within the test dataset. This demonstrates the profound capability of deep learning for estimating perspective motion in ultra-high-resolution fMRI, confirming its efficacy and robustness in handling complex motion estimation tasks. Conclusions and Future Work The subject matter described herein includes an AI approach, i.e., dual- stream deep network (DSDN), for direct perspective motion estimation in aliased simultaneous multislice fMRI, eliminating the need for time-consuming image reconstruction and co-registration in conventional motion correction frameworks. By leveraging a dual-stream deep network architecture, our method accurately estimates the motion parameters from aliased fMRIs, offering significant improvements in processing time and computational efficiency. Experimental results suggest the precision and efficiency of DSDN, highlighting its potential to enhance real-time ultrahigh-resolution fMRI acquisition applications. In this work, we rely on T2-weighted MRI without motion as the reference image. It is interesting to explore the usage of different types of reference volume (e.g., T1-weighted MRI) to test the effectiveness of the proposed AI model, which will be our future work. In addition, to further explore the generalizability and robustness of the proposed AI model, we plan to generate more simulated SMS images with more complex motion patterns for performance testing. Furthermore, we will explore the usage of more advanced attention mechanisms to further enhance the efficiency of model training and the performance of the proposed method in prospective motion estimation. Alternate Deep Neural Network Architecture and Overall Methodology Figure 6 is a block diagram illustrating another example architecture for a deep neural network for prospective motion estimation and correction of fMRI images where the reference image is a T1-weighted MRI image and the neural network includes 3 branches – one for learning features of the T1-weighted reference images, one for learning features of aliased simultaneous multi-slice images, and one for learning features of single-slice fMRI images. Figures 7A and 7B are graphs of training error for the architecture illustrated in Figure 6. In Figure 6, the middle and lower branches are used as reference images. More particularly, the images from the middle branch are used to learn the differences between the FoV of the T1-weighted image and the FoV of the fMRI scanner. The upper branch is used to extract features from the multi- slice SMS images for which motion parameters are being predicted. The middle branch's input is the very first slice acquired in a session. Since our goal is to align the images over time, we do not need to perform motion correction on this initial image. From the second image onward, motion correction is performed with respect to the first image. Figure 8 is a block diagram illustrating a computing platform including a neural network trained to predict motion parameters suitable for performing prospective motion correction of fMRI images. Referring to Figure 8, a computing platform 800 includes at least one processor 802 and memory 804. Computing platform 800 further includes a trained multi-branch feature extractor CNN 806 that receives fMRI images and a reference image from an fMRI scanner 805 and generates feature vectors as an output. Trained multi- branch feature extractor CNN 806 may be of the design illustrated in Figure 1 or Figure 6. Computing platform 800 further includes a trained motion estimator 808 that estimates motion parameters of the subject from the fMRI images and the reference image. Trained motion estimator 808 may also be of the design illustrated in Figure 1 or Figure 6. Trained motion estimator 808 may output the values of the motion parameters to an fMRI image acquisition controller 810 of fMRI scanner 805, which adjusts the fields of view of the MRI coils of fMRI scanner 805 based on the estimated motion parameters. Trained multi-branch feature extractor CNN 806 and trained motion estimator 808 may be implemented using computer executable instructions stored in memory 804 and executed by processor 802. Figure 9 is a flow chart illustrating an exemplary process for prospective motion estimation and correction of fMRI images. Referring to Figure 9, in step 900, the process includes inputting aliased simultaneous multi-slice (SMS) fMRI images into a first branch of a convolutional neural network (CNN) trained to generate feature vectors from the SMS fMRI images. For example, aliased SMS images of a subject may be input to one branch of a multi-branch CNN, such as that illustrated in Figure 1 or Figure 6, that is trained to extract features from the images. In step 902, the process further includes inputting a reference image into a second branch of the CNN trained to generate a feature vector from the reference image. For example, a T1-weighted or T2-weighted MR reference image may be input into another branch of a multi-branch CNN, such as that illustrated in Figure 1 or Figure 6. In step 904, the process further includes inputting the feature vectors into a motion estimator trained to predict motion parameters of the SMS fMRI images relative to the reference image and receiving, as output, values for the predicted motion parameters of the SMS fMRI images. For example, the feature vectors generated by the feature extraction branches may be input into a motion estimation branch, which predicts motion parameters that indicate shifts and rotations in the aliased SMS fMRI images relative to the reference image. In step 906, the process further includes outputting the values of the predicted motion parameters to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to obtain the fMRI scan. For example, the motion estimator may output the values of the predicted motion parameters to an fMRI image acquisition controller of an fMRI scanner, which adjusts the fields of view of magnetic coils used to obtain the aliased SMS images such that the coil fields of view are shifted and rotated by the amounts indicated by the values of the motion parameters. The disclosure of each of the following references is hereby incorporated herein by reference in its entirety. References 1. Ugurbil, K.: Ultrahigh field and ultrahigh resolution fMRI. Current Opinion in Biomedical Engineering 18 (2021) 100288 2. Feinberg, D.A., Beckett, A.J., Vu, A.T., Stockmann, J., Huber, L., Ma, S., Ahn, S., Setsompop, K., Cao, X., Park, S., et al.: Next-generation MRI scanner designed for ultra-high-resolution human brain imaging at 7 Tesla. Nature Methods 20(12) (2023) 2048–2057 3. Zaitsev, M., Akin, B., LeVan, P., Knowles, B.R.: Prospective motion correction in functional MRI. Neuroimage 154 (2017) 33–42 4. Thesen, S., Heid, O., Mueller, E., Schad, L.R.: Prospective acquisition correction for head motion with image-based tracking for real-time fMRI. Magnetic Resonance in Medicine: An Official Journal of the International Society for Magnetic Resonance in Medicine 44(3) (2000) 457–465 5. Hoinkiss, D.C., Porter, D.A.: Prospective motion correction in 2D multishot MRI using EPI navigators and multislice-to-volume image registration. Magnetic Resonance in Medicine 78(6) (2017) 2127–2135 6. LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature 521(7553) (2015) 436–444 7. Chen, X., Wang, X., Zhang, K., Fung, K.M., Thai, T.C., Moore, K., Mannel, R.S., Liu, H., Zheng, B., Qiu, Y.: Recent advances and clinical applications of deep learning in medical image analysis. Medical Image Analysis 79 (2022) 102444 8. Gore, J.C.: Artificial intelligence in medical imaging (2020) 9. Barth, M., Breuer, F., Koopmans, P.J., Norris, D.G., Poser, B.A.: Simultaneous multislice (SMS) imaging techniques. Magnetic Resonance in Medicine 75(1) (2016) 63–81 10. Risk, B.B., Kociuba, M.C., Rowe, D.B.: Impacts of simultaneous multislice acquisition on sensitivity and specificity in fmri. NeuroImage 172 (2018) 538–553 11. Uecker, M., Lai, P., Murphy, M.J., Virtue, P., Elad, M., Pauly, J.M., Vasanawala, S.S., Lustig, M.: ESPIRiT—an eigenvalue approach to autocalibrating parallel MRI: where SENSE meets GRAPPA. Magnetic Resonance in Medicine 71(3) (2014) 990–1001 12. Maclaren, J., Speck, O., Stucht, D., Schulze, P., Hennig, J., Zaitsev, M.: Navigator accuracy requirements for prospective motion correction. Magnetic Resonance in Medicine: An Official Journal of the International Society for Magnetic Resonance in Medicine 63(1) (2010) 162–170 13. Stucht, D., Danishad, K.A., Schulze, P., Godenschweger, F., Zaitsev, M., Speck, O.: Highest resolution in vivo human brain MRI using prospective motion correction. PLOS ONE 10(7) (2015) 1–17 14. Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR). (2018) 7132–7141 15. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations (ICLR). 16. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. Advances in Neural Information Processing Systems 30 (2017) It will be understood that various details of the subject matter described herein may be changed without departing from the scope of the subject matter described herein. Furthermore, the foregoing description is for the purpose of illustration only, and not for the purpose of limitation, as the subject matter described herein is defined by the claims as set forth hereinafter.

Claims

CLAIMS What is claimed is:

1. A method for prospective motion estimation and correction in functional magnetic resonance imaging (fMRI) using deep neural networks, the method comprising: inputting aliased simultaneous multi-slice (SMS) fMRI images into a first branch of a convolutional neural network (CNN) trained to generate a feature vector from the aliased SMS fMRI images and receiving, as output, a feature vector for the aliased SMS fMRI images; inputting a reference image into a second branch of the CNN trained to generate a feature vector from the reference image and receiving, as output, a feature vector for the reference image; inputting the feature vectors into a motion estimator trained to predict motion parameters of the SMS fMRI images relative to the reference image and receiving, as output, values for the predicted motion parameters of the aliased SMS fMRI images; and outputting the values for the predicted motion parameters to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to perform the fMRI scan.

2. The method of claim 1 wherein the first branch of the CNN includes a plurality of CNN blocks, each including a 2D convolution layer, a batch normalization layer, a rectilinear unit (ReLU) activation layer, and a max pooling layer.

3. The method of claim 1 wherein the second branch of the CNN includes a plurality of CNN blocks, each including a 3D convolution layer, a batch normalization layer, a rectilinear unit (ReLU) activation layer, and a max pooling layer.

4. The method of claim 1 wherein the motion estimator is configured to concatenate the feature vectors and reduce dimensionality of the concatenated feature vectors to produce the values of the predicted motion parameters.

5. The method of claim 1 wherein the values for the predicted motion parameters include predicted values of shift and rotation parameters of the aliased SMS fMRI images relative to the reference image.

6. The method of claim 5 comprising adjusting the fields of view of the magnetic coils by shifting and rotating the fields of view of the magnetic coils using the predicted values of the shift and rotation parameters.

7. The method of claim 1 wherein the reference image comprises a T2- weighted MR image.

8. The method of claim 1 wherein the reference image comprises a T1- weighted MR image.

9. The method of claim 1 comprising inputting aliased single slice fMRI images into a third branch of the convolutional neural network trained to generate a feature vector from the aliased single slice fMRI images.

10. The method of claim 9 wherein inputting the feature vectors into the motion estimator comprises inputting the feature vectors from the aliased single slice fMRI images into the motion estimator and the motion estimator is configured to predict motion parameters of the aliased single slice images relative to the reference image and receiving, as output, values for the predicted motion parameters of the SMS fMRI images.

11. A system for prospective motion estimation and correction in functional magnetic resonance imaging (fMRI) using neural networks, the system comprising: a computing platform including at least one processor and a memory; a trained multi-branch feature extractor convolutional neural network (CNN) implemented by the at least one processor and including a first branch for receiving, as input, aliased simultaneous multi-slice (SMS) fMRI images and generating, as output, a feature vector from the aliased SMS fMRI images and a second branch for receiving, as input, a reference image and generating, as output, a feature vector from the reference image; anda trained motion estimator for receiving the feature vectors as input and generating, as output, predicted values for motion parameters of the SMS fMRI images relative to the reference image, and providing the values of the predicted motion parameters of the SMS fMRI images to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to perform the fMRI scan.

12. The system of claim 11 wherein the first branch of the multi-branch feature extractor CNN includes a plurality of CNN blocks, each including a 2D convolution layer, a batch normalization layer, and a rectilinear unit (ReLU) activation layer, and a max pooling layer.

13. The system of claim 11 wherein the second branch of the multi-branch feature extractor CNN includes a plurality of CNN blocks, each including a 3D convolution layer, a batch normalization layer, and a rectilinear unit (ReLU) activation layer, and a max pooling layer.

14. The system of claim 11 wherein the trained motion estimator is configured to concatenate the feature vectors and reduce dimensionality of the concatenated feature vectors to produce the values of the predicted motion parameters.

15. The system of claim 11 wherein the values for the predicted motion parameters include predicted values of shift and rotation parameters of the aliased SMS fMRI images relative to the reference image.

16. The system of claim 11 wherein the fMRI image acquisition controller is configured to use the values of the predicted motion parameters to adjust the fields of view of the coils by shifting and rotating the fields of view of the coils using the predicted values of the shift and rotation parameters.

17. The system of claim 11 wherein the reference image comprises a T2- weighted MR image.

18. The system of claim 11 wherein the reference image comprises a T1- weighted MR image.

19. The system of claim 11 wherein the trained feature extractor CNN comprises a third branch for receiving aliased single slice fMRI imagesand generating a feature vector from the aliased single slice fMRI images.

20. A non-transitory computer readable medium having stored thereon computer-executable instructions that when executed by a processor of a computer control the computer to perform steps comprising: inputting aliased simultaneous multi-slice (SMS) functional magnetic resonance imaging (fMRI) images into a first branch of a convolutional neural network (CNN) trained to generate a feature vector from the aliased SMS fMRI images and receiving, as output, a feature vector for the aliased SMS fMRI images; inputting a reference image into a second branch of the CNN trained to generate a feature vector from the reference image and receiving, as output, a feature vector for the reference image; inputting the feature vectors into a motion estimator trained to predict motion parameters of the SMS fMRI images relative to the reference image and receiving, as output, values for the predicted motion parameters of the aliased SMS fMRI images; and outputting the values for the predicted motion parameters to an fMRI image acquisition controller for adjusting, during an fMRI scan, fields of view of magnetic coils used to perform the fMRI scan.

Citation Information

Patent Citations

  • Apparatus and method for deep learning to mitigate artifacts arising in simultaneous multi slice (SMS) magnetic resonance imaging (MRI)

    US20200300954A1

  • Real-time motion monitoring using deep neural network

    US20210057083A1

  • Multi-slice magnetic resonance imaging method and device based on long-distance attention model reconstruction

    US20230097417A1

  • Simultaneous Multi-Slice Protocol and Deep-Learning Magnetic Resonance Imaging Reconstruction

    US20230298230A1

  • Detecting motion artifacts from k-space data in segmentedmagnetic resonance imaging

    US20230337987A1