Motion artifact self-adaptive correction method based on multi-mode MRI (Magnetic Resonance Imaging) fusion
Through the motion artifact adaptive correction method of multimodal MRI fusion, the artifact problem caused by complex head movements in MRI images is solved, high-precision image correction and feature fusion are achieved, and image quality and reliability of clinical diagnosis are improved.
Patent Information
- Application Number
- CN202510780671.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing MRI motion artifact correction technology cannot fully cope with various complex head movements, resulting in the lack of information between image modes and insufficient utilization of features, especially in structural magnetic resonance images and functional magnetic resonance images, which affects clinical diagnosis and research.
Adaptive motion artifact correction method based on multimodal MRI fusion is adopted to pre-process structural magnetic resonance images and functional magnetic resonance images through a cross-modal feature fusion model, feature interaction fusion and multi-scale reconstruction, combining dynamic spatiotemporal attention correction and three-level loss function optimization to achieve artifact detection and correction.
It improves the registration accuracy and data quality of MRI images, significantly improves the feature representation ability of multimodal data, ensures high-fidelity image output, and overcomes the problem of artifact correction in complex head movements.
Smart Images

Figure CN120298448A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing heat, and particularly relates to a method for adaptively correcting motion artifacts based on multi-modal MRI fusion. Background Art
[0002] In magnetic resonance images, head motion artifacts are a common and unavoidable problem, especially in structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI), where head motion artifacts are more significant. In severe cases, it may lead to the inability to effectively extract the structural and functional information of the brain, thus affecting clinical diagnosis and research.
[0003] Current MRI motion artifact correction techniques mainly detect and compensate for motion artifacts based on K-space or use deep learning to remove artifacts in the image domain. Although an inter-domain data consistency loss is adopted to maintain the consistency between K-space and the spatial domain, this approach still faces certain challenges. Especially in the case of severe artifacts, how to achieve a high degree of consistency without losing image details remains a difficult point. Existing technologies usually perform motion artifact simulation in the preliminary stage and perform correction through frequency domain and spatial domain generators, but these methods cannot fully handle various complex head motion situations in clinical practice, resulting in insufficient robustness and adaptability of the trained model in actual applications, and ultimately may cause residual artifacts or image blurring.
[0004] In addition, most current technologies focus on single-modal MRI images, ignoring the correlation between modalities, resulting in missing information dimensions (sMRI lacks temporal dynamic information, fMRI lacks spatial high-frequency information), and insufficient feature utilization.
[0005] Therefore, how to improve the MRI motion artifact correction process in the prior art to overcome the disadvantage of being unable to fully handle various complex head motion situations in the prior art and solve the technical problems of missing information between image modalities and insufficient feature utilization. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for adaptively correcting motion artifacts based on multi-modal MRI fusion to overcome the disadvantage of being unable to fully handle various complex head motion situations in the prior art and solve the problems of missing information between image modalities and insufficient feature utilization.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for adaptively correcting motion artifacts based on multi-modal MRI fusion, comprising the following steps: S1: Obtain the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI of the patient, and preprocess the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI; S2: Create a cross-modal feature fusion model, and input the preprocessed structural magnetic resonance image (sMRI) and functional magnetic resonance image (fMRI) into the cross-modal feature fusion model for dual-branch feature extraction; S3: The cross-modal feature fusion model performs cross-modal feature interaction and fusion on the extracted dual-branch features, and conducts multi-scale feature reconstruction; S4: Detect image artifacts based on the reconstructed multi-scale features, and perform dynamic spatio-temporal attention correction on the detected image artifacts; S5: Establish a three-level loss function system for multi-task joint optimization.
[0008] Preferably, in step S1, the preprocessing of the structural magnetic resonance image (sMRI) and functional magnetic resonance image (fMRI) is implemented by a multi-modal data preprocessing module. The multi-modal data preprocessing module includes a motion artifact simulation unit and a multi-modal registration unit. The specific process of the preprocessing is as follows: S11: Apply an elastic deformation field with a specified displacement and rotation angle to the structural magnetic resonance image (sMRI) through the motion artifact simulation unit; S12: Use time-varying motion parameters for the functional magnetic resonance image (fMRI), and perform image enhancement processing by randomly rotating a specified angle, scaling by a specified multiple, and injecting Rician noise; S13: Implement a hierarchical registration strategy through the multi-modal registration unit to achieve the registration of the structural magnetic resonance image (sMRI) and functional magnetic resonance image (fMRI).
[0009] Preferably, the specific process of implementing the hierarchical registration strategy through the multi-modal registration unit in step S13 is as follows: S131: Perform normalization processing on the structural magnetic resonance image (sMRI) and functional magnetic resonance image (fMRI) to unify their gray-scale ranges, resolutions, and spatial dimensions; S132: Segment the structural magnetic resonance image (sMRI) and functional magnetic resonance image (fMRI), extract key anatomical structures as regions of interest, and perform morphological processing on the regions of interest to provide a specified number of registration feature points; S133: Define six-degree-of-freedom registration parameters. The six-degree-of-freedom registration parameters include three translation parameters, namely the translation amounts along the x, y, and z axes, and three rotation parameters, namely the Euler angles around the x, y, and z axes; S134: Use mutual information or normalized mutual information to measure the statistical dependence relationship between images, and use mean square error or normalized cross-correlation processing for same-modal images; S135: Set the initial transformation of the structural magnetic resonance image (sMRI) and the functional magnetic resonance image (fMRI) to the identity transformation, downsample the images to 1 / 4 or 1 / 8 of the original size, and quickly obtain rough registration parameters to avoid falling into local optima. S136: Gradually restore to the original resolution, use the low-resolution results as initial values, refine the registration parameters, select at least ten pairs of landmark points in the region of interest, such as blood vessel bifurcation points and cerebral sulcus endpoints, calculate their spatial distance errors and direction deviations, and determine whether the anatomical structures are aligned. S137: Construct a deformation field using cubic B-spline interpolation, set the grid node spacing to 2 mm, ensure that local details are retained while avoiding overfitting, add a bending energy penalty term for deformation field regularization, constrain the smoothness of the deformation, prevent non-physiological deformation, continue the mutual information of the rigid registration stage, or introduce local mutual information to enhance the registration accuracy of local structures.
[0010] Preferably, the specific process of the dual-branch feature extraction by the cross-modal feature fusion model in step S2 is as follows: S21: Convolve the structural magnetic resonance image (sMRI) through the initial convolutional layer of the first branch of the cross-modal feature fusion model. The specific convolution calculation formula is as follows: ; where is the input structural magnetic resonance image (sMRI), stride is the step size, and 7×7×7 is the convolution kernel size; represents the feature map obtained after processing by the initial convolutional layer of the first branch, which is the output result of the convolution operation on the structural magnetic resonance image (sMRI). represents a three-dimensional convolution operation (3D Convolution) with a convolution kernel size of 7×7×7 (in three-dimensional space, each dimension has a size of 7), and stride (step size) is 2, that is, the convolution kernel moves 2 pixels each time when sliding on the image (in each dimension of the three-dimensional space); S22: Process through the residual block hierarchy of the first branch of the cross-modal feature fusion model. The specific calculation formula is as follows: ; where is the feature map processed by the previous (l−1)-th residual block and serves as the input to the current l-th residual block; R l is an improved 3D residual block with channel numbers {64, 128, 256, 512} in sequence, and the output feature resolution is {8 mm 3 , 4 mm 3 , 2 mm 3 , 1 mm 3}}。
[0011] S23: Convolve the functional magnetic resonance image fMRI through the 3D convolutional layer of the second branch of the cross-modal feature fusion model. The specific calculation formula is as follows: ; where, represents the feature map obtained after being processed by the 3D convolutional layer of the second branch and is the output result of the convolutional operation on the functional magnetic resonance image (fMRI); represents a three-dimensional convolutional operation. The size of the convolutional kernel is 3×3×3, that is, the size of the convolutional kernel in each dimension of the three-dimensional space is 3, and the default stride is 1. is the input functional magnetic resonance image data. Here, t:t + 9 means selecting a series of fMRI images from time t to time t + 9 in the time dimension as the input to capture the dynamic features of the functional magnetic resonance image over a period of time.
[0012] S24: Perform LSTM temporal modeling. The calculation formula is as follows: ; where, represents the fMRI feature representation of the functional magnetic resonance image obtained after being processed by the long short-term memory network (LSTM) at time t; represents an LSTM layer with 128 hidden units. The number of its hidden units determines the memory capacity and processing ability of the LSTM layer. 128 hidden units means that this layer can maintain a 128-dimensional hidden state when processing sequence data; is the feature sequence input to the LSTM layer. It is composed of the feature maps of 10 time steps from time t - 9 to time t in the time dimension of the feature map output by the previous 3D convolutional layer, and is used to capture the dynamic change information of the fMRI image over a period of time.
[0013] S25: Perform hierarchical feature alignment on the outputs of the two branches, and perform channel dimension concatenation at 5 scales. The specific formula is as follows: ; where, is the sMRI feature map of the k th layer, is the fMRI feature map of the k th layer, is the output concatenated feature, which represents the feature map obtained by concatenating the channels at a certain scale after hierarchical feature alignment at time t; It is a channel dimension concatenation operation, that is, two feature maps with the same spatial dimensions (height, width, depth, here in the three-dimensional image scenario) but different numbers of channels are concatenated in the channel dimension; It represents values from 1 to 5, representing 5 different scales.
[0014] S26: Generate dynamic attention, and the specific process is as follows: ; Among them, , represents the channel-adaptive attention map generated at the k-th scale, and its element values are between [0,1]. The larger the value, the more important the features at the corresponding position or channel; σ is the Sigmoid function, and the formula is It maps the input value to the interval [0,1] for generating attention weights; is the convolution kernel weight (convolution kernel), with a shape of 1×1×1×( C sMRI + C fMRI )×1; C sMRI and C fMRI are respectively and 's number of channels. Through this convolution kernel and the concatenated feature map perform a convolution operation to calculate the attention weights.
[0015] Preferably, the cross-modal feature fusion model in step S3 performs cross-modal feature interaction fusion on the extracted dual-branch features and the specific process of multi-scale feature reconstruction is as follows: S31: Perform weighted feature fusion of the dual branches, and the specific formula is as follows: ; Among them, ⊙ represents element-wise multiplication, is the output fusion feature, , achieving the dynamic balance of anatomical and functional features; is the channel-adaptive attention map generated at the k-th scale, and its element values are between [0,1], used to measure F sMRI the importance of features.
[0016] S32: Perform multi-scale feature reconstruction, inverse pyramid aggregation, and finally output a resolution of 1mm³, and the specific formula is as follows: ; Among them, , Wk is the kernel parameter. Among them, represents the output feature after multi-scale feature reconstruction and inverse pyramid aggregation; is the k -th fused feature map; Here, the features fused at 5 different scales from k = 1 to k = 5 are U k calculated ([[]] k = 1, resolution 8mm 3 , k = 5, resolution 1mm 3 ); U k is a 3×3×3 transposed convolution upsampling operation.
[0017] Preferably, in step S4, based on the reconstructed multi-scale features, image artifacts are detected, and the specific process of dynamically correcting the spatio-temporal attention for the detected image artifacts is as follows: S41: Calculate the gradient magnitude through a Sobel operator with a kernel size of 3×3×3, perform bilinear interpolation processing, and generate a spatial attention map. The calculation formula is: ; Among them, F is a three-dimensional feature map; represents the generated spatial attention map, which is used to indicate the importance of different spatial positions in the image for artifact detection and correction. Norm is a normalization operation, and its purpose is to map the calculated values to the appropriate range [0, 1] interval, which is convenient for subsequent processing and as an attention weight; , respectively represent the gradients of the feature map F in the x direction and the z direction.
[0018] S42: Generate temporal attention based on LSTM motion modeling. The specific formula is as follows: ; ; Among them, are the 6-degree-of-freedom head motion parameters respectively, is the output time-varying attention weight; represents a long short-term memory network (LSTM) layer with 64 hidden units; m t is a feature vector, which clarifies that the unit of the displacement change amount Δxt is millimeter (mm), and the unit of the rotation angle is radian (rad). This vector is used to be input into the LSTM to characterize the motion information.
[0019] S43: The generator model based on the U-Net3+ architecture configures multi-scale skip connections (5 scales) and a temporal convolutional branch (kernel size 3×3×3), and outputs the corrected structural magnetic resonance image sMRI and functional magnetic resonance image fMRI. The module can adaptively identify the spatio-temporal distribution characteristics of head motion artifacts; S44: Perform multi-scale feature skip connections, ; Among them, represents the feature map obtained after the multi-scale feature skip connection operation at the k-th scale; represents a three-dimensional transposed convolution (also called a transposed convolution) operation with a convolution kernel size of 1×1×1; Concat is a concatenation operation, which here concatenates different feature maps in the channel dimension or other appropriate dimensions to fuse various feature information; E k is the feature map at the k-th scale in the encoder path, T k is the feature map related to time (from the temporal convolutional branch) at the k-th scale, D k+1 is the feature map at the (k + 1)-th scale in the decoder path; U is upsampling by a factor of 2, k Aggregate multi-scale features reversely from 5 to 1.
[0020] Preferably, after performing the multi-scale feature skip connection in S44, the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI are also reconstructed; The reconstruction formula for the structural magnetic resonance image sMRI is as follows: ; Among them, represents the reconstructed structural magnetic resonance image, which is the corrected sMRI image output by the model; S1 is the multi-scale fusion feature with a resolution of 1 mm³ and the number of channels C = 64, w s is the convolution kernel, b s is the bias, tanh is the activation function, and tanh constrains the output value range to [−1, 1] to adapt to the MRI signal range and outputs the corrected sMRI image; The reconstruction formula for the functional magnetic resonance image fMRI is as follows: ; Among them, represents the reconstructed functional magnetic resonance image, which is the corrected fMRI image output by the model; Denotes a summation operation, which accumulates the relevant calculation results for the 5 time steps from time t−5 to time t−1; is the feature map at time point τ, wf is the convolution kernel, and Wτ is the time-varying weight, satisfying ∑τWτ = 1; is the bias term, which is used to adjust the result of the convolution operation and increase the flexibility of the model fitting.
[0021] Preferably, the specific process of establishing a three-level loss function system for multi-task joint optimization in step S5 is as follows: S51: Calculate the anatomical fidelity loss, and the specific formula is as follows: Among them, denotes the anatomical fidelity loss, which is used to measure the similarity in anatomical structure between the reconstructed image (the corrected image) and the real image. The smaller the value, the closer the reconstructed image is to the real anatomical structure; calculates the structural similarity between the corrected image Icor and the real image Igt, and its value range is [0,1]. The closer the value is to 1, the more similar the structure is; is to calculate the corrected image I cor and the real image I gt The L1 norm of the gradient. ∇ represents the gradient operator, and the L1 norm is the sum of the absolute values of each element in the vector, which is used to measure the image gradient difference and reflect the difference in structural information such as image edges.
[0022] S52: Calculate the preserved structural similarity, and the specific formula is as follows: ; Among them, denotes the structural similarity index, which is used to measure the similarity in structure between images. Here, (x,y,z) corresponds to the spatial coordinates of the image. In the context of magnetic resonance images, it represents the three dimensions of depth, height, and width; μ is the local mean, and are respectively the means of the corrected image I cor and the real image I gt , reflecting the overall brightness level of the image; σ is the local standard deviation, and are respectively the standard deviations of the corrected image I cor and the real image I gt , reflecting the degree of dispersion of the image pixel values, that is, the contrast situation of the image; is the corrected image I cor and the real image I gtThe covariance measures the correlation of the pixel value changes between the two; C1 and C2 are two constants used to avoid the denominator being zero. C 1=(0.01 L ) 2 , C 2=(0.03 L ) 2 ; L is the signal range.
[0023] S53: Enhance the edge sharpness through gradient constraint, and the calculation formula is as follows: ; Among them, represents the gradient magnitude of image I, which is used to measure the change intensity of the image in space and reflects the structural information such as the edges of the image; , , are the partial derivatives of image I in the x, y, and z directions respectively, representing the change rate of the image in the corresponding directions. ∇ is the Sobel operator; S54: Calculate the functional preservation loss, and the calculation formula is as follows: ; Among them, represents the functional preservation loss, which is used to measure the difference between the function-related features after processing (such as correction) and the true function features. The smaller the loss value, the closer the processed features are to the true situation in terms of function; ρ is a metric function, indicating FC cor and FC gt the degree of difference between; represents the function-related features after correction or transformation, etc.; represents the true function-related features, that is, the function features used as the reference standard.
[0024] S55: Calculate the cross-modal consistency loss, and the calculation formula is as follows: ; Among them, represents the cross-modal consistency loss, which is used to ensure that during the cross-modal conversion process, the image can be restored to the original state as much as possible after cyclic conversion. The smaller the loss value, the better the consistency of the cyclic conversion; is a generator network, and its function is to first convert the structural magnetic resonance image (sMRI) into a functional magnetic resonance image (fMRI), and then convert the obtained fMRI back to sMRI; is the original structural magnetic resonance image, which serves as the input for the entire cyclic conversion process; Represents the square of the L2 norm.
[0025] S56: Create a cyclic generator and perform pre-training and joint training. The pre-training formula is as follows: ; Where, Represents the model initialization parameters; argmin: is a mathematical operator whose function is to find the parameter value that minimizes the value of the following expression; Represents the anatomical fidelity loss, which is used to constrain the generated image to be consistent with the real situation in terms of anatomical structure, and measures the similarity between the reconstructed image and the real image in terms of anatomical structure; Represents the function preservation loss, which is used to ensure that the function-related features in the generation process are retained, and measures the difference between the processed function-related features and the real function features. Its constraint conditions are: freeze the parameters of the cross-modal attention module, train independently for each single modality, and only optimize the sMRI branch for L anatomy , and only optimize the fMRI branch for L function ; The joint training formula is as follows: ; Where, Represents the parameters finally optimized through joint training of the model; argmin : The function of this operator is to find the parameter value that can minimize the value of the subsequent expression; Represents the cross-modal consistency loss, which ensures the consistency of the image after cross-modal conversion (such as from sMRI to fMRI and then back to sMRI) with the original image. The smaller the loss, the higher the consistency of the conversion; Represents the anatomical fidelity loss, which is used to constrain the generated image to be consistent with the real situation in terms of anatomical structure, and measures the similarity between the reconstructed image and the real image in terms of anatomical structure; Represents the function preservation loss, which is used to ensure that the function-related features in the generation process are retained, and measures the difference between the processed function-related features and the real function features.
[0026] S57: Enable all modules to achieve end-to-end optimization.
[0027] The beneficial effects of the present invention include: The motion artifact adaptive correction method based on multi-modal MRI fusion provided by the present invention acquires the patient's sMRI and fMRI, preprocesses the sMRI and fMRI, creates a cross-modal feature fusion model, and inputs the preprocessed sMRI and fMRI into the cross-modal feature fusion model for dual-branch feature extraction; the cross-modal feature fusion model performs cross-modal feature interaction fusion on the extracted dual-branch features and conducts multi-scale feature reconstruction; detects image artifacts based on the reconstructed multi-scale features, and performs dynamic spatio-temporal attention correction on the detected image artifacts; establishes a three-level loss function system for multi-task joint optimization. By integrating sMRI and fMRI information and based on the above image processing and deep learning, it overcomes the shortcoming in the prior art that it cannot fully handle various complex head motion situations, and solves the technical problems of information loss between image modalities and insufficient utilization of features.
[0028] First, by adopting a motion artifact simulation unit and a multi-modal registration unit, and a multi-modal data preprocessing and enhancement strategy of physical simulation and data augmentation technology, the problems of motion artifacts and registration errors in sMRI and fMRI data are solved, and the registration accuracy and data quality are improved.
[0029] Second, through the cross-modal feature fusion network architecture of a dual-branch feature extractor and a cross-modal attention mechanism, the spatio-temporal feature fusion of sMRI and fMRI is realized, significantly improving the feature representation ability of multi-modal data.
[0030] Third, through the dynamic spatio-temporal attention correction mechanism that combines an artifact detection and generator network with spatio-temporal attention weights, the head motion artifacts are innovatively corrected dynamically, improving the data quality and ensuring the output of high-fidelity sMRI and fMRI.
[0031] Finally, by establishing a three-level loss function system and a multi-task joint optimization strategy that combines anatomical fidelity, functional preservation, and cross-modal consistency losses, the model is optimized during the multi-stage training process to ensure the best performance and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic flow chart of the motion artifact adaptive correction method based on multi-modal MRI fusion of the present invention.
[0033] Figure 2 It is a schematic diagram of the cross-modal feature fusion network architecture of the present invention.
[0034] Figure 3 It is a schematic diagram of the architecture of the dynamic spatio-temporal attention correction module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following further elaborates on the present invention in conjunction with the attached Figures 1 to 3 drawings: Example 1 See the appendix Figure 1 As shown, the motion artifact adaptive correction method based on multimodal MRI fusion includes the following steps: S1: Obtain the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI of the patient, and preprocess the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI; S2: Create a cross-modal feature fusion model, and input the preprocessed structural magnetic resonance image sMRI and functional magnetic resonance image fMRI into the cross-modal feature fusion model for dual-branch feature extraction; S3: The cross-modal feature fusion model performs cross-modal feature interaction fusion on the extracted dual-branch features and conducts multi-scale feature reconstruction; S4: Detect image artifacts based on the reconstructed multi-scale features, and perform dynamic spatio-temporal attention correction on the detected image artifacts; S5: Establish a three-level loss function system for multi-task joint optimization.
[0036] In this embodiment, the preprocessing of the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI in step S1 is implemented by a multimodal data preprocessing module. The multimodal data preprocessing module includes a motion artifact simulation unit and a multimodal registration unit. The specific process of preprocessing is as follows: S11: Apply an elastic deformation field with a specified displacement and rotation angle to the structural magnetic resonance image sMRI through the motion artifact simulation unit; S12: Use time-varying motion parameters for the functional magnetic resonance image fMRI, and perform image enhancement processing by randomly rotating a specified angle, scaling by a specified multiple, and injecting Rician noise; S13: Implement a hierarchical registration strategy through the multimodal registration unit to achieve the registration of the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI.
[0037] The motion artifact simulation unit uses physical model simulation technology to apply an elastic deformation field with a maximum displacement of 3 mm and a rotation angle of ±5° to the sMRI, and uses time-varying motion parameters for the fMRI (maximum displacement of 1.5 mm / TR when TR = 2 s). Data augmentation uses random rotation (±10°), scaling (0.9 - 1.1 times), and Rician noise injection (SNR = 20 - 30 dB). The multimodal registration unit implements a hierarchical registration strategy, first performing 6-degree-of-freedom rigid registration, and then performing B-spline non-linear registration (grid spacing 2 mm) to ensure that the registration accuracy reaches the technical indicators of a translation error <0.5 mm and a rotation error <0.5°.
[0038] The specific process of implementing the hierarchical registration strategy through the multimodal registration unit in step S13 is as follows: S131: Standardize the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI to unify their gray-scale ranges, resolutions, and spatial dimensions.
[0039] Traverse each pixel point of the sMRI and fMRI images, and statistically analyze the gray-scale value distribution. Calculate the minimum value of the image gray-scale value I min and the maximum value I max , and map the original gray-scale value I (x,y,z) through the formula: I′ (x,y,z)=255×( I (x,y,z)− I min ) / ( I max − I min ) to the standard gray-scale range of [0, 255]. If there are abnormal peaks in the image gray-scale distribution, use adaptive histogram equalization to divide the image into multiple sub-blocks, and perform histogram equalization separately within each sub-block to avoid over-enhancing noise.
[0040] Calculate the spatial dimension difference between the sMRI and fMRI images, and determine the translation amount and scaling ratio of the images. Use the affine transformation matrix to perform translation and scaling operations on the images. For translation, if the sMRI image needs to be translated t x units to the right in the x-axis direction, t y units upward in the y-axis direction, and t z units forward in the z-axis direction, then the affine transformation matrix T is: .
[0041] Multiply the coordinate vector of the image by the transformation matrix to obtain the transformed image coordinates, and complete the unification of spatial dimensions.
[0042] S132: Segment the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI, extract the key anatomical structures as the regions of interest, and perform morphological processing on the regions of interest to provide a specified number of registration feature points.
[0043] Using a deep learning segmentation algorithm, the U-Net network is adopted in this embodiment. First, a large number of labeled sMRI and fMRI image datasets are collected and divided into a training set, a validation set, and a test set. During the training process, the images are input into the U-Net network. The network extracts image features through convolutional layers, restores the image resolution through upsampling and skip connections, and outputs a probability map for each pixel belonging to different anatomical structure categories. The cross-entropy loss function is used to calculate the difference between the prediction result and the labeled image, and the network parameters are updated through the backpropagation algorithm to continuously optimize the model. After training, the sMRI and fMRI images are input into the trained model to obtain the segmented images, and different anatomical structures are segmented out.
[0044] Perform morphological operations on the region of interest mask image. First, use the dilation operation. By defining a structuring element, which can be a square or a circle, traverse the mask image. When the center of the structuring element is located at a foreground pixel point of the mask image, all the pixel points covered by the structuring element are set as foreground, enhancing the continuity of the structure and filling in small holes. Then use the erosion operation. When the center of the structuring element is located at a foreground pixel point of the mask image, only when all the pixel points covered by the structuring element are foreground, the central pixel point is retained as foreground, removing noise and isolated pixel points, making the boundary of the region of interest clearer and providing a specified number of registration feature points.
[0045] S133: Define 6-degree-of-freedom registration parameters. The 6-degree-of-freedom registration parameters include 3 translation parameters, namely the translation amounts along the x, y, and z axes. The value range is determined according to the image size and registration requirements and can be set to [-10mm, 10mm], and 3 rotation parameters, namely the Euler angles around the x, y, and z axes, with the value range being [-π, π]. In the subsequent registration process, the spatial transformation of the image is achieved by adjusting the degree-of-freedom registration parameters.
[0046] S134: Use mutual information or normalized mutual information to measure the statistical dependence relationship between images. For same-modal images, use mean square error or normalized cross-correlation processing.
[0047] For the sMRI and fMRI heterogeneous images, use mutual information (MI) or normalized mutual information (NMI) to measure the statistical dependence relationship between images. Mutual information is based on information theory and calculates the joint probability distribution p ( I sMRI ,I fMRI ) and marginal probability distribution p ( I sMRI )、 p ( IfMRI ) and use the formula MI = H ( I sMRI ) + H ( I fMRI ) - H ( I sMRI , I fMRI ) for calculation, where H is the information entropy. The normalized mutual information is to normalize the mutual information, and the formula is NMI = H ( I sMRI ) H ( I fMRI )MI, making the similarity metric value between [0, 1], which is more convenient for comparison and optimization. During the calculation process, use the histogram to statistically analyze the distribution of the image gray values and approximately calculate the probability distribution.
[0048] S135: Set the initial transformation of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI as the identity transformation, downsample the image to 1 / 4 or 1 / 8 of the original size, and quickly obtain the rough registration parameters to avoid falling into the local optimum.
[0049] S136: Gradually restore to the original resolution, use the low-resolution result as the initial value, refine the registration parameters, select at least ten pairs of landmark points in the region of interest, such as blood vessel bifurcation points and cerebral sulcus endpoints, calculate the spatial distance error and direction deviation of them, and judge whether the anatomical structures are aligned.
[0050] Gradually restore the image resolution, doubling the image resolution each time until the original resolution is restored. Use the registration parameters obtained in the previous low-resolution stage as the initial value of the current high-resolution stage, continue to use the optimization algorithm, within a smaller parameter search range, based on the similarity metric, further adjust the 6-degree-of-freedom registration parameters to further improve the similarity metric value between the images. In the region of interest, manually or automatically select at least ten pairs of landmark points, such as blood vessel bifurcation points, cerebral sulcus endpoints, etc. Calculate the spatial distance error and direction deviation of these landmark points in the registered image. If the spatial distance error is less than the set threshold and the direction deviation is less than the set angle, it is considered that the anatomical structures are aligned and the fine registration is completed; otherwise, continue to adjust the registration parameters until the requirements are met.
[0051] S137: Construct a deformation field using cubic B-spline interpolation. Set the grid node spacing to 2 mm to ensure avoiding overfitting while retaining local details. Add a bending energy penalty term to regularize the deformation field, constrain the smoothness of the deformation, prevent non-physiological deformation, continue the mutual information from the rigid registration stage, or introduce local mutual information to enhance the registration accuracy of local structures.
[0052] Construct a deformation field using cubic B-spline interpolation. Divide the image space into a grid with a grid node spacing of 2 mm. For each pixel in the image, construct cubic B-spline basis functions based on its surrounding grid nodes, and calculate the displacement of the pixel in the deformation field through a linear combination of the basis functions. To ensure avoiding overfitting while retaining local details in the deformation field, add a bending energy penalty term to regularize the deformation field. The bending energy penalty term constrains the smoothness of the deformation and prevents non-physiological deformation by calculating the sum of the squares of the second derivatives of the deformation field. During the optimization process, continue the mutual information used in the rigid registration stage, or introduce local mutual information. Divide the image into multiple local regions, calculate the mutual information within each local region respectively, and enhance the registration accuracy of local structures. Use an optimization algorithm to continuously adjust the deformation field parameters to minimize the objective function containing the similarity metric term and the bending energy penalty term, obtain the final deformation field, and complete image registration.
[0053] Example 2 Based on Example 1, the specific process of performing dual-branch feature extraction by the cross-modal feature fusion model in step S2 is as follows: S21: Perform convolution processing on the structural magnetic resonance image sMRI through the initial convolution layer of the first branch of the cross-modal feature fusion model. The specific convolution calculation formula is as follows: ; where represents the feature map obtained after processing by the initial convolution layer of the first branch, which is the output result of performing a convolution operation on the input structural magnetic resonance image (sMRI) and is used for subsequent feature extraction and processing; is the input structural magnetic resonance image sMRI; represents a three-dimensional convolution operation. The convolution kernel size is 7×7×7 (in three-dimensional space, the size in each dimension is 7), and the stride is 2, which means that when the convolution kernel slides on the image, it moves 2 pixels in each dimension each time, and the output resolution is , where C1 = 64.
[0054] S22: Process through the residual block level of the first branch of the cross-modal feature fusion model. The specific calculation formula is as follows: ; Among them, is the feature map processed by the previous layer (the (l - 1)-th layer) residual block, serving as the input of the current l-th layer residual block, carrying the feature information extracted from the previous level. l = 2, 3, 4, 5 correspond to the output of residual blocks at different levels respectively, and are used for further feature extraction and transformation; R l is an improved 3D residual block, with the number of channels being {64, 128, 256, 512} in sequence, and the output feature resolution being {8mm 3 , 4mm 3 , 2mm 3 , 1mm 3}, representing the l-th residual block. It is a special neural network module. By introducing a skip connection, the network can learn the residual relationship between the input and the output, effectively alleviating the vanishing gradient problem in deep neural networks, and enhancing the expression ability and training stability of the network.
[0055] S23: Perform convolution processing on the functional magnetic resonance image fMRI through the 3D convolutional layer of the second branch of the cross-modal feature fusion model. The specific calculation formula is as follows: ; Among them, This is the feature map obtained after being processed by the 3D convolutional layer of the second branch; represents a three-dimensional convolution operation, where the size of the convolution kernel is 3×3×3; represents the input functional magnetic resonance image data. Here, t:t + 9 indicates a series of fMRI images from time t to time t + 9 in the time dimension. is the input functional magnetic resonance image fMRI, , the window length of the time window is 10TR, TR = 2s; the convolution kernel is 3×3×3, the stride is 1, and the padding mode is "SAME"; the output feature map: resolution: , where the number of channels C = 64.
[0056] S24: Perform LSTM time series modeling. The calculation formula is as follows: ; Among them, represents the feature representation of the functional magnetic resonance image (fMRI) obtained after being processed by the long short-term memory network (LSTM) at time t. It synthesizes the information of multiple previous time steps and is a kind of feature refinement of the fMRI time series data at the current moment, and is used for subsequent feature fusion or task analysis; Represents an LSTM layer with 128 hidden units. The number of hidden units determines the memory capacity and processing ability of the LSTM layer. 128 hidden units mean that this layer can maintain a 128-dimensional hidden state when processing sequential data, and can capture relatively complex time series dependencies; Is the feature sequence input to the LSTM layer, which is composed of the feature maps output by the previous 3D convolutional layer at 10 time steps from time t−9 to time t in the time dimension. The number of LSTM units is 128, input sequence: sliding time window t -9: t (window length 10TR, step size 5TR), output: ; H , W , D Is the spatial dimension, T Is the number of time points, total duration .
[0057] S25: Perform hierarchical feature alignment on the outputs of the two branches, and perform channel dimension concatenation at 5 scales. The specific formula is as follows: ; Among them, Represents the feature map obtained by concatenating the channel dimensions at a certain scale after hierarchical feature alignment at time t, ; Is the channel dimension concatenation operation; Is the k-th layer sMRI feature map, ; Is the k-th layer fMRI feature map.
[0058] S26: Perform dynamic attention generation. The specific process is as follows: ; Among them, Is the feature map obtained by concatenating the channel dimensions at the k-th scale, which integrates the feature information of the two branches (sMRI and fMRI) and serves as the input for calculating the attention map; Is the bias term, which adds a learnable offset to the convolution operation result to improve the model fitting ability; Represents the channel-adaptive attention map generated at the k-th scale, whose element values are between [0,1]. The larger the value, the more important the features at the corresponding position or channel; σ is the Sigmoid function, and the formula is It maps the input value to the [0,1] interval for generating attention weights; Is the convolution kernel weight, , the convolution kernel shape is 1×1×1×( CsMRI + C fMRI ) × 1; C sMRI and C fMRI are respectively and of the number of channels. Through the convolution kernel and the spliced feature map perform a convolution operation to calculate the attention weights.
[0059] See Figure 2 , the cross-modal feature fusion model sets up double-branch processing channels. The first-branch processing channel includes a first input layer, an initial convolution layer, a residual block level, and a feature pyramid layer. The second-branch processing channel includes a second output layer, a 3D convolution layer, a spatio-temporal feature extraction time window, an LSTM layer, and a feature tensor output layer. The cross-modal feature fusion model also includes a feature splicing channel, a convolution layer, an activation layer, a generated attention map layer, and a weighted fusion layer. The feature splicing channel is respectively connected to the feature pyramid layer of the first-branch processing channel and the feature tensor output layer of the second-branch processing channel.
[0060] Embodiment 3 On the basis of Embodiment 1 or Embodiment 2, the specific process of the cross-modal feature fusion model for cross-modal feature interaction fusion and multi-scale feature reconstruction of the extracted double-branch features is as follows: S31: Perform weighted feature fusion of the two branches. The specific formula is as follows: ; where ⊙ represents element-wise multiplication; is to generate a spatial / channel adaptive attention map; is the k-th layer sMRI feature map; is the k-th layer fMRI feature map; is the output fusion feature, , to achieve the dynamic balance of anatomical and functional features; S32: Perform multi-scale feature reconstruction, inverse pyramid aggregation, and finally output a resolution of 1mm³. The specific formula is as follows: ; where represents the final feature map obtained after multi-scale feature reconstruction and inverse pyramid aggregation. This feature map synthesizes the fused feature information at 5 scales and is the final output of the model in the feature processing stage; is the k-th layer fused feature map, k = 1, resolution 8mm 3 , k = 5, resolution 1mm 3 , , U k is a 3×3×3 deconvolution upsampling operation, and Wk is the kernel parameter. . Stride: 25−k (stride is 16 when k = 1, and stride is 11 when k = 5), output resolution: 1mm 3 ×Cout (Cout = 64).
[0061] In this embodiment, referring to Figure 3 the architecture of the dynamic spatio-temporal attention correction module, in step S4, the image artifacts are detected based on the reconstructed multi-scale features, and the specific process of dynamically correcting the spatio-temporal attention of the detected image artifacts is as follows: S41: Calculate the gradient magnitude through the Sobel operator with a kernel size of 3×3×3 in the artifact detection unit of the dynamic spatio-temporal attention correction module architecture, perform bilinear interpolation processing, and generate a spatial attention map. The calculation formula is: ; Among them, represents the generated spatial attention map, and its element values reflect the importance of different positions in the image in space. The larger the value, the more attention should be paid to the corresponding position; represents the normalization operation, and the purpose is to map the calculated gradient magnitude to a suitable range; , are the partial derivatives of the feature map F in the x direction and the z direction respectively; F is a three-dimensional feature map, , with a resolution of 2mm 3 , sobel operator: three-direction convolution kernel (taking the x direction as an example), , (3×3 kernel, extended to 3×3×3 three-dimensional), gradient calculation: , , resolution improvement: bilinear interpolation (actually trilinear interpolation: , among them, represents the high-resolution spatial attention map, with a resolution of 0.5mm 3 , output spatial attention map: represents the trilinear interpolation operation, and here ×2 means doubling the resolution of the original image in the three dimensions of depth, height, and width; represents the generated spatial attention map, and its element values reflect the importance of different positions in the image in space. The larger the value, the more attention should be paid to the corresponding position, .
[0062] S42: Generate time attention based on LSTM motion modeling. The specific formula is as follows: ; ; Among them, represents a long short-term memory network (LSTM) layer with 64 hidden units, are respectively the head movement parameters of 6 degrees of freedom, is the output time-varying attention weight; represents the motion feature vector at time t.
[0063] S43: The generator model based on the U-Net3+ architecture configures multi-scale skip connections (5 scales) and a temporal convolutional branch (kernel size 3×3×3), and outputs the corrected structural magnetic resonance image sMRI and functional magnetic resonance image fMRI. The module can adaptively identify the spatio-temporal distribution characteristics of head motion artifacts.
[0064] The encoder structure formula is: ; Among them, represents the feature map after the l-th layer passes through the convolution operation in the generator model based on the U-Net+ architecture; represents a three-dimensional convolution operation, with a convolution kernel size of 3×3×3 and a stride of 2; is the feature map of the previous layer (the l−1-th layer), which serves as the input for the current l-th layer convolution operation and carries the feature information extracted from the previous level; the input is , with a convolution kernel of 3×3×3, a stride of 2, and the number of output channels is C l , , and the output resolution is .
[0065] Temporal convolutional branch: ; Parameter description: Tl: represents the feature map obtained after the l-th layer passes through the convolution operation in the temporal convolutional branch; Here is a three-dimensional convolution operation, with a convolution kernel size of 3×3×3; is the channel dimension concatenation operation. The time attention weight , concatenation operation: , ⊕ represents the channel concatenation of the feature and the temporal weight, and the 3D temporal convolution kernel slides in the spatial dimension and weights in the time dimension; S44: Perform multi-scale feature skip connections, ; Among them, represents the fused feature map obtained after specific operations during the multi-scale feature skip connection process; That is, a three-dimensional deconvolution (upsampling convolution) operation with a convolution kernel size of 1×1×1, U is 2x upsampling, and k aggregates multi-scale features from 5 to 1 in reverse; Concat represents a channel dimension concatenation operation. It concatenates the feature maps E k , T k and D k+1 along the channel dimension to fuse multi-faceted feature information; E k : is the feature map at the k-th layer in the generator model based on the U-Net3+ architecture, containing the spatial feature information of the image at this scale; T k : is the feature map obtained after convolution operation at the k-th layer in the temporal convolution branch, fusing temporal attention information and spatial feature information, reflecting the dynamic features in the time dimension; D k+1 : is the feature map from a deeper layer (the k+1-th layer), with more abstract semantic information and different spatial dimensions (after downsampling).
[0066] After multi-scale feature skip connections are made in S44, structural magnetic resonance image sMRI and functional magnetic resonance image fMRI reconstruction are also performed; the reconstruction formula for the structural magnetic resonance image sMRI is as follows: ; where, represents the reconstructed structural magnetic resonance image (sMRI); represents the hyperbolic tangent activation function, whose function values are mapped between [-1, 1], introducing non-linearity to enhance the model's expressive ability, adapt to the MRI signal range, and output the corrected sMRI image; : resolution 1mm³, number of channels C =64, a certain feature map obtained after multi-scale feature skip connection operation (here K takes the value of 1); : weight matrix (convolution kernel), used to perform a linear transformation on the input feature , : bias term, increasing the model's fitting ability and adjusting the result of the linear transformation.
[0067] The reconstruction formula for the functional magnetic resonance image fMRI is as follows: ; where, represents the reconstructed functional magnetic resonance image (fMRI); represents the weight coefficient, which is used to weight the calculation results at different time steps (or scale-related) during the summation process; weight matrix (convolution kernel), performing a linear transformation on the input feature to learn the mapping relationship between the features and the reconstructed functional magnetic resonance image; It is the representation of the feature map S1 obtained by multi-scale feature skip connection at different time steps τ (value range: t to t - 5). Denotes the bias term, which increases the fitting ability of the model and adjusts the result of the linear transformation. And the formula satisfies ∑ τ W τ = 1. (Feature map at time point τ ), convolution kernel , bias , time-varying weight W τ ∈ [0, 1] (attention weight from LSTM, satisfying ∑ τ W τ = 1), output the corrected fMRI time point signal.
[0068] Example 4 Based on Example 1 or Example 2 or Example 3, the specific process of establishing a three-level loss function system for multi-task joint optimization in step S5 is as follows: S51: Calculate the anatomical fidelity loss, and the specific formula is as follows: ; Among them, represents the anatomical fidelity loss, which is used to measure the similarity of the reconstructed or corrected image and the real image in anatomical structure. The lower the loss value, the closer the anatomical structure of the reconstructed image is to the real situation; SSIM: Structural Similarity Index, which is used to evaluate the similarity of two images in structure. Its value range is between [-1, 1], and the closer the value is to 1, the more similar the image structures are; Icor: Corrected image, which is the output image processed by the model; Igt: Real image, which is used as a reference standard to evaluate the quality of the corrected image; ∇: Gradient operator, which is used to calculate the gradient of the image; represents the L1 norm between the gradients of the corrected image and the real image, that is, the sum of the absolute values of the differences in the corresponding pixel gradient values, which reflects the differences in details such as the edges of the image.
[0069] S52: Calculate the preservation of structural similarity, and the specific formula is as follows: ; Among them, Structural Similarity Index, which is used to quantify the similarity of two images (here x, y, z are regarded as the spatial coordinates of the image) in structure, and its value range is between -1 and 1. The closer the value is to 1, the more similar the image structures are; are the means of the reconstructed image Icor and the real image Igt respectively, which reflect the overall brightness level of the image. μ is the local mean, σ is the local standard deviation, C1=(0.01 L ) 2 , C 2=(0.03 L ) 2 , L is the signal range, S53: Enhance the edge sharpness through gradient constraint, and the calculation formula is as follows: ; where ∇ is the Sobel operator; , , are the partial derivatives of the image I in the x, y, and z directions respectively, and these partial derivatives measure the change rate of the pixel values of the image in the corresponding directions.
[0070] S54: Calculate the functional preservation loss, and the calculation formula is as follows: ; where represents the functional preservation loss; is the Euclidean distance; represents the feature map after correction and transformation; represents the ground-truth feature map, that is, the feature map used as the reference standard.
[0071] S55: Calculate the cross-modal consistency loss, and the calculation formula is as follows: ; where represents the cycle-consistency loss, which is used to measure the difference degree between the image after cyclic transformation and the original image during the cross-modal transformation process; is a generator, and its function is to first convert the structural magnetic resonance image (sMRI) into a functional magnetic resonance image (fMRI), and then convert the obtained fMRI back to sMRI; represents the square of the L2 norm, and the square of the L2 norm is calculated for the vector composed of the image pixel values.
[0072] The cyclic generator G contains 3 levels of residual blocks, and the structure is: .
[0073] where G(x): represents the output result of the cyclic generator G for the input x; x is the original image data input to the cyclic generator G; F1 and F2 represent the convolutional operations or other non-linear transformation operations in the residual blocks; It means to first perform the F1 transformation on the input x, and then use the result after the F1 transformation as the input to perform the F2 transformation.
[0074] S56: Create a cyclic generator and perform pre-training and joint training. The pre-training formula is as follows: ; Among them, represents the initialization parameters, which are the initial parameter values optimized under the constraint of a specific loss function and are used for the subsequent training of the generator; argmin is an operator whose function is to find the parameter value that minimizes the value of the subsequent expression. represents the loss related to the anatomical structure, which is used to constrain the generated image by the generator to be consistent with the real situation in terms of anatomical structure. For example, in medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense. represents the function preservation loss. Its constraint condition is: freeze the parameters of the cross-modal attention module, perform independent training for single modalities, only optimize Lanatomy for the sMRI branch, and only optimize Lfunction for the fMRI branch. The joint training formula is as follows: Among them, represents the finally optimized parameters, which are the optimal parameter values found under the constraint of multiple loss functions and are used to determine the final model parameters of the cyclic generator. represents the cross-modal consistency loss. represents the loss related to the anatomical structure, which is used to constrain the generated image by the generator to be consistent with the real situation in terms of anatomical structure. For example, in medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense. represents the function preservation loss.
[0075] S57: Enable end-to-end optimization for all modules. Optimizer configuration: Adam β1 = 0.9, β2 = 0.999, ϵ = 10−8; Initial learning rate: 10−4, decay × 0.5 every 20k steps; Batch size: 8 pairs of sMRI-fMRI data.
[0076] The motion artifact adaptive correction method based on multi-modal MRI fusion of the present invention inputs paired sMRI-fMRI data with artifacts, generates a multi-modal data cube through registration; then extracts anatomical and functional features in parallel and conducts feature interaction at 4 scales; then dynamically generates spatio-temporal attention maps and reconstructs them region by region; finally outputs artifact-free sMRI and high-fidelity fMRI, and simultaneously generates a quality control report including the comparison before and after correction. The key parameter configuration is: the number of convolutional kernels for feature extraction [64, 128, 256, 512], the number of LSTM units is 128, and the learning rate decays by 0.5 times every 20k steps.
[0077] In summary, the motion artifact adaptive correction method based on multi-modal MRI fusion provided by the present invention acquires the patient's sMRI and fMRI, preprocesses the sMRI and fMRI, creates a cross-modal feature fusion model, and inputs the preprocessed sMRI and fMRI into the cross-modal feature fusion model for double-branch feature extraction; the cross-modal feature fusion model performs cross-modal feature interaction and fusion on the extracted double-branch features, and conducts multi-scale feature reconstruction; detects image artifacts based on the reconstructed multi-scale features, and performs dynamic spatio-temporal attention correction on the detected image artifacts; establishes a three-level loss function system for multi-task joint optimization. By integrating sMRI and fMRI information and based on the above image processing and deep learning, it overcomes the shortcoming in the prior art that it cannot fully handle various complex head motion situations, and solves the technical problems of information loss between image modalities and insufficient utilization of features.
[0078] By adopting a motion artifact simulation unit and a multi-modal registration unit, and a multi-modal data preprocessing and enhancement strategy of physical simulation and data augmentation technology, it solves the problems of motion artifacts and registration errors in sMRI and fMRI data, and improves the registration accuracy and data quality. Through the cross-modal feature fusion network architecture of a double-branch feature extractor and a cross-modal attention mechanism, it realizes the spatio-temporal feature fusion of sMRI and fMRI, and significantly improves the feature representation ability of multi-modal data. Through the dynamic spatio-temporal attention correction mechanism that combines an artifact detection and generator network with spatio-temporal attention weights, it innovatively corrects head motion artifacts dynamically, improves the data quality, and ensures the output of high-fidelity sMRI and fMRI. By establishing a three-level loss function system and combining a multi-task joint optimization strategy of anatomical fidelity, functional preservation, and cross-modal consistency loss, it optimizes the model during the multi-stage training process to ensure the best performance and accuracy.
Claims
1. An adaptive correction method for motion artifacts based on multimodal MRI fusion, characterized in that Including the following steps: S1: Obtain the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI of the patient, and preprocess the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI; S2: Create a cross-modal feature fusion model, and input the preprocessed structural magnetic resonance image sMRI and functional magnetic resonance image fMRI into the cross-modal feature fusion model for dual-branch feature extraction; S3: The cross-modal feature fusion model performs cross-modal feature interaction fusion on the extracted dual-branch features and conducts multi-scale feature reconstruction; S4: Detect image artifacts based on the reconstructed multi-scale features, and perform dynamic spatio-temporal attention correction on the detected image artifacts; S5: Establish a three-level loss function system for multi-task joint optimization.
2. The motion artifact adaptive correction method based on multi-modal MRI fusion according to claim 1, characterized in that, In step S1, the preprocessing of the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI is implemented by a multi-modal data preprocessing module. The multi-modal data preprocessing module includes a motion artifact simulation unit and a multi-modal registration unit. The specific process of the preprocessing is as follows: S11: Apply an elastic deformation field with a specified displacement and rotation angle to the structural magnetic resonance image sMRI through the motion artifact simulation unit; S12: Use time-varying motion parameters for the functional magnetic resonance image fMRI, and perform image enhancement processing by randomly rotating a specified angle, scaling by a specified multiple, and injecting Rician noise; S13: Implement a hierarchical registration strategy through the multi-modal registration unit to achieve the registration of the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI.
3. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 2, wherein, The specific process of implementing the hierarchical registration strategy through the multi-modal registration unit in step S13 is as follows: S131: Perform normalization processing on the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI to unify their gray value ranges, resolutions, and spatial dimensions; S132: Segment the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI, extract key anatomical structures as regions of interest, and perform morphological processing on the regions of interest to provide a specified number of registration feature points; S133: Define 6-degree-of-freedom registration parameters including 3 translation parameters and 3 rotation parameters; S134: Measure the statistical dependence relationship between images, and use mean square error or normalized cross-correlation processing for same-modal images; S135: Set the initial transformation of the structural magnetic resonance image sMRI and functional magnetic resonance image fMRI as the identity transformation, and downsample the images to a specified proportion of the original size; S136: Gradually restore to the original resolution, use the low-resolution result as the initial value, select at least ten pairs of landmark points within the region of interest, calculate their spatial distance error and direction deviation, and determine whether the anatomical structures are aligned; S137: Construct a deformation field, add a bending energy penalty term for deformation field regularization, continue with the mutual information in the rigid registration stage or introduce local mutual information to enhance the registration accuracy of local structures.
4. The motion artifact adaptive correction method based on multi-modal MRI fusion according to claim 1, wherein, The specific process of the cross-modal feature fusion model performing dual-branch feature extraction in step S2 is as follows: S21: Convolve the structural magnetic resonance image sMRI through the initial convolutional layer of the first branch of the cross-modal feature fusion model; S22: Process through the residual block hierarchy of the first branch of the cross-modal feature fusion model; S23: Convolve the functional magnetic resonance image fMRI through the 3D convolutional layer of the second branch of the cross-modal feature fusion model; S24: Perform LSTM temporal modeling, and the modeling formula is as follows: ; Among them, represents the feature sequence of 10 time steps from t−9 to t extracted by the CNN at time step t. The CNN is used to extract features from the spatial dimension of fMRI data; Represents an LSTM layer with 128 hidden units. Using 128 hidden units means that the dimensions of structures such as the internal memory units in the LSTM are 128, enabling it to learn more complex temporal features; Is the fMRI feature representation finally obtained at time step t after being processed by the LSTM, which is used for subsequent analysis. The number of LSTM units is 128. t -9: t Is the sliding time window. Is the output. H , W , D Is the spatial dimension. T Is the number of time points. S25: Perform hierarchical feature alignment on the outputs of the two branches, and perform channel dimension concatenation at 5 scales, and the specific formula is as follows: ; Among them, is the k layer sMRI feature map; is the k layer fMRI feature map; is the output concatenated feature; represents values from 1 to 5, representing 5 different scales; refers to the operation of tensor concatenation in the channel dimension in deep learning; S26: Generate dynamic attention, and the specific process is as follows: ; Among them, denotes The weight matrix has five dimensions and is responsible for linearly transforming the input features; among them, 1×1×1 means that the convolution kernel weights of the first three dimensions are all 1; and the convolution kernel weight of the fourth dimension denotes the sum of the number of channels of structural magnetic resonance imaging features ( ) and the number of channels of functional magnetic resonance imaging features ( ); the convolution kernel weight of the last dimension is also 1; is to generate a spatial / channel adaptive attention map; σ is the Sigmoid function, and the output value range is [0, 1]; is the output concatenated feature; is the bias term, which is used to increase the fitting ability of the model.
5. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 4, wherein The specific process of the cross-modal feature interaction and fusion of the double-branch features extracted by the cross-modal feature fusion model in step S3 and the multi-scale feature reconstruction is as follows: S31: Perform weighted feature fusion of the two branches, and the specific formula is as follows: ; Among them, represents the dimensional information of the k-th feature map after fusion; among them, represents the size of the feature map in the height direction; represents the size of the feature map in the width direction; represents the size of the feature map in the depth direction; represents the number of channels of structural magnetic resonance imaging features; is to generate a spatial / channel adaptive attention map; is the k layer of sMRI feature maps; the k layer of fMRI feature maps; ⊙ represents element-wise multiplication; is the k-th feature map after fusion, achieving a dynamic balance between anatomical and functional features; S32: Perform multi-scale feature reconstruction, inverse pyramid aggregation, and finally output a resolution of 1 mm³, and the specific formula is as follows: ; Among them, represents the output feature finally after multi-scale feature reconstruction and inverse pyramid aggregation; is the k layer fusion feature map; Here, the features after fusion at 5 different scales from k = 1 to k = 5 are U k calculated, k = 1, resolution 8mm 3 , k = 5, resolution 1mm 3 ; U k is the 3×3×3 deconvolution upsampling operation.
6. The motion artifact adaptive correction method based on multi-modal MRI fusion according to claim 1, characterized in that In step S4, detect image artifacts based on the reconstructed multi-scale features, and the specific process of dynamically correcting the spatio-temporal attention of the detected image artifacts is as follows: S41: Calculate the gradient magnitude through a Sobel operator with a kernel size of 3×3×3, perform bilinear interpolation processing, and generate a spatial attention map; S42: Generate temporal attention based on LSTM motion modeling; S43: Configure a multi-scale skip connection and a temporal convolutional branch based on the generator model of the U-Net3+ architecture, and output the corrected structural magnetic resonance image sMRI and functional magnetic resonance image fMRI. The module can adaptively identify the spatio-temporal distribution characteristics of head motion artifacts; S44: Perform multi-scale feature skip connection, and the specific formula is as follows: ; Among them, represents the feature map obtained after the multi-scale feature skip connection operation; U is upsampling by a factor of 2; k takes values from 5 to 1, representing the reverse aggregation of multi-scale features from 5 to 1; represents a three-dimensional convolution operation with a convolution kernel size of 1×1×1, which adjusts the number of channels and performs linear transformation on the input feature map, changing the number of channels without changing the spatial dimensions (length, width, depth) of the feature map; represents concatenating the E k , T k , D k+1 three feature maps in the channel dimension, where E k , T k , D k+1 are feature maps from different scales, different processing stages, or different branches.
7. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 6, characterized in that, After performing the multi-scale feature skip connection in S44, the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI are also reconstructed; The reconstruction formula of the structural magnetic resonance image sMRI is as follows: ; Among them, represents the reconstructed structural magnetic resonance image (sMRI); represents the hyperbolic tangent activation function, whose function values are mapped between [-1, 1], introducing non-linear factors to enhance the expression ability of the model; : a certain feature map obtained after the multi-scale feature skip connection operation (here K takes the value of 1); : the weight matrix (convolution kernel) used for the input features , to perform a linear transformation; : the bias term, which increases the fitting ability of the model and adjusts the result of the linear transformation; The reconstruction formula of the functional magnetic resonance image fMRI is as follows: ; Among them, represents the reconstructed functional magnetic resonance image (fMRI); represents the weight coefficient, which is used to weight the calculation results related to different time steps or scales during the summation process; Weight matrix, for the input features perform a linear transformation to learn the mapping relationship between the features and the reconstructed functional magnetic resonance image; is the representation of the feature map S1 obtained by multi-scale feature skip connection at different time steps τ; the value range of τ is from t to t - 5; represents the bias term, which increases the fitting ability of the model, adjusts the result of the linear transformation, and satisfies ∑ τ W τ = 1.
8. The motion artifact adaptive correction method based on multi-modal MRI fusion according to claim 1, wherein The specific process of establishing a three-level loss function system for multi-task joint optimization in step S5 is as follows: S51: Calculate the anatomical fidelity loss, and the specific formula is as follows: ; Among them, represents the anatomical fidelity loss, which is used to measure the similarity of the reconstructed image and the real image in terms of anatomical structure. The smaller the loss value, the closer the reconstructed image is to the real image; represents the structural similarity index, which is used to measure the structural similarity between two images. The value range is limited to [-1, 1]. The closer it is to 1, the more similar the image structures are. In SSIM( I cor , I gt ), I cor is the reconstructed image, I gt is the real reference image; ∇ represents the gradient operator; calculates the I cor 1-norm of the gradient of the reconstructed image I gt and the real image, L which is used to measure the difference in structural information such as image edges; S52: Calculate the preservation of structural similarity, and the specific formula is as follows: ; Among them, The structural similarity index, which is used to quantify the structural similarity degree between two images, ranges from -1 to 1, and the closer the value is to 1, the more similar the image structures are; are the reconstructed images I cor and the real images I gt 's mean values, which reflect the overall brightness level of the images; μ is the local mean value, σ is the local standard deviation, C 1=(0.01 L ) 2 , C 2=(0.03 L ) 2 , L is the signal range; S53: Enhance the edge sharpness through gradient constraint, and the calculation formula is as follows: ; where ∇ is the Sobel operator; , , are the partial derivatives of the image I in the x, y, and z directions respectively, and the partial derivative measures the change rate of the pixel values of the image in the corresponding direction; S54: Calculate the functional preservation loss, and the calculation formula is as follows: ; Among them, represents the Functional Preservation Loss; is the Euclidean distance; represents the feature map after calibration and transformation; represents the ground truth feature map, that is, the feature map used as a reference standard; S55: Calculate the cross-modal consistency loss, and the calculation formula is as follows: ; Among them, represents the cross-modal consistency loss, which is used to measure the degree of difference between the image after cyclic transformation and the original image during the cross-modal transformation process; is a generator whose function is to first convert the structural magnetic resonance image (sMRI) into the functional magnetic resonance image fMRI, and then convert the obtained fMRI back to sMRI; represents the square of the L2 norm, and calculates the square of the L2 norm for the vector composed of the image pixel values; S56: Create a cyclic generator and perform pre-training and joint training, and the pre-training formula is as follows: ; Among them, represents the initialization parameter, which is the initial parameter value optimized under a specific loss function constraint and is used for the subsequent training of the generator; argmin is an operator used to find the parameter value that minimizes the value of the following expression; represents the loss related to anatomical structure, which is used to constrain the generated image by the generator to conform to the real situation in terms of anatomical structure. In medical image generation, it is necessary to ensure that the shape, position, etc. of the organs conform to anatomical common sense; represents the function preservation loss; The constraints are as follows: Freeze the parameters of the cross-modal attention module, perform single-modal independent training, and only optimize the sMRI branch L anatomy , and only optimize the fMRI branch L function ; The joint training formula is as follows: ; Among them, represents the parameters finally optimized, which are the optimal parameter values found under the constraints of multiple loss functions and are used to determine the final model parameters of the cyclic generator; represents the cross-modal consistency loss; represents the loss related to anatomical structure, which is used to constrain the generated images by the generator to conform to the real situation in terms of anatomical structure. In medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense; represents the function preservation loss; S57: Enable end-to-end optimization for all modules, optimizer configuration: Adam β 1 = 0.9, β 2 = 0.999, ϵ = 10 −8 ; Initial learning rate: 10 −4 , decay by ×0.5 every 20k steps; Batch size: 8 pairs of sMRI-fMRI data.
Citation Information
Patent Citations
Cancer accompanying depression identification method based on multi-mode magnetic resonance data
CN113705680A
Deep network model for accelerating multi-modal MR imaging
CN114049408A
Brain function fusion analysis method based on multi-modal registration
CN114419015A
Deep learning electroencephalogram noise reduction method based on time-frequency domain information fusion
CN114947883A
Alzheimer's disease cerebral atrophy model construction method, prediction method and device
CN117457222A
Cited By
Method for identifying and processing artifact region in nuclear medicine imaging image
CN120598823A
Throat structure and detail high-resolution image enhancement method based on deep learning
CN120746872A
Multi-center mental image data enhancement method and system based on federal learning
CN120876297A
Image enhancement method and device, electronic equipment and storage medium
CN120876804A