Adaptive motion artifact correction method based on multimodal MRI fusion

Through the motion artifact adaptive correction method of multimodal MRI fusion, the problem of missing information and insufficient feature utilization in the prior art is solved in the absence of information and insufficient feature utilization in complex head movements, and high-fidelity MRI image correction is achieved, which improves image quality and robustness.

CN120298448BActive Publication Date: 2025-08-22川北医学院附属医院 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510780671.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-22
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing MRI motion artifact correction technology cannot fully cope with various complex cephalomotor situations, resulting in the lack of information between image modalities and insufficient utilization of features, affecting clinical diagnosis and research.

Method used

Adaptive motion artifact correction method based on multimodal MRI fusion is adopted, and sMRI and fMRI are preprocessed, feature extraction, interactive fusion and multi-scale reconstruction are performed through a cross-modal feature fusion model, and dynamic spatiotemporal attention correction is performed, and multi-task optimization is performed in combination with three-level loss function.

Benefits of technology

The data quality and robustness of MRI images are improved, the fidelity and consistency of image details under complex head movements are ensured, and the problems of missing and insufficient feature utilization between information modes are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298448B_ABST
    Figure CN120298448B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology, and specifically relates to a method for adaptive motion artifact correction based on multimodal MRI fusion. The method involves obtaining a patient's sMRI and fMRI images, preprocessing the sMRI and fMRI images, creating a cross-modal feature fusion model, and inputting the preprocessed sMRI and fMRI images into the cross-modal feature fusion model for dual-branch feature extraction. The cross-modal feature fusion model then interactively fuses the extracted dual-branch features and reconstructs multi-scale features. Image artifacts are detected based on the reconstructed multi-scale features, and dynamic spatiotemporal attention correction is performed on the detected image artifacts. Finally, a three-level loss function system is established for multi-task joint optimization. By integrating sMRI and fMRI information and based on the aforementioned image processing and deep learning, the method overcomes the shortcomings of existing technologies, which are unable to fully cope with various complex head motion situations, and addresses the technical issues of missing information between image modalities and insufficient feature utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of thermal image processing technology, and in particular relates to a motion artifact adaptive correction method based on multimodal MRI fusion. Background Art

[0002] In magnetic resonance imaging, head motion artifact is a common and difficult-to-avoid problem, especially in structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI). Head motion artifact is more significant. In severe cases, it may lead to the inability to effectively extract structural and functional information of the brain, thereby affecting clinical diagnosis and research.

[0003] Current MRI motion artifact correction technologies are mainly based on detecting and compensating motion artifacts in K-space or using deep learning to remove artifacts in the image domain. Although inter-domain data consistency loss is used to maintain consistency between K-space and the spatial domain, this approach still faces certain challenges, especially in cases where artifacts are more severe. How to achieve high consistency without losing image details remains a difficult point. Existing technologies usually simulate motion artifacts in the preliminary stage and correct them through frequency domain and spatial domain generators, but these methods cannot fully cope with various complex head movements in clinical practice, resulting in insufficient robustness and adaptability of the trained models in practical applications, which may ultimately cause residual artifacts or image blur.

[0004] In addition, current technologies mostly focus on single-modality MRI images, ignoring the correlation between modalities, resulting in missing information dimensions (sMRI lacks temporal dynamic information, fMRI lacks spatial high-frequency information) and insufficient feature utilization.

[0005] Therefore, how to improve the MRI motion artifact correction process in the existing technology to overcome the shortcomings of the existing technology that it cannot fully cope with various complex head movement situations and solve the technical problems of missing information between image modalities and insufficient feature utilization. Summary of the Invention

[0006] The purpose of the present invention is to provide a motion artifact adaptive correction method based on multimodal MRI fusion to overcome the shortcomings of the existing technology that it cannot fully cope with various complex head motion situations and solve the problems of missing information and insufficient feature utilization between image modalities.

[0007] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0008] The method for adaptive correction of motion artifacts based on multimodal MRI fusion includes the following steps:

[0009] S1: acquiring a structural magnetic resonance image (sMRI) and a functional magnetic resonance image (fMRI) of a patient, and preprocessing the structural magnetic resonance image (sMRI) and the functional magnetic resonance image (fMRI);

[0010] S2: creating a cross-modal feature fusion model, and inputting the pre-processed structural magnetic resonance image sMRI and functional magnetic resonance image fMRI into the cross-modal feature fusion model for dual-branch feature extraction;

[0011] S3: The cross-modal feature fusion model performs cross-modal feature interactive fusion on the extracted dual-branch features and performs multi-scale feature reconstruction;

[0012] S4: Detect image artifacts based on the reconstructed multi-scale features and perform dynamic spatiotemporal attention correction on the detected image artifacts;

[0013] S5: Establish a three-level loss function system for multi-task joint optimization.

[0014] Preferably, the preprocessing of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI in step S1 is implemented by a multimodal data preprocessing module, and the multimodal data preprocessing module includes a motion artifact simulation unit and a multimodal registration unit. The specific process of the preprocessing is as follows:

[0015] S11: Applying an elastic deformation field with a specified displacement and rotation angle to the structural magnetic resonance image (sMRI) through a motion artifact simulation unit;

[0016] S12: Use time-varying motion parameters on functional magnetic resonance images (fMRI) and perform image enhancement processing by randomly rotating them at a specified angle, scaling them by a specified multiple, and injecting Rician noise.

[0017] S13: Implementing a hierarchical registration strategy through a multimodal registration unit to achieve registration of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI.

[0018] Preferably, the specific process of implementing the hierarchical registration strategy by the multimodal registration unit in step S13 is as follows:

[0019] S131: Standardize structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI) to unify their grayscale range, resolution, and spatial size;

[0020] S132: Segmenting the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI, extracting key anatomical structures as regions of interest, and performing morphological processing on the regions of interest to provide a specified number of registration feature points;

[0021] S133: Define 6-DOF registration parameters, which include three translation parameters, i.e., x-axis translation, y-axis translation, and three rotation parameters, i.e., Euler angles around x-axis, y-axis, and z-axis.

[0022] S134: using mutual information or normalized mutual information to measure the statistical dependence between images, and using mean square error or normalized cross correlation for same-modality images;

[0023] S135: setting the initial transformation of the structural magnetic resonance image (sMRI) and the functional magnetic resonance image (fMRI) to the identity transformation, downsampling the images to 1 / 4 or 1 / 8 of their original size, and quickly obtaining rough registration parameters to avoid falling into local optimality;

[0024] S136: Gradually restore the image to the original resolution, use the low-resolution result as the initial value, refine the registration parameters, select at least ten pairs of landmark points in the region of interest, such as vascular bifurcation points and brain sulcus endpoints, calculate their spatial distance error and direction deviation, and determine whether the anatomical structure is aligned;

[0025] S137: Cubic B-spline interpolation is used to construct the deformation field. The grid node spacing is set to 2 mm to ensure that overfitting is avoided while preserving local details. A bending energy penalty term is added to regularize the deformation field to constrain the smoothness of the deformation and prevent non-physiological deformation. The mutual information in the rigid registration stage is continued, or local mutual information is introduced to enhance the registration accuracy of local structures.

[0026] Preferably, the specific process of performing dual-branch feature extraction using the cross-modal feature fusion model in step S2 is as follows:

[0027] S21: performing convolution processing on the structural magnetic resonance image sMRI through the initial convolution layer of the first branch of the cross-modal feature fusion model. The specific convolution calculation formula is as follows:

[0028] ;

[0029] in, The input structural magnetic resonance image sMRI, stride is the step size, and 7×7×7 is the convolution kernel size; It represents the feature map obtained after processing by the initial convolutional layer of the first branch, which is the output result of the convolution operation on the structural magnetic resonance image (sMRI). Represents a 3D convolution operation (3DConvolution), with a convolution kernel size of 7×7×7 (in three-dimensional space, the size of each dimension is 7), and a stride of 2, that is, the convolution kernel moves 2 pixels each time when sliding on the image (in each dimension of the three-dimensional space);

[0030] S22: Processing is performed at the residual block level of the first branch of the cross-modal feature fusion model. The specific calculation formula is as follows:

[0031] ;

[0032] in, It is the feature map processed by the residual block of the previous layer (the l-1 layer), which serves as the input of the current l-layer residual block; R l For the improved 3D residual block, the number of channels is {64, 128, 256, 512}, and the output feature resolution is {8mm 3 , 4mm 3 , 2mm 3 , 1mm 3}.

[0033] S23: Convolution processing is performed on the functional magnetic resonance image (fMRI) through the 3D convolution layer of the second branch of the cross-modal feature fusion model. The specific calculation formula is as follows:

[0034] ;

[0035] in, It represents the feature map obtained after processing by the second branch 3D convolution layer, which is the output result of the functional magnetic resonance image (fMRI) convolution operation; Represents a three-dimensional convolution operation with a convolution kernel size of 3×3×3, meaning the kernel size is 3 in each dimension of the three-dimensional space. The default step size is 1. is the input functional magnetic resonance image data. Here, t:t+9 indicates that a series of fMRI images from time t to time t+9 are selected as input in the time dimension to capture the dynamic characteristics of the functional magnetic resonance image over a period of time.

[0036] S24: Perform LSTM time series modeling. The calculation formula is as follows:

[0037] ;

[0038] in, It represents the functional magnetic resonance image (fMRI) feature representation obtained after long short-term memory network (LSTM) processing at time t; Represents an LSTM layer with 128 hidden units. The number of hidden units determines the memory capacity and processing power of the LSTM layer. 128 hidden units means that the layer can maintain a 128-dimensional hidden state when processing sequence data. It is the feature sequence input to the LSTM layer. It is composed of the feature maps of the previous 3D convolutional layer output in the time dimension from time t−9 to time t, which are 10 time steps. It is used to capture the dynamic change information of the fMRI image over a period of time.

[0039] S25: Perform hierarchical feature alignment on the output of the two branches and perform channel dimension splicing at five scales. The specific formula is as follows:

[0040] ;

[0041] in, For the k Layer sMRI feature maps, For the k Layer fMRI feature maps, is the output splicing feature, which represents the feature map obtained by splicing the channel dimension at a certain scale after hierarchical feature alignment at time t; It is a channel dimension splicing operation, which connects two feature maps with the same spatial dimensions (height, width, depth, in the 3D image scenario) but different number of channels in the channel dimension; The representative value ranges from 1 to 5, representing 5 different scales.

[0042] S26: Generate dynamic attention. The specific process is as follows:

[0043] ;

[0044] in, , It represents the channel-adaptive attention map generated at the kth scale, whose element values ​​are between [0,1]. The larger the value, the more important the feature of the corresponding position or channel. σ is the Sigmoid function, and the formula is It maps the input value to the [0,1] interval to generate attention weights; Is the convolution kernel weight (convolution kernel), the shape is 1×1×1×( C sMRI + C fMRI )×1; C sMRI and C fMRI They are and The number of channels is obtained by combining the convolution kernel with the concatenated feature map. Perform convolution operation to calculate attention weights.

[0045] Preferably, the cross-modal feature fusion model in step S3 performs cross-modal feature interactive fusion on the extracted dual-branch features and performs multi-scale feature reconstruction in the following specific process:

[0046] S31: Perform weighted feature fusion of the two branches. The specific formula is as follows:

[0047] ;

[0048] Among them, ⊙ represents element-by-element multiplication, To output the fusion features, , achieving a dynamic balance between anatomical and functional characteristics; It is a channel-adaptive attention map generated at the kth scale, whose element values ​​are between [0,1], used to measure F sMRI The importance of the feature.

[0049] S32: Perform multi-scale feature reconstruction and reverse pyramid aggregation, with a final output resolution of 1mm³. The specific formula is as follows:

[0050] ;

[0051] in, , W k is the kernel parameter. Represents the final output features after multi-scale feature reconstruction and inverse pyramid aggregation; For the k Layer fusion feature map; Here is the fusion of features from k=1 to k=5 at 5 different scales. U k calculate( k =1, resolution 8mm 3 , k =5, resolution 1mm 3 ); U k It is a 3×3×3 deconvolution upsampling operation.

[0052] Preferably, in step S4, the image artifacts are detected based on the reconstructed multi-scale features, and the specific process of performing dynamic spatiotemporal attention correction on the detected image artifacts is as follows:

[0053] S41: Calculate the gradient amplitude through the Sobel operator with a kernel size of 3×3×3, perform bilinear interpolation, and generate a spatial attention map. The calculation formula is:

[0054] ;

[0055] in,F is a three-dimensional feature map; Represents the generated spatial attention map, which is used to indicate the importance of different spatial locations in the image for artifact detection and correction. Norm is a normalization operation, the purpose of which is to map the calculated values ​​to the appropriate range [0,1] interval for subsequent processing and as attention weights; 、 They represent the gradients of the feature map F in the x and z directions respectively.

[0056] S42: Temporal attention generation based on LSTM motion modeling. The specific formula is as follows:

[0057] ; ;

[0058] in, are the 6-DOF head motion parameters, is the output time-varying attention weight; Represents a long short-term memory (LSTM) layer with 64 hidden units; m t Is a characteristic vector, which clarifies that the unit of displacement change Δxt is millimeter (mm), and the rotation angle The unit is radians (rad). This vector is used as input to LSTM to represent motion information.

[0059] S43: A generator model based on the U-Net3+ architecture is configured with multi-scale skip connections (5 scales) and temporal convolution branches (kernel size 3×3×3). It outputs rectified structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI). The module can adaptively identify the spatiotemporal distribution characteristics of head motion artifacts.

[0060] S44: Perform multi-scale feature jump connection,

[0061] ;

[0062] in, Represents the feature map obtained after the multi-scale feature jump connection operation at the kth scale; Represents a three-dimensional deconvolution (also called transposed convolution) operation with a convolution kernel size of 1×1×1; Concat is a concatenation operation, where different feature maps are connected in the channel dimension or other appropriate dimensions to fuse multiple feature information; E k is the feature map at the kth scale in the encoder path, T k is the time-related feature map (from the temporal convolution branch) at the kth scale, Dk+1 is the feature map at the k+1th scale in the decoder path; U is 2 times upsampling, k Inversely aggregate multi-scale features from 5 to 1.

[0063] Preferably, after performing multi-scale feature jump connection in S44, structural magnetic resonance image sMRI and functional magnetic resonance image fMRI reconstruction are also performed;

[0064] The reconstruction formula of structural magnetic resonance image sMRI is as follows:

[0065] ;

[0066] in, Represents the reconstructed structural magnetic resonance image, which is the corrected sMRI image output by the model; S1 is the multi-scale fusion feature with a resolution of 1mm³ and a number of channels C =64, w s is the convolution kernel, b s is the bias, tanh is the activation function, tanh constrains the output value range to [−1,1], adapts to the MRI signal range, and outputs the corrected sMRI image;

[0067] The functional magnetic resonance image fMRI reconstruction formula is as follows:

[0068] ;

[0069] in, Represents the reconstructed functional magnetic resonance image, which is the corrected fMRI image output by the model; represents the summation operation, which accumulates the relevant calculation results of the five time steps from time t−5 to time t−1; is the feature map at time point τ, wf is the convolution kernel, Wτ is the time-varying weight, and satisfies ∑τWτ=1; It is a bias term used to adjust the result of the convolution operation and increase the flexibility of model fitting.

[0070] Preferably, the specific process of establishing a three-level loss function system for multi-task joint optimization in step S5 is as follows:

[0071] S51: Calculate the anatomical fidelity loss. The specific formula is as follows:

[0072]

[0073] in, Represents the anatomical fidelity loss, which is used to measure the similarity between the reconstructed image (corrected image) and the real image in terms of anatomical structure. The smaller the value, the closer the reconstructed image is to the real anatomical structure. The structural similarity between the corrected image Icor and the real image Igt is calculated. Its value range is [0,1]. The closer the value is to 1, the more similar the structure is. The image after calculation I cor With real images I gt The L1 norm of the gradient. ∇ represents the gradient operator. The L1 norm is the sum of the absolute values ​​of each element in the vector. It is used to measure image gradient differences and reflect differences in structural information such as image edges.

[0074] S52: Calculate the structural similarity. The specific formula is as follows:

[0075] ;

[0076] in, It represents the structural similarity index, which is used to measure the structural similarity of images. Here, (x, y, z) corresponds to the spatial coordinates of the image. In the magnetic resonance image scenario, it represents the three dimensions of depth, height, and width. μ is the local mean. and They are the corrected images I cor and real image I gt The mean value of reflects the overall brightness level of the image; σ is the local standard deviation, and The corrected images are I cor and real images I gt The standard deviation reflects the discrete degree of the image pixel value, that is, the contrast of the image; is the corrected image I cor and real image I gt The covariance of the two pixels is used to measure the correlation between the pixel value changes; C1 and C2 are two constants used to avoid the situation where the denominator is zero. C 1=(0.01 L ) 2 , C 2=(0.03 L ) 2 ; L is the signal range.

[0077] S53: Enhance edge sharpness through gradient constraint. The calculation formula is as follows:

[0078] ;

[0079] in, Represents the gradient amplitude of image I, which is used to measure the intensity of image changes in space and reflect the structural information of the image such as the edge; 、 、 are the partial derivatives of image I in the x, y, and z directions, respectively, indicating the rate of change of the image in the corresponding directions. ∇ is the Sobel operator;

[0080] S54: Calculate the function retention loss. The calculation formula is as follows:

[0081] ;

[0082] in, Represents function preservation loss, which is used to measure the difference between the function-related features after processing (such as correction) and the true function features. The smaller the loss value, the closer the processed features are to the true situation in terms of function. ρ is a metric function, indicating FC cor and FC gt the degree of difference between Indicates function-related features after correction or conversion; Represents the true function-related features, that is, the functional features used as reference standards.

[0083] S55: Calculate the cross-modal consistency loss. The calculation formula is as follows:

[0084] ;

[0085] in, Represents the cross-modal consistency loss, which is used to ensure that during the cross-modal conversion process, the image can be restored to its original state as much as possible after the cyclic conversion. The smaller the loss value, the better the consistency of the cyclic conversion. It is a generator network that converts structural magnetic resonance images (sMRI) into functional magnetic resonance images (fMRI) and then converts the obtained fMRI back to sMRI. is the original structural magnetic resonance image, which serves as the input of the entire cyclic conversion process; Represents the square of the L2 norm.

[0086] S56: Create a loop generator and perform pre-training and joint training. The pre-training formula is as follows:

[0087] ;

[0088] in, Represents the model initialization parameters; argmin: is a mathematical operator whose function is to find the parameter value that minimizes the value of the subsequent expression; Represents the anatomical fidelity loss, which is used to constrain the generated image to be consistent with the actual situation in terms of anatomical structure, and measures the similarity between the reconstructed image and the real image in terms of anatomical structure; Represents the function preservation loss, which is used to ensure that the function-related features are preserved during the generation process and measure the difference between the processed function-related features and the true function features. Its constraints are: freezing the cross-modal attention module parameters, single-modality independent training, and the sMRI branch only optimizing L anatomy , the fMRI branch only optimizes L function ;

[0089] The joint training formula is as follows:

[0090] ;

[0091] in, Represents the parameters of the model finally obtained through joint training optimization; argmin : The function of this operator is to find the parameter value that can minimize the value of the subsequent expression; Represents the cross-modal consistency loss, which ensures the consistency of the image with the original image after cross-modal conversion (such as from sMRI to fMRI and then back to sMRI). The smaller the loss, the higher the consistency of the conversion; Represents the anatomical fidelity loss, which is used to constrain the generated image to be consistent with the actual situation in terms of anatomical structure, and measures the similarity between the reconstructed image and the real image in terms of anatomical structure; It represents the function preservation loss, which is used to ensure that the function-related features are retained during the generation process and measures the difference between the processed function-related features and the true function features.

[0092] S57: Enable all modules to achieve end-to-end optimization.

[0093] The beneficial effects of the present invention include:

[0094] The present invention provides a method for adaptive motion artifact correction based on multimodal MRI fusion. The method obtains a patient's sMRI and fMRI images, preprocesses the sMRI and fMRI images, creates a cross-modal feature fusion model, and inputs the preprocessed sMRI and fMRI images into the cross-modal feature fusion model for dual-branch feature extraction. The cross-modal feature fusion model then interactively fuses the extracted dual-branch features and reconstructs multi-scale features. Image artifacts are detected based on the reconstructed multi-scale features, and dynamic spatiotemporal attention correction is performed on the detected image artifacts. Finally, a three-level loss function system is established for multi-task joint optimization. By integrating sMRI and fMRI information and based on the aforementioned image processing and deep learning, the method overcomes the shortcomings of existing technologies, which are unable to fully cope with various complex head motion situations, and addresses the technical issues of missing information between image modalities and insufficient feature utilization.

[0095] First, by adopting a motion artifact simulation unit and a multimodal registration unit, and through a multimodal data preprocessing and enhancement strategy of physical simulation and data enhancement technology, the motion artifact and registration error problems in sMRI and fMRI data were solved, and the registration accuracy and data quality were improved.

[0096] Secondly, through the cross-modal feature fusion network architecture of the dual-branch feature extractor and the cross-modal attention mechanism, the spatiotemporal feature fusion of sMRI and fMRI is achieved, significantly improving the feature representation ability of multimodal data.

[0097] Thirdly, through the dynamic spatiotemporal attention correction mechanism of artifact detection and generator network combined with spatiotemporal attention weights, head motion artifacts are innovatively corrected dynamically, which improves data quality and ensures high-fidelity sMRI and fMRI output.

[0098] Finally, by establishing a three-level loss function system and combining a multi-task joint optimization strategy of anatomical fidelity, function preservation, and cross-modal consistency loss, the model is optimized during a multi-stage training process to ensure optimal performance and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 Schematic diagram of the process of the motion artifact adaptive correction method based on multimodal MRI fusion of the present invention.

[0100] Figure 2 Schematic diagram of the cross-modal feature fusion network architecture of the present invention.

[0101] Figure 3 Schematic diagram of the dynamic spatiotemporal attention correction module architecture of the present invention. DETAILED DESCRIPTION

[0102] The following is combined with Figures 1 to 3 The present invention is described in further detail:

[0103] Example 1

[0104] See attached Figure 1 As shown, the motion artifact adaptive correction method based on multimodal MRI fusion includes the following steps:

[0105] S1: acquiring a structural magnetic resonance image (sMRI) and a functional magnetic resonance image (fMRI) of a patient, and preprocessing the structural magnetic resonance image (sMRI) and the functional magnetic resonance image (fMRI);

[0106] S2: creating a cross-modal feature fusion model, and inputting the pre-processed structural magnetic resonance image sMRI and functional magnetic resonance image fMRI into the cross-modal feature fusion model for dual-branch feature extraction;

[0107] S3: The cross-modal feature fusion model performs cross-modal feature interactive fusion on the extracted dual-branch features and performs multi-scale feature reconstruction;

[0108] S4: Detect image artifacts based on the reconstructed multi-scale features and perform dynamic spatiotemporal attention correction on the detected image artifacts;

[0109] S5: Establish a three-level loss function system for multi-task joint optimization.

[0110] In this embodiment, the preprocessing of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI in step S1 is implemented by a multimodal data preprocessing module. The multimodal data preprocessing module includes a motion artifact simulation unit and a multimodal registration unit. The specific process of the preprocessing is as follows:

[0111] S11: Applying an elastic deformation field with a specified displacement and rotation angle to the structural magnetic resonance image (sMRI) through a motion artifact simulation unit;

[0112] S12: Use time-varying motion parameters on functional magnetic resonance images (fMRI) and perform image enhancement processing by randomly rotating them at a specified angle, scaling them by a specified multiple, and injecting Rician noise.

[0113] S13: Implementing a hierarchical registration strategy through a multimodal registration unit to achieve registration of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI.

[0114] The motion artifact simulation unit uses physical modeling technology to apply an elastic deformation field with a maximum displacement of 3mm and a rotation angle of ±5° for sMRI. For fMRI, time-varying motion parameters (maximum displacement of 1.5mm / TR at TR = 2s) are used. Data augmentation utilizes random rotation (±10°), scaling (0.9-1.1x), and Rician noise injection (SNR = 20-30dB). The multimodal registration unit implements a hierarchical registration strategy, first performing 6-degree-of-freedom rigid registration, followed by B-spline nonlinear registration (with a grid spacing of 2mm), ensuring registration accuracy of <0.5mm for translation and <0.5° for rotation.

[0115] The specific process of implementing the hierarchical registration strategy by the multimodal registration unit in step S13 is as follows:

[0116] S131: Standardize the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI to unify their grayscale range, resolution, and spatial size.

[0117] Traverse each pixel of sMRI and fMRI images and calculate the grayscale value distribution. Calculate the minimum grayscale value of the image I min and maximum value I max , the original gray value I (x,y,z) by the formula: I′ (x,y,z)=255×( I (x,y,z)− I min ) / ( I max − I min ) is mapped to the standard grayscale range of [0, 255]. If there are abnormal peaks in the image grayscale distribution, adaptive histogram equalization is used to divide the image into multiple sub-blocks and perform histogram equalization within each sub-block to avoid excessive noise enhancement.

[0118] Calculate the spatial size difference between sMRI and fMRI images and determine the image translation and scaling ratio. Use the affine transformation matrix to translate and scale the image. For translation, if the sMRI image needs to be translated to the right in the x-axis direction, t x units, shift upward along the y-axis t y units, moving forward along the z axis t z units, then the affine transformation matrix T is:

[0119] .

[0120] Multiply the image coordinate vector by the transformation matrix to obtain the transformed image coordinates and complete the spatial size unification.

[0121] S132: Segment the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI, extract key anatomical structures as regions of interest, and perform morphological processing on the regions of interest to provide a specified number of registration feature points.

[0122] A deep learning segmentation algorithm is used, and in this embodiment, a U-Net network is employed. First, a large number of annotated sMRI and fMRI image datasets are collected and divided into training, validation, and test sets. During training, the images are fed into the U-Net network. The network extracts image features through convolutional layers, restores image resolution through upsampling and skip connections, and outputs a probability map indicating that each pixel belongs to a different anatomical structure category. A cross-entropy loss function is used to calculate the difference between the predicted results and the annotated images. The network parameters are updated using a backpropagation algorithm to continuously optimize the model. After training is complete, the sMRI and fMRI images are fed into the trained model to obtain segmented images, separating the different anatomical structures.

[0123] Morphological operations are performed on the mask image of the region of interest. First, a dilation operation is used to define a structuring element, which can be square or circular. The mask image is traversed. When the center of the structuring element is located at a foreground pixel in the mask image, all pixels covered by the structuring element are set as foreground pixels, enhancing structural continuity and filling small holes. Then, an erosion operation is used to center the structuring element at a foreground pixel in the mask image. Only when all pixels covered by the structuring element are foreground pixels is the center pixel retained as foreground, removing noise and isolated pixels, making the boundary of the region of interest clearer and providing a specified number of registration feature points.

[0124] S133: Define the 6-DOF registration parameters. These parameters include three translation parameters (x, y, and z axis translations), whose value range is determined by the image size and registration requirements and can be set to [-10mm, 10mm], and three rotation parameters (Euler angles around the x, y, and z axes), whose value range is [-π, π]. In the subsequent registration process, the spatial transformation of the image is achieved by adjusting the DOF registration parameters.

[0125] S134: Mutual information or normalized mutual information is used to measure the statistical dependence between images. For homomodal images, mean square error or normalized cross correlation processing is used.

[0126] For heteromodal images of sMRI and fMRI, mutual information (MI) or normalized mutual information (NMI) is used to measure the statistical dependence between images. Mutual information is based on information theory and is calculated by calculating the joint probability distribution of the grayscale values ​​of the two images. p ( I sMRI ,I fMRI ) and marginal probability distribution p ( I sMRI ), p ( I fMRI ), using the formula MI = H ( I sMRI )+ H ( I fMRI )− H ( I sMRI ,I fMRI ), where H is the information entropy. Normalized mutual information is the normalized mutual information, and the formula is NMI = H ( I sMRI ) H ( I fMRI )MI, so that the similarity metric value is between [0, 1], which is easier to compare and optimize. During the calculation process, the histogram is used to statistically analyze the distribution of image grayscale values ​​and approximate the probability distribution.

[0127] S135: The initial transformation of the structural magnetic resonance image (sMRI) and the functional magnetic resonance image (fMRI) is set to the identity transformation, and the images are downsampled to 1 / 4 or 1 / 8 of the original size to quickly obtain rough registration parameters to avoid falling into a local optimum.

[0128] S136: Gradually restore to the original resolution, use the low-resolution result as the initial value, refine the registration parameters, select at least ten pairs of landmark points in the region of interest, such as vascular bifurcation points and brain sulcus endpoints, calculate their spatial distance error and direction deviation, and determine whether the anatomical structure is aligned.

[0129] The image resolution is gradually restored, doubling the image resolution each time until it returns to the original resolution. The registration parameters obtained in the previous low-resolution stage are used as the initial values ​​for the current high-resolution stage. The optimization algorithm is continued to be used. Based on the similarity metric, the 6-DOF registration parameters are further adjusted within a smaller parameter search range to further improve the similarity metric between images. Within the region of interest, at least ten pairs of landmark points are manually or automatically selected, such as vascular bifurcations and sulcus endpoints. The spatial distance error and directional deviation of these landmark points in the registered image are calculated. If the spatial distance error is less than the set threshold and the directional deviation is less than the set angle, the anatomical structure is considered aligned and the refined registration is completed; otherwise, the registration parameters are continued to be adjusted until the requirements are met.

[0130] S137: Cubic B-spline interpolation is used to construct the deformation field. The grid node spacing is set to 2 mm to ensure that overfitting is avoided while preserving local details. A bending energy penalty term is added to regularize the deformation field to constrain the smoothness of the deformation and prevent non-physiological deformation. The mutual information in the rigid registration stage is continued, or local mutual information is introduced to enhance the registration accuracy of local structures.

[0131] A deformation field is constructed using cubic B-spline interpolation. The image space is divided into a grid with a grid node spacing of 2 mm. For each pixel in the image, a cubic B-spline basis function is constructed based on the surrounding grid nodes. The pixel's displacement in the deformation field is calculated using a linear combination of the basis functions. To ensure that the deformation field preserves local details while avoiding overfitting, a bending energy penalty term is added to regularize the deformation field. This penalty term calculates the sum of the squares of the second-order derivatives of the deformation field to constrain the smoothness of the deformation and prevent non-physiological deformations. During the optimization process, the mutual information used in the rigid registration stage is continued, or local mutual information is introduced. The image is divided into multiple local regions, and the mutual information is calculated within each local region to enhance the registration accuracy of local structures. Using an optimization algorithm, the deformation field parameters are continuously adjusted to minimize the objective function consisting of the similarity measure and the bending energy penalty term. The final deformation field is obtained, completing the image registration.

[0132] Example 2

[0133] Based on Example 1, the specific process of performing dual-branch feature extraction using the cross-modal feature fusion model in step S2 is as follows:

[0134] S21: performing convolution processing on the structural magnetic resonance image sMRI through the initial convolution layer of the first branch of the cross-modal feature fusion model. The specific convolution calculation formula is as follows:

[0135] ;

[0136] in, It represents the feature map obtained after the initial convolution layer of the first branch. It is the output result of the convolution operation on the input structural magnetic resonance image (sMRI) and is used for subsequent feature extraction and processing. The input structural magnetic resonance image sMRI; Represents a three-dimensional convolution operation with a kernel size of 7×7×7 (in three-dimensional space, the size of each dimension is 7) and a stride of 2, which means that when the kernel slides across the image, it moves 2 pixels in each dimension at a time, resulting in an output resolution of , where C1=64.

[0137] S22: Processing is performed at the residual block level of the first branch of the cross-modal feature fusion model. The specific calculation formula is as follows:

[0138] ;

[0139] in, It is the feature map processed by the residual block of the previous layer (the l-1 layer). As the input of the current l-layer residual block, it carries the feature information extracted by the previous layer. l = 2, 3, 4, 5 correspond to the output of the residual block of different layers, which are used for further feature extraction and transformation. l For the improved 3D residual block, the number of channels is {64, 128, 256, 512}, and the output feature resolution is {8mm 3 , 4mm 3 , 2mm 3 , 1mm 3}, represents the lth residual block, which is a special neural network module. By introducing skip connections, the network can learn the residual relationship between input and output, effectively alleviate the gradient vanishing problem in deep neural networks, and enhance the network's expression ability and training stability.

[0140] S23: Convolution processing is performed on the functional magnetic resonance image (fMRI) through the 3D convolution layer of the second branch of the cross-modal feature fusion model. The specific calculation formula is as follows:

[0141] ;

[0142] in, This is the feature map obtained after processing by the second branch 3D convolution layer; Represents a three-dimensional convolution operation, where the size of the convolution kernel is 3×3×3; Represents the input functional magnetic resonance image data. Here, t:t+9 indicates a series of fMRI images from time t to time t+9 in the time dimension. is the input functional magnetic resonance image fMRI, , the window length of the time window is 10TR, TR=2s; the convolution kernel is 3×3×3, the step size is 1, and the filling mode is "SAME"; the output feature map: resolution: , where the number of channels C=64.

[0143] S24: Perform LSTM time series modeling. The calculation formula is as follows:

[0144] ;

[0145] in, It represents the functional magnetic resonance image (fMRI) feature representation obtained after long short-term memory network (LSTM) processing at time t. It integrates the information of multiple previous time steps and is a feature extraction of fMRI time series data at the current moment, which is used for subsequent feature fusion or task analysis. Represents an LSTM layer with 128 hidden units. The number of hidden units determines the memory capacity and processing power of the LSTM layer. 128 hidden units means that the layer can maintain a 128-dimensional hidden state when processing sequence data, enabling it to capture more complex time series dependencies. It is the feature sequence input to the LSTM layer. It is composed of the feature maps of the previous 3D convolutional layer output from time t-9 to time t in the time dimension. The number of LSTM units is 128, and the input sequence is: sliding time window t -9: t (window length 10TR, step length 5TR), output: ; H , W , D is the spatial dimension, T is the number of time points, total duration .

[0146] S25: Perform hierarchical feature alignment on the output of the two branches and perform channel dimension splicing at five scales. The specific formula is as follows:

[0147] ;

[0148] in, It represents the feature map obtained by splicing the channel dimension at a certain scale after hierarchical feature alignment at time t. ; It is a channel dimension splicing operation; is the k-th layer sMRI feature map, ; is the k-th layer fMRI feature map.

[0149] S26: Generate dynamic attention. The specific process is as follows:

[0150] ;

[0151] in, It is the feature map obtained by splicing the channel dimension at the kth scale, which integrates the feature information of the two branches (sMRI and fMRI) as the input for calculating the attention map; It is a bias term that adds a learnable offset to the convolution operation result to improve the model fitting ability; It represents the channel-adaptive attention map generated at the kth scale, whose element values ​​are between [0,1]. The larger the value, the more important the feature of the corresponding position or channel. σ is the Sigmoid function, and the formula is It maps the input value to the [0,1] interval to generate attention weights; is the convolution kernel weight, , the convolution kernel shape is 1×1×1×( C sMRI + C fMRI )×1; C sMRI and C fMRI They are and The number of channels is obtained by combining the convolution kernel with the concatenated feature map. Perform convolution operation to calculate attention weights.

[0152] See also Figure 2 The cross-modal feature fusion model sets a dual-branch processing channel, the first branch processing channel includes a first input layer, an initial convolution layer, a residual block level, and a feature pyramid layer, the second branch processing channel includes a second output layer, a 3D convolution layer, a spatiotemporal feature extraction time window, an LSTM layer, and a feature tensor output layer, the cross-modal feature fusion model also includes a feature splicing channel, a convolution layer, an activation layer, a generated attention layer, and a weighted fusion layer, and the feature splicing channel is respectively connected to the feature pyramid layer of the first branch processing channel and the feature tensor output layer of the second branch processing channel.

[0153] Example 3

[0154] On the basis of Example 1 or Example 2, the specific process of the cross-modal feature fusion model in step S3 performing cross-modal feature interactive fusion on the extracted dual-branch features and performing multi-scale feature reconstruction is as follows:

[0155] S31: Perform weighted feature fusion of the two branches. The specific formula is as follows:

[0156] ;

[0157] Among them, ⊙ represents element-by-element multiplication; To generate spatial / channel adaptive attention maps; is the k-th layer sMRI feature map; is the k-th layer fMRI feature map; To output the fusion features, , achieving a dynamic balance between anatomical and functional characteristics;

[0158] S32: Perform multi-scale feature reconstruction and reverse pyramid aggregation, with a final output resolution of 1mm³. The specific formula is as follows:

[0159] ;

[0160] in, Represents the final feature map obtained after multi-scale feature reconstruction and reverse pyramid aggregation. This feature map integrates the fused feature information at five scales and is the final output of the model in the feature processing stage; is the k-th layer fusion feature map, k=1, resolution 8mm 3 , k=5, resolution 1mm 3 , , U k is a 3×3×3 deconvolution upsampling operation, Wk is the kernel parameter, Step size: 25-k (16 steps when k=1, 11 steps when k=5), output resolution: 1mm 3 × Cout (Cout=64).

[0161] In this embodiment, see Figure 3 The dynamic spatiotemporal attention correction module architecture is as follows: In step S4, image artifacts are detected based on the reconstructed multi-scale features, and the specific process of dynamic spatiotemporal attention correction of the detected image artifacts is as follows:

[0162] S41: The gradient amplitude is calculated by the Sobel operator with a kernel size of 3×3×3 in the artifact detection unit in the dynamic spatiotemporal attention correction module architecture, and bilinear interpolation is performed to generate a spatial attention map. The calculation formula is:

[0163] ;

[0164] in, Represents the generated spatial attention map, whose element values ​​reflect the importance of different positions in the image in space. The larger the value, the more attention the corresponding position deserves; Represents the normalization operation, the purpose of which is to map the calculated gradient amplitude to a suitable range; 、 are the partial derivatives of the feature map F in the x direction and the z direction respectively; F is a three-dimensional feature map, , resolution 2mm 3 , Sobel operator: three-directional convolution kernel (taking the x direction as an example), , (3×3 kernel, extended to 3×3×3 three dimensions), gradient calculation: , , resolution improvement: bilinear interpolation (actually trilinear interpolation: ,in, Represents a high-resolution spatial attention map with a resolution of 0.5mm 3 , output spatial attention map: Represents a trilinear interpolation operation, where ×2 means increasing the resolution of the original image by a factor of 2 in depth, height, and width. Represents the generated spatial attention map, whose element values ​​reflect the importance of different positions in the image in space. The larger the value, the more attention the corresponding position deserves. .

[0165] S42: Temporal attention generation based on LSTM motion modeling. The specific formula is as follows:

[0166] ;

[0167] ;

[0168] in, represents a long short-term memory (LSTM) layer with 64 hidden units, are the 6-DOF head motion parameters, is the output time-varying attention weight; represents the motion feature vector at time t.

[0169] S43: The generator model based on the U-Net3+ architecture is configured with multi-scale skip connections (5 scales) and temporal convolution branches (kernel size 3×3×3), and outputs corrected structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI). The module can adaptively identify the spatiotemporal distribution characteristics of head motion artifacts.

[0170] The encoder structure formula is:

[0171] ;

[0172] in, Represents the feature map of the lth layer after convolution operation in the generator model based on the U-Net+ architecture; Represents a three-dimensional convolution operation with a kernel size of 3×3×3 and a stride of 2; is the feature map of the previous layer (the l-1 layer), which serves as the input of the current l-th layer convolution operation and carries the feature information extracted by the previous layer; the input is , convolution kernel 3×3×3, stride 2, the number of output channels is C l , , output resolution .

[0173] Temporal convolution branch:

[0174] ;

[0175] Parameter description: Tl: represents the feature map obtained after the convolution operation of the lth layer in the temporal convolution branch; This is a three-dimensional convolution operation with a kernel size of 3×3×3; It is the channel dimension splicing operation. Temporal attention weight , splicing operation:

[0176] , ⊕ represents the channel splicing of features and time domain weights, the 3D temporal convolution kernel slides in the spatial dimension and is weighted in the time dimension;

[0177] S44: Perform multi-scale feature jump connection,

[0178] ;

[0179] in, Represents the fused feature map obtained after specific operations in the process of multi-scale feature jump connection; That is, a three-dimensional deconvolution (up-sampling convolution) operation, with a convolution kernel size of 1×1×1, U being a 2-fold upsampling, and k being reversely aggregated from 5 to 1 for multi-scale features; Concat represents a channel dimension concatenation operation. It combines feature maps E from different scales k 、T k and D k+1 Splicing along the channel dimension to integrate multiple feature information; E k : It is the feature map of the kth layer in the generator model based on the U-Net3+ architecture, which contains the spatial feature information of the image at that scale; T k : It is the feature map obtained after the convolution operation of the kth layer in the temporal convolution branch, which integrates the temporal attention information and spatial feature information, and reflects the dynamic characteristics in the time dimension; D k+1: It is a feature map from a deeper layer (layer k+1), with more abstract semantic information and different spatial dimensions (after downsampling).

[0180] After S44 performs multi-scale feature jump connection, structural magnetic resonance image sMRI and functional magnetic resonance image fMRI reconstruction are also performed; the reconstruction formula of structural magnetic resonance image sMRI is as follows:

[0181] ;

[0182] in, represents the reconstructed structural magnetic resonance image (sMRI); represents the hyperbolic tangent activation function, whose function value is mapped to the range [-1, 1], introducing nonlinear factors, enhancing the expressive ability of the model, adapting to the MRI signal range, and outputting the corrected sMRI image; : Resolution 1mm³, number of channels C =64, a feature map obtained after the multi-scale feature jump connection operation (here K is 1); : Weight matrix (convolution kernel), used to adjust the input features , perform linear transformation; : Bias term, which increases the fitting ability of the model and adjusts the results of linear transformation.

[0183] The functional magnetic resonance image fMRI reconstruction formula is as follows:

[0184] ;

[0185] in, represents the reconstructed functional magnetic resonance image (fMRI); Represents the weight coefficient, which is used to weight the calculation results of different time steps (or scale-related) during the summation process; Weight matrix (convolution kernel), for input features Perform linear transformation to learn the mapping relationship between features and reconstructed functional magnetic resonance images; It is the representation of the feature map S1 obtained by multi-scale feature jump connection at different time steps τ (value range: t to t-5); Represents the bias term, increases the model's fitting ability, and adjusts the result of the linear transformation. And the formula satisfies ∑ τ W τ =1. (Time point τ feature map), convolution kernel , bias , time-varying weight W τ ∈[0,1] (attention weights from LSTM, satisfying ∑τ W τ =1), and output the corrected fMRI time point signal.

[0186] Example 4

[0187] On the basis of Example 1 or Example 2 or Example 3, the specific process of establishing a three-level loss function system for multi-task joint optimization in step S5 is as follows:

[0188] S51: Calculate the anatomical fidelity loss. The specific formula is as follows:

[0189] ;

[0190] in, =SSIM stands for anatomical fidelity loss, which is used to measure the degree of similarity between the anatomical structure of the reconstructed or corrected image and the real image. The lower the loss value, the closer the anatomical structure of the reconstructed image is to the real situation. SSIM: Structural Similarity Index, which is used to evaluate the structural similarity between two images. Its value range is between [−1, 1]. The closer the value is to 1, the more similar the image structures are. Icor: Corrected image, which is the output image after model processing. Igt: Real image, which serves as a reference standard for evaluating the quality of the corrected image. ∇: Gradient operator, used to calculate the gradient of the image. It represents the L1 norm between the calculated corrected image and the true image gradient, that is, the sum of the absolute values ​​of the corresponding pixel gradient value differences, reflecting the difference in detail information such as image edges.

[0191] S52: Calculate the structural similarity. The specific formula is as follows:

[0192] ;

[0193] in, The structural similarity index is used to quantify the structural similarity between two images (here x, y, and z are regarded as the spatial coordinates of the images). The value range is between -1 and 1. The closer the value is to 1, the more similar the image structures are. It is the mean of the reconstructed image Icor and the real image Igt, reflecting the overall brightness level of the image. μ is the local mean, σ is the local standard deviation, C 1=(0.01 L ) 2 , C 2=(0.03 L ) 2 , L is the signal range,

[0194] S53: Enhance edge sharpness through gradient constraint. The calculation formula is as follows:

[0195] ;

[0196] Where ∇ is the Sobel operator; 、 、 are the partial derivatives of image I in the x, y, and z directions, respectively. These partial derivatives measure the rate of change of the pixel values ​​of the image in the corresponding directions.

[0197] S54: Calculate the function retention loss. The calculation formula is as follows:

[0198] ;

[0199] in, Indicates Functional Preservation Loss; is the Euclidean distance; Represents the corrected and transformed feature map; Represents the ground-truth feature map, which is the feature map used as the reference standard.

[0200] S55: Calculate the cross-modal consistency loss. The calculation formula is as follows:

[0201] ;

[0202] in Cycle-Consistency Loss represents the cross-modal consistency loss, which is used to measure the difference between the image after cyclic conversion and the original image during the cross-modal conversion process; It is a generator that converts the structural magnetic resonance image (sMRI) into a functional magnetic resonance image (fMRI), and then converts the obtained fMRI back to sMRI; It represents the square of the L2 norm, which is calculated for the vector of image pixel values.

[0203] The loop generator G contains 3 levels of residual blocks with the following structure:

[0204] .

[0205] Where G(x) represents the output of the loop generator G for the input x; x is the original image data input to the loop generator G; F1 and F2 represent the convolution operation or other nonlinear transformation operations in the residual block; It means that the input x is first transformed by F1, and then the result of F1 transformation is used as input for F2 transformation.

[0206] S56: Create a loop generator and perform pre-training and joint training,

[0207] The pre-training formula is as follows:

[0208] ;

[0209] in, Represents the initialization parameter, which is the initial parameter value obtained by optimization under the constraints of a specific loss function and is used for subsequent generator training; argmin is an operator whose function is to find the parameter value that minimizes the value of the subsequent expression; Represents the loss related to anatomical structure, which is used to constrain the image generated by the generator to be consistent with the actual anatomical structure. For example, in medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense. Denotes function preservation loss. Its constraints are: freezing the cross-modal attention module parameters, independent training of a single modality, optimizing only Lanatomy in the sMRI branch and only Lfunction in the fMRI branch;

[0210] The joint training formula is as follows:

[0211]

[0212] in, The parameters obtained by final optimization are the optimal parameter values ​​found under the constraints of multiple loss functions, which are used to determine the final model parameters of the loop generator. represents the cross-modal consistency loss; Represents the loss related to anatomical structure, which is used to constrain the image generated by the generator to be consistent with the actual anatomical structure. For example, in medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense. Indicates loss of functionality.

[0213] S57: Enable end-to-end optimization for all modules. Optimizer configuration: Adam β1=0.9, β2=0.999, ϵ=10−8; initial learning rate: 10−4, decay ×0.5 every 20k steps; batch size: 8 sMRI-fMRI data pairs.

[0214] This method, based on multimodal MRI fusion, adaptively corrects motion artifacts. Paired artifact-laden sMRI and fMRI data are input and registered to generate a multimodal data cube. Anatomical and functional features are then extracted in parallel, with feature interactions performed at four scales. A spatiotemporal attention map is then dynamically generated and reconstructed by region. Finally, artifact-free sMRI and high-fidelity fMRI images are output, along with a quality control report containing pre- and post-correction comparisons. Key parameters are: the number of convolution kernels for feature extraction (64, 128, 256, 512), the number of LSTM units (128), and a learning rate decay of 0.5 every 20k steps.

[0215] In summary, the present invention provides a method for adaptive correction of motion artifacts based on multimodal MRI fusion. The method obtains the patient's sMRI and fMRI, preprocesses the sMRI and fMRI, creates a cross-modal feature fusion model, and inputs the preprocessed sMRI and fMRI into the cross-modal feature fusion model for dual-branch feature extraction. The cross-modal feature fusion model performs cross-modal feature interactive fusion on the extracted dual-branch features and performs multi-scale feature reconstruction. Image artifacts are detected based on the reconstructed multi-scale features, and dynamic spatiotemporal attention correction is performed on the detected image artifacts. A three-level loss function system is established for multi-task joint optimization. By integrating sMRI and fMRI information, based on the above-mentioned image processing and deep learning, the shortcomings of the existing technology that cannot fully cope with various complex head movement situations are overcome, and the technical problems of missing information between image modalities and insufficient feature utilization are solved.

[0216] By employing a motion artifact simulation unit and a multimodal registration unit, and through a multimodal data preprocessing and enhancement strategy combining physical simulation and data augmentation techniques, the team addressed motion artifacts and registration errors in sMRI and fMRI data, improving registration accuracy and data quality. A cross-modal feature fusion network architecture, combining a dual-branch feature extractor with a cross-modal attention mechanism, enabled the fusion of spatiotemporal features from sMRI and fMRI, significantly enhancing the feature representation capabilities of multimodal data. A dynamic spatiotemporal attention correction mechanism, combining artifact detection and generator networks with spatiotemporal attention weights, innovatively corrected head motion artifacts dynamically, improving data quality and ensuring high-fidelity sMRI and fMRI output. By establishing a three-level loss function system and combining a multi-task joint optimization strategy that integrates anatomical fidelity, functional preservation, and cross-modal consistency losses, the team optimized the model over multiple training stages to ensure optimal performance and accuracy.

Claims

1. A motion artifact adaptive correction method based on multimodal MRI fusion, characterized in that: The following steps are involved: S1: acquiring a structural magnetic resonance image (sMRI) and a functional magnetic resonance image (fMRI) of a patient, and preprocessing the structural magnetic resonance image (sMRI) and the functional magnetic resonance image (fMRI); S2: creating a cross-modal feature fusion model, and inputting the pre-processed structural magnetic resonance image sMRI and functional magnetic resonance image fMRI into the cross-modal feature fusion model for dual-branch feature extraction; S3: The cross-modal feature fusion model performs cross-modal feature interactive fusion on the extracted dual-branch features and performs multi-scale feature reconstruction; S4: Detect image artifacts based on the reconstructed multi-scale features and perform dynamic spatiotemporal attention correction on the detected image artifacts; S5: Establish a three-level loss function system for multi-task joint optimization; The specific process of preprocessing is as follows: S11: Applying an elastic deformation field with a specified displacement and rotation angle to the structural magnetic resonance image (sMRI) through a motion artifact simulation unit; S12: Use time-varying motion parameters on functional magnetic resonance images (fMRI) and perform image enhancement processing by randomly rotating them at a specified angle, scaling them by a specified multiple, and injecting Rician noise. S13: implementing a hierarchical registration strategy through a multimodal registration unit to achieve registration of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI; In step S4, the image artifacts are detected based on the reconstructed multi-scale features, and the specific process of dynamic spatiotemporal attention correction of the detected image artifacts is as follows: S41: Calculate the gradient magnitude using the Sobel operator with a kernel size of 3×3×3, perform bilinear interpolation, and generate a spatial attention map. S42: Temporal attention generation based on LSTM motion modeling; S43: A generator model based on the U-Net3+ architecture is configured with multi-scale skip connections and temporal convolution branches to output corrected structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI). The module can adaptively identify the spatiotemporal distribution characteristics of head motion artifacts. S44: Perform multi-scale feature jump connection. The specific formula is as follows: ; in, Represents the feature map obtained after the multi-scale feature jump connection operation; u 2x upsampling; k The value 5-1 means reverse aggregation of multi-scale features from 5 to 1; Represents a three-dimensional convolution operation with a convolution kernel size of 1×1×1. It adjusts the number of channels and performs linear transformation on the input feature map, changing the number of channels without changing the spatial dimensions (length, width, and depth) of the feature map. Indicates that E k 、T k 、D k+1 The three feature maps are concatenated in the channel dimension, E k 、T k 、D k+1 These are feature maps from different scales, different processing stages, or different branches.

2. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 1, characterized in that: The preprocessing of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI in step S1 is implemented by a multimodal data preprocessing module, and the multimodal data preprocessing module includes a motion artifact simulation unit and a multimodal registration unit.

3. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 1, characterized in that: The specific process of implementing the hierarchical registration strategy by the multimodal registration unit in step S13 is as follows: S131: Standardize structural magnetic resonance images (sMRI) and functional magnetic resonance images (fMRI) to unify their grayscale range, resolution, and spatial size; S132: Segmenting the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI, extracting key anatomical structures as regions of interest, and performing morphological processing on the regions of interest to provide a specified number of registration feature points; S133: defining 6-DOF registration parameters including 3 translation parameters and 3 rotation parameters; S134: Measure the statistical dependence between images, using mean square error or normalized cross-correlation processing for same-modality images; S135: setting the initial transformation of the structural magnetic resonance image sMRI and the functional magnetic resonance image fMRI to the identity transformation, and downsampling the images to a specified ratio of the original size; S136: gradually restore to the original resolution, use the low-resolution result as the initial value, select at least ten pairs of landmark points in the region of interest, calculate their spatial distance error and direction deviation, and determine whether the anatomical structure is aligned; S137: Construct a deformation field, add a bending energy penalty term to regularize the deformation field, continue the mutual information of the rigid registration stage or introduce local mutual information to enhance the registration accuracy of local structures.

4. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 1, characterized in that: The specific process of performing dual-branch feature extraction using the cross-modal feature fusion model in step S2 is as follows: S21: performing convolution processing on the structural magnetic resonance image sMRI through the initial convolution layer of the first branch of the cross-modal feature fusion model; S22: Processing through the residual block level of the first branch of the cross-modal feature fusion model; S23: performing convolution processing on the functional magnetic resonance image (fMRI) through the 3D convolution layer of the second branch of the cross-modal feature fusion model; S24: Perform LSTM time series modeling. The modeling formula is as follows: ; in, represents the feature sequence of 10 time steps from t−9 to t extracted by CNN at time step t. CNN is used to extract features from the spatial dimension of fMRI data; Represents an LSTM layer with 128 hidden units. Using 128 hidden units means that the dimension of the LSTM internal memory unit and other structures is 128, which can learn more complex time series features; It is the final fMRI feature representation obtained at time step t after LSTM processing, which is used for subsequent analysis. The number of LSTM units is 128. t -9: t is a sliding time window, is the output; H , W , D is the spatial dimension, T is the time point; S25: Perform hierarchical feature alignment on the output of the two branches and perform channel dimension splicing at five scales. The specific formula is as follows: ; in, For the k Layer sMRI feature map; For the k Layer fMRI feature maps; To output the splicing features; The representative value ranges from 1 to 5, representing 5 different scales; Refers to the operation of tensor splicing in the channel dimension in deep learning; S26: Generate dynamic attention. The specific process is as follows: ; in, express The weight matrix has five dimensions and is responsible for linearly transforming the input features. 1×1×1 means that the convolution kernel weights of the first three dimensions are all 1, and the convolution kernel weights of the fourth dimension are all 1. represents the number of structural magnetic resonance imaging feature channels ( ) and the number of functional magnetic resonance imaging feature channels ( ); the convolution kernel weight of the last dimension is also 1; To generate spatial / channel adaptive attention maps; σ Sigmoid function, output value range [0, 1]; To output the splicing features; is a bias term used to increase the fitting ability of the model.

5. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 4, characterized in that: The specific process of the cross-modal feature fusion model in step S3 performing cross-modal feature interactive fusion on the extracted dual-branch features and performing multi-scale feature reconstruction is as follows: S31: Perform weighted feature fusion of the two branches. The specific formula is as follows: ; in, express The dimension information of the k-th feature map after fusion; among them, Represents the size of the feature map in the height direction; Represents the size of the feature map in the width direction; Represents the size of the feature map in the depth direction; represents the number of structural MRI feature channels; To generate spatial / channel adaptive attention maps; For the k Layer sMRI feature map; No. k Layer fMRI feature map; ⊙ represents element-wise multiplication; The kth feature map after fusion achieves a dynamic balance between anatomical and functional features; S32: Perform multi-scale feature reconstruction and reverse pyramid aggregation, with a final output resolution of 1mm³. The specific formula is as follows: ; in, Represents the final output features after multi-scale feature reconstruction and inverse pyramid aggregation; For the k Layer fusion feature map; Here is the fusion of features from k=1 to k=5 at 5 different scales. u k calculate, k =1, resolution 8mm 3 , k =5, resolution 1mm 3 ; u k It is a 3×3×3 deconvolution upsampling operation.

6. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 1, characterized in that: After S44 performs multi-scale feature jump connection, structural magnetic resonance image (sMRI) and functional magnetic resonance image (fMRI) reconstruction are also performed; The reconstruction formula of structural magnetic resonance image sMRI is as follows: ; in, represents the reconstructed structural magnetic resonance image (sMRI); Represents the hyperbolic tangent activation function, whose function value is mapped to [-1, 1], introducing nonlinear factors and enhancing the expressive power of the model; : A feature map obtained after the multi-scale feature jump connection operation (here K is 1); : Weight matrix (convolution kernel), used to adjust the input features , perform linear transformation; : Bias term, which increases the fitting ability of the model and adjusts the results of linear transformation; The functional magnetic resonance image fMRI reconstruction formula is as follows: ; in, represents the reconstructed functional magnetic resonance image (fMRI); Represents the weight coefficient, which is used to weight the calculation results related to different time steps or scales during the summation process; Weight matrix, for input features Perform linear transformation to learn the mapping relationship between features and reconstructed functional magnetic resonance images; It is the representation of the feature map S1 obtained by multi-scale feature jump connection at different time steps τ; the value range of τ is: t to t-5; Represents the bias term, increases the model's fitting ability, adjusts the result of the linear transformation, and satisfies ∑ τ w τ =1.

7. The motion artifact adaptive correction method based on multimodal MRI fusion according to claim 1, characterized in that: The specific process of establishing a three-level loss function system for multi-task joint optimization in step S5 is as follows: S51: Calculate the anatomical fidelity loss. The specific formula is as follows: ; in, Represents the anatomical fidelity loss, which is used to measure the similarity between the reconstructed image and the real image in anatomical structure. The smaller the loss value, the closer the reconstructed image is to the real image. It stands for structural similarity index, which is used to measure the structural similarity between two images. Its value range is limited to [-1,1]. The closer it is to 1, the more similar the image structures are. SSIM ( I cor , I gt )middle, I cor is the reconstructed image, I gt is the true reference image; ∇ represents the gradient operator; The calculation is to reconstruct the image I cor With real images I gt Gradient L 1 norm, used to measure the difference in structural information such as image edges; S52: Calculate the structural similarity. The specific formula is as follows: ; in, The structural similarity index is used to quantify the structural similarity between two images. The value range is between -1 and 1. The closer the value is to 1, the more similar the image structures are. It is the mean of the reconstructed image Icor and the real image Igt, reflecting the overall brightness level of the image. u is the local mean, σ is the local standard deviation, C 1=(0.01 L ) 2 , C 2=(0.03 L ) 2 , L is the signal range; S53: Enhance edge sharpness through gradient constraint. The calculation formula is as follows: ; Where ∇ is the Sobel operator; 、 、 are the partial derivatives of image I in the x, y, and z directions, respectively. The partial derivatives measure the rate of change of the pixel values ​​of the image in the corresponding directions; S54: Calculate the function retention loss. The calculation formula is as follows: ; in, Indicates Functional Preservation Loss; is the Euclidean distance; Represents the corrected and transformed feature map; Represents the true value feature map, that is, the feature map used as the reference standard; S55: Calculate the cross-modal consistency loss. The calculation formula is as follows: ; in, Represents the cross-modal consistency loss, which is used to measure the difference between the image after cyclic transformation and the original image during the cross-modal transformation process; It is a generator that converts the structural magnetic resonance image (sMRI) into the functional magnetic resonance image fMRI, and then converts the obtained fMRI back to sMRI; Indicates the square of the L2 norm, which is calculated for the vector of image pixel values; S56: Create a loop generator and perform pre-training and joint training. The pre-training formula is as follows: ; in, Represents the initialization parameter, which is the initial parameter value obtained by optimization under the constraints of a specific loss function and is used for subsequent generator training; argmin is an operator used to find the parameter value that minimizes the value of the subsequent expression; Represents the loss related to anatomical structure, which is used to constrain the image generated by the generator to be consistent with the actual anatomical structure. In medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense; Indicates loss of functionality; The constraints are: freezing the cross-modal attention module parameters, independent training of the single modality, and optimizing the sMRI branch only. L anatomy , the fMRI branch only optimizes L function ; The joint training formula is as follows: ; in, The parameters obtained by final optimization are the optimal parameter values ​​found under the constraints of multiple loss functions, which are used to determine the final model parameters of the loop generator. represents the cross-modal consistency loss; Represents the loss related to anatomical structure, which is used to constrain the image generated by the generator to be consistent with the actual anatomical structure. In medical image generation, it is necessary to ensure that the shape and position of organs conform to anatomical common sense; Indicates loss of functionality; S57: Enable end-to-end optimization of all modules, optimizer configuration: Adam β 1=0.9, β 2=0.999, ϵ =10 −8 ; Initial learning rate: 10 −4 , decay × 0.5 every 20k steps; batch size: 8 pairs of sMRI-fMRI data.

Citation Information

Patent Citations

  • Cancer accompanying depression identification method based on multi-mode magnetic resonance data

    CN113705680A

  • Deep learning electroencephalogram noise reduction method based on time-frequency domain information fusion

    CN114947883A