Medical image registration method based on multilayer perceptron
By constructing a medical image registration method of a multi-layer perceptron, using multi-scale feature extraction and correlation perception registration decoder, the existing method solves the problem of capturing long-distance dependence and representing spatial correspondence at full resolution, and achieves a high accuracy and robust medical image registration effect.
Patent Information
- Application Number
- CN202510194071.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing medical image registration methods based on convolutional neural networks are difficult to capture long-distance dependence of fine-grained size at full resolution, and lack of effective models to represent the spatial correspondence between mobile and fixed images, limiting their application in clinical practice.
Using a medical image registration method based on a multi-layer perceptron, an unstructured correlation-aware multi-layer perceptron is designed to capture the multi-scale dependency dependence of local and global features by constructing a multi-scale feature extraction encoder and an associated perception registration decoder, and using the correlation between the multi-scale feature map and multiple registration steps as complementary information.
Effectively capture the long-distance dependence of fine-grained size at full resolution, improves the accuracy of feature matching, enhances the robustness of the model, ensures that each registration step can utilize the results of the previous step, and solves the problem of local deformation of traditional methods when dealing with large-scale or high-resolution images.
Smart Images

Figure CN120047504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of deep learning and medical image registration, and particularly relates to a medical image registration method based on a multi - layer perceptron. Background Art
[0002] Medical image registration is a key technology in the field of medical image analysis. When registering two images, one image (the moving image) is mapped to another image (the fixed image) so that points with the same anatomical meaning on the two images correspond to each other, thereby realizing the matching and fusion of information. In clinical diagnosis, doctors often need to compare multiple images of a patient, such as computed tomography (CT) images, magnetic resonance imaging (MRI) images, positron emission tomography (PET) images, and ultrasound images, to obtain more comprehensive information. Through medical image registration, doctors can accurately compare the positions and morphological characteristics of lesions, organ structures, or functional regions in the images, so as to better understand the development of diseases and their impacts.
[0003] In the past decade, convolutional neural networks (CNNs) have become the focus of medical image research. Mok T.C et al. proposed a fast symmetric diffeomorphic image registration method based on convolutional neural networks, which achieved efficient medical image registration while maintaining the symmetry and diffeomorphic characteristics of the displacement field. Cao et al. used a convolutional neural network to map each input image patch of a pair of 3D brain MR volume data to its respective displacement vector, and then added these displacement vectors to obtain the final registration displacement field. Ito M et al. used a convolutional neural network to learn reasonable deformations to generate a realistic displacement field. DeVos B.D et al. used an unsupervised convolutional neural network (CNN) method for deformable image registration, which not only calculates the displacement field of the image but also ensures the consistency of anatomical structures during image deformation. Deep registration methods based on convolutional neural networks (CNNs) have been widely applied to fast end - to - end registration. These methods learn the mapping from image pairs to spatial transformations based on a set of training data and outperform traditional methods in terms of registration performance. Golland P et al. proposed DiffPose, a self - supervised framework for differentiable registration of intraoperative 2D X - ray images and preoperative CT scans, and finally achieved sub - millimeter registration accuracy on a surgical dataset. Li et al. proposed a modality - independent structural representation learning method, which uses deep neighborhood self - similarity (DNS) and contrast learning to learn discriminative and contrast - invariant deep - structure image representations and performs well in multi - modality registration tasks. Liu Lei et al. used a convolutional neural network for feature detection and description and introduced a graph neural network GNN for feature matching, and this method achieved excellent performance in registration tasks with large - range rotation and scale changes.
[0004] Traditional image registration methods often struggle to maintain the accuracy of significant structural boundaries or discontinuous regions when processing images. Chen et al. proposed a Deep Discontinuity-Preserving Image Registration Network (DDIR), which effectively captures and retains the details and boundary information in images by extracting image features at multiple scales and levels, focusing on boundaries and significant feature regions, and avoiding blurring and distorting these important regions during the registration process. Balakrishnan G et al. proposed VoxelMorph based on the U-Net framework for deformable registration of brain MRI. VoxelMorph consists of an encoder-decoder structure, where the encoder is responsible for extracting high-level features of the image, and the decoder maps these features to the displacement field. However, existing registration methods based on convolutional neural networks lack an effective model to represent the spatial correspondence between the moving image and the fixed image, limiting the wide application of related methods in clinical practice.
[0005] In recent years, vision transformers have effectively addressed the deficiencies of convolutional neural networks in this regard. The Transformer architecture provides a larger receptive field for the model, helping to capture long-range dependency information and achieve higher accuracy. Chen et al. proposed ViT-V-Net, which integrates the vision transformer (ViT) into a V-Net style convolutional network. The ViT-V-Net model consists of a hybrid architecture of convolutional layers and Transformer layers. To effectively propagate information, it uses long skip connections between the encoder and the decoder. The output of the decoder is a dense displacement field to achieve the deformation of anatomical structures. To improve the limitations of traditional convolutional neural networks (CNNs), such as local receptive fields and difficulties in processing long-range dependencies. Ma M et al. proposed a new symmetric Transformer network called SymTrans, which adds a multi-head self-attention mechanism to the convolutional neural network and constructs a symmetric Transformer network architecture to model long-range spatial associations across images. Yang et al.'s GraformerDIR uses graph convolution to effectively model the non-Euclidean spatial structure in images and combines the long-range dependency capture ability of the Transformer to address the trade-off between registration accuracy and computational complexity in existing methods. However, although Transformer-based medical image registration methods can capture long-range dependencies between image features, due to the computational and memory consumption issues of the self-attention mechanism, they usually can only operate on low-resolution features and cannot capture fine-grained long-range dependencies at full resolution. This poses a challenge to the precise non-linear deformable registration of complex organs such as the cerebral cortex. Secondly, existing methods mainly focus on the matching between the features of two images and do not fully explore the association between the coarse-to-fine registration steps. Summary of the Invention
[0006] The object of the present invention is to provide a multi-layer perceptron-based medical image registration method to solve the above technical problems.
[0007] To this end, the technical solution of the present invention is as follows:
[0008] A multi-layer perceptron-based medical image registration method is realized by inputting a medical image into a trained deep learning-based single-modal and multi-modal medical image registration network; the deep learning-based single-modal and multi-modal medical image registration network is composed of a multi-scale feature extraction encoder and a correlation-aware registration decoder; wherein,
[0009] The multi-scale feature extraction encoder is composed of a first coding module, a second coding module, a third coding module, and a fourth coding module connected in sequence; each coding module is composed of a first convolutional layer, a second convolutional layer, a Leaky ReLU activation function, an instance normalization module, and a max pooling layer connected in sequence;
[0010] The correlation-aware registration decoder is composed of a first decoding module, a second decoding module, a third decoding module, and a fourth decoding module; the first decoding module is composed of a UCA-MLP, an R-Head, a displacement field module, an upsampling module, and a transformation module connected in sequence; the second decoding module and the third decoding module are both composed of a first UCA-MLP, a second UCA-MLP, an R-Head, a summation module, a displacement field module, an upsampling module, and a transformation module connected in sequence; the fourth decoding module is composed of a first UCA-MLP, a second UCA-MLP, an R-Head, a summation module, and a displacement field module connected in sequence;
[0011] The output ends of the R-Head, upsampling module, and transformation module of the first decoding module are respectively connected to the input ends of the second UCA-MLP, summation module, and first UCA-MLP of the second decoding module; the output ends of the R-Head, upsampling module, and transformation module of the second decoding module are respectively connected to the input ends of the second UCA-MLP, summation module, and first UCA-MLP of the third decoding module; the output ends of the R-Head, upsampling module, and transformation module of the third decoding module are respectively connected to the input ends of the second UCA-MLP, summation module, and first UCA-MLP of the fourth decoding module;
[0012] The output ends of the first encoding module are respectively connected to the conversion module of the third decoding module and the input ends of the first UCA-MLP of the fourth decoding module. The output ends of the second encoding module are respectively connected to the conversion module of the second decoding module and the input ends of the first UCA-MLP of the third decoding module. The output ends of the third encoding module are respectively connected to the conversion module of the first decoding module and the input ends of the first UCA-MLP of the second decoding module. The output end of the fourth encoding module is connected to the input end of the UCA-MLP of the first decoding module.
[0013] Further, in the multi-scale feature extraction encoder, the convolution kernel sizes of the first convolutional layer and the second convolutional layer of each encoding module are set to 3×3×3; the parameter of the Leaky ReLU activation function is set to 0.2; the window size of the max pooling layer is set to 2×2×2.
[0014] Further, in the correlation-aware registration decoder, each UCA-MLP is composed of a correlation layer, a layer splicing module, a third convolutional layer, a layer normalization module, a first perception module, a second perception module, a third perception module, a first summation module, a local cross-channel attention mechanism, and a second summation module. Among them, the correlation layer, the layer splicing module, the third convolutional layer, and the layer normalization module are connected in sequence. The output end of the layer normalization module is respectively connected to the input ends of the first perception module, the second perception module, the third perception module, and the second summation module. The output ends of the first perception module, the second perception module, and the third perception module are respectively connected to the input end of the first summation module. The output end of the first summation module is connected to the input end of the local cross-channel attention mechanism. The output end of the local cross-channel attention mechanism is connected to the input end of the second summation module.
[0015] Further, in each UCA-MLP, the input ends of the layer splicing module are respectively connected to the output end of the correlation layer and the output end of the module connected to the input end of the correlation layer, so as to obtain a spliced image formed by sequentially splicing the fixed image, the moving image, and the correlation image obtained by processing the fixed image and the moving image. The convolution kernel of the third convolutional layer is set to 2×2×2, and its output end is also connected to the input end of the second summation module. The first perception module is composed of a 3×3 sliding window, a gMLP module, and a Region Merge module connected in sequence. The second perception module is composed of a 5×5 sliding window, a gMLP module, and a Region Merge module connected in sequence. The third perception module is composed of a 7×7 sliding window, a gMLP module, and a Region Merge module connected in sequence. In the local cross-channel attention mechanism, K is set to 3.
[0016] Further, the specific implementation steps of the medical image registration method based on the multi-layer perceptron are as follows:
[0017] Step 1: Obtain a medical image set for training, which consists of several medical images with the same organ or the same structure;
[0018] Step 2: Preprocess all the medical images obtained in Step 1 to have the same image specifications, and divide them into a training set, a test set, and a validation set. Moreover, the medical images in each image set are formed in pairs of images. In each pair of images, one image serves as the moving image and the other image serves as the fixed image;
[0019] Step 3: Construct a deep learning-based single-modal and multi-modal medical image registration network;
[0020] Step 4: Use the medical image pairs processed in Step 2 to train the deep learning-based single-modal and multi-modal medical image registration network constructed in Step 3;
[0021] Step 5: Input the medical images to be registered into the deep learning-based single-modal and multi-modal medical image registration network trained in Step 4, and output the registered result images.
[0022] Furthermore, in Step 1, the medical images are CT images and / or MRI images.
[0023] Furthermore, in Step 2, the preprocessing steps of the medical images are: perform affine transformation, resampling, and cropping on the medical images in sequence to preprocess them into images with the same specifications as the MNI-152 meninges version.
[0024] Furthermore, in Step 4, define the loss function as an unsupervised loss function L, and its expression is:
[0025] L = λ 1 ×L sim +λ 2 ×L reg ,
[0026] In the formula, λ 1 and λ 2 are adjustment factors,
[0027] L sim is the similarity penalty loss, and its expression is:
[0028]
[0029] In the formula, represents the average intensity of the image I at the local position P, and p i represents the iteration of n 3 neighborhood regions around P;
[0030] L regis the regularization loss, and its expression is:
[0031]
[0032] In the formula, represents the spatial gradient operator.
[0033] Furthermore, in the unsupervised loss function L of step S4, λ 1 is set to 0.8, and λ 2 is set to 0.2.
[0034] Compared with the prior art, the medical image registration method based on the multi-layer perceptron uses the correlation between multi-scale feature maps and multiple registration steps as complementary information to provide key context information for each registration step. Secondly, in order to capture the multi-scale dependencies of local and global features during the registration process, an unstructured correlation-aware multi-layer perceptron is designed to improve the accuracy of feature matching without losing important details, solving the problem that existing methods cannot capture fine-grained long-range dependencies at full resolution. At the same time, the added correlation information between multi-step registrations further enhances the robustness of the model, ensuring that each registration step can utilize the results of the previous step; finally, to solve the problem of local deformation of traditional registration methods when dealing with large-scale or high-resolution images, UCA-MLP introduces an unstructured correlation-aware module that can effectively capture local deformation and multi-scale dependencies and uses a specific partitioning strategy to obtain local information under different window sizes. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of the medical image registration method based on the multi-layer perceptron of the present invention;
[0036] Figure 2 is an overall framework diagram of the unstructured correlation-aware network based on the multi-layer perceptron of the present invention;
[0037] Figure 3 is a structural schematic diagram of the unstructured correlation-aware multi-layer perceptron module of the present invention;
[0038] Figure 4 is a comparison diagram of registration results of different methods on the SR-Reg dataset in the embodiment of the present invention;
[0039] Figure 5 is a comparison diagram of registration error results of different methods on six brain datasets in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the following embodiments are by no means any limitation to the present invention.
[0041] See Figure 1 , the specific implementation steps of the medical image registration method based on the multi-layer perceptron are described as follows.
[0042] Step 1: Obtain a medical image set for training, which consists of several medical images. The medical images are specifically CT images and / or MRI images of the same organ or the same structure.
[0043] In practical applications, the medical images in Step 1 can be collected by oneself or directly apply existing medical image datasets. In this embodiment, the medical image set uses an existing dataset, specifically including the SR-Reg dataset, ADNI dataset, ABIDE dataset, ADHD dataset, IXI dataset, Mindboggle dataset, and Buckner dataset; among them, the SR-Reg dataset is a dataset containing MRI images and CT images of 3D brains, and the remaining datasets are datasets containing only MRI images of 3D brains.
[0044] Based on this, two image registration tasks are designed in this embodiment. One task is the multi-modal medical image registration task of 3D brain MRI images and CT images of the same patient, and the other task is the single-modal medical image registration task of 3D brain MRI images among the same patients.
[0045] Step 2: Preprocess all the medical images obtained in Step 1 to have the same image specifications, and divide them into a training set, a test set, and a validation set;
[0046] The specific implementation steps of this Step 2 are as follows:
[0047] Step 2.1: Perform affine transformation, resampling, and cropping on the medical images in sequence to preprocess them into images with the same specifications as the MNI-152 meninges version;
[0048] In this embodiment, the SR-Reg dataset is used for the multi-modal medical image registration task. The original size of the medical images selected from the SR-Reg dataset is 192×208×176 voxels, and the resolution is 1×1×1mm 3 , based on this, first resample the voxels to 128×128×128 through affine transformation and resampling to align with the MNI-152 brain template, with isotropic voxels of 1mm 3 , and then crop to the size of 144×192×160 to be consistent with the size of the MNI-152 brain template;
[0049] The ADNI dataset, ABIDE dataset, ADHD dataset, IXI dataset, Mindboggle dataset, and Buckner dataset are used for the single-modal medical image registration task. The medical images selected from these six datasets are also aligned to the MNI-152 brain template through affine transformation and resampling in sequence as described above, with isotropic voxels of 1mm 3 and then cropped to the final size of 144×192×160;
[0050] Step 2.2: Divide the preprocessed images into a training set, a validation set, and a test set;
[0051] In this embodiment, for the multi-modal medical image registration task, 180 cases are randomly selected. Among them, 150 cases are used for training, 10 cases are used for validation, and the remaining 20 cases are used for testing. Specifically, each case consists of CT images and MRI images of the 3D brain of the same patient. For the single-modal medical image registration task, 2,656 brain MRI images are randomly selected from the ADNI dataset, ABIDE dataset, ADHD dataset, and IXI dataset as the training set, and 100 pairs of images are randomly selected from the Mindboggle dataset and Buckner dataset to form 200 test image pairs as the test set;
[0052] Step 2.3: Pair the preprocessed images in the above two training sets. In each pair of preprocessed images, one image is used as the moving image I m , and the other image is used as the fixed image I f ; In this embodiment, the CT image in the multi-modal medical image registration task is used as the moving image, and the MRI image is used as the fixed image.
[0053] Step 3: Construct a deep learning-based single-modal and multi-modal medical image registration network UCANet, including a multi-scale feature extraction encoder and an unstructured coarse-to-fine correlation-aware registration decoder;
[0054] See Figure 2 and Figure 3 , and the specific implementation steps of this step 3 are as follows:
[0055] Step 3.1: Construct a multi-scale feature extraction encoder;
[0056] Specifically, the multi-scale feature extraction encoder is composed of a first encoding module S1, a second encoding module S2, a third encoding module S3, and a fourth encoding module S4 connected in sequence; each encoding module is composed of a first convolutional layer, a second convolutional layer, a Leaky ReLU activation function, an instance normalization module, and a max pooling layer connected in sequence; in this embodiment, the convolutional kernels of the first convolutional layer and the second convolutional layer of each encoding module have a size of 3×3×3 and a stride of 1; the parameter of the Leaky ReLU activation function is set to 0.2; the window size of the max pooling layer is set to 2×2×2.
[0057] The processing steps of this multi-scale feature extraction encoder are as follows: Input the paired moving image I m (Moving Image) and the fixed image I f (Fixed Image) into the multi-scale feature extraction encoder, and pass through four encoding modules to sequentially generate four multi-scale features from the moving image I m and the fixed image I f ; as Figure 2 shown, F m 1 and are respectively the first-scale features of the moving image and the fixed image generated by the first encoding module, F m 2 and are respectively the second-scale features of the moving image and the fixed image generated by the second encoding module, and F m 3 are respectively the third-scale features of the moving image and the fixed image generated by the third encoding module, F m 4 and are respectively the fourth-scale features of the moving image and the fixed image generated by the fourth encoding module.
[0058] Step 3.2: Construct an association-aware registration decoder:
[0059] This association-aware registration decoder uses UCA-MLP (multi-layer perception mechanism) to perform unstructured correlation perception, so as to achieve coarse-to-fine registration; specifically, the association-aware registration decoder is composed of a first decoding module, a second decoding module, a third decoding module, and a fourth decoding module; among them,
[0060] The first decoding module is composed of a UCA-MLP, an R-Head, a displacement field module, an upsampling module, and a transformation module connected in sequence; the second decoding module and the third decoding module are both composed of a first UCA-MLP, a second UCA-MLP, an R-Head, a summing module, a displacement field module, an upsampling module, and a transformation module connected in sequence; the fourth decoding module is composed of a first UCA-MLP, a second UCA-MLP, an R-Head, a summing module, and a displacement field module connected in sequence; wherein, the output end of the R-Head of the first decoding module is further connected to the input end of the second UCA-MLP of the second decoding module, the output end of its upsampling module is further connected to the input end of the summing module of the second decoding module, and the output end of its transformation module is connected to the input end of the first UCA-MLP of the second decoding module; the output end of the R-Head of the second decoding module is further connected to the input end of the second UCA-MLP of the third decoding module, the output end of its upsampling module is further connected to the input end of the summing module of the third decoding module, and the output end of its transformation module is connected to the input end of the first UCA-MLP of the third decoding module; the output end of the R-Head of the third decoding module is further connected to the input end of the second UCA-MLP of the fourth decoding module, the output end of its upsampling module is further connected to the input end of the summing module of the fourth decoding module, and the output end of its transformation module is connected to the input end of the first UCA-MLP of the fourth decoding module.
[0061] In each decoding module, each UCA-MLP has the same structural composition. Each UCA-MLP is specifically composed of a Correlation Layer, a layer splicing module, a third convolutional layer, a LayerNorm module, a first perception module, a second perception module, a third perception module, a first summation module, a local cross-channel attention mechanism, and a second summation module. The Correlation Layer, the layer splicing module, the third convolutional layer, and the LayerNorm module are connected in sequence. The output end of the LayerNorm module is respectively connected to the input ends of the first perception module, the second perception module, the third perception module, and the second summation module. The output ends of the first perception module, the second perception module, and the third perception module are respectively connected to the input end of the first summation module. The output end of the first summation module is connected to the input end of the local cross-channel attention mechanism. The output end of the local cross-channel attention mechanism is connected to the input end of the second summation module. Among them, in addition to being connected to the output end of the Correlation Layer, the input end of the layer splicing module is also connected to the input end of the Correlation Layer to the same output end, so as to obtain a spliced image formed by sequentially splicing a fixed image, a moving image, and a correlation image obtained by processing the fixed image and the moving image. The convolutional kernel of the third convolutional layer is 2×2×2, the stride is 1, and its output end is also connected to the input end of the second summation module. The first perception module is composed of a 3×3 sliding window, a gMLP module, and a Region Merge module connected in sequence. The second perception module is composed of a 5×5 sliding window, a gMLP module, and a Region Merge module connected in sequence. The third perception module is composed of a 7×7 sliding window, a gMLP module, and a Region Merge module connected in sequence. In the local cross-channel attention mechanism, K is set to 3.
[0062] In each decoding module, a displacement field function is set in the displacement field module to parameterize the deformable registration problem into a displacement field. The formula of the displacement field function is: where ψ is the displacement field, and θ is a learnable parameter, which is obtained through unsupervised learning during network training.
[0063] See Figure 3 , the specific processing steps of each UCA-MLP for the input fixed image F A and the moving image F B are as follows: First, input F A and F B into the Correlation Layer, which can calculate the local correlation between the two feature maps in the convolutional feature space and emphasize the local details in the depth feature representation. The correlation calculation formula between the two is:
[0064]
[0065] In the formula, Pi A and P i B respectively represent the central voxels sampled from F A and F B ; i and j respectively represent the 2D central coordinates of the central voxel; n A , n B ∈[-k, k] 3 indicates that n A , n B iterates respectively within a three-dimensional neighborhood centered on P i A and P i B with a size of [-k, k]×[-k, k]×[-k, k]; in this embodiment, k is set to 1;
[0066] Furthermore, a set of points is sampled within a 3D region of size (d×d×d), which can be achieved through 3D convolution; in this embodiment, d = 3 is set to calculate the local correlation within the 3D neighborhood, thereby outputting the feature correlation feature map F C , whose shape is the same as that of the feature maps F A and F B , and the number of channels d 3 = 27;
[0067] Subsequently, F A , F B and F C are concatenated by channels and passed through a third convolutional layer of 3×3×3 to generate the correlation perception feature map F corr ; then, after F corr is processed by LayerNorm, it is input into a multi-layer perception module with different window sizes to capture multi-scale dependencies; specifically, in this embodiment, sliding windows of sizes 3×3, 5×5, and 7×7 are respectively used to divide the feature map into non-overlapping regions, denoted as F RS , which helps to capture fine-grained local information, enhance the sensitivity of the model to features of different scales, and perform local interaction operations on the feature map obtained through region division; the gMLP module is used to process the output of the unstructured layer, and the Region Merge module is used for region integration, and then the features from different regions are fused through weighted summation of channels;
[0068] Next, the feature fusion results of different regions interact using the local cross-channel attention mechanism to establish dependencies between different channels, thereby extracting more representative feature representations; specifically, a one-dimensional convolution of size K along the channel dimension creates dependencies between different channels; subsequently, the Sigmoid activation function is applied to generate the weights of the local channel attention, and its expression is:
[0069] F O = F M × σ(C1D K (F M )),
[0070] In the formula, C1D K represents a one-dimensional convolution with a convolution kernel size of K, F M represents the feature map obtained after region merging, and the final output F O is achieved by multiplying the attention weights element-wise with the original input feature map, and the result is an attention-weighted feature map with updated weights; as the number of channels increases, the range of local cross-channel interactions should also be expanded accordingly; the non-linear expression of this relationship is:
[0071] C = φ(K) = 2 η*K-b ,
[0072] In the formula, K represents the kernel size of the one-dimensional convolution, and η and b are the scaling and offset parameters of the linear transformation;
[0073] Given the total number of channels C, the kernel size K can be adaptively determined as follows:
[0074]
[0075] In this embodiment, |.| odd is defined as rounding to the nearest odd number, η is set to 2, and b is set to 1; then K = 3 is calculated;
[0076] Furthermore, through each UCA-MLP module in the decoding module, the potential correspondence between them is explored through region partitioning and an unstructured deformation layer to specifically capture the dependencies related to deformable medical image registration.
[0077] Step 3.3, construct a network model UCANet for single-modal and multi-modal medical image registration based on deep learning, which is composed of the multi-scale feature extraction encoder constructed in Step 3.1 and the association-aware registration decoder constructed in Step 3.2; among them, the output end of the first encoding module is respectively connected to the input ends of the conversion module of the third decoding module and the first UCA-MLP of the fourth decoding module to input the first-scale feature
[0078] F m 1 of the moving image into the conversion module of the third decoding module, while the first-scale feature The first UCA-MLP input to the fourth decoding module; the output ends of the second encoding module are respectively connected to the conversion module of the second decoding module and the input end of the first UCA-MLP of the third decoding module to transmit the second-scale feature F of the moving image m 2 are input to the conversion module of the second decoding module, while the second-scale feature of the fixed image is input to the first UCA-MLP of the third decoding module; the output ends of the third encoding module are respectively connected to the conversion module of the first decoding module and the input end of the first UCA-MLP of the second decoding module to transmit the third-scale feature F of the moving image m 3 are input to the conversion module of the first decoding module, while the third-scale feature of the fixed image is input to the first UCA-MLP of the second decoding module; the output end of the fourth encoding module is connected to the input end of the UCA-MLP of the first decoding module to transmit the fourth-scale feature F of the moving image m 4 and the fourth-scale feature of the fixed image are input to the first UCA-MLP of the first decoding module.
[0079] In practical applications, the specific processing steps of the network model UCANet for single-modal and multi-modal medical image registration based on deep learning are as follows:
[0080] F output by the fourth encoding module m 4 and are input into the UCA-MLP of the first decoding module, which explores their spatial correspondence relationships to obtain relationship features and then inputs them into the deformable registration head (R-Head). The output results are respectively input to the displacement field module of the first decoding module to map out the displacement field ψ 4 , and this displacement field ψ 4 after upsampling and F output by the third encoding module m 3 are input to the transformation module, enabling F m 3 to be represented in the form of the displacement field ψ 4 : for subsequent guidance of the second decoding module;
[0081] Then, output by the third encoding module It is input into the first UCA-MLP of the second decoding module, and the obtained relational features are then jointly input into the second UCA-MLP of the second decoding module together with the output result of the R-Head of the first decoding module, and then input into the R-Head of the second decoding module. The output result is combined with the upsampled displacement field ψ in the first decoding module 4 Input to the input summation module for summation processing, and then input to the displacement field module of the second decoding module to map out the displacement field ψ 3 This displacement field ψ 3 After upsampling and the F output by the second encoding module m 2 Input to the input transformation module, and the output of which Continues to be used to guide the third decoding module;
[0082] Next, the output of the second encoding module And the output by the second decoding module Are input into the first UCA-MLP of the third decoding module. The obtained relational features are then jointly input into the second UCA-MLP of the third decoding module together with the output result of the R-Head of the second decoding module, and then input into the R-Head of the third decoding module. The output result is combined with the upsampled displacement field ψ in the second decoding module 3 Input to the input summation module for summation processing, and then input to the displacement field module of the third decoding module to map out the displacement field ψ 2 This displacement field ψ 2 After upsampling and the F output by the first encoding module m 1 Input to the input transformation module, and the output of which Continues to be used to guide the fourth decoding module;
[0083] Finally, the F output by the first encoding module f 1 And the output by the third decoding module Are input into the first UCA-MLP of the fourth decoding module. The obtained relational features are then jointly input into the second UCA-MLP of the fourth decoding module together with the output result of the R-Head of the third decoding module, and then input into the R-Head of the fourth decoding module. The output result is combined with the upsampled displacement field ψ in the third decoding module 2 Input to the input summation module for summation processing, and then input to the displacement field module of the fourth decoding module to map out the displacement field ψ 1 This displacement field ψ 1 Aligns I m With I f And its output result is the final registration result.
[0084] Step 4: Use the medical images obtained from Step 2 to train the single-modal and multi-modal medical image registration network based on deep learning constructed in Step 3.
[0085] The specific implementation steps of this Step 4 are as follows:
[0086] Step 4.1: Define the loss function as an unsupervised loss function L, which is composed of a similarity penalty loss L sim and a regularization loss L reg ; among them,
[0087] The similarity penalty loss L sim is used to measure the similarity between the deformed image and the fixed image F f , and its expression is:
[0088]
[0089] In the formula, represents the average intensity of the image I at the local position P, and p i represents the iteration of n 3 neighborhood regions around P;
[0090] The regularization loss L reg is used to impose regularization on the deformable registration transformation ψ to promote smooth and realistic transformations in the physical space, and its expression is:
[0091]
[0092] In the formula, represents the spatial gradient operator;
[0093] Furthermore, the unsupervised loss function L adopts the weighted sum of L sim and L reg , and its expression is:
[0094] L = λ 1 × L sim + λ 2 × L reg ,
[0095] In the formula, λ 1 and λ 2 are adjustment factors used to balance the registration accuracy and smoothness; in this embodiment, λ 1 = 0.8 and λ 2 = 0.2;
[0096] Step 4.2: During the training process, use the unsupervised loss function L to train and optimize the medical image registration network model UCANet until the unsupervised loss function L reaches the minimum value and remains stable.
[0097] Step 5: Input two medical images to be registered (one as the moving image I m , and the other as the fixed image I f ) into the deep learning-based single-modal and multi-modal medical image registration network UCANet trained in Step 4, and the registration results of the two images can be output.
[0098] To further prove the accuracy and usability of the medical image registration method based on multi-layer perceptron of the present invention in terms of registration effect, the Dice similarity coefficient (DSC) and the normalized Jacobian determinant (NJD) are used as evaluation indicators.
[0099] DSC is used to compare the similarity between the segmented region after automatic registration and the true segmented region, while NJD measures the reversibility and smoothness of the local deformation of the image by calculating the Jacobian determinant of the displacement field, and evaluates whether the displacement field after registration maintains the topological structure; the 95% Hausdorff distance (HD95) is used to evaluate the registration accuracy, and the percentage of non-positive Jacobian determinant (P|J| <= 0) is calculated to evaluate the differential characteristics of the displacement field; Mem (GB) is used to represent the memory consumed during training and testing, Par (M) is used to measure the number of parameters of the model, and time represents the time consumption during the execution of the model.
[0100] On the SR-Reg dataset, the performance comparison results of different registration methods are shown in Table 1 below.
[0101] Table 1:
[0102]
[0103]
[0104] It can be seen from the comparison of the performance results in Table 1 that the proposed UCANet is superior to other methods in terms of Dice, HD95, and P metrics, and the best performance is achieved in all three metrics. Specifically, UCANet reaches the highest Dice score of 83.63%, indicating its superior overlap degree between the prediction and the true segmentation result. In addition, UCANet shows the lowest HD95 value (1.33), reflecting its more accurate boundary prediction than VoxelMorph, TransMorph, and VMambaMorph;
[0105] In addition, although the memory usage and model parameters of UCANet (23.72 million) are slightly higher, its significant performance improvement justifies the additional computational cost. The running time of UCANet (0.49 seconds) is higher than that of other methods, but considering the significant improvement in accuracy and precision, this trade-off is reasonable.
[0106] As Figure 4 shown in the schematic diagram of the visual comparison of the registration results of this application and two other methods in Table 1 on the SR-Reg dataset; among them, in actual use, the CT image in the SR-Reg dataset is used as the fixed image, and the nuclear magnetic resonance (MR) image is used as the moving image; from Figure 4 the visual comparison results shown, it can be seen that in the specially marked red and green areas, UCANet is significantly closer to the real situation, while other methods have problems of over-registration and under-registration.
[0107] The performance comparison results of different methods on six brain datasets are shown in Table 2 below.
[0108] Table 2:
[0109]
[0110] From the comparison of the performance results in Table 2, it can be seen that on the Mindboggle dataset, the DSC of UCANet reached 0.645%, which is the highest among all methods, indicating its high registration accuracy. Its NJD percentage is 1.915%, showing the smallest topological distortion in the deformation field; on the Buckner dataset, UCANet continues to outperform other methods, with a DSC of 0.659% and an NJD of 1.825%. On the GPU, its average running time per registration is 0.51 seconds, close to LapIRN (0.52 seconds) and Dual-PRNet++ (0.49 seconds), but slightly slower than TransMorph (0.32 seconds). Although it makes a slight compromise in running time, it is reasonable because it has a significant improvement in registration accuracy and smoothness, as shown by higher DSC and lower NJD; In summary, the above comparison results clearly demonstrate the effectiveness of UCANet proposed by the present invention in coping with the challenges of deformable medical image registration.
[0111] As Figure 5 shown in the comparison of the registration error results of different methods on six brain datasets, which shows the registration error results of different methods in the single-modal brain image registration experiment. The more black areas, the greater the error; obviously, UCANet proposed by the present invention significantly reduces the error, showing excellent accuracy and stability.
[0112] Although the present invention has been described above with reference to the embodiments, various modifications can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the features in the embodiments disclosed in the present invention can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
[0113] References:
[0114] [1] Balakrishnan, G.; Zhao, A.; Sabuncu, M. R.; Guttag, J.; Dalca, A. V. Voxelmorph: a learning framework for deformable medical image registration. IEEE transactions on medical imaging 2019, 38, 1788 - 1800.
[0115] [2] Chen, J.; Frey, E. C.; He, Y.; Segars, W. P.; Li, Y.; Du, Y. Transmorph: Transformer for unsupervised medical image registration. Medical image analysis 2022, 82, 102615.
[0116] [3] Guo, T.; Wang, Y.; Meng, C. Mambamorph: a mamba - based backbone with contrastive feature learning for deformable mr - ct registration. arXiv preprint arXiv:2401.13934 2024.
[0117] [4] Wang, Z.; Zheng, J. Q.; Ma, C.; Guo, T. Vmambamorph: A visual mamba-based framework with cross-scan module for deformable 3d image registration. arXiv preprint arXiv:2404.05105, 2024.
[0118] [5] Zhu, Y.; Lu, S. Swin-voxelmorph: A symmetric unsupervised learning model for deformable medical image registration using swin transformer. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2022, pp. 78 - 87.
[0119] [6] Mok, T. C.; Chung, A. C. Large deformation diffeomorphic image registration with Laplacian pyramid networks. In Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2020: 23rd International Conference, Lima, Peru, October 4 - 8, 2020, Proceedings, Part III 23. Springer, 2020, pp. 211 - 221
[0120] [7]Shu, Y.; Wang, H.; Xiao, B.; Bi, X.; Li, W. Medical image registration based on uncoupled learning and accumulative enhancement. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2021, pp. 3 - 13
[0121] [8]Kang, M.; Hu, X.; Huang, W.; Scott, M. R.; Reyes, M. Dual-stream pyramid registration network. Medical image analysis 2022, 78, 102379.
[0122] [9]Zhou, S.; Hu, B.; Xiong, Z.; Wu, F. Self-distilled hierarchical network for unsupervised deformable image registration. IEEE Transactions on Medical Imaging 2023, 42, 2162 - 2175.
[0123]
[10] Meng, M.; Bi, L.; Feng, D.; Kim, J. Non-iterative coarse-to-fine registration based on single-pass deep cumulative learning. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2022, pp. 88 - 97
[0124]
[11] Meng, M.; Bi, L.; Fulham, M.; Feng, D.; Kim, J. Non-iterative coarse-to-fine transformer networks for joint affine and deformable image registration. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2023, pp. 750 - 760.
Claims
1. A medical image registration method based on a multi-layer perceptron, characterized in that: This is achieved through a single-modality and multi-modality medical image registration network based on deep learning; The encoder is composed of a first encoding module, a second encoding module, a third encoding module and a fourth encoding module connected in sequence; each encoding module is composed of a first convolutional layer, a second convolutional layer, a Leaky ReLU activation function, an instance normalization module and a maximum pooling layer connected in sequence; The decoder is composed of a first decoding module, a second decoding module, a third decoding module and a fourth decoding module connected in sequence; the first decoding module is composed of a UCA-MLP, an R-Head, a displacement field module, an up-sampling module and a transformation module connected in sequence; the second decoding module and the third decoding module are both composed of a first UCA-MLP, a second UCA-MLP, an R-Head, a summing module, a displacement field module, an up-sampling module and a transformation module connected in sequence; the fourth decoding module is composed of a first UCA-MLP, a second UCA-MLP, an R-Head, a summing module and a displacement field module connected in sequence; the output end of the R-Head and the up-sampling module of the first decoding module is also connected to the second UCA-MLP and the input end of the summing module of the second decoding module respectively; the output end of the R-Head and the up-sampling module of the second decoding module is also connected to the second UCA-MLP and the input end of the summing module of the third decoding module respectively; the output end of the R-Head and the up-sampling module of the third decoding module is also connected to the second UCA-MLP and the input end of the summing module of the fourth decoding module respectively; The output end of the first encoding module is respectively connected to the conversion module of the third decoding module and the first UCA-MLP input end of the fourth decoding module, the output end of the second encoding module is respectively connected to the conversion module of the second decoding module and the first UCA-MLP input end of the third decoding module, the output end of the third encoding module is respectively connected to the conversion module of the first decoding module and the first UCA-MLP input end of the second decoding module, and the output end of the fourth encoding module is connected to the UCA-MLP input end of the first decoding module.
2. The medical image registration method based on multi-layer perceptron according to claim 1, characterized in that: In the multi-scale feature extraction encoder, the convolution kernel size of the first convolution layer and the second convolution layer of each encoding module is set to 3×3×3; the parameter of the Leaky ReLU activation function is set to 0.2; and the window size of the maximum pooling layer is set to 2×2×2.
3. The medical image registration method based on multi-layer perceptron according to claim 1, characterized in that: In the association-aware registration decoder, each UCA-MLP is composed of a correlation layer, a layer splicing module, a third convolutional layer, a layer normalization module, a first perception module, a second perception module, a third perception module, a first summation module, a local cross-channel attention mechanism and a second summation module; wherein the correlation layer, the layer splicing module, the third convolutional layer and the layer normalization module are connected in sequence, the output end of the layer normalization module is respectively connected to the input end of the first perception module, the second perception module, the third perception module and the second summation module, the output end of the first perception module, the second perception module and the third perception module are respectively connected to the input end of the first summation module, the output end of the first summation module is connected to the input end of the local cross-channel attention mechanism, and the output end of the local cross-channel attention mechanism is connected to the input end of the second summation module.
4. The medical image registration method based on multi-layer perceptron according to claim 1, characterized in that: In each UCA-MLP, the input end of the layer splicing module is connected to the output end of the relevant layer and the output end of the module connected to the input end of the relevant layer, so as to obtain a spliced image formed by sequentially splicing a fixed image, a moving image, and a related image obtained by processing the fixed image and the moving image; the convolution kernel of the third convolutional layer is set to 2×2×2, and its output end is also connected to the input end of the second summation module; the first-layer perception module is composed of a 3×3 sliding window, a gMLP module, and a RegionMerge module connected in sequence; the second-layer perception module is composed of a 5×5 sliding window, a gMLP module, and a Region Merge module connected in sequence; the third-layer perception module is composed of a 7×7 sliding window, a gMLP module, and a Region Merge module connected in sequence; in the local cross-channel attention mechanism, K is set to 3.
5. The medical image registration method based on multi-layer perceptron according to claim 1, characterized in that: Here are the steps: Step 1, obtaining a medical image set for training, which consists of several medical images with the same organ or the same structure; Step 2, preprocessing all medical images obtained in step 1 to have the same image specifications, and dividing them into a training set, a test set and a validation set, and the medical images in each image set are composed of image pairs, in each image pair, one image is used as a moving image, and the other image is used as a fixed image; Step 3: Construct a single-modality and multi-modality medical image registration network based on deep learning; Step 4: using the medical images processed in step 2 to train the single-modality and multi-modality medical image registration network based on deep learning constructed in step 3; Step 5: Input the medical image to be registered into the single-modality and multi-modality medical image registration network based on deep learning trained in step 4, and output the registration result map.
6. The medical image registration method based on multi-layer perceptron according to claim 5, characterized in that: In step 1, the medical image is a CT image and / or an MRI image.
7. The medical image registration method based on multi-layer perceptron according to claim 5, characterized in that: In step 2, the preprocessing steps of the medical image are: performing affine transformation, resampling and cropping on the medical image in sequence to preprocess it into an image with the same specifications as the MNI-152 meningeal plate.
8. The medical image registration method based on multi-layer perceptron according to claim 5, characterized in that: In step 4, the loss function is defined as the unsupervised loss function L, which is expressed as: L=λ1×L sim +λ2×L reg , In the formula, λ1 and λ2 are adjustment factors, L sim is the similarity penalty loss, which is expressed as: In the formula, represents the average intensity of image I at local position P, p i represents the number of n around P 3 Iteration of neighborhood areas; L reg is the regularization loss, and its expression is: In the formula, represents the spatial gradient operator.
9. The medical image registration method based on multi-layer perceptron according to claim 8, characterized in that: In the unsupervised loss function L in step S4, λ1 is set to 0.8 and λ2 is set to 0.2.
Citation Information
Patent Citations
Brain medical image registration method and system based on local-global information cooperation
CN116309754A
Medical image segmentation method based on boundary perception and attention mechanism
CN117078930A
Deformable medical image registration model and registration method based on Swin Transform
CN117274330A
Brain nuclear magnetic resonance image registration method and device based on deep self-attention network
CN117372484A
Medical image segmentation method based on global and local feature joint learning and multi-scale feature fusion
CN118840548A