Transformer-Based Medical Image Registration Method and System

By combining convolutional networks and Transformer structures in medical image registration, the problems of limited receptive fields and loss of detailed information in the prior art are solved, and more efficient image registration effect and global modeling capabilities are achieved.

CN115170622BActive Publication Date: 2025-05-30FUDAN UNIVERSITY +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210515128.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2025-05-30
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

The existing medical image registration method based on deep learning limits the global modeling capability of the model due to the limited receptive field and downsampling operations of the convolutional layer.

Method used

Using the medical image registration method based on Transformer, a high-level semantic feature map is generated through convolutional network downsampling, and the feature map is expanded into a one-dimensional vector and added to the position embedding vector. Enter the Transformer encoder and decoder module to capture global context information and generate deformation field.

Benefits of technology

It improves the effect of image registration, enhances the fusion of local features and global features, improves the sensitivity to local and global deformation, reduces the amount of calculation and retains image detail information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170622B_ABST
    Figure CN115170622B_ABST
Patent Text Reader

Abstract

The present invention provides a medical image registration method and system based on a transformer, including: Step S1: Obtain a pair of medical images, including a floating image and a fixed image, and perform preprocessing; Step S2: Generate a feature map by downsampling the preprocessed floating image in three stages; Step S3: Construct a transformer network model, expand the feature map generated by downsampling in each stage into a one-dimensional vector, and add a position embedding vector to send it into the transformer encoder module; Step S4: The fixed image generates a vector according to Step S2 and Step S3, and sends it into the transformer structure decoder module; Step S5: Upsample and fuse the features of the vector output by the transformer network model to generate a deformation field; Step S6: Register the medical image to be registered according to the generated deformation field. The present invention can enhance the connection between local information and global information of the image, thereby improving the effect of image registration under unsupervised learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image registration, and in particular, to a medical image registration method and system based on transformer. Background Art

[0002] With the development of medical imaging devices, it is not only possible to acquire images containing accurate anatomical information, but also images containing functional information. Observing images of these modalities requires spatial imagination and subjective experience. By using image registration methods, multi-modal information can be fused into the same image, enabling doctors to observe lesions and results more accurately from various angles. At the same time, by registering dynamic images acquired at different times, the changes in lesions and organs can be quantitatively analyzed, making medical diagnosis, surgical planning, and radiotherapy planning more accurate and reliable.

[0003] Currently, most registration methods based on deep learning use deep convolutional neural networks to generate image deformation fields. Shan et al. used a fully convolutional network to input the entire image to be registered and output a global deformation field, overcoming the limitations of image patches. Ferrante et al. studied an efficient registration network based on the U-Net structure. Li et al. and Jiang et al. guided the training of the registration network by calculating the similarity between the deformed image and the reference image at multiple scales. In order to obtain the global deformation of the image, these methods stack convolutional layers deep enough to cover the global region of the image. In this process, the effective receptive field of the network is limited, severely restricting the global modeling ability of the model. At the same time, continuous downsampling operations also cause the loss of image detail information. With the application of transformer in the field of computer vision, it establishes long-range dependencies from different dimensions to capture global context information, improving the global modeling ability while ensuring efficiency and being more sensitive to global deformation.

[0004] The invention patent with the publication number CN105765626B discloses a registration of medical images, including: receiving a 2D X-ray image acquired by a medical 2D imaging device in a first observation direction; filtering the 2D X-ray image so that the high-frequency components of the 2D X-ray image are enhanced relative to the low-frequency components of the 2D X-ray image; receiving a 3D image acquired by a medical 3D imaging device; generating a 2D projection image according to the 3D image, wherein the 2D projection image is generated in a second observation direction; superimposing the filtered 2D X-ray image and the 2D projection image; providing a function for changing the second observation direction so that the 2D projection image is registered with the filtered 2D X-ray image. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a medical image registration method and system based on Transformer.

[0006] According to a medical image registration method and system based on Transformer provided by the present invention, the solution is as follows:

[0007] In the first aspect, a medical image registration method based on Transformer is provided, and the method includes:

[0008] Step S1: Obtain a pair of medical images, including a floating image and a fixed image, and preprocess the floating image and the fixed image to obtain a preprocessed floating image and a preprocessed fixed image;

[0009] Step S2: Generate feature maps of three different scales and resolutions by downsampling the preprocessed floating image through three stages;

[0010] Step S3: Construct a Transformer network model, including an encoder module and a decoder module, unfold the feature maps generated by each stage of downsampling into one-dimensional vectors, and add position embedding vectors and send them into the Transformer encoder module;

[0011] Step S4: Generate vectors for the preprocessed fixed image according to Steps S2 and S3, and send them into the decoder module of the Transformer structure;

[0012] Step S5: Perform upsampling feature fusion on the vectors output by the Transformer network model to generate a deformation field;

[0013] Step S6: Register the medical image to be registered according to the generated deformation field.

[0014] Preferably, in Step S1, the preprocessing of the floating image and the fixed image is performed by setting the floating image and the fixed image to the same size.

[0015] Preferably, for the three-stage downsampling in Step S2, each stage of downsampling is composed of different numbers of 3D depth residual modules. The height of the floating image x m is H, the width is w, and the thickness is D. The feature map is generated through the first-stage downsampling. C is the number of channels of the feature map. This stage of downsampling is composed of three 3D residual convolutional blocks, and the height and width are reduced to 1 / 4 of the original size, and the thickness is reduced to 1 / 8 of the original thickness;

[0016] The Feature map. The downsampling in this stage consists of two 3D residual convolutional blocks. The height and width of the feature map are reduced to 1 / 8 of the original size, and the thickness is reduced to 1 / 16 of the original thickness;

[0017] Generated through the third-stage downsampling Feature map. This stage consists of two 3D residual convolutional blocks. The size of the feature map is reduced to 1 / 32 of the original size, and the thickness is reduced to 1 / 32 of the original thickness;

[0018] The feature map F after each downsampling CNN Is expressed as:

[0019]

[0020] Among them, l = 3 represents the downsampling of three stages, θ represents the parameters of the residual network in each stage, C represents the dimension of the 3D feature map after downsampling; R represents the set of feature map dimensions; x m Represents the floating image.

[0021] Preferably, the step S3 includes: In the encoder module, it includes a multi-head self-attention layer and a feed-forward neural network layer, and multiple network blocks are built in the entire encoder module; The decoder is similar to the encoder, except that an additional multi-head attention layer is added in one network block of the decoder;

[0022] The expansion of the feature map into a one-dimensional vector includes:

[0023] Encode the position information of the three dimensions, merge the encoded information and the one-dimensional vector and send them into the multi-head self-attention layer; Use the sin and cos functions to encode the positions of the three dimensions to generate position embedding vectors;

[0024] The calculation formula of the position encoding is as follows:

[0025]

[0026] Among them, i ∈ {D, H, W} represents the three dimensions of the feature map, which are thickness, height, and width respectively, where Merge PE D 、PE H And PE W The position embedding vectors on the dimensions and expand them into a one-dimensional vector; Add the position vector to the one-dimensional vector of the expanded feature map and send it into the multi-head self-attention layer.

[0027] Preferably, the step S5 includes: Divide the sequence output by the Transformer into three parts, adjust the size of each part of the sequence to correspond to the size of the feature map in each stage, and adopt a top-down fusion strategy;

[0028] The feature map of the first sequence is bilinearly upsampled by a factor of two and fused with the feature map of the second sequence;

[0029] Then, the feature map of the fused second sequence is upsampled to generate a feature map twice the size, which is fused with the feature map of the third sequence;

[0030] Finally, a 3-channel output corresponding to the fixed image size is established, which respectively corresponds to the displacements of each voxel point in the X, Y, and Z directions of the image

[0031] In a second aspect, a medical image registration system based on a transformer is provided, and the system includes:

[0032] Module M1: Obtain a pair of medical images, including a floating image and a fixed image, and preprocess the floating image and the fixed image to obtain a preprocessed floating image and a preprocessed fixed image;

[0033] Module M2: Downsample the preprocessed floating image through three stages to generate feature maps of three different scales and resolutions;

[0034] Module M3: Construct a transformer network model, including an encoder module and a decoder module, expand the feature map generated by each stage of downsampling into a one-dimensional vector, and add a positional embedding vector to send it into the transformer encoder module;

[0035] Module M4: The preprocessed fixed image generates a vector according to Module M2 and Module M3, and sends it into the transformer structure decoder module;

[0036] Module M5: Upsample the feature fusion of the vector output by the transformer network model to generate a deformation field;

[0037] Step S6: Register the medical image to be registered according to the generated deformation field.

[0038] Preferably, in the module M1, the preprocessing of the floating image and the fixed image is performed by: setting the floating image and the fixed image to the same size.

[0039] Preferably, for the three-stage downsampling in the module M2, each stage of downsampling is composed of different numbers of 3D depth residual modules. The height of the floating image x m is H, the width is w, and the thickness is D. The feature map is generated through the first-stage downsampling. C is the number of channels of the feature map. This stage of downsampling is composed of three 3D residual convolutional blocks, and the height and width are reduced to 1 / 4 of the original size, and the thickness is reduced to 1 / 8 of the original thickness;

[0040] Generated by the downsampling in the second stage Feature maps. The downsampling in this stage consists of two 3D residual convolutional blocks. The height and width of the feature maps are reduced to 1 / 8 of the original size, and the thickness is reduced to 1 / 16 of the original thickness;

[0041] Generated by the downsampling in the third stage Feature maps. This stage consists of two 3D residual convolutional blocks. The size of the feature maps is reduced to 1 / 32 of the original size, and the thickness is reduced to 1 / 32 of the original thickness;

[0042] The feature maps F after each downsampling CNN Are expressed as:

[0043]

[0044] Where l = 3 represents the downsampling in three stages, θ represents the parameters of the residual network in each stage, C represents the dimension of the 3D feature maps after downsampling; R represents the set of feature map dimensions; x m Represents the floating image.

[0045] Preferably, the module M3 includes: In the encoder module, it includes a multi-head self-attention layer and a feed-forward neural network layer, and multiple network blocks are built in the entire encoder module; The decoder is similar to the encoder, except that an additional multi-head attention layer is added in one network block of the decoder;

[0046] The expansion of the feature maps into one-dimensional vectors includes:

[0047] Encoding the position information in three dimensions, merging the encoded information and the one-dimensional vector and sending them into the multi-head self-attention layer; Using the sin and cos functions to encode the positions in three dimensions to generate position embedding vectors;

[0048] The calculation formula of the position encoding is as follows:

[0049]

[0050] Where i ∈ {D, H, W} represents the three dimensions of the feature maps, namely thickness, height and width, where Merging PE D 、Pe H And Pe W The position embedding vectors on the dimensions and expanding them into one-dimensional vectors; Adding the position vectors to the one-dimensional vectors expanded from the feature maps and sending them into the multi-head self-attention layer.

[0051] Preferably, the module M5 includes: Dividing the sequence output by the Transformer into three parts, adjusting the size of each part of the sequence to correspond to the size of the feature maps in each stage, and adopting a top-down fusion strategy;

[0052] The feature map of the first sequence is bilinearly upsampled by a factor of two and fused with the feature map of the second sequence;

[0053] Then, the feature map of the fused second sequence is upsampled to generate a feature map twice the size and fused with the feature map of the third sequence;

[0054] Finally, a 3-channel output corresponding to the fixed image size is established, corresponding to the displacements of each voxel point in the X, Y, and Z directions of the image.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] 1. The present invention generates feature maps with high-level semantics through downsampling of the convolutional network, and at the same time, the reduction in spatial resolution brings a reduction in the computational complexity of the transformer;

[0057] 2. In the present invention, the transformer structure is used to establish long-range dependencies to capture global context information, enabling the full fusion of local features and global features, being more sensitive to local and global deformations, and greatly improving the effect of image registration. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0059] Figure 1 It is a flowchart of a medical image registration method based on Transformer provided by an embodiment of the present invention;

[0060] Figure 2 It is a flowchart of downsampling of the convolutional network provided by an embodiment of the present invention;

[0061] Figure 3 It is a structural diagram of a residual module provided by an embodiment of the present invention;

[0062] Figure 4 It is a structural diagram of a transformer provided by an embodiment of the present invention;

[0063] Figure 5 It is a flowchart of upsampling feature fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0065] An embodiment of the present invention provides a medical image registration method based on a transformer, which combines a convolutional neural network and a transformer structure for medical image registration. The convolutional neural network can effectively extract local features, but lacks the ability of global modeling; the transformer can establish long-distance dependencies of context features and strengthen the connection of local features, thereby enhancing the ability of global modeling. First, the floating image and the fixed image are downsampled by a convolutional neural network to generate feature maps with different scales and resolutions. Then, the feature maps are unfolded into one-dimensional vectors and position embedding vectors are added as the input of the transformer. Among them, the vector generated by the floating image is used as the input of the transformer encoder, and the vector generated by the fixed image is used as the input of the transformer decoder. Finally, the feature vector output by the transformer is upsampled by a convolutional network to generate a deformation field, and the registered image is generated after the floating image acts on the deformation field. Refer to Figure 1 As shown, the specific steps of this method are as follows:

[0066] Step S1: Obtain a pair of medical images, including a floating image and a fixed image, and preprocess the floating image and the fixed image to obtain the preprocessed floating image and the preprocessed fixed image. In this step, the preprocessing of the floating image and the fixed image is as follows: set the floating image and the fixed image to the same size.

[0067] Step S2: Downsample the preprocessed floating image through three stages to generate three feature maps with different scales and resolutions.

[0068] The floating image first passes through a common convolutional module and then undergoes three stages of downsampling. Each stage of downsampling is composed of different numbers of 3D depth residual modules. The floating image x m has a height of H, a width of w, and a thickness of D. After the first stage of downsampling, the feature map is generated. C is the number of channels of the feature map. This stage of downsampling is composed of three 3D residual convolutional blocks, and the height and width are reduced to 1 / 4 of the original size, and the thickness is reduced to 1 / 8 of the original thickness; after the second stage of downsampling, the feature map is generated. This stage of downsampling is composed of two 3D residual convolutional blocks, and the height and width of the feature map are reduced to 1 / 8 of the original size, and the thickness is 1 / 16 of the original thickness; after the third stage of downsampling, the feature map is generated. This stage is composed of two 3D residual convolutional blocks, and the size of the feature map is reduced to 1 / 32 of the original size, and the thickness is 1 / 32 of the original thickness. Therefore, the feature map F CNN after each downsampling can be expressed as:

[0069]

[0070] Among them, l = 3 represents downsampling in three stages, θ represents the parameters of the residual network in each stage, C represents the dimension of the 3D feature map after downsampling, and R represents the set of feature map dimensions; x m represents the floating image.

[0071] Step S3: Construct a Transformer network model, including an encoder module and a decoder module. Unfold the feature map generated by downsampling in each stage into a one-dimensional vector, and add a positional embedding vector and send it into the Transformer encoder module.

[0072] Limited by the receptive field of the convolutional network, the depth convolutional network has a limited range of learning feature information and cannot establish long-range dependencies between local features. Therefore, a Transformer structure is combined for long-distance context modeling, which consists of an encoder module and a decoder module. In a network block of the encoder, it is composed of a multi-head self-attention layer and a feed-forward neural network layer, and multiple network blocks are built in the entire encoder. The decoder is similar to the encoder, except that an additional multi-head attention layer is added in a network block of the decoder. To better optimize the depth network, the entire network uses residual connections and normalization layers.

[0073] The Transformer encoder does not support three-dimensional feature vectors as input, and the feature map generated in each stage needs to be unfolded into a one-dimensional feature vector. However, the positional information between features is lost during the process of unfolding the one-dimensional vector. Therefore, the positional information of the three dimensions is encoded, and finally the encoded information and the one-dimensional vector are merged and sent into the self-attention layer. The sine and cosine functions are used to encode the positions of the three dimensions to generate positional embedding vectors. The calculation formula for the above-mentioned positional encoding is as follows:

[0074]

[0075] where i ∈ {D, H, W} represents the three dimensions of the feature map, namely thickness, height, and width, where Merge PE D 、PE H and PE W The positional embedding vectors on the dimensions and unfold them into a one-dimensional vector. The one-dimensional vector obtained by unfolding the feature map plus the positional vector is sent into the self-attention layer module.

[0076] The multi-head self-attention layer consists of multiple dot-product attention sub-layers, which can learn the expression information of different positions from different representation sub-spaces. The multi-head self-attention layer obtains the Quarry vector Q, Key vector K, and Value vector V according to the input vector. After Q, K, and V pass through multiple linear projections, the dot-product attention layer calculation is performed. The dot-product attention layer can calculate a score score = q * k for each vector. To make the training more stable, the score is normalized, then passed through the softmax activation function, and then multiplied by V to obtain the score of each input vector. After the scores are added up, the output result of this module is obtained. The specific formula is as follows:

[0077]

[0078] Finally, the attention scores output by multiple dot-product attention layers are concatenated as the output result of multi-head attention. Denote as:

[0079] MultiHead(Q, K, V) = Concat(Attention 1 , …, Attention n )

[0080] The feed-forward neural network consists of two layers. The first layer uses the ReLU activation function to achieve the purpose of non-linear change, and the latter layer uses a linear function. Before the output of the self-attention layer enters the feed-forward neural network, it passes through a residual module and a normalization module. The residual module sums the output of the self-attention layer and the input of the encoder and then performs a Dropout operation to reduce redundant information. The normalization module uses the mean and standard deviation on a single sample data to continuously adjust the intermediate output of the neural network, so that the numerical values of the intermediate outputs of each layer of the entire neural network are more stable, and at the same time has a regularization effect.

[0081] Step S4: The preprocessed fixed image generates a vector according to Step S2 and Step S3 and is sent into the decoder module of the transformer structure.

[0082] The fixed image passes through another branch and undergoes three downsamplings to generate three scale feature maps. After the three feature maps are unfolded and added with position encoding information, they are input into the multi-head attention layer to generate feature vectors. The input of the decoder module is divided into two items, one is the feature vector output by the encoder, and the other is the feature vector generated by the fixed image. The decoder structure is also stacked by a multi-head attention layer and a feed-forward neural layer, and the network performance is enhanced through parallel operations. Similar to the encoder, a residual network is used to connect each sub-layer, followed by a regularization layer, and then through the feed-forward network layer. The feature vector after decoding passes through the fully connected layer of the softmax activation function to obtain the output vector.

[0083] Step S5: Upsample the features of the vectors output by the transformer network model for feature fusion to generate a deformation field.

[0084] The sequence output by the Transformer is divided into three parts, and the size of each part of the sequence is adjusted to correspond to the size of the feature map at each stage. To enhance the interaction between features of different layers, a top-down fusion strategy is adopted. The feature map of the first sequence is subjected to a bilinear upsampling operation by a factor of two and fused with the feature map of the second sequence. Then, the fused feature map of the second sequence is also upsampled to generate a feature map twice the size, which is fused with the feature map of the third sequence. Finally, a 3-channel output with a fixed image size is established, corresponding to the displacements of each voxel point in the X, Y, and Z directions of the image respectively.

[0085] Step S6: Register the medical image to be registered according to the generated deformation field.

[0086] The vectors output by the transformer are fused through upsampled features to generate a deformation field, and the floating image generates a registered image through the deformation field.

[0087] Next, a more specific description of the present invention will be given.

[0088] A medical image registration method based on a transformer provided by an embodiment of the present invention specifically includes the following steps:

[0089] Step 1: Obtain a medical data set, where each sample includes a pair of reference images and floating images. The reference image and the floating image in each sample are respectively denoted as x f and x m , and preprocess the reference image and the floating image so that the sizes of the reference image and the floating image are the same, and the size is 1×96×96×96.

[0090] Step 2: The floating image generates feature maps of three different scales and resolutions through three stages of downsampling.

[0091] The image first passes through a common convolution module and then undergoes three stages of convolutional downsampling. Each stage of downsampling is composed of different numbers of 3D depth residual modules. The flow chart of the convolutional downsampling is as Figure 2 shown. The height of the floating image x m is H, the width is w, and the thickness is D. The feature map is generated through the first-stage downsampling. C is the number of channels of the feature map. This stage of downsampling is composed of three 3D residual convolutional blocks. The width and height of the feature map are reduced to 1 / 4 of the original size, and the thickness is reduced to 1 / 8 of the original thickness. The Feature map. The downsampling in this stage consists of two 3D residual convolutional blocks. The height and width of the feature map are reduced to 1 / 8 of the original size, and the thickness is reduced to 1 / 16 of the original thickness. The feature map generated through the downsampling in the third stage Feature map. This stage consists of two 3D residual convolutional blocks. The height and width of the feature map size are reduced to 1 / 16 of the original size, and the thickness is reduced to 1 / 32 of the initial thickness. Therefore, the feature map F after each downsampling CNN Can be expressed as:

[0092]

[0093] Where L = 3, representing the downsampling in three stages, θ represents the parameters of the residual network in each stage, and C represents the number of channels of the 3D feature map after downsampling.

[0094] By using the 3D depth residual module for downsampling, on the one hand, highly local features are extracted, and on the other hand, the resolution of the feature map is reduced. The structure of the 3D depth residual module is as Figure 3 Shown. In residual learning, one or more layers can be skipped to perform identity mapping. This method can avoid problems such as vanishing gradients caused by the deepening of the network and enable the network to be better optimized. The residual block consists of two convolutional layers with a kernel size of 3×3×3 and a dropout layer. Since the PReLU (Parametric Rectified Linear Unit) activation function can enable the network to filter out irrelevant information faster and is more suitable for training, in the residual network, PReLU is used as the activation function of the convolutional layer. The PReLU activation function is expressed as:

[0095] g(x) = max(0, x) + a i min(0, x)

[0096] Where, a i Represents the coefficient that controls the slope of the negative part of the activation function. To accelerate the convergence of the network, batch normalization (BN) processing is also added to the convolutional layer. Introducing a dropout layer (with a dropout probability set to 0.2) in the residual network can inactivate some neurons during training, avoid overfitting of parameters, and improve the generalization ability of the network.

[0097] Floating image x mThe size is 1×96×96×96. After passing through the first convolutional block with a convolutional kernel size of 3×3×3 and a convolutional stride of 2, the height and width are reduced to 1 / 2 of the original size, the dimension is 64, and the size of the feature map is 64×48×48×48. The downsampling consists of 3 residual blocks. The convolutional stride in the first residual block is 2, and the strides of the remaining residual blocks are 1. Therefore, after the first-stage downsampling, the size of the feature map is reduced by 1 / 4, the dimension of the feature map is 192, the thickness is 12, and the size of the feature map is 192×12×24×24. The second-stage downsampling consists of 3 residual blocks. The convolutional stride in the first residual block is 2, and the strides of the remaining residual blocks are 1. The size after the second-stage downsampling is 192×6×12×12, the size is reduced by 1 / 8, and the feature dimension remains 192, with a thickness of 6. The third-stage downsampling consists of 2 residual blocks. The convolutional stride in the first residual block is 2, and the convolutional stride of the second residual block is 1. After the third-stage downsampling, the size is 192×3×6×6, the feature dimension does not change, the thickness is 3, and the height and width are reduced to 1 / 16 of the original size.

[0098] Step 3: Unfold the feature maps generated by each stage of downsampling into one-dimensional vectors, add position embedding vectors, and feed them into the transformer encoder module.

[0099] Global context modeling is a long-range dependence modeling method, and its role is to understand the image from a global perspective as much as possible. Limited by the receptive field of the convolutional network, the depth convolutional network has a limited range of learning feature information and cannot establish long-range dependence relationships between local features. Therefore, the transformer structure is used for long-distance context modeling, which consists of an encoder module and a decoder module. In a network block of the encoder, it is composed of a multi-head self-attention layer and a feed-forward neural network layer, and multiple network blocks are built in the entire encoder. The structure diagram of the Transformer is as Figure 4 shown. The decoder is similar to the encoder, except that a multi-head sub-attention layer is added in a network block of the decoder. The Transformer encoder does not support three-dimensional feature vectors as input, and the feature maps generated in each stage need to be unfolded into one-dimensional feature vectors. However, the position information between features is lost during the process of unfolding the one-dimensional vectors. Therefore, the position information of the three dimensions is encoded, and finally the encoded information and the one-dimensional vector are merged and fed into the self-attention layer. The sine and cosine functions are used to encode the positions of the three dimensions to generate position embedding vectors. The calculation formula for position encoding is as follows:

[0100]

[0101] where i ∈ {D, H, W} represents the three dimensions of the feature map, namely thickness, height, and width, where Merge PE D , PE H and PE W The position embedding vectors on the dimension are combined and expanded into a one-dimensional vector. The vector size of the first-stage feature map after expansion is 192×6912, the position embedding vector is 192×6912, and the feature vector of the first stage after adding the position vector is 192×6912; the vector size of the second-stage feature map after expansion is 192×864, the position embedding vector is 192×864, and the feature vector of the second stage after adding the position vector is 192×864; the vector after the third-stage feature map is expanded is 384×108, the position embedding vector is 192×108, and the feature vector of the third stage after adding the position vector is 192×108. The vectors of the three stages are combined, and the finally generated vector is 192×7884.

[0102] The generated vector is fed into the self-attention layer module. The multi-head self-attention layer is composed of multiple dot-product attention sub-layers, which can learn the expression information of different positions from different representation sub-spaces. The input of the multi-head attention layer consists of q, K, and V. d = 512 is the dimension of the input embedding, and q = 8 is the number of dot-product attention layers. The multi-head self-attention layer obtains the Quarry vector Q, Key vector K, and Value vector V according to the input vector. After Q, K, and V are linearly projected 8 times, the dot-product attention layer calculation is performed. The dot-product attention layer can calculate a score score = q*k for each vector. To make the training more stable, the score is normalized, then passed through the softmax activation function, and then multiplied by V to obtain the score of each input vector. After the scores are added, the output result of this module is obtained. The specific formula is as follows:

[0103]

[0104] Finally, the attention scores output by multiple dot-product attention layers are connected as the output result of the multi-head attention. Denote as:

[0105] MultiHead(Q,K,V) = Concat(Attention 1 , …, Attention 8 )

[0106] The feed-forward neural network consists of two layers. The first layer uses the ReLU activation function to achieve the purpose of non-linear change, and the latter layer uses a linear function. Before the output of the self-attention layer enters the feed-forward neural network, it passes through a residual module and a normalization module. The residual module sums the output of the self-attention layer and the input of the encoder, and then performs a Dropout operation to reduce redundant information.

[0107] Step 4: Fix the image. Generate vectors according to Steps 2 and 3 and feed them into the Transformer decoder module.

[0108] The fixed image x f has a size of 1×96×96×96. Feature maps of different sizes are generated through three stages of downsampling by the convolutional network. The feature map generated in the first stage is 192×12×24×24, the feature map generated in the second stage is 192×6×12×12, and the feature map generated in the third stage is 192×3×6×6. Then, position vector encoding is performed on the generated feature maps, unfolded and merged, and finally a 192×7884-dimensional vector is generated and input into the multi-head attention layer in the decoder. The input of the decoder module is divided into two items. One is the feature vector of the floating image x m output by the encoder, and the other is the fixed image x f to generate the feature vector. The decoder structure is also stacked by the multi-head attention layer and the feed-forward neural layer, and the network performance is enhanced through parallel operations. Similar to the encoder, residual networks are used to connect each sub-layer, followed by a regularization layer and then passed through the feed-forward network layer. The feature vector after decoding passes through the fully connected layer of the softmax activation function to obtain the output vector.

[0109] Step 5: The vector output by the Transformer is upsampled and feature-fused to generate the deformation field.

[0110] The vector output by the Transformer is 192×7884. The vector is divided into sequence one vector 192×108, sequence two vector 192×864, and sequence three vector 192×6912. Adjust the sizes of each sequence vector to correspond to the sizes of the feature maps at each stage. The vector 192×108 is adjusted to the size of the third-stage feature map 192×3×6×6; the vector 192×864 is adjusted to the size of the second-stage feature map 192×6×12×12; the vector 192×6912 is adjusted to the size of the first-stage feature map 192×12×24×24. In order to enhance the interaction between features of different layers, a top-down fusion strategy is adopted. The flowchart of upsampling feature fusion is as Figure 5 shown. The sequence one vector is subjected to bilinear upsampling operation twice and fused with the second sequence vector. Then, the fused second sequence vector is also subjected to upsampling operation to generate a feature vector twice the size and fused with the third sequence vector. Finally, an output with the same dimension as the fixed image is established, corresponding to the displacements of each voxel point in the X, Y, and Z directions of the image.

[0111] Step 6: The floating image generates the registered image through the deformation field.

[0112] The vectors output by the transformer are upsampled and feature-fused to generate a deformation field, and the floating image is warped by the deformation field to generate the registered image.

[0113] An embodiment of the present invention provides a medical image registration method and system based on a transformer. A pair of medical images, including a floating image and a fixed image, is obtained and preprocessed; a transformer model, including an encoder and a decoder, is constructed; the floating image is downsampled by a convolutional neural network to obtain feature maps of different scales. Each scale of feature map is unfolded into a one-dimensional vector, and a position embedding vector is added as the input of the encoder; at the same time, the same operations are performed on the fixed image, including sampling by a convolutional network, transforming the size of the feature map, and embedding the position vector, as the input of the decoder; after passing through the transformer structure, a registration deformation field is generated, and the floating image is warped by the deformation field to generate the registered image. By combining the convolutional neural network and the transformer structure for image registration, the connection between local information and global information of the image is enhanced, thereby improving the effect of image registration under unsupervised learning.

[0114] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be regarded as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as both software modules for implementing the method and the structure within the hardware component.

[0115] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. A medical image registration method based on Transformer, characterized in that, it includes: Step S1: Obtain a pair of medical images, including a floating image and a fixed image, and preprocess the floating image and the fixed image to obtain a preprocessed floating image and a preprocessed fixed image; Step S2: Downsample the preprocessed floating image through three stages to generate feature maps of three different scales and resolutions; Step S3: Construct a Transformer network model, including an encoder module and a decoder module, expand the feature maps generated by each stage of downsampling into one-dimensional vectors, and add position embedding vectors to send them into the Transformer encoder module; Step S4: The preprocessed fixed image generates vectors according to Step S2 and Step S3 and sends them into the Transformer structure decoder module; Step S5: Upsample and fuse the vectors output by the Transformer network model to generate a deformation field; Step S6: Register the medical image to be registered according to the generated deformation field; The said Step S5 includes: Divide the sequence output by Transformer into three parts, adjust the size of each part of the sequence to correspond to the size of the feature maps of each stage, and adopt a top-down fusion strategy; Perform bilinear upsampling operation twice on the feature maps of the first sequence and fuse them with the feature maps of the second sequence; Then perform an upsampling operation on the fused feature maps of the second sequence to generate feature maps twice the size and fuse them with the feature maps of the third sequence; Finally, establish a 3-channel output with the same size as the fixed image, corresponding to the displacements of each voxel point in the X, Y, and Z directions of the image.

2. The medical image registration method based on Transformer according to claim 1, characterized in that, In the said Step S1, the preprocessing of the floating image and the fixed image adopts: Set the floating image and the fixed image to the same size.

3. The medical image registration method based on Transformer according to claim 1, characterized in that, The downsampling in the three stages in step S2 is composed of different numbers of 3D depth residual modules for each stage, and the floating image x m has a height of H, a width of w, and a thickness of D, and is generated through the first-stage downsampling feature map, where C is the number of channels of the feature map. The downsampling in this stage is composed of three 3D residual convolutional blocks, and the height and width are reduced to 1 / 4 of the original size, and the thickness is reduced to 1 / 8 of the original thickness; Generated by downsampling in the second stage Feature map. The downsampling in this stage consists of two 3D residual convolutional blocks. The height and width of the feature map are reduced to 1 / 8 of the original size, and the thickness is reduced to 1 / 16 of the original thickness; Generated through the third-stage downsampling feature map, which consists of two 3D residual convolutional blocks in this stage. The size of the feature map is reduced to 1 / 32 of the original size, and the thickness is reduced to 1 / 32 of the original thickness; The feature map F after each downsampling CNN is expressed as: Among them, l = 3 represents downsampling in three stages, θ represents the parameters of the residual network in each stage, C represents the dimension of the 3D feature map after downsampling; R represents the set of feature map dimensions; x m represents the floating image.

4. The medical image registration method based on Transformer according to claim 1, characterized in that, The said Step S3 includes: In the encoder module, it includes a multi-head self-attention layer and a feed-forward neural network layer, and multiple network blocks are built in the entire encoder module; The decoder is similar to the encoder, except that an additional multi-head attention layer is added in one network block of the decoder; The expansion of the feature map into a one-dimensional vector includes: Encode the position information of three dimensions, merge the encoded information and the one-dimensional vector and send them into the multi-head self-attention layer; Encode the positions of three dimensions using sine and cosine functions to generate position embedding vectors; Encode the positions of three dimensions, and the calculation formula is as follows: Among them, i ∈ {D, H, W} represents the three dimensions of the feature map, namely thickness, height, and width, where Merge PE D , PE H and PE W The position embedding vectors on the dimension are combined and expanded into a one-dimensional vector; the one-dimensional vector obtained by expanding the feature map is added to the position vector and fed into the multi-head self-attention layer.

5. A medical image registration system based on Transformer, characterized in that, it includes: Module M1: Obtain a pair of medical images, including a floating image and a fixed image, and preprocess the floating image and the fixed image to obtain a preprocessed floating image and a preprocessed fixed image; Module M2: Generate feature maps of three different scales and resolutions by downsampling the preprocessed floating image in three stages; Module M3: Construct a transformer network model, including an encoder module and a decoder module. Unfold the feature maps generated by downsampling in each stage into one-dimensional vectors, add positional embedding vectors, and send them into the transformer encoder module; Module M4: Generate vectors for the preprocessed fixed image according to Module M2 and Module M3, and send them into the decoder module of the transformer structure; Module M5: Upsample and fuse the features of the vectors output by the transformer network model to generate a deformation field; Step S6: Register the medical image to be registered according to the generated deformation field; The said Module M5 includes: Divide the sequence output by the Transformer into three parts, adjust the size of each part of the sequence to correspond to the size of the feature map in each stage, and adopt a top-down fusion strategy; Perform bilinear upsampling operation with a factor of two on the feature map of the first sequence and fuse it with the feature map of the second sequence; Then perform upsampling operation on the fused feature map of the second sequence to generate a feature map with twice the size, and fuse it with the feature map of the third sequence; Finally, establish a 3-channel output with the same size as the fixed image, corresponding to the displacements of each voxel point in the X, Y, and Z directions of the image.

6. The transformer-based medical image registration system according to claim 5, wherein, In the said Module M1, the preprocessing of the floating image and the fixed image is carried out by: setting the floating image and the fixed image to the same size.

7. The transformer-based medical image registration system according to claim 5, wherein, The downsampling in the three stages of the module M2, where the downsampling in each stage is composed of different numbers of 3D depth residual modules, and the floating image x m has a height of H, a width of w, and a thickness of D, and generates a feature map through the first-stage downsampling. C is the number of channels of the feature map. The downsampling in this stage is composed of three 3D residual convolutional blocks, and the height and width are reduced to 1 / 4 of the original size, and the thickness is reduced to 1 / 8 of the original thickness; Generated by downsampling in the second stage Feature map. The downsampling in this stage consists of two 3D residual convolutional blocks. The height and width of the feature map are reduced to 1 / 8 of the original size, and the thickness is reduced to 1 / 16 of the original thickness; Generated by the third-stage downsampling feature map. This stage consists of two 3D residual convolutional blocks. The size of the feature map is reduced to 1 / 32 of the original size, and the thickness is reduced to 1 / 32 of the original thickness; The feature map F after each downsampling CNN is expressed as: Among them, l = 3 represents downsampling in three stages, θ represents the parameters of the residual network in each stage, C represents the dimension of the 3D feature map after downsampling; R represents the set of feature map dimensions; x m represents the floating image.

8. The transformer-based medical image registration system according to claim 5, wherein, The said Module M3 includes: In the encoder module, it includes a multi-head self-attention layer and a feed-forward neural network layer, and multiple network blocks are built in the entire encoder module; The decoder is similar to the encoder, except that an additional multi-head attention layer is added in one network block of the decoder; The unfolding of the said feature map into a one-dimensional vector includes: Encode the position information of the three dimensions, and merge the encoded information and the one-dimensional vector and send them into the multi-head self-attention layer; Use the sine and cosine functions to encode the positions of the three dimensions to generate positional embedding vectors; Encode the positions of the three dimensions, and the calculation formula is as follows: where \(i\in\{D, H, W\}\) represents the three dimensions of the feature map, namely thickness, height, and width, where Merge PE D , PE H and PE W The position embedding vectors on the dimension are merged and expanded into a one-dimensional vector; the one-dimensional vector obtained by expanding the feature map is added to the position vector and fed into the multi-head self-attention layer.

Citation Information

Patent Citations

  • Medical image registration

    CN105765626B

  • Unsupervised three-dimensional medical image registration method and system based on neural network

    CN110599528A

  • Unsupervised learning medical image registration method and system

    CN113763441A