Medical image registration method and system based on multi-dimensional loss function
By using a two-step registration method of multi-dimensional loss function and transformer deep learning network in medical image registration, the problem of long calculation time and difficult to compatible with the accuracy adaptation range in the prior art is solved, and efficient and accurate medical image registration is achieved.
Patent Information
- Application Number
- CN202311589364.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-24
AI Technical Summary
The existing medical image registration technology has the problem of long calculation time and poor results in processing complex medical images. Deep learning requires high computing power and time-consuming in three-dimensional data processing, and a single loss function is difficult to compatible with registration accuracy and adaptation range.
Using a medical image registration method based on multidimensional loss function, it provides a transformer-based deep learning network model, which is divided into two steps: coarse registration and fine registration, and uses multidimensional adaptive loss function to optimize fine registration transformation parameters to improve registration accuracy and adaptation range.
It achieves higher medical image registration accuracy, improves detection efficiency, maximizes accuracy under the same computer computing power, and is compatible with large-scale adaptation range and high-precision registration.
Smart Images

Figure CN120047500A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image registration processing. Specifically, it relates to a medical image registration method and system based on a multi-dimensional loss function, and also relates to a corresponding computer terminal and a computer-readable storage medium. Background Art
[0002] Medical images refer to the technology and processing process of obtaining internal tissue images of the human body or a part of the human body in a non-invasive manner for medical treatment or medical research, including X-rays, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), etc. In the medical field, image registration is a very important task. Medical images usually contain a large amount of information, including tumors, lesions, organs, etc., and this information can be used to obtain more accurate detection information through image registration.
[0003] In image registration, it is necessary to align two or more sets of images so that the corresponding pixel positions are the same in the same coordinate system. The goal of image registration is to minimize the differences between images so that they can be compared and fused in subsequent processing and analysis. However, traditional medical image registration technologies still have certain limitations, such as long calculation time and poor effects in processing complex medical images. With the continuous development of deep learning and computer vision technologies, the research on medical image registration has become more in-depth, and applying deep learning to medical image registration problems can provide better support for medical diagnosis and treatment.
[0004] Compared with traditional medical image registration solutions, medical image registration technologies based on deep learning have some significant advantages: they can perform adaptive feature extraction without the need to manually design a feature extractor; they can process more complex data; they have higher accuracy, better robustness, faster speed, etc. Therefore, they are widely used in the fields of medical image registration, segmentation, detection, etc.
[0005] Miao, et al constructed a CNN regressor to directly estimate the transformation parameters (Miao, Shun, Z. Jane Wang, and Rui Liao. "A CNN regression approach for real-time 2D / 3D registration." IEEE transactions on medical imaging 35.5 (2016): 1352-1363.). Liao, et al. implemented more reliable 2D / 3D registration by building the POINT^2 network through tracking the points of interest (POI) and triangulation layers (Liao, Haofu, et al. "Multiview 2D / 3D rigid registration via a point-of-interest network for tracking and triangulation." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019.). Kausch, et al. proposed an end-to-end automatic registration process for C-arm pose automatic estimation based on sequential deep learning, and made the network more robust by the constraints of segmented images and fiducial points (Kausch, Lisa, et al. "C-arm positioning for standard projections during spinal implant placement." Medical Image Analysis 81 (2022): 102557.). Czolbe, et al introduced semantic similarity evidence to make the transformation matrix smoother (Czolbe, Steffen, et al. "Semantic similarity metrics for image registration." Medical Image Analysis 87 (2023): 102830.). Zhu, et al used multi-scale feature extraction and similarity attention model to further obtain more accurate and robust registration results (Zhu, Fei, et al. "Similarity attention-based CNN for robust 3D medical image registration." Biomedical Signal Processing and Control 81 (2023): 104403.).
[0006] Although deep learning has been widely used in the field of medical image registration, there are still many deficiencies after its introduction: First, the calculation of three-dimensional data requires high computing power of the computer. Directly processing the original data requires huge computing power and is time-consuming, while processing such as downsampling or dimensionality reduction loses accuracy and some semantic information, and the two need to be balanced. In addition, existing direct registration algorithms generally use a single loss function, and problems that the registration accuracy and the adaptation range cannot be compatible often occur during training. Summary of the Invention
[0007] In view of the above deficiencies in the prior art, the present invention provides a medical image registration method and system based on a multi-dimensional loss function, and also provides a corresponding computer terminal and a computer-readable storage medium.
[0008] According to one aspect of the present invention, there is provided a medical image registration method based on a multi-dimensional loss function, including:
[0009] Preprocess the 3D-CT medical image data to obtain a training data set and the medical image data to be registered;
[0010] Provide a deep learning network model based on a transformer, where the deep learning network model includes a coarse registration network part and a fine registration network part;
[0011] Use the coarse registration network part to perform coarse registration on the training data set to obtain coarse registration transformation parameters;
[0012] Based on the coarse registration transformation parameters, use the fine registration network part to perform fine registration on the training data set to obtain fine registration transformation parameters;
[0013] Use a multi-dimensional adaptive loss function to optimize the fine registration transformation parameters, and train to obtain a registration model;
[0014] Use the registration model to perform registration processing on the medical image data to be registered.
[0015] Preferably, the preprocessing of the 3D-CT medical image data to obtain a training data set includes:
[0016] Perform image format conversion, image size normalization, and simulated floating data processing on the 3D-CT medical image data to obtain preprocessing data; wherein, the preprocessing data includes: a fixed image and a floating image;
[0017] Divide the preprocessing data into a training set, a validation set, and a test set to obtain a training data set.
[0018] Preferably, both the coarse registration network part and the fine registration network part include: a transformer module and a classification head module; where:
[0019] The transformer module is used to extract the features of the input data to obtain attention scores;
[0020] The classification head module includes: two consecutive multi-layer perceptron layers and one hyperbolic tangent activation function layer; taking the attention score as the input of the classification head module, and finally outputting the corresponding transformation parameter result after linear mapping and activation function.
[0021] Preferably, using the coarse registration network part to perform coarse registration on the training data set to obtain coarse registration transformation parameters includes:
[0022] Performing dimensionality reduction and block processing on the fixed image and the floating image in the training data set respectively to form corresponding low-resolution 3D feature maps;
[0023] Taking the low-resolution 3D feature map as the input of the transformer module of the coarse registration network part to generate low-resolution patch attention scores;
[0024] Outputting the coarse registration transformation parameters by passing the low-resolution patch attention scores through the classification head module of the coarse registration network part.
[0025] Preferably, based on the coarse registration transformation parameters, using the fine registration network part to perform fine registration on the training data set to obtain fine registration transformation parameters includes:
[0026] Combining the coarse registration transformation parameters with the floating image in the training data set to obtain the image after coarse registration transformation;
[0027] Performing dimensionality reduction and block processing on the image after coarse registration transformation and the fixed image in the training data set respectively to form corresponding high-resolution 3D feature maps;
[0028] Taking the high-resolution 3D feature map as the input of the transformer module of the fine registration network part to generate high-resolution patch attention scores;
[0029] Outputting the fine registration transformation parameters by passing the high-resolution patch attention scores through the classification head module of the fine registration network part.
[0030] Preferably, using the multi-dimensional adaptive loss function to optimize the fine registration transformation parameters and training to obtain the registration model includes:
[0031] Optimize the fine registration transformation parameters by using a multi-dimensional adaptive loss function that includes a parameter domain and an image domain, where the multi-dimensional adaptive loss function L total is as follows:
[0032] L total = L MSE + μL NCC
[0033] In the formula, L MSE is the loss function for pose estimation, μ is a sigmoid-like function, and L NCC is the inter-pixel loss function;
[0034] Among them, the loss function L MSE for pose estimation is as follows:
[0035]
[0036] In the formula, is the predicted value, is the standard value, and n is the number of fine registration transformation parameters;
[0037] The inter-pixel loss function L MCC is as follows:
[0038]
[0039] In the formula, x is the original image, x * is the predicted image, Cov(·) is the covariance of two images, and Var(·) is the variance of the image itself.
[0040] Preferably, using the registration model to perform registration processing on the medical image data to be registered includes:
[0041] Perform rigid body transformation on the medical image data to be registered through the registration model to generate corresponding fixed medical image data, and complete the registration of the medical image to be registered.
[0042] According to another aspect of the present invention, a medical image registration system based on a multi-dimensional loss function is provided, including:
[0043] A data processing module, which is used to preprocess 3D-CT medical image data to obtain a training data set and medical image data to be registered respectively;
[0044] A model training module, which is used to provide a deep learning network model based on transformer, and the deep learning network model includes a coarse registration network part and a fine registration network part; where:
[0045] The rough registration network part is used to perform rough registration on the training data set to obtain rough registration transformation parameters;
[0046] The fine registration network part, based on the rough registration transformation parameters, performs fine registration on the training data set to obtain fine registration transformation parameters;
[0047] The fine registration transformation parameters are optimized by using a multi-dimensional adaptive loss function, and a registration model is trained;
[0048] A registration module, which is used to perform registration processing on the medical image data to be registered.
[0049] According to the third aspect of the present invention, there is provided a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method described in any one of the above of the present invention, or run the system described in any one of the above of the present invention.
[0050] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method described in any one of the above of the present invention, or run the system described in any one of the above of the present invention.
[0051] Due to the adoption of the above technical solutions, compared with the prior art, the present invention has at least one of the following beneficial effects:
[0052] The medical image registration method and system based on a multi-dimensional loss function provided by the present invention realize a two-step 3D registration technology, use a deep learning method to replace the traditional method to realize 3D registration of medical images, improve the registration accuracy, and thus improve the detection efficiency.
[0053] The medical image registration method and system based on a multi-dimensional loss function provided by the present invention combine rough and fine registration based on multi-scale input, and maximize the accuracy under the condition of the same computer computing power.
[0054] The medical image registration method and system based on a multi-dimensional loss function provided by the present invention adopt a multi-dimensional adaptive loss function to comprehensively consider pose estimation and loss between pixels, and realize the simultaneous compatibility of a large adaptation range and high-precision registration. Description of the Drawings
[0055] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0056] Figure 1This is the workflow diagram of the medical image registration method based on a multi-dimensional loss function in an embodiment of the present invention.
[0057] Figure 2 This is the working schematic diagram of the medical image registration model in a preferred embodiment of the present invention.
[0058] Figure 3 This is the working schematic diagram of the scaled dot-product attention module in a preferred embodiment of the present invention.
[0059] Figure 4 This is the schematic diagram of the multi-head attention mechanism in a preferred embodiment of the present invention.
[0060] Figure 5 This is the schematic diagram of two groups of registration results in a preferred embodiment of the present invention; among them, (a) and (b) are the schematic diagrams of one group of registration results, and (c) and (d) are the schematic diagrams of another group of registration results.
[0061] Figure 6 This is the schematic diagram of the component modules of the medical image registration system based on a multi-dimensional loss function in an embodiment of the present invention. Detailed implementation manners
[0062] The embodiments of the present invention will be described in detail below: These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
[0063] An embodiment of the present invention provides a medical image registration method based on a multi-dimensional loss function. This method comprehensively uses deep learning network technology, a two-stage registration structure from coarse to fine based on the transformer deep learning network framework, and introduces a multi-dimensional adaptive loss function, while improving the adaptation range and registration accuracy of registration.
[0064] Specifically, as Figure 1 shown, the medical image registration method based on a multi-dimensional loss function provided by this embodiment may include the following operations:
[0065] S1. Preprocess the 3D-CT medical image data to obtain a training data set and the medical image data to be registered;
[0066] S2. Provide a deep learning network model based on the transformer. This deep learning network model includes a coarse registration network part and a fine registration network part;
[0067] S3. Coarsely register the training data set using the coarse registration network part and obtain the coarse registration transformation parameters;
[0068] S4. Based on the rough registration transformation parameters, use the fine registration network part to perform fine registration on the training data set to obtain the fine registration transformation parameters;
[0069] S5. Use a multi-dimensional adaptive loss function to optimize the fine registration transformation parameters and train to obtain a registration model;
[0070] S6. Use the registration model to perform registration processing on the medical image data to be registered.
[0071] In some preferred embodiments, the above S1, preprocess the 3D-CT medical image data to obtain a training data set, which may further include the following operations:
[0072] S11. Perform image format conversion, image size normalization, and simulated floating data processing on the 3D-CT medical image data to obtain preprocessed data; wherein, the preprocessed data includes: fixed images and floating images;
[0073] S12. Divide the preprocessed data into a training set, a validation set, and a test set to obtain a training data set.
[0074] In some preferred embodiments, the above S2, both the rough registration network part and the fine registration network part may further include: a transformer module and a classification head module; wherein:
[0075] The transformer module is used to extract the features of the input data to obtain attention scores;
[0076] The classification head module includes: two consecutive multi-layer perceptron layers and one layer of hyperbolic tangent activation function layer; use the attention scores as the input of the classification head module, and finally output the corresponding transformation parameter results after linear mapping and activation function.
[0077] In some preferred embodiments, the above S3, use the rough registration network part to perform rough registration on the training data set to obtain the rough registration transformation parameters, which may further include the following operations:
[0078] S31. Perform dimensionality reduction and block processing on the fixed image and the floating image in the training data set respectively to form corresponding low-resolution 3D feature maps;
[0079] S32. Use the low-resolution 3D feature maps as the input of the transformer module of the rough registration network part to generate low-resolution patch attention scores;
[0080] S33. Output the rough registration transformation parameters (6DOF rigid transformation parameters) through the classification head module of the rough registration network part for the low-resolution patch attention scores.
[0081] In some preferred embodiments, in step S4, based on the rough registration transformation parameters, the fine registration network part is used to perform fine registration on the training data set to obtain the fine registration transformation parameters, and the following operations may further be included:
[0082] S41, combining the rough registration transformation parameters with the floating image in the training data set to obtain the image after rough registration transformation;
[0083] S42, respectively performing dimensionality reduction and block processing on the image after rough registration transformation and the fixed image in the training data set to form corresponding high-resolution 3D feature maps;
[0084] S43, using the high-resolution 3D feature maps as the input of the transformer module of the fine registration network part to generate high-resolution patch attention scores;
[0085] S44, outputting the fine registration transformation parameters (6DOF rigid transformation parameters) through the classification head module of the fine registration network part for the high-resolution patch attention scores.
[0086] In some preferred embodiments, in step S5, a multi-dimensional adaptive loss function is used to optimize the fine registration transformation parameters, and a registration model is trained, and the following operations may further be included:
[0087] Using a multi-dimensional adaptive loss function including a parameter domain and an image domain to optimize the fine registration model, where the multi-dimensional adaptive loss function including a parameter domain and an image domain is:
[0088] L total = L MSE + μL NCC
[0089] In the formula, L MSE is the loss function for pose estimation, μ is a sigmoid-like function, and L NCC is the inter-pixel loss function.
[0090] In some preferred embodiments, the above-mentioned loss function L MSE for pose estimation is:
[0091]
[0092] In the formula, is the predicted value, is the standard value, and n is the number of fine registration transformation parameters (6DOF rigid transformation parameters, including: 3 displacement parameters and 3 rotation parameters);
[0093] The inter-pixel loss function L NCC is:
[0094]
[0095] In the formula, x is the original image, and x * is the predicted image, Cov(·) is the covariance of two images, and Var(·) is the variance of the image itself.
[0096] In some preferred embodiments, in S6 above, when using a registration model to perform registration processing on the medical image data to be registered, it may further include:
[0097] Performing a rigid body transformation on the medical image data to be registered through the registration model to generate corresponding fixed medical image data, and completing the registration of the medical image to be registered.
[0098] Next, in combination with a preferred embodiment, the technical solution provided in the above embodiments of the present invention will be further described in detail.
[0099] The medical image registration method provided by this preferred embodiment, based on deep learning technology and a multi-dimensional adaptive loss function, realizes precise registration of medical images through two-step training.
[0100] As Figure 1 shown, the medical image registration method based on a multi-dimensional loss function provided by this preferred embodiment mainly includes the following steps:
[0101] Step 1, acquisition and preprocessing of 3D-CT data;
[0102] Step 2, performing rough registration on the 3D-CT data using the rough registration network part of a deep learning architecture based on transformer;
[0103] Step 3, further performing fine registration through the fine registration network part of a deep learning architecture based on transformer by updating parameters based on the rough registration result;
[0104] Step 4, further optimizing the network architecture using a multi-dimensional adaptive loss function, training to obtain a registration model for obtaining the optimal registration result.
[0105] In Step 1, the acquisition and preprocessing of 3D-CT data includes:
[0106] Selecting an experimental data set and uniformly saving it in NifTI format, performing image size normalization, simulating floating data processing, and dividing the obtained data into a training set, a validation set, and a test set.
[0107] In Step 2, performing rough registration on the 3D data using a deep learning architecture based on transformer includes:
[0108] Step 2.1: First, downsample the standard image (fixed image) and the floating image to obtain low-resolution images and input them into the registration network. The registration network is based on the ViT architecture under the transformer and is a coarse registration network. Specifically, the convolutional idea is also applied to the ViT architecture in this registration network. For the input low-resolution 3D feature map As the input of the first stage, learn a function mapping relationship f. f is a 3D convolution with a convolutional kernel size of s×s×s, a stride of s-o, and zero-padding of p. It maps the low-resolution 3D feature map x 0 to a new feature with C channels 1 f(x 0 ). Subsequently, expand f(x 0 ) to H 1 W 1 D 1 ×C 1 . Finally, after layer normalization of the features, input them into the subsequent transformer network.
[0109] Step 2.2: In the transformer network, capture and model the misalignment and global relationship between the fixed image and the moving image through the similarity between the projected query-key pairs, so as to generate attention scores for each patch embedding, and finally adopt the multi-head attention mechanism.
[0110] Step 2.3: Finally, a classification head is attached at the end of each stage. The classification head is implemented by two consecutive multi-layer perceptron (MLP) layers and uses the hyperbolic tangent (Tanh) activation function. The classification head takes the features averaged by the transformer module as the input, and finally outputs the rigid body transformation parameters after linear mapping and activation function, which are the coarse registration 6DOF transformation parameters.
[0111] In Step 3, further refine the registration of the coarse registration result update parameters through a deep learning architecture based on the transformer, including:
[0112] Step 3.1: After the end of the first stage, the output rigid body transformation parameters are converted into a rigid body transformation matrix through a formula and then used to transform the images of the next stage. The images of the second stage are the high-resolution images of the standard image (fixed image) and the floating image obtained by downsampling. The input of the second stage is the transformed image obtained by combining the high-resolution fixed image, the coarse registration result, and the high-resolution floating image.
[0113] Step 3.2: For the 3D feature map of the input transformed image as the input of the registration network in the second stage, learn a function mapping relationship f. f is a 3D convolution with a convolutional kernel size of s×s×s, a stride of s-o, and zero-padding of p. It maps the feature x1 Mapped to channel C 2 of the new feature f(X 1 ), and then f(X 1 ) is expanded to H 2 W 2 D 2 ×C 2 . Finally, layer normalization is performed on the features.
[0114] Step 3.3: Feed the new feature into the transformer network in the second stage. The structure of the transformer network in the second stage is the same as that in Step 2.2.
[0115] Step 3.4: The classification head in the second stage is also implemented by two consecutive multi-layer perceptron (MLP) layers, using the hyperbolic tangent (Tanh) activation function. The classification head takes the features averaged by the transformer module as input, and finally outputs the rigid body transformation parameters after linear mapping and activation function. Finally, the corresponding rigid body transformation parameters and floating CT data are used to generate the corresponding standard CT data through rigid body transformation.
[0116] In Step 4, the loss function of the network is composed of the 6DOF error in the pose estimation dimension and the pixel loss in the image domain. The rigid body transformation parameters can measure the error between the predicted pose and the standard pose. The most common way is to calculate the difference of each parameter and then minimize the difference to make the predicted value close to the standard value. In this preferred embodiment, MSE is used as the loss function L MSE for pose estimation. Further, the predicted pose is combined with the original 3D-CT to obtain the predicted 3D-CT, and the pixel-level error between it and the standard 3D-CT is calculated. The pixel-wise loss function L NCC is used to measure the difference between the predicted image and the target image. Finally, an adaptive multi-dimensional loss function is proposed, and the specific formula is as follows:
[0117] L total = L MSE + μL NCC
[0118] where μ is a sigmoid-like function, which adaptively allocates the weights of the Loss based on 6 dimensions and the Loss based on the image domain according to the total number of network training rounds set, so that the network focuses on the dimension error at the beginning of training and the image domain error in the later stage. In this way, compatibility is achieved while adapting to a wide range and high-precision registration.
[0119] The following combines a specific application example to further detail the technical solutions provided in the above embodiments of the present invention.
[0120] The medical image registration method based on a multi-dimensional loss function adopted in this specific application example uses computer-aided clinical multi-modal image data for registration, and then realizes the registration of three-dimensional CT medical images.
[0121] As Figure 1 shown, it is the workflow diagram of the medical image registration method based on a multi-dimensional loss function adopted in this specific application example, including the following steps:
[0122] Step 1, acquisition and preprocessing of 3D-CT data;
[0123] Step 2, use the rough registration network part of the deep learning architecture based on transformer to perform rough registration on 3D-CT data;
[0124] Step 3, update the parameters based on the rough registration result and further perform fine registration through the fine registration network part of the deep learning architecture based on transformer to obtain a registration model;
[0125] Step 4, use the multi-dimensional adaptive loss function to further optimize the network architecture, train to obtain a registration model, and obtain the optimal registration result.
[0126] 3D-CT registration is to match a 3D standard image with a 3D floating image. Registration is divided into rigid registration and non-rigid registration. This specific application example is based on rigid registration to study the rigid registration of a 3D standard image and a 3D floating image. First, starting from the standard image F CT , a corresponding floating image M CT is made through a rigid body transformation. The registration target is to learn the rigid body transformation matrix from M CT to F CT . Specifically, the designed network is used to parameterize the rigid body registration problem as the function f θ (M CT , F CT ) = T, where θ is a set of rigid body transformation parameters, and T represents the predicted rigid body transformation matrix. The purpose is to optimize the matrix T to make it closer to the true rigid body transformation matrix, so that the error between the DOF parameters after the transformation of the matrix T and the actual DOF parameters is minimized.
[0127] In geometric transformation, rigid body transformation is a transformation that keeps the shape and size unchanged. It includes two basic operations: translation and rotation. Rigid body transformation can be represented by six parameters, which describe the components of translation and rotation. Specifically, the six parameters of rigid body transformation are the displacement components of translation in three axial directions (t x , t y , t z ) and the rotation components around three axial directions (r x , r y , rz )。These parameters describe the translation and rotation operations of a rigid body transformation in three-dimensional space. The rigid body transformation matrix represents the rigid body transformation as a 4x4 matrix, which transforms a point from its initial position to a new position. In the present invention, the rigid body transformation matrix is usually denoted as T and has the following form:
[0128]
[0129] where t is a 3x1 translation vector and R is a 3x3 rotation matrix, which is obtained by multiplying three rotation matrices R x , R y and R z to get R = R x * R y * R z , where
[0130]
[0131]
[0132]
[0133] In step one, the collection and preprocessing of the experimental data set include:
[0134] Select the currently largest publicly available annotated spinal CT data set CTSpine1K as the experimental data set. All data are uniformly saved in NIfTI format, with a slice size of 512*512. The number of slices is center-cropped to 512, so the overall size of all data is 5123. Select 279 cases from CTSpine1K as the training set and 81 cases as the test set.
[0135] For each case of CT data, 6 different combinations of parameters are randomly generated within the ranges of t x,y,z ∈ (-45mm, 45mm) and r x,y,z ∈ (-45°, 45°). Subsequently, the DOF parameters are sequentially converted into R x , R y , R z and t according to the above formula and combined into a 4x4 rigid body transformation matrix. Finally, the corresponding floating CT data are generated by applying the rigid body transformation to the different rigid body transformation parameters and the original CT data using the affine_transform method in the ndimage library of the scipy toolkit, resulting in a total of 1674 groups of data. Considering the actual clinical situation, the parameter ranges in the test set are t x,y,z ∈ (-20±15mm, 20±15mm) and r x,y,z∈(-20±15°,20±15°), 4 sets of different combined data are generated for each case, totaling 324 sets of data.
[0136] This specific application example solves the 3D-CT registration problem from coarse to fine, which can be specifically divided into two stages i = 1, 2. When i = 1, it is coarse registration, and when i = 2, it is fine registration.
[0137] In step two, a deep learning architecture based on transformer is used for coarse registration of 3D data:
[0138] Step 2.1: First, downsample the standard image and floating image with a size of 512^3 to obtain low-resolution images with a size of 128^3 and input them into the coarse registration network. The registration network is based on the vit architecture under transformer. Specifically, the convolutional idea is also applied to the ViT architecture in the network. For the input 3D feature map As the input of stage one, learn a function mapping relationship f. f is a 3D convolution with a convolution kernel size of s×s×s, a stride of s - o, and zero-padding of p. At this time It maps the feature x 0 to a new feature f(x 1 ) with a channel of C 0 . The size of the new feature is Subsequently, expand f(x 0 ) to H 1 W 1 D 1 ×C 1 . In this method, H 1 W 1 D 1 is 4096, and C 1 is 256. Finally, after layer normalization of the features, they are input into the transformer module of the stage.
[0139] Step 2.2: In the transformer network as shown in Figure 2 , capture and model the misalignment and global relationship between the fixed image and the moving image through the similarity between the projected query-key pairs, so as to generate attention scores for each patch embedding. Specifically, given the input embedding patch block as X = {x 1 , x 2 ... x n-1 , x n}, where n is the number of patch blocks.
[0140] Q(x k ) = x k W Qk = 1, 2…n - 1, n
[0141] K(x k ) = x k W K k = 1, 2…n - 1, n
[0142] V(x k ) = x k W V k = 1, 2…n - 1, n
[0143] Among them, Q(x k ), K(x k ), and V(x k ) represent the query, key, and value of the k-th input patch respectively. W Q , W K , and W V are learnable parameter matrices used to map the input patch to the query, key, and value spaces. Then, Q(x k ), K(x k ), and V(x k ) are input into the scaled dot-product attention module shown in Figure 3 to calculate the attention scores as follows:
[0144]
[0145] Among them, d h is the dimension of the key space. Finally, a multi-head attention mechanism with h = 2 is adopted. As shown in Figure 4 , the outputs of all scaled dot-product attention modules are combined together on the channel and linearly projected through the matrix W O ∈ R d×d . In this study, d = 256.
[0146] Step 2.3: Finally, a classification head is attached at the end of each stage. The classification head is implemented by two consecutive multi-layer perceptron (MLP) layers and uses the hyperbolic tangent (Tanh) activation function. The classification head takes the features averaged by the transformer module as input and finally outputs 6 rigid body transformation parameters after linear mapping and activation function.
[0147] In step three, the parameters of the coarse registration result are further refined through the registration network for fine registration:
[0148] Step 3.1: After the end of the first stage, the output rigid body transformation parameters are converted into a rigid body transformation matrix through a formula and then used to transform the image of the next stage. The image of the second stage is obtained by downsampling the standard image and the floating image with a size of 512 3 to a size of 256 3For the high-resolution image, the input in the second stage is the "moved" image combined with the high-resolution image and the coarse registration rigid body transformation matrix.
[0149] Step 3.2: For the input high-resolution "moved" 3D feature map As the input of the second stage, learn a function mapping relationship f, where f is a 3D convolution with a convolution kernel size of s×s×s, a stride of s - o, and zero-padding of p. At this time It maps the feature x 1 to a new feature f(x 2 ) with C channels 1 . The size of the new feature is Subsequently, expand f(x 1 ) to H 2 W 2 D 2 ×C 2 . In this method, H 2 W 2 D 2 is 4096, C 2 is 256. Finally, layer normalization is performed on the features.
[0150] Step 3.3: Feed the new feature into the transformer module in the second stage. The structure of the transformer module in the second stage is the same as that in Step 2.2.
[0151] Step 3.4: The classification head in the second stage is also implemented by two consecutive multi-layer perceptron (MLP) layers, using the hyperbolic tangent (Tanh) activation function. The classification head takes the features averaged by the transformer module as the input, and finally outputs 6 rigid body transformation parameters after linear mapping and activation function. Finally, the corresponding standard CT data is generated by performing rigid body transformation on the rigid body transformation parameters and the floating CT data according to the affine_transform method in the ndimage library of the scipy toolkit.
[0152] Compared with the requirement of the initial registration error range for one-time registration, multi-resolution registration has higher robustness and universality for large-range registration, and features may not be easily matched at a single scale. By performing registration in two stages at different scales, the robustness of matching can be improved, and the possibility of false matching can be reduced. And multi-scale registration can capture more detailed information in the image, thereby improving the accuracy of registration. Performing rough large-range registration at a smaller scale and then performing detailed small-range registration at a larger scale can ensure that the final registration result is more accurate and reduce the computational amount. The two stages of the present invention adopt multi-scale resolution registration. Specifically, the input of the first stage is 128 3The low-resolution image, and the second stage inputs 256 3 The high-resolution image.
[0153] In step four of this specific application example, the loss function for measuring the network is composed of the 6DOF error in the pose estimation dimension and the pixel loss in the image domain.
[0154] The 6 rigid body transformation parameters can measure the error between the predicted pose and the standard pose. The most common way is to calculate the difference of each parameter and then minimize this difference to make the predicted value close to the standard value In this experiment, MSE is used as the loss function L for pose estimation MSE :
[0155]
[0156] In this specific application example, the obtained estimated pose is further combined with the original 3D-CT to obtain the predicted 3D-CT, and the pixel-level error between it and the standard 3D-CT is calculated. The inter-pixel loss is used to measure the difference between the predicted image and the target image, and the algorithm L is optimized by making the predicted image closer to the target image NCC :
[0157]
[0158] In the formula, x represents the original image, and x * represents the predicted image, Cov(·) represents the covariance of the two images, and Var(·) represents the variance of the image itself.
[0159] According to the conclusions of previous work and experimental experience, the loss based on 6 dimensions is more robust for the registration of two images with large errors. When the error between the two images is small, the error value in each dimension is negligible compared to the initial error value, which will cause the network to be difficult to continue gradient descent after a certain time, ultimately resulting in underfitting. On the contrary, the loss based on the image is very sensitive to the images to be registered, and a large initial error will probably cause the network to be unable to fit or fall into a local minimum. When the error is small, each pixel will superimpose the error on the image-based Loss, which can further guide the gradient descent during network training and obtain a more optimized result. Based on the above, a multi-dimensional adaptive loss function is proposed in this specific application example, and the specific formula is as follows:
[0160] L total = L MSE + μL NCC
[0161] Where μ is a sigmoid-like function that adaptively allocates the weights of the Loss based on 6 dimensions and the Loss based on the image domain according to the total number of network training rounds set, so that the network emphasizes the dimensional error at the initial stage of training and the image domain error at the later stage. In this way, it can achieve compatibility while achieving a wide range of adaptation and high-precision registration.
[0162] In summary, the medical image registration method provided by the above embodiments of the present invention inputs the moving position 3D-CT and outputs the rigid body transformation matrix from the moving position to the standard position. Further, the standard position 3D-CT can be obtained, and the three-dimensional effect diagram of the final registration result is as Figure 5 shown. It can be seen that compared with the large misalignment in space before registration, the method provided by the above embodiments of the present invention can achieve spatial alignment, thereby realizing the rigid registration of three-dimensional image data to guide clinicians to make relevant judgments and operations.
[0163] An embodiment of the present invention provides a medical image registration system based on a multi-dimensional loss function, as Figure 6 shown. The system may include the following modules:
[0164] A data processing module for preprocessing 3D-CT medical image data to obtain a training data set and the medical image data to be registered respectively;
[0165] A model training module for providing a deep learning network model based on transformer. The deep learning network model includes a coarse registration network part and a fine registration network part; where:
[0166] The coarse registration network part is used to perform coarse registration on the training data set to obtain coarse registration transformation parameters;
[0167] The fine registration network part is based on the coarse registration transformation parameters to perform fine registration on the training data set to obtain fine registration transformation parameters;
[0168] A multi-dimensional adaptive loss function is used to optimize the fine registration transformation parameters, and a registration model is trained;
[0169] A registration module for performing registration processing on the medical image data to be registered.
[0170] Among them, when selecting the loss function in the model training module, according to the conclusions of previous work and experimental experience, the loss based on 6 dimensions will have greater robustness for the registration of two images with large errors. When the error between the two images is small, the error value in each dimension is negligible relative to the initial error value, which will cause the network to be difficult to continue gradient descent after a certain period of time, ultimately resulting in underfitting. On the contrary, the loss based on the image is very sensitive to the images to be registered, and a large initial error will probably cause the network to be unable to fit or fall into a local minimum. When the error is small, each pixel will superimpose the error on the image-based Loss, which can further guide the gradient descent during network training and obtain a more optimized result. Therefore, an adaptive multi-dimensional loss function that comprehensively considers pose estimation and pixel-wise loss is designed.
[0171] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the method to implement the composition of the system. That is, the embodiments in the method can be understood as the preferred examples for constructing the system, which will not be elaborated here.
[0172] An embodiment of the present invention provides a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method of any one of the above embodiments of the present invention, or run the system of any one of the above embodiments of the present invention.
[0173] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviation: RAM), such as static random access memory (English: static random-access memory, abbreviation: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviation: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be stored in one or more memories in a partitioned manner. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0174] The computer programs, computer instructions, etc. described above can be stored in partitions in one or more memories. And the computer programs, computer instructions, data, etc. described above can be called by the processor.
[0175] A processor, configured to execute the computer program stored in the memory to implement each step in the method involved in the above embodiments or each module of the system. For details, reference can be made to the relevant descriptions in the foregoing method and system embodiments.
[0176] The processor and the memory can be of an independent structure or an integrated structure integrated together. When the processor and the memory are of an independent structure, the memory and the processor can be coupled and connected through a bus.
[0177] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method of any one of the above embodiments of the present invention, or to run the system of any one of the above embodiments of the present invention.
[0178] The medical image registration method and system based on a multi-dimensional loss function provided in the above embodiments of the present invention implement a two-step 3D registration technology, use a deep learning method to replace the traditional method to achieve three-dimensional registration of medical images, improve the registration accuracy, and thus improve the detection efficiency; based on the combination of coarse and fine registration with multi-scale input, the accuracy is maximally guaranteed under the condition of the same computer computing power; a multi-dimensional adaptive loss function is adopted to comprehensively consider pose estimation and loss between pixels, realizing the simultaneous compatibility of a large adaptation range and high-precision registration.
[0179] Those skilled in the art know that in addition to implementing the system and its various devices provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices provided by the present invention can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices for implementing various functions can also be regarded as both software modules for implementing the method and structures within the hardware component.
[0180] Matters not described in detail in the above embodiments of the present invention are all well-known technologies in the art.
[0181] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A medical image registration method based on a multi-dimensional loss function, characterized in that, it includes: Preprocess the 3D-CT medical image data to obtain a training data set and the medical image data to be registered; Provide a deep learning network model based on transformer, and the deep learning network model includes a coarse registration network part and a fine registration network part; Use the coarse registration network part to perform coarse registration on the training data set to obtain coarse registration transformation parameters; Based on the coarse registration transformation parameters, use the fine registration network part to perform fine registration on the training data set to obtain fine registration transformation parameters; Use a multi-dimensional adaptive loss function to optimize the fine registration transformation parameters and train a registration model; Use the registration model to perform registration processing on the medical image data to be registered.
2. The medical image registration method based on a multi-dimensional loss function according to claim 1, characterized in that, The preprocessing of the 3D-CT medical image data to obtain a training data set includes: Perform image format conversion, image size normalization and simulated floating data processing on the 3D-CT medical image data to obtain preprocessing data; wherein, the preprocessing data includes: fixed image and floating image; Divide the preprocessing data into a training set, a validation set and a test set to obtain a training data set.
3. The medical image registration method based on a multi-dimensional loss function according to claim 1, characterized in that, Both the coarse registration network part and the fine registration network part include: a transformer module and a classification head module; wherein: The transformer module is used to extract the features of the input data to obtain attention scores; The classification head module includes: two consecutive multi-layer perceptron layers and one hyperbolic tangent activation function layer; use the attention scores as the input of the classification head module, and finally output the corresponding transformation parameter results after linear mapping and activation function.
4. The medical image registration method based on a multi-dimensional loss function according to claim 3, characterized in that, The use of the coarse registration network part to perform coarse registration on the training data set to obtain coarse registration transformation parameters includes: Perform dimensionality reduction and block processing on the fixed image and the floating image in the training data set respectively to form corresponding low-resolution 3D feature maps; Use the low-resolution 3D feature maps as the input of the transformer module of the coarse registration network part to generate low-resolution patch attention scores; Output the coarse registration transformation parameters through the classification head module of the coarse registration network part with the low-resolution patch attention scores.
5. The medical image registration method based on a multi-dimensional loss function according to claim 3, characterized in that, Based on the coarse registration transformation parameters, using the fine registration network part to perform fine registration on the training data set to obtain fine registration transformation parameters includes: Combine the coarse registration transformation parameters with the floating image in the training data set to obtain a coarsely registered transformed image; Perform dimensionality reduction and block processing on the image after the rough registration transformation and the fixed image in the training dataset respectively to form corresponding high-resolution 3D feature maps; Use the high-resolution 3D feature maps as the input of the transformer module in the fine registration network part to generate high-resolution patch attention scores; Output the fine registration transformation parameters through the classification head module in the fine registration network part with the high-resolution patch attention scores.
6. The medical image registration method based on a multi-dimensional loss function according to claim 1, characterized in that, The multi-dimensional adaptive loss function is used to optimize the fine registration transformation parameters, and a registration model is trained, including: Optimize the fine registration transformation parameters by using a multi-dimensional adaptive loss function that includes a parameter domain and an image domain, where the multi-dimensional adaptive loss function L totaL is as follows: L total = L MSE + μL NCC where L MSE is the loss function for pose prediction, μ is the sigmoid-like function, and L NCC is the inter-pixel loss function; Among them, the loss function L of the pose estimation MSE is as follows: Wherein, is the predicted value, is the standard value, and n is the number of fine registration transformation parameters; The inter-pixel loss function L NCC is as follows: where x is the original image, and x * is the predicted image, Cov(·) is the covariance of two images, and Var(·) is the variance of the image itself.
7. The medical image registration method based on a multi-dimensional loss function according to claim 1, characterized in that, The registration model is used to perform registration processing on the medical image data to be registered, including: Perform a rigid body transformation on the medical image data to be registered through the registration model to generate corresponding fixed medical image data, and complete the registration of the medical image to be registered.
8. A medical image registration system based on a multi-dimensional loss function, characterized in that, including: A data processing module, which is used to preprocess 3D-CT medical image data to obtain a training dataset and medical image data to be registered respectively; A model training module, which is used to provide a deep learning network model based on a transformer. The deep learning network model includes a rough registration network part and a fine registration network part; where: The rough registration network part is used to perform rough registration on the training dataset to obtain rough registration transformation parameters; The fine registration network part is based on the rough registration transformation parameters to perform fine registration on the training dataset to obtain fine registration transformation parameters; Use a multi-dimensional adaptive loss function to optimize the fine registration transformation parameters, and train to obtain a registration model; A registration module, which is used to perform registration processing on the medical image data to be registered.
9. A computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it can be used to execute the method described in any one of claims 1-7, or, run the system described in claim 8.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it can be used to execute the method described in any one of claims 1-7, or, run the system described in claim 8.
Citation Information
Patent Citations
A medical image registration algorithm based on a convolutional neural network
CN109584283A
2D medical image registration method fusing residual image information
CN115457020A
Magnetic resonance image unsupervised cascade registration method based on artificial intelligence
CN116228823A
Medical image registration method based on LOFTR network model and improved particle swarm algorithm
CN116883462A
Image registration method and model training method thereof
US20210390716A1
Cited By
Cardiovascular atlas calculation and visualization system based on generative model
CN120510310A
Cardiovascular atlas calculation and visualization system based on generative model
CN120510310B