Multi-modal medical image registration method independent of segmentation labels

Through modal conversion and feature-based thickness registration methods, the shortcomings of multimodal medical image registration relying on segmented labels in the prior art are solved, and high-precision registration in complex anatomical structures and large displacement scenarios are achieved.

CN120235918APending Publication Date: 2025-07-01NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510331403.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing multimodal medical image registration methods rely on segmentation labels, making it difficult to effectively register in complex anatomical structures and large displacement scenarios, and are limited by the performance of segmentation models.

Method used

A multimodal medical image registration method that does not rely on segmentation labels is designed to unify image modalities through modal conversion and combine feature-based coarse registration and fine registration models to achieve fine alignment of images.

Benefits of technology

It realizes high-precision multimodal medical image registration without segmentation labels, avoids the negative impact of segmentation errors, and improves the robustness and reliability of registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235918A_ABST
    Figure CN120235918A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal medical image registration method independent of segmentation labels, and relates to the technical field of medical image processing. The method comprises the following steps: firstly, collecting a multi-modal medical image data set, inputting the multi-modal data set into a modal transformation model, and carrying out training and modal transformation to realize modal unification; inputting the converted moving image and the fixed image into a coarse registration model, and realizing coarse registration by using a feature-based coarse registration method, so that each feature of the moving image is roughly aligned with the fixed image; and finally, inputting the moving image and the fixed image subjected to coarse matching into a fine registration model for training to realize pixel-level fine registration. According to the method, a multi-modal medical image registration frame which does not depend on any segmentation label is designed, the effect of a registration model is prevented from being affected by the performance of an image segmentation model, it is ensured that the structure of a converted image generated by a modal conversion model is consistent before and after conversion, and meanwhile it is ensured that the same effect is achieved in a registration scene with large displacement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a multimodal medical image registration method that does not rely on segmentation labels. Background Art

[0002] In medical diagnosis, it is often necessary to register medical images of different modalities to obtain more comprehensive and accurate information. For example, images of different modalities such as magnetic resonance imaging (MRI), computed tomography (CT), and positron emission tomography (PET) can reflect different physiological structures and pathological characteristics of the human body. However, images of different modalities differ in grayscale, resolution, contrast, etc., and the complexity and individual differences of human organ tissues also increase the difficulty of registration. In the field of medical diagnosis, accuracy and efficiency are two crucial goals. Manual registration requires professional doctors to compare the details of different modal images one by one to obtain effective registration results. This process often consumes a lot of time and effort.

[0003] At this time, a multimodal medical image registration model with good performance is needed. The registration algorithm can automatically register the input images of different modalities and accurately align different types of medical images (such as CT, MRI, etc.), so that doctors can observe and analyze the details of the lesion area more comprehensively. For example, in the diagnosis of lung diseases, doctors can refer to CT and MRI image data at the same time to locate and analyze the lesions more accurately. This technology can not only increase the accuracy of diagnosis, but also greatly save doctors' time and efforts, allowing doctors to develop appropriate treatment plans for patients more quickly. In the prevention, monitoring and early diagnosis of major diseases, multimodal medical image registration technology has shown its irreplaceable importance.

[0004] With the continuous popularization and development of artificial intelligence technology, artificial intelligence has achieved excellent results in the field of image registration, which can significantly improve the registration effect and speed of machines and replace manual work. Machine learning is an important branch of artificial intelligence, and deep learning is the most important algorithm in machine learning. It aims to extract high-level abstract features of data through multi-layer nonlinear transformations, learn the potential distribution laws of data, and thus acquire the ability to make reasonable judgments or predictions on new data. With the help of big data, deep learning has begun to emerge in various fields with its powerful fitting ability, especially in the field of medical image registration.

[0005] In the registration method described in the Chinese patent "CN116402865A Multimodal Image Registration Method, Device and Medium Using Diffusion Model", the registration method uses a discriminator-free generative model based on the diffusion idea, which helps to reduce the inconsistency and artifacts of the generated images, improve the result of multimodal registration, and improve the quality of the generated images. However, when encountering registration images with a large displacement space, as the displacement increases, the positioning of common feature points between the two modalities becomes more difficult, which is not conducive to the training of the pixel-level registration model. Therefore, this invention has limitations.

[0006] In the registration method described in the Chinese patent "CN114387317A Registration Method and Device for CT Images and MRI Three-Dimensional Images", the registration method collaborates through a multi-stage deep learning network, effectively improving the accuracy and efficiency of CT and MRI image registration, and providing strong technical support for medical image analysis and clinical diagnosis. However, this method faces challenges when processing images with complex anatomical structures, such as fundus images. Specifically, the vascular structure in fundus images is complex, and using a modality conversion generator may cause offset and deformation of the vascular structure. This difference may affect the performance of this registration algorithm, thereby leading to a decrease in registration accuracy and limiting its accuracy and reliability in clinical applications.

[0007] For the multimodal medical image registration problem, existing research all converts the multimodal medical image registration problem into a single-modal medical image registration problem by means of modality conversion or segmentation labels. In the registration process of organs such as the brain and lungs, since the contours of these organs are clear and the segmentation labels are easy to obtain, using segmentation labels is appropriate; however, for the registration of complex anatomical structures such as fundus images, using segmentation labels is not very wise because its effect is limited by the performance of the segmentation model, and the evaluation method based on the Dice coefficient cannot accurately reflect the quality of the registration effect. In addition, due to the limited pixel action range of deformable registration, it is difficult to achieve good results in registration scenarios with large displacements. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a multimodal medical image registration method that does not rely on segmentation labels in view of the above-mentioned deficiencies of the prior art, design a multimodal medical image registration framework that does not rely on any segmentation labels, avoid the influence of the performance of the image segmentation model on the effect of the registration model, ensure that the structure of the converted image generated by the modality conversion model is consistent before and after conversion, and at the same time ensure that it is also effective in registration scenarios with large displacements.

[0009] To solve the above technical problems, the technical solutions adopted by the present invention are:

[0010] A multi-modal medical image registration method that does not rely on segmentation labels. First, collect a multi-modal medical image dataset, input the multi-modal dataset into a modality conversion model, train and perform modality conversion to achieve modality unification. Then, input the converted moving image and fixed image into a coarse registration model, and use a feature-based coarse registration method to achieve coarse registration, so that each feature of the moving image is roughly aligned with the fixed image. Finally, input the coarsely registered moving image and fixed image into a fine registration model for training to achieve pixel-level fine registration.

[0011] Further, the modality conversion model is constructed using the CycleGAN architecture in the generative adversarial network, and a structural similarity constraint is introduced into the loss function. Feature maps are extracted from the generator G_A and the generator G_B respectively, and then the structural differences between the corresponding feature maps are calculated. The formula is as follows:

[0012] L translation =L CycleGAN +L struct ;

[0013]

[0014] Among them, L translation is the loss of the modality conversion model, L CycleGAN is the original loss of the CycleGAN model, L struct is the structural similarity constraint, N represents the number of generator layers, and are the i-th layer features of the generator G_A and the generator G_B respectively, represents taking the average of the sum of the L1 norms between each layer of features;

[0015] Set the proportionality coefficients of the loss function to 1, so that the modality conversion model treats each constraint equally.

[0016] Further, before the modality conversion, the collected multi-modal medical image dataset is divided into a training set and a test set, and each dataset includes two typical modality images; the dataset is screened, and images with high brightness and clear structure are selected for training; the training set is used to train the modality conversion model. After training, use the model to perform modality conversion on one of the modality images, and use it as the moving image, and the corresponding one as the fixed image and input them into the coarse registration model together.

[0017] Further, in the coarse registration, a feature - based coarse registration method is used to globally transform the moving image with excessive displacement deviation so as to be roughly aligned with the fixed - image structure; the feature - based coarse registration method includes feature detection and description, initial matching, and outlier rejection; the SuperPoint network is used as the feature detector and descriptor, and the SuperGlue network is used as the feature matcher and outlier screening tool; the SuperPoint network and the SuperGlue network are pre - trained networks, and there is no need to retrain the entire network. The pre - trained weights of the two networks are directly applied to the coarse registration task.

[0018] Further, the SuperPoint network includes an encoder and two decoders, which output the key - point positions and visual descriptors respectively; the SuperPoint network is pre - trained on a dataset formed by geometric transformations to directly extract the key points in the image.

[0019] The SuperGlue network is pre - trained on the large - scale indoor scene dataset ScanNet and includes an attention graph neural network and an optimal matching layer; in the attention graph neural network, the key - point encoder uses self - attention and cross - attention layers to create more powerful representations; the optimal matching layer achieves strong matching by using the Sinkhorn algorithm; finally, according to the strong matching relationship output by the SuperGlue network, the transformation matrix for affine transformation is estimated, and the affine transformation is performed using this transformation matrix. The formula is as follows:

[0020]

[0021] where \(I\) is the moving image, and \(I\) ′ is the moving image after affine transformation; is the matrix of affine transformation; \(a\) 11 and \(a\) 22 represent the scaling coefficients, controlling the scaling of the image in the horizontal and vertical directions; \(a\) 12 and \(a\) 21 represent the shearing coefficients, controlling the horizontal and vertical shearing of the image; \(t\) x and \(t\) y represent the translation coefficients, controlling the translation of the image in the horizontal and vertical directions; the row of 0, 0, 1 is used to extend the two - dimensional coordinates to the homogeneous coordinate system so that the affine transformation can be uniformly represented by matrix multiplication.

[0022] Further, the specific process of applying the SuperPoint network and the SuperGlue network to the coarse registration task is as follows:

[0023] Input the motion image and the fixed image after modal conversion into the coarse registration model. The coarse registration method extracts key points in the images, gives the matching relationship, and generates an estimation matrix. Use the estimation matrix to globally distort the entire motion image to obtain an image with the main structures roughly aligned.

[0024] Furthermore, the fine registration model is an unsupervised deformable registration network based on UNet, including two encoders that respectively extract the features of the motion image and the fixed image, a decoder, an STN network for deformation, a cross-fusion transformer module (CFT module), and a cross-attention feature fusion module (CAFF module);

[0025] Two encoders are used to respectively extract the features of the motion image and the fixed image, that is, in the way of dual-stream input;

[0026] The CFT module and the CAFF module are used to solve the channel-level semantic gap problem brought by skip connections; the CFT module performs cross-channel feature fusion on the encoder side, and the CAFF module enhances the effect of feature fusion on the decoder side, and the two work together;

[0027] The STN network is used to perform pixel-level deformation, apply the deformation field output by the decoder to the motion image, and output the deformed image;

[0028] Input the motion image and the fixed image pair after coarse registration into the fine registration model for training. Set the initial learning rate, batch size, and number of training epochs, and monitor the changes of various performance indicators during the training process at all times. Use the Adam optimizer to adjust the model parameters and retain the model weights with the best effect; during the training process, use the normalized cross-correlation loss function and the smoothing loss function to ensure that the registered image and the fixed image have high similarity and the deformation field is smooth. The loss formula is as follows:

[0029] L Def =λ NCC L NCC +λ sm L sm ;

[0030] Among them, L Def is the loss of the fine registration model, L NCC is the normalized cross-correlation loss, L sm is the smoothing loss, λ NCC and λ sm are weight coefficients;

[0031] After the training is completed, use the trained fine registration model to perform pixel-level deformation on the coarsely registered image pair and output the final registration result.

[0032] The beneficial effects of adopting the above technical solution are as follows: A multi-modal medical image registration method that does not rely on segmentation labels provided by the present invention completely does not rely on any segmentation labels. Instead, it unifies the modalities through image-to-image conversion and combines coarse registration and fine deformation registration to complete the task, avoiding the negative impact brought by segmentation errors. Experiments have proved that even in the absence of segmentation labels, the method of the present invention can still achieve excellent visual effects and state-of-the-art registration accuracy, proving its feasibility and superiority in practical applications. The cross-fusion transformer module and cross-attention feature fusion module introduced by the present invention can effectively handle the challenges brought by the semantic differences between different modalities through cross-channel and multi-scale feature fusion, greatly enhancing the effect of fine registration. The ablation experiment results show that the evaluation metrics of the model (Baseline + CFT + CAFF) using these modules on two datasets are significantly higher than those of the baseline model without using these modules, highlighting the importance of multi-scale and multi-channel feature fusion in improving the registration performance of the encoder-decoder structure registration model. The images obtained by the method of the present invention have higher image similarity and lower errors, so they can provide more accurate visual effects in practical applications and provide more reliable support for clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flowchart of a multi-modal medical image registration method that does not rely on segmentation labels provided by an embodiment of the present invention;

[0034] Figure 2 It is a visualization image of experimental verification on the CF-FA dataset provided by an embodiment of the present invention;

[0035] Figure 3 It is a visualization image of experimental verification on the OASIS-2d dataset provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Next, with reference to the drawings and embodiments, the specific embodiments of the present invention will be described in further detail. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0037] This embodiment provides a multi-modal medical image registration method that does not rely on segmentation labels, and focuses on the construction of modality conversion and deformable registration networks. The multi-modal medical image registration process is as follows: First, the moving image and the fixed image are input into the modality conversion network to achieve modality unification; then the moving image is globally distorted so that its contour, anatomical structure, etc. are roughly aligned with the fixed image; finally, a fine registration model is relied on to achieve pixel-level fine registration.

[0038] First, design a modality conversion model to unify and convert modalities, solve the differences in grayscale, resolution, contrast, etc. among different modality images, and provide possibilities for subsequent coarse registration and fine registration. At the same time, it is necessary to ensure that the anatomical structure of the converted image does not shift or deform.

[0039] Secondly, introduce a coarse registration model. Add a coarse registration model before fine registration to perform rough alignment in advance. Roughly align the structures of the images to be registered, solve the problem that the deformable registration model at the pixel level is difficult to train due to the large displacement deviation between the moving graph and the fixed image, and improve the robustness and generalization of the overall registration framework.

[0040] Finally, according to the disadvantages of the traditional fine registration model UNet structure, design a pixel-level fine registration model to solve the semantic gap between the moving image and the fixed image, overcome the negative impact brought by skip connections, and at the same time introduce an attention mechanism to greatly improve the registration effect.

[0041] As Figure 1 shown, the method of this embodiment is as follows.

[0042] Step 1: Modality conversion.

[0043] Generative adversarial networks (GANs) are often used for modality conversion of medical images, especially the CycleGAN architecture, because it can be trained without paired labeled data, avoiding the difficulty of manually labeling data. This embodiment also uses CycleGAN to build a model that can convert one modality image to another modality image. However, when facing images with complex structures, the cycle consistency loss in CycleGAN cannot effectively guarantee the structural similarity between the converted image and the real image, affecting the effect of the subsequent registration model.

[0044] Therefore, this embodiment introduces a structural similarity constraint into the loss function, extracts feature maps from the generator G_A and the generator G_B respectively, and then calculates the structural differences between the corresponding feature maps to ensure that the converted image not only undergoes modality conversion but also retains the structural features of the original image. Set the weight coefficients of the loss function to 1, so that the modality conversion model treats each constraint equally.

[0045] Divide the collected multi-modal dataset into a training set and a test set, and each dataset contains two typical modality images. To improve the effect of the images generated by the generator, it is necessary to screen the dataset, select images with high brightness and clear structures for training, and avoid being interfered by low-quality images. Use the training set to train the model. After training, use the model to convert one modality of the images, and use the converted images as the moving images, and the corresponding fixed images are jointly input into the coarse registration network.

[0046] Step 2: Coarse registration.

[0047] Due to large displacement deviations in some data, it is difficult to achieve good results directly using a pixel-based registration model. Therefore, in this embodiment, a feature-based coarse registration method is first used to globally transform the moving image with excessive displacement deviation so that it is roughly aligned with the fixed image structure, providing a good basis for pixel-based fine registration. The feature-based coarse registration method includes feature detection and description, initial matching, and outlier rejection. In this embodiment, the SuperPoint network is used as the feature detector and descriptor, and the SuperGlue network is used as the feature matcher and outlier screening tool.

[0048] The SuperPoint network consists of an encoder and two decoders, which output the key point positions and visual descriptors respectively. The model is pre-trained on a dataset formed by geometric transformations and can directly extract the key points in the image.

[0049] The SuperGlue network is pre-trained on the large-scale indoor scene dataset ScanNet and consists of an attention graph neural network and an optimal matching layer. In the attention graph neural network, the key point encoder uses self-attention and cross-attention layers to create more powerful representations. The optimal matching layer achieves strong matching by using the Sinkhorn algorithm. Finally, the transformation matrix for affine transformation is estimated according to the strong matching relationship output by the SuperGlue network, and the affine transformation is performed using this transformation matrix. The formula is as follows:

[0050]

[0051] where \(I\) is the moving image, \(I'\) ′ is the moving image after affine transformation, is the matrix of affine transformation, \(a_{11}\) 11 and \(a_{21}\) 22 represent the scaling coefficients, controlling the scaling of the image in the horizontal and vertical directions, \(a_{12}\) 12 and \(a_{22}\) 21 represent the shearing coefficients, controlling the horizontal and vertical shearing of the image, \(t_1\) x and \(t_2\) y represent the translation coefficients, controlling the translation of the image in the horizontal and vertical directions, and the row of 0, 0, 1 is used to extend the two-dimensional coordinates to the homogeneous coordinate system so that the affine transformation can be uniformly represented by matrix multiplication.

[0052] Different from the other two steps, both networks used in this step are pre-trained. Instead of retraining the entire SuperPoint or SuperGlue network, their pre-trained weights are directly applied to the coarse registration task. The pre-trained model can perform rough registration on the image pairs after modality conversion. The converted moving image and fixed image are input into the coarse registration model. The coarse registration method extracts key points in the images, gives the matching relationship, and generates an estimated matrix. The entire moving image is globally distorted using the matrix to obtain an image with the main structure roughly aligned.

[0053] Step 3: Fine registration.

[0054] After coarse registration, there are still very small areas in the local region of the moving image that are difficult to correct through affine transformation. Therefore, a pixel-level registration model is required to finely adjust individual pixels. In this embodiment, an unsupervised deformable registration network is designed to achieve pixel-level registration. Structurally, a UNet structure similar to traditional pixel-level registration models is adopted, and the problems brought by the traditional UNet structure are solved.

[0055] Since there are still semantic differences in features between the moving image and the fixed image after modality conversion, the model uses two encoders to extract the respective features of the moving image and the fixed image, that is, the dual-stream input method. In addition, the model introduces a cross-fusion transformer module (CFT module) and a cross-attention feature fusion module (CAFF module) to solve the channel-level semantic gap problem brought by skip connections. The CFT module performs cross-channel feature fusion on the encoder side, while the CAFF module enhances the effect of feature fusion on the decoder side. The two work together to achieve effective fusion of multi-scale and multi-channel features, improving the final registration effect and accuracy.

[0056] The moving and fixed images are input, and the output of the fine registration model is a deformation field of the same size as the input. The STN network is used for pixel-level deformation, and the deformation field output by the decoder is applied to the moving image to output the deformed image.

[0057] The moving image and fixed image pair after coarse registration are input into the fine registration model for training. The initial learning rate is set to 0.001, the batch size is 2, and the number of training epochs is 1000. Monitor the changes in various performance indicators during the training process at all times. Use the Adam optimizer to adjust the model parameters and retain the model weights with the best effect. During the training process, the normalized cross-correlation loss function and the smoothness loss function are used to ensure high similarity between the registered images and a smooth deformation field. The loss formula is as follows:

[0058] L Def =λ NCC LNCC +λ sm L sm ;

[0059] where L NCC is the normalized cross - correlation loss, L sm is the smoothness loss, and λ NCC and λ sm are weight coefficients.

[0060] After the training is completed, the trained fine - registration model is used to perform pixel - level deformation on the coarsely registered image pair, and the final registration result is output.

[0061] After obtaining the final registration result through the fine - registration model, the registered moving image is compared with the fixed image for qualitative and quantitative tests. In terms of qualitative tests, in order to intuitively reflect the superiority of the method of this embodiment, the moving images and fixed images registered by each comparison method are cropped into grids and spliced together to compare and observe the visual registration effects of each method. In terms of quantitative tests, multiple metrics are used to evaluate the model, and these metrics include SSIM, MSE, Dice, and hd95. For datasets with complex structures or without segmentation labels, SSIM and MSE can be used for evaluation. For datasets with simple structures or with segmentation labels, the segmentation labels can be registered synchronously, and then Dice and hd95 are used to evaluate the registration effect.

[0062] Two publicly available medical registration datasets, the CF - FA dataset and the OASIS - 2d dataset, are selected to experimentally verify the method of this embodiment and compare it with other advanced methods. In the experiment on the CF - FA dataset, experiments are carried out under two conditions: without coarse registration and with coarse registration, and SSIM and MSE are used as evaluation metrics. In the experiment on the OASIS - 2d dataset, Dice and hd95 are used as evaluation metrics.

[0063] The cross - fusion transformer module and cross - attention feature fusion module introduced by the method of this embodiment can effectively handle the challenges brought by the semantic differences between different modalities through cross - channel and multi - scale feature fusion, greatly enhancing the effect of fine registration. The ablation experiment results show that the evaluation metrics of the model (Baseline + CFT + CAFF) using these modules on the two datasets are significantly higher than those of the baseline model without using these modules. As shown in Table 1, it highlights the importance of multi - scale and multi - channel feature fusion in improving the registration performance of the encoder - decoder structure registration model.

[0064] Table 1 Ablation experiments on the CF - FA dataset and the OASIS - 2d dataset

[0065]

[0066] The method proposed in this embodiment is superior to a variety of existing advanced methods in quantitative evaluation metrics (such as SSIM, MSE), as shown in Table 2. For example, on the CF-FA dataset, the method of this embodiment achieved an SSIM value of 0.528 and an MSE value of 326, higher than other methods such as SyN (SSIM: 0.411, MSE: 2182) and VoxelMorph (SSIM: 0.483, MSE: 395).

[0067] Table 2 Quantitative comparison on CF-FA dataset and OASIS-2d dataset

[0068]

[0069]

[0070] The comparison results also show that the method of this embodiment is superior to the case without coarse registration when there is coarse registration, verifying that coarse registration can indeed solve the problem that excessive displacement is not conducive to fine registration. In addition, the method of this embodiment also compares with the method that can use diffeomorphism (the part with +diff in Table 2) and obtains good results on both datasets. The visualization images of the experimental results also verify the above experimental conclusions, such as Figure 2 and Figure 3 shown.

[0071] In addition, the method of this embodiment has a low algorithm time complexity. The software and hardware environment it is based on is as follows: The hardware environment includes a CPU (Intel(R) Core(TM) i7-10700 CPU@2.90GHz), a GPU (NVIDIA Quadro RTX 6000), the memory capacity is 24.00GB, and the hard disk capacity is 1TB. The software environment includes the operating system Windows11, and the algorithm development language is Python. In the above software and hardware environment, using the pre-trained registration model for registration operations, the average algorithm time consumption is 10s, which is much less than the time consumption of manual registration.

[0072] In summary, traditional multi-modal fundus image registration methods usually rely on segmentation labels, but this method is limited by the low segmentation accuracy caused by the complex anatomical structures of medical images. In contrast, the method of this embodiment does not rely on any segmentation labels at all. Instead, it unifies the modalities through image-to-image conversion and combines coarse registration and fine deformation registration to complete the task. This innovation avoids the negative impact brought by segmentation errors. Experiments have proved that even without segmentation labels, the method of this embodiment can still achieve excellent visual effects and state-of-the-art registration accuracy, demonstrating its feasibility and superiority in practical applications. The images obtained by the method of the present invention have higher image similarity and lower errors, thus being able to provide more accurate visual effects in practical applications and providing more reliable support for clinical diagnosis.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A multimodal medical image registration method that does not rely on segmentation labels, characterized by: The method first collects a multimodal medical image dataset, inputs the multimodal dataset into a modality conversion model, performs training and modality conversion, and realizes modality unification; then the converted moving image and fixed image are input into a coarse registration model, and a feature-based coarse registration method is used to realize coarse registration, so that each feature of the moving image is roughly aligned with the fixed image; finally, the coarsely registered moving image and fixed image are input into a fine registration model for training, so as to realize fine registration at the pixel level.

2. According to claim 1, a multimodal medical image registration method that does not rely on segmentation labels is characterized by: The modality transfer model is constructed using the CycleGAN architecture in the generative adversarial network, and a structural similarity constraint is introduced in the loss function. Feature maps are extracted from generator G_A and generator G_B respectively, and then the structural differences between the corresponding feature maps are calculated. The formula is as follows: L translation =L CycleGAN +L struct ; Among them, L translation is the loss of the modal conversion model, L CycleGAN is the original loss of the CycleGAN model, L struct is the structural similarity constraint, N represents the number of generator layers, and are the i-th layer features of generator G_A and generator G_B respectively, It means taking the average of the L1 norm sum between each layer of features; The proportional coefficients of the loss function are all set to 1, so that the mode conversion model treats each restriction constraint equally.

3. The multimodal medical image registration method according to claim 2, which is independent of segmentation labels, is characterized in that: Before the modality conversion, the collected multimodal medical image data set is divided into a training set and a test set, each data set includes typical two-modality images; the data set is screened to select images with high brightness and clear structure for training; the training set is used to train the modality conversion model, and after the training is completed, the model is used to perform modality conversion on the image of one of the modalities and input it as a moving image and the corresponding fixed image into the coarse registration model.

4. The multimodal medical image registration method according to claim 3 that does not rely on segmentation labels is characterized by: In the coarse registration, a feature-based coarse registration method is used to globally transform the moving image with excessive displacement deviation so as to roughly align it with the fixed image structure; The feature-based coarse registration method includes feature detection and description, initial matching and outlier rejection; the SuperPoint network is used as a feature detector and descriptor, and the SuperGlue network is used as a feature matcher and outlier screening tool; the SuperPoint network and the SuperGlue network are pre-trained networks, and there is no need to retrain the entire network. The pre-trained weights of the two networks are directly applied to the coarse registration task.

5. The multimodal medical image registration method according to claim 4 that does not rely on segmentation labels is characterized by: The SuperPoint network includes an encoder and two decoders, which respectively output key point positions and visual descriptors; the SuperPoint network is pre-trained on a data set formed by geometric transformations to directly extract key points in an image; The SuperGlue network is pre-trained on the large-scale indoor scene dataset ScanNet, including an attention graph neural network and an optimal matching layer; in the attention graph neural network, the key point encoder uses self-attention and cross-attention layers to create a more powerful representation; the optimal matching layer achieves strong matching by utilizing the Sinkhorn algorithm; finally, the transformation matrix for affine transformation is estimated based on the strong matching relationship output by the SuperGlue network, and the transformation matrix is ​​used for affine transformation, and the formula is as follows: Where, I is the moving image, and I′ is the moving image after affine transformation; is the matrix of affine transformation; a 11 and a 22 Represents the zoom factor, which controls the image zoom in the horizontal and vertical directions; a 12 and a 21 Represents the shear coefficient, which controls the horizontal and vertical shear of the image; t x and t y Represents the translation coefficient, which controls the horizontal and vertical translation of the image; the row of 0, 0, 1 is used to expand the two-dimensional coordinate system to the homogeneous coordinate system so that the affine transformation can be uniformly represented by matrix multiplication.

6. The multimodal medical image registration method according to claim 5, which is independent of segmentation labels, is characterized in that: The specific process of applying the SuperPoint network and SuperGlue network to the rough registration task is as follows: The modal-converted moving image and fixed image are input into the coarse registration model. The coarse registration method extracts the key points in the image, gives the matching relationship and generates an estimation matrix. The estimation matrix is ​​used to globally warp the entire moving image to obtain an image with the main structure roughly aligned.

7. The multimodal medical image registration method according to claim 6, which is independent of segmentation labels, is characterized in that: The fine registration model is an unsupervised deformable registration network based on UNet, including two encoders for extracting the features of the moving image and the fixed image respectively, a decoder, an STN network for deformation, a cross-fusion transformer module (CFT module) and a cross-attention feature fusion module (CAFF module); Two encoders are used to extract the features of moving images and fixed images respectively, that is, the dual-stream input method; The CFT module and the CAFF module are used to solve the channel-level semantic gap problem caused by the skip connection; the CFT module performs cross-channel feature fusion on the encoder side, and the CAFF module enhances the effect of feature fusion on the decoder side, and the two work together; The STN network is used to perform pixel-level deformation, apply the deformation field output by the decoder to the moving image, and output the deformed image; The moving image and fixed image pairs after coarse registration are input into the fine registration model for training. The initial learning rate, batch size and training rounds are set, and the changes of various performance indicators during the training process are monitored at all times. The Adam optimizer is used to adjust the model parameters and retain the model weights with the best effect. The normalized cross-correlation loss function and the smoothing loss function are used during the training process to ensure that the registered image and the fixed image have a high degree of similarity and the deformation field is smooth. The loss formula is as follows: L Def =λ NCC L NCC +λ sm L sm ; Among them, L Def is the fine registration model loss, L NCC is the normalized cross-correlation loss, L sm is the smoothing loss, λ NCC and λ sm is the weight coefficient; After the training is completed, the trained fine registration model is used to perform pixel-level deformation on the coarsely registered image pairs and output the final registration result.

Citation Information

Patent Citations

  • Registration method and device for CT image and MRI three-dimensional image

    CN114387317A

  • Multimodal image registration method and device using diffusion model, and medium

    CN116402865A

Cited By

  • Three-dimensional medical image segmentation method and system for thyroid-related eye diseases

    CN121640056A