Abdominal medical image registration method and device based on structure guidance and medium

By introducing structural guidance and swin-transformer networks into medical image registration, the problems of low registration accuracy and organ contour distortion of abdominal image are solved, and higher registration accuracy and more accurate organ contour matching are achieved.

CN120182170APending Publication Date: 2025-06-20FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311766651.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing medical image registration technology has poor accuracy when processing abdominal images, and cannot effectively utilize anatomical structure information. The convolutional network is limited by the receptive field and is difficult to deal with long-distance deformation, resulting in distortion of organ contours.

Method used

The structure-guided abdominal medical image registration method is adopted, and the images are linear and nonlinear transformed by acquiring the abdominal multimodal medical source image and target image with segmentation labels, and the registration accuracy is improved by using structural information and region weights. The registration network is constructed using swin-transformer to handle long-distance deformation.

Benefits of technology

It significantly improves the accuracy of abdominal medical image registration, reduces the distortion of the organ profile, and can individually improve the registration accuracy of a certain organ or part of the area to meet actual medical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182170A_ABST
    Figure CN120182170A_ABST
Patent Text Reader

Abstract

The invention relates to an abdominal medical image registration method and device based on structure guidance and a medium, and the method comprises the following steps: obtaining a source image with a segmentation label and a target image pair, and inputting the source image and the target image pair into a linear registration network to obtain a linear transformation matrix; inputting the source image and the target image pair subjected to linear transformation by using the linear transformation matrix into a nonlinear transformation network to obtain a nonlinear deformation field; performing nonlinear deformation on the linearly transformed source image and the segmentation label thereof according to a nonlinear deformation field; setting a region weight, and performing structural similarity evaluation and pixel similarity evaluation based on the target image, the source image after nonlinear deformation, the segmentation label of the source image and the segmentation label of the source image; and training the linear registration network and the nonlinear transformation network based on the structural similarity evaluation result and the pixel similarity evaluation result, iteratively updating network parameters, and performing image registration after training is completed. Compared with the prior art, the method has the advantage of high registration precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image registration, and in particular to a method, device and medium for abdominal medical image registration based on structure guidance. Background Art

[0002] Medical image registration refers to a class of technologies that transform images of the same organ or body region from different times, different sensors, and different perspectives into the same coordinate system to match the image content at corresponding positions. This technology originated in the 1990s. With the maturity and clinical popularity of medical imaging technologies such as computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography (PET), there is an urgent need to fuse different modalities of medical images to make full use of the complementary information provided therein, so as to better diagnose and treat diseases. Medical image registration technology is the basis of image fusion. Different images often come from different devices, and there are differences in the spatial coordinate reference and scale of these devices. Even for the same device, there are differences in the organ morphology among different individuals. To solve this problem, it is necessary to use registration-related algorithms to find a set of spatial transformations to make the corresponding anatomical structures in two or more images spatially consistent.

[0003] The technical solution proposed in "VoxelMorph: A Learning Framework for Deformable Medical Image Registration" utilizes the learning ability of a deep neural network. Through the training process on the target dataset, a general registration method is learned, so that the deformation field and the deformed image to be registered can be directly output through a single inference process, avoiding the time-consuming iterative optimization process in traditional registration methods and greatly improving the execution efficiency of the algorithm. At the same time, this solution proposes a deformation field regularization scheme to reduce the unreasonable deformations in the deformation field by reducing the Euclidean norm (L2 norm) of the spatial gradient of the deformation field, ensuring the anatomical credibility of the deformed image. However, this technical solution has the following disadvantages:

[0004] 1. This solution only uses the gray information of the images to evaluate the similarity between the moving image and the fixed image after registration, without using anatomical structure information, resulting in poor accuracy when processing abdominal images.

[0005] 2. The convolutional network used in this solution is limited by its own receptive field and has an upper limit on the deformation distance estimation. When there is a long-distance correspondence relationship between the source image and the target image, the accuracy decreases or the organ contour is distorted.

[0006] 3. This solution lacks the ability to focus on partial areas required in actual medical scenarios and cannot improve the registration accuracy of partial areas alone. Summary of the Invention

[0007] The purpose of the present invention is to provide a structure-guided abdominal medical image registration method, device, and medium, which use structural information to improve the registration accuracy in scenarios with complex deformations such as abdominal medical images and reduce the distortion of organ contours.

[0008] The purpose of the present invention can be achieved through the following technical solutions:

[0009] A structure-guided abdominal medical image registration method includes the following steps:

[0010] S1. Obtain a pair of abdominal multi-modal medical source images and target images with segmentation labels, input them into a linear registration network to obtain a linear transformation matrix, and perform a linear transformation on the source image and its segmentation label based on the linear transformation matrix;

[0011] S2. Input the linearly transformed source image and target image pair into a non-linear transformation network to obtain a non-linear deformation field;

[0012] S3. Perform a non-linear deformation on the linearly transformed source image and its segmentation label according to the non-linear deformation field;

[0013] S4. Perform a structural similarity evaluation based on the segmentation label of the target image and the segmentation label of the non-linearly deformed source image to determine the regional weight, and perform a pixel similarity evaluation based on the target image, the non-linearly deformed source image, and the corresponding regional weight;

[0014] S5. Train the linear registration network and the non-linear transformation network based on the structural similarity evaluation result and the pixel similarity evaluation result, and iteratively update the network parameters;

[0015] S6. Based on the trained linear registration network and non-linear transformation network, perform a registration operation on the pair of images to be registered, and output the registered image pair.

[0016] The linear registration network is composed of multiple layers of swin-transformer.

[0017] The linear transformation matrix includes a total of 12 parameters describing four operations of rotation, translation, scaling, and shearing.

[0018] In step S1, the MINE method is used to estimate the mutual information value between the target image and the linearly transformed source image, and the overall alignment is achieved by maximizing the mutual information value.

[0019] In step S3, the STN is used to apply the non - linear deformation field obtained in step S2 to the source image after linear transformation to obtain the final output image. The process of using the STN for transformation includes the sampling and interpolation processes of the pixel values of the source image. Among them, the trilinear interpolation algorithm is used for interpolation.

[0020] The structural similarity evaluation includes the calculation of the Dice loss and the edge loss between the segmentation label of the target image and the segmentation label of the source image after non - linear deformation. Among them, the edge loss is the distance between the corresponding organ contour edge points. During the network training process, the Dice loss and the edge loss cooperate with each other.

[0021] The pixel - level similarity evaluation includes the following steps:

[0022] Extract the modality - independent nearest - neighbor descriptors (MIND) of the target image and the source image after non - linear deformation respectively. The MIND constructs a high - dimensional vector using the gray - scale change information in the area around the pixel.

[0023] Calculate the distance between the MIND vectors of the target image and the source image after non - linear deformation to obtain the image difference at the pixel level.

[0024] Determine the regional weight based on the segmentation label of the image, multiply the regional weight by the image difference and take the average to obtain the final similarity evaluation result.

[0025] The method for determining the regional weight is as follows: preset the region of interest (ROI) in the image, set the regional weight of the pixels in the ROI to the maximum value of 1, and the regional weights of other pixels decrease linearly according to the distance from the pixel to the ROI.

[0026] A structural - guidance - based abdominal medical image registration device includes a memory, a processor, and a program stored in the memory. When the processor executes the program, the above - described method is implemented.

[0027] A storage medium stores a program, and when the program is executed, the above - described method is implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The present invention extracts structural information from the organ segmentation template and uses the structural information to guide the registration network to improve the registration accuracy in scenarios with complex deformations such as abdominal medical images.

[0030] 2. The present invention uses swin - transformer as the basic structure to construct the registration network, improves the network's ability to handle long - distance deformations, and avoids the problem of organ contour distortion.

[0031] 3. The present invention uses organ segmentation labels to construct regional weights, which can improve the registration accuracy of a certain organ or part of the region alone without affecting the overall registration accuracy, meeting the actual medical needs. Description of the Drawings

[0032] Figure 1 is the flowchart of the method of the present invention;

[0033] Figure 2 is the network structure diagram of the registration method of the present invention;

[0034] Figure 3 is the schematic diagram of the VoxelMorph network structure;

[0035] Figure 4 is the schematic diagram of the pixel-level similarity evaluation process. Detailed Embodiment

[0036] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0037] This embodiment provides a method for registering abdominal medical images based on structure guidance, as Figure 1 shown, including the following steps:

[0038] S1. Obtain a pair of abdominal multi-modal medical source images and target images with segmentation labels, input them into a linear registration network to obtain a linear transformation matrix a, and perform a linear transformation on the source image and its segmentation labels based on the linear transformation matrix a.

[0039] Since there are differences in morphological characteristics such as height and weight among different individuals, and there are also differences in the shooting angles and scaling scales of medical devices, it is necessary to perform linear registration before non-linear registration to make the images to be registered as similar as possible as a whole. Therefore, a linear registration network is constructed.

[0040] As Figure 2 shown, the linear registration network inputs an image pair [I m , I f composed of the source image and the target image, sends it into a linear registration network composed of multiple layers of swin-transformer, and outputs a linear transformation matrix, which includes a total of 12 parameters describing four operations of rotation, shift, scaling, and shear. Thereafter, use this linear transformation matrix to perform an affine transformation on the source image (I m ) to obtain a linearly deformed source image (I m ′). Then, use the MINE method to estimate the target image (I f ) and I mThe mutual information value of the ′ image. By maximizing this mutual information value, the two images are made as similar as possible to achieve the purpose of overall alignment.

[0041] S2. The linearly transformed source image and target image pair are input into the non - linear transformation network to obtain the non - linear deformation field φ.

[0042] Such as Figure 2 shown, the non - linear transformation network completes the local deformation of the image. Its input is the [I f , I m ′] image pair, and the output is the non - linear deformation field (φ).

[0043] Figure 3 Figure 16 shows the U - NET network structure adopted by Voxelmorph. The subscript fraction in each box represents the ratio of the feature map size at this stage to the size of the original picture, and the number marked in the box is the number of feature map channels. The non - linear transformation network structure adopted in this embodiment is the same as that of Voxelmorph. However, the feature extraction method adopted by VoxelMorph is a convolutional network, and this invention uses swin - transformer for layer - by - layer feature extraction (encode) and finally obtains the deformation vector field (DVF).

[0044] The non - linear deformation field is expressed as:

[0045] φ = Id + u

[0046] Where, φ is the non - linear deformation field; Id is the identity transformation field, expressed in tensor form, representing the coordinates of each pixel in the linearly transformed source image; u is the direct output of the non - linear transformation network, representing the non - linear DVF, which is composed of the displacement vectors of each pixel. Specifically, for any voxel p belonging to an n - dimensional source image, u(p) gives an n - dimensional displacement vector.

[0047] S3. According to the non - linear deformation field φ, the linearly transformed source image and its segmentation label are non - linearly deformed.

[0048] In this embodiment, STN (Spatial Transformer Networks) is used to apply the deformation field to the I m ′ image to obtain the final output image (I m ″), which includes the sampling and interpolation process of the original image pixel values. The interpolation algorithm uses trilinear interpolation.

[0049] S4. Based on the segmentation labels of the target image and the segmentation labels of the source image after non-linear deformation, perform structural similarity evaluation to determine the regional weights, and perform pixel similarity evaluation based on the target image, the source image after non-linear deformation, and the corresponding regional weights.

[0050] To determine the regional weights, it is necessary to manually set the region of interest (ROI), which is usually the whole or part of an organ. The regional weight represents the importance of the similarity of the pixel position to the similarity of the overall image. The weights of the pixels in the ROI are set to the maximum value of 1, and the weights of other pixels decrease linearly according to their distance from the ROI. In practical applications, the difference between the minimum weight value and the maximum weight value generally does not exceed 0.2.

[0051] Pixel similarity evaluation and structural similarity evaluation are key steps in introducing structural information.

[0052] Among them, the structural similarity evaluation directly calculates the segmentation labels, including two aspects. One is to calculate the Dice loss between the segmentation labels of the target image and the segmentation labels of the source image after non-linear transformation, and gradually increase the proportion of the overlapping area to the total area to align the organ structure. The other is to calculate the boundary loss, that is, the distance between the corresponding organ contour edge points, and try to reduce the gap between the edge points to align the organ contour. The Dice loss and the boundary loss cooperate with each other, with the Dice loss being the main one in the early stage of training and the boundary loss being the main one in the later stage. Such a design can avoid drastic fluctuations in the evaluation results caused by noise or human errors in the segmentation labels of small-volume organs and make the training more stable.

[0053] Pixel-level similarity evaluation calculates for I f and I m ″ images, as Figure 4 shown, including the following steps:

[0054] Respectively extract the modality-independent nearest neighbor descriptors MIND of the target image I f and the source image I m ″ after non-linear deformation. MIND constructs a high-dimensional vector using the gray-scale change information of the area around the pixel;

[0055] Calculate the distance between the MIND vectors of the target image I f and the source image I m ″ after non-linear deformation to obtain the pixel-level image difference;

[0056] Based on the segmentation labels of the image, determine the regional weights, multiply the regional weights by the image difference, and take the average to obtain the final similarity evaluation result.

[0057] This method generates regional weights from segmentation labels to guide the algorithm to focus on the region of interest (ROI), achieving the purpose of segmentation-guided registration.

[0058] S5. Based on the structural similarity evaluation result and the pixel similarity evaluation result, train the linear registration network and the non-linear transformation network, and iteratively update the network parameters.

[0059] S6. Based on the trained linear registration network and non-linear transformation network, perform a registration operation on the pair of images to be registered, and output the registered pair of images.

[0060] Embodiment 2

[0061] This embodiment further illustrates the application of the method described in Embodiment 1 above in the construction of a liver digital twin model.

[0062] Note: The liver digital twin model can be used in scenarios such as the diagnosis and treatment of liver-related diseases, and surgical navigation. The key to constructing a liver twin model is not only to reflect the morphological changes of the liver in terms of structure but also to reflect the functional changes of the liver. To achieve this goal, image fusion technology is required to fuse the CT image reflecting the liver structure and the arterial phase Primovist MRI image reflecting the liver function. The key to fusion is to use registration technology to establish the position correspondence between voxels in the two images. The present invention can be used to complete this task.

[0063] Implementation steps: 1. Collect CT images and arterial phase Primovist MRI images; 2. Establish a dataset by unifying the image size and pixel size; 3. Use the manual segmentation method to obtain image segmentation labels and delimit the liver as the ROI; 4. Use the method described in Embodiment 1 to train the linear registration network and the non-linear transformation network, and adjust the hyperparameters until the model effect reaches the optimal; 5. Directly input the pair of images to be registered into the trained network and output the registered arterial phase Primovist MRI image; 6. Use the deformed image for image fusion to establish a liver twin model.

[0064] Embodiment 3

[0065] This embodiment further illustrates the application of the method described in Embodiment 1 above in liver function assessment.

[0066] Description: The absorption and decomposition ability of the liver for specific chemical agents is an important indicator to measure liver function. By comparing the Gd-EOB-DTPA MRI images in the arterial phase and the hepatobiliary phase, the response of any region of the liver to the agent can be visually observed, which can better complete tasks such as regional liver function assessment and prediction of the risk of postoperative liver failure. However, during the process of obtaining images in the arterial phase and the hepatobiliary phase, due to the effects of respiration and heartbeat, the movement of the liver position is inevitably brought about. Therefore, it is necessary to use the method proposed in the present invention to align the images in the arterial phase and the hepatobiliary phase.

[0067] Implementation steps: 1. Collect Gd-EOB-DTPA MRI images in the arterial phase and the hepatobiliary phase; 2. Establish a data set by unifying the image size and pixel size; 3. Use the manual segmentation method to obtain the segmentation labels of abdominal organs, and define the liver as the ROI; 4. Use the method described in Example 1 to train the linear registration network and the non-linear transformation network, and adjust the hyperparameters until the model effect reaches the optimal; 5. Directly input the image pairs to be registered into the trained network, and output the registered MRI image pairs; 6. Observe and calculate the image change curve of any region of the liver to complete the liver function assessment.

[0068] Example 4

[0069] This embodiment provides a structure-guided abdominal medical image registration device, including a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements the method described in Example 1 above.

[0070] Example 5

[0071] This embodiment provides a storage medium, on which a program is stored. When the program is executed, it implements the method described in Example 1 above.

[0072] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0073] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in this technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A method for abdominal medical image registration based on structure guidance, characterized in that, It includes the following steps: S1. Obtain the abdominal multi-modal medical source image and target image pair with segmentation labels, input them into the linear registration network to obtain the linear transformation matrix, and perform linear transformation on the source image and its segmentation labels based on the linear transformation matrix; S2. Input the linearly transformed source image and target image pair into the non-linear transformation network to obtain the non-linear deformation field; S3. Perform non-linear deformation on the linearly transformed source image and its segmentation labels according to the non-linear deformation field; S4. Perform structural similarity evaluation based on the segmentation labels of the target image and the segmentation labels of the non-linearly deformed source image to determine the regional weights, and perform pixel similarity evaluation based on the target image, the non-linearly deformed source image, and the corresponding regional weights; S5. Train the linear registration network and the non-linear transformation network based on the structural similarity evaluation result and the pixel similarity evaluation result, and iteratively update the network parameters; S6. Based on the trained linear registration network and non-linear transformation network, perform a registration operation on the pair of images to be registered, and output the registered image pair.

2. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, The linear registration network is composed of multiple layers of swin-transformer.

3. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, The linear transformation matrix includes a total of 12 parameters describing four operations of rotation, translation, scaling, and shear.

4. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, In step S1, the MINE method is used to estimate the mutual information value between the target image and the linearly transformed source image, and the overall alignment is achieved by maximizing the mutual information value.

5. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, In step S3, the STN is used to apply the non-linear deformation field obtained in step S2 to the linearly transformed source image to obtain the final output image. The process of using STN for transformation includes the sampling and interpolation process of the pixel values of the source image. Among them, the trilinear interpolation algorithm is used for interpolation.

6. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, The structural similarity evaluation includes the calculation of the Dice loss and the edge loss between the segmentation label of the target image and the segmentation label of the non-linearly deformed source image. Among them, the edge loss is the distance between the corresponding organ contour edge points. During the network training process, the Dice loss and the edge loss cooperate with each other.

7. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, The pixel-level similarity evaluation includes the following steps: Extract the modality-independent nearest neighbor descriptors MIND of the target image and the non-linearly deformed source image respectively. MIND constructs a high-dimensional vector using the gray-scale change information in the surrounding area of the pixel; Calculate the distance between the MIND vectors of the target image and the non-linearly deformed source image to obtain the image difference at the pixel level; Determine the regional weights based on the segmentation labels of the image, multiply the regional weights by the image difference, and take the average to obtain the final similarity evaluation result.

8. The method for abdominal medical image registration based on structure guidance according to claim 1, characterized in that, The method for determining the regional weights is as follows: Preset the region of interest ROI in the image, set the regional weight of the pixels in the ROI to the maximum value of 1, and the regional weights of other pixels decrease linearly according to the distance from the pixel to the ROI.

9. An apparatus for abdominal medical image registration based on structure guidance, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1-8.

10. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the method described in any one of claims 1-8.