Cross-modal Large Deformation Image Registration Method Based on Semantic Mask
Through a multi-level registration framework based on semantic masks, the problem of difficulty in extracting the internal structure information of organs in cross-modal large deformation medical image registration is solved, and more accurate and smooth image registration results are achieved, which is suitable for the fusion of large deformation cross-modal medical images.
Patent Information
- Application Number
- CN202210896469.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-28
AI Technical Summary
In the prior art, in cross-modal medical image registration, especially in large deformations, it is difficult to accurately extract the internal structure information of the organ, resulting in unsmooth registration results and cannot meet the needs of doctors.
A multi-level registration framework based on semantic masks is adopted, including Mask-RCNN network segmentation organs, and a mean semantic mask is constructed. Through affine registration and deformation registration, the semantic mask is used to provide image information at different stages to improve registration accuracy.
It improves the accuracy and smoothness of cross-modal large deformation image registration, maintains the internal structure of the organ, adapts to the information differences of different modal images, and improves the image registration effect.
Smart Images

Figure CN115222780B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and mainly relates to an image registration method, which can be used to register medical images in cross-modal and large deformation scenarios. Background Art
[0002] Medical image registration is one of the most challenging tasks in the field of image processing. Whether in radiotherapy or puncture surgery, the role of cross-modal image registration technology is very crucial. With the progress of medical imaging equipment, for the same patient, images containing accurate anatomical information and images containing functional information can be collected. Doctors usually compare images of different modalities based on their subjective experience and knowledge to obtain different lesion information. A correct registration method can fuse the image information of different modalities, making it more convenient and accurate for doctors to observe lesions. There are mainly two problems to be solved in the image registration tasks for assisting radiotherapy and surgery, namely the image cross-modal problem and the large deformation problem of organs.
[0003] Currently, for cross-modal image registration with small deformations, registration methods based on deep features and generative models are mainly used. These methods are mainly applied to images with fixed structures such as the brain and lungs. For images of the abdomen with large deformations and unfixed structures, these methods cannot accurately register. At the same time, existing methods have better registration effects only on images with a small span of gray information, such as between T1 and T2 of magnetic resonance MRI, and between high-quality CT and low-quality CT. For images with a large span of gray information such as CT and magnetic resonance MRI, the registration effects of existing methods are very poor.
[0004] For the problem of large deformation image registration, currently, a multi-layer registration framework is mainly used to solve it. The commonly used frameworks include two steps: first rigid registration and then non-rigid registration. However, these frameworks are difficult to capture the information of cross-modal images. Some methods use the image segmentation results as auxiliary supervision information. However, the segmentation results can only carry the shape and position information of the corresponding organs, and cannot carry the internal structure information of the organs, resulting in non-smooth registration results, and the internal structure of the organs is damaged, which cannot meet the needs of doctors. Summary of the Invention
[0005] The purpose of the present invention is to propose a cross-modal large deformation image registration method based on semantic masks in view of the above-mentioned deficiencies of the prior art, so as to maintain the internal structure of organs during the large deformation registration process, improve the smoothness of the results, and improve the cross-modal registration effect.
[0006] To achieve the above purpose, the technical solution of the present invention includes:
[0007] (1) Using the cross-modal large deformation medical image dataset as the original data, it is divided into a training set and a test set according to a ratio of 4:1. Both the training set and the test set include two parts: labels and images. Among them, the images in the training set are marked as I V , and the images in the test set are marked as I T ;
[0008] (2) Input the images in the training set into the Mask-RCNN network, and use the labels in the training set and the output results of the Mask-RCNN network to iteratively update the network parameters. After 10,000 iterations, a trained Mask-RCNN network is obtained;
[0009] (3) Input the images in the test set and the training set into the trained Mask-RCNN network to obtain the segmentation result D T of the test set and the segmentation result DV of the training set;
[0010] (4) Use the segmentation results to construct a multi-level registration framework, and use the multi-level registration framework to obtain multi-level registration results:
[0011] (4a) Multiply the segmentation result D T of the test set and the segmentation result DV of the training set by a fixed coefficient L to obtain the mean semantic mask M T of the test set and the mean semantic mask M V of the training set;
[0012] (4b) Use the mean semantic mask M V of the training set to construct an affine registration network S A ′ trained with the semantic mask, and use S A ′ to obtain the affine registration result from the mean semantic mask M T of the test set;
[0013] (4c) Use the affine registration result to generate the centroid principal axis semantic mask MP θV of the training set and the centroid principal axis semantic mask M PθT of the test set;
[0014] (4d) Use the centroid principal axis semantic mask M PθV of the training set to construct a deformation registration network S B ′ trained with the semantic mask, and use S B ′ to obtain the deformation registration result from the centroid principal axis semantic mask M PθT of the test set. The obtained deformation registration result is the multi-level registration result.
[0015] The present invention has the following advantages compared with the prior art:
[0016] 1. The present invention solves the problem that general semantic information cannot be extracted during the registration of cross-modal images by introducing a method for generating semantic masks and generating different semantic masks according to the characteristics of organs.
[0017] 2. The present invention improves the accuracy of large deformation registration by introducing a multi-level registration framework, first performing affine registration and then deformation registration.
[0018] 3. The present invention improves the registration ability for organs with complex internal structures and unfixed positional structures by introducing semantic masks adapted to different registration stages and providing different image information for different registration stages. Description of the Drawings
[0019] Figure 1 is a flowchart of the implementation of the present invention.
[0020] Figure 2 is a comparison diagram of the registration results of lower abdominal magnetic resonance and CT data using the method of the present invention and five existing medical image registration algorithms respectively. Detailed Embodiment
[0021] In combination with the drawings, the embodiments and effects of the present invention will be further described in detail.
[0022] Refer to Figure 1 , and the implementation steps of this example are as follows:
[0023] Step 1, use the Mask-RCNN network to obtain the segmentation results of the training set and the test set.
[0024] One of the difficulties in cross-modal image processing is the problem of inconsistent gray levels. Using segmentation labels for supervised learning or information representation is one of the common methods. At present, it is difficult to obtain medical image segmentation labels, so this example uses the Mask-RCNN network of weak supervised learning for organ segmentation, and the specific implementation is as follows:
[0025] 1.1) Use the cross-modal large deformation medical image dataset as the original data, and divide it into a training set and a test set according to a ratio of 4:1. Both the training set and the test set include two parts: labels and images. Among them, the images in the training set are marked as I V , and the images in the test set are marked as I T ;
[0026] 1.2) Calculate its loss value L mask according to the Mask-RCNN network loss function:
[0027] L mask =-∑ i (y c log(p c ))
[0028] Among them, y c is the pixel value in the label, p c is the pixel value in the network output result, and ∑ i represents the summation of the calculation results for each point in the image;
[0029] 1.3) Use the backpropagation algorithm to calculate the parameter gradients of the Mask - RCNN network according to the loss value;
[0030] 1.4) Set the learning rate to 0.0005, and update the parameters of the Mask - RCNN network using the stochastic gradient descent method according to the parameter gradients of the Mask - RCNN network.
[0031] 1.5) After performing steps 1.2 to 1.4 iteratively 10,000 times, obtain the trained Mask - RCNN network;
[0032] 1.6) Input the images in the test set and the training set into the trained Mask - RCNN network to obtain the segmentation result D T of the test set and the segmentation result D V of the training set.
[0033] Step 2, for the segmentation result D T of the test set and the segmentation result D V of the training set, obtain the mean mask M T of the test set and the mean mask M V of the training set.
[0034] Since there are fewer parameters in the affine registration process, the registration algorithm has little deformation on the floating image. Therefore, the information required for the affine registration step is only the position, shape, and size of the target organ. Although the distance map can better represent the boundary information, the negative values inside its contour are not applicable in the process of neural network backpropagation. Therefore, the mean map that only carries the position, shape, and size information is used as the mask. The specific implementation is as follows:
[0035] 2.1) According to the average number of pixels N l of the organ to be registered and the total number of pixels N of the image, set the fixed coefficient L:
[0036]
[0037] 2.2) Multiply the segmentation result D V of the training set by the fixed coefficient L to obtain the mean mask M V of the training set;
[0038] 2.3) Multiply the segmentation result D T of the test set by the fixed coefficient L to obtain the mean mask M T of the test set.
[0039] Step 3: Use the mean semantic mask M of the training set V to construct an affine registration network S trained with the semantic mask A ′, and use S A ′ to obtain the affine registration result from the mean semantic mask M of the test set T
[0040] Using multi-step registration is a common method to solve the large deformation problem, where registration is divided into affine registration and deformation registration. Affine transformation includes linear transformation and translation, which can greatly change the shape of the image as a whole. Taking affine registration can greatly change the position and size of the image to be registered, preventing large deformation registration from damaging the image structure.
[0041] In this step, the affine registration network is used to calculate the affine registration parameters of the image to be registered, and the specific implementation is as follows:
[0042] 3.1) Input the mean mask M of the training set V into the affine registration network S A to obtain the affine registration parameter H of the training set V ;
[0043] 3.2) According to the affine registration parameter H of the training set V , perform coordinate mapping on the segmentation result D V and the image I V of the training set;
[0044] 3.3) Perform bilinear interpolation on the blank pixels after coordinate mapping of the segmentation result D V and the image I V of the training set, that is, weight and sum the adjacent pixel values according to the ratio of the distance between the blank pixel and the adjacent pixels, to obtain the pixel values of the blank pixels of the segmentation result D V and the image I V of the training set
[0045] 3.4) Use the pixel values of the blank pixels of the segmentation result D V and the image I V of the training set to fill the blank pixels in the coordinate mapping result of the segmentation result D V and the image I V of the training set, to obtain the affine registration result D V ′ of D V and the affine registration result I′ V of I V ;
[0046] 3.5) According to the affine registration network loss function, calculate the determinant loss value L of the affine registration network det and the orthogonal loss value L ortho :
[0047] L det = (-1 + det(H V + I)) 2
[0048]
[0049] where I is the identity matrix and λ is the singular value of H V + I;
[0050] 3.6) Use the backpropagation algorithm to calculate the parameter gradients of the affine registration network according to the loss value;
[0051] 3.7) Set the learning rate to 0.0001, and update the parameters of the affine registration network using the adaptive moment estimation method according to the parameter gradients of the affine registration network.
[0052] 3.8) After iterating steps 3.5 to 3.7 for 20,000 times, obtain the affine registration network S A ' trained with the semantic mask;
[0053] 3.9) Input the mean mask M T of the test set into the affine registration network S A ' trained with the semantic mask to obtain the affine registration parameters H T of the test set;
[0054] 3.10) According to the affine registration parameters H T of the test set, perform coordinate mapping on the segmentation result D T of the test set and the image I T ;
[0055] 3.11) Perform bilinear interpolation on the blank pixels after coordinate mapping of the segmentation result D T of the test set and the image I T , that is, obtain the pixel values of the blank pixels of the segmentation result D T of the test set and the image I T by calculating the linear interpolation results in two directions;
[0056] 3.12) Use the obtained pixel values of the blank pixels of the segmentation result D T of the test set and the image I T to fill the blank pixels in the coordinate mapping results of the segmentation result D T of the test set and the image I T to obtain the affine registration result D T ' of D T and the affine registration result I T ' of I T′.
[0057] Step 4: Generate the centroid principal axis semantic mask M of the training set using the affine registration result PθV and the centroid principal axis semantic mask M of the test set PθT .
[0058] Although the fixed-value mask can provide information about position, shape, and size, since the internal structure and texture information of the organ also need to be concerned during the registration process, it is very important to construct a mask that contains the internal structure information of the organ. For a single organ, although the imaging gray values of the same tissue in different modality images are different, a single modality is the same for the same tissue imaging. Therefore, the centroid feature constructed based on the mean value on the organ of a single image is highly adaptable to different modalities. If the centroid mask has poor effect on constraining the rotation transformation of the structure, it is easy to cause errors and insufficient information in the polar coordinate space. Therefore, the present invention also introduces the principal axis information as an additional constraint, and constructs a centroid principal axis mask based on this principle. The specific implementation is as follows:
[0059] 4.1) Multiplication and cutting: Multiply the affine registration result D V ′ of D V and the affine registration result I′ V of I V to obtain the cut training set image C V ;
[0060] 4.2) Calculate the centroid P V of the cut training set image C V :
[0061]
[0062]
[0063] where x and y are the coordinates of the pixels in the picture, and b V (x, y) is the pixel value of the cut training set image C V at the coordinate (x, y), is the coordinate of the centroid P V of the image to be obtained;
[0064] 4.3) Calculate the double angle 2θ V of the principal axis of the cut training set image C V :
[0065]
[0066] where a, b, and c are intermediate variables, and the calculation formulas are as follows:
[0067] a = ∫∫x V ′2 b V (x V ′, y V ′)dx V ′dy V ′
[0068] b = 2∫∫x V ′y V ′b V (x V ′, y V ′)dx V ′dy V ′
[0069] c = ∫∫y V ′ 2 b V (x V ′, y V ′)dx V ′dy V ′
[0070] where is the centroid coordinate of the image, (x, y) is the coordinate of the pixel in the image;
[0071] 4.4) According to the double angle 2θ of the major axis of the training set image C V calculate the major axis θ V : V
[0072]
[0073] 4.5) According to the major axis θ V calculate the coordinate of another point on the major axis except the centroid:
[0074]
[0075] where 1 is the abscissa of this point, is the ordinate of the point;
[0076] 4.6) Calculate the Euclidean distance d from each point of the segmented training set image to the major axis according to the two - point formula av ;
[0077]
[0078] 4.7) Calculate the Euclidean distance d from each point of the segmented training set image to the centroid according to the centroid coordinate cv ;
[0079]
[0080] 4.8) According to dav and d cv Calculate the centroid principal axis mask M of the training set PθV :
[0081]
[0082] where x and Y are the length and width of the training set image, and m″ xyv is based on d av and d cv Calculate the centroid principal axis mask M PθV The pixel value of the x-th row and y-th column, the formula is as follows:
[0083] m″ xyv = 255 - d cv c p - d av a p
[0084] In the formula, a p is the principal axis attenuation coefficient, and c p is the centroid attenuation coefficient;
[0085] 4.9) Multiplication and cutting: Multiply the affine registration result D T of D T ' and the affine registration result I T of I T ' to obtain the cut test set image C T ;
[0086] 4.10) Calculate the centroid P T of the cut test set image C T :
[0087]
[0088]
[0089] where x and y are the coordinates of the pixels in the picture, and b T (x, y) is the pixel value of the cut training set image C T at the coordinate (x, y), is the coordinate of the centroid P T of the image to be found.
[0090] 4.11) Calculate the double angle 2θ T of the principal axis of the cut test set image C T :
[0091]
[0092] where a, b, and c are intermediate variables, and the calculation formulas are as follows:
[0093] a = ∫∫x T ′ 2 b T (x T ′, y T ′)dx T ′dy T ′
[0094] b = 2∫∫x T ′y T ′b T (x T ′, y T ′)dx T ′dy T ′
[0095] c = ∫∫y T ′ 2 b T (x T ′, y T ′)dx T ′dy T ′
[0096] where are the centroid coordinates of the image, (x, y) are the coordinates of the pixels in the image;
[0097] 4.12) Calculate the major axis θ of the test set image C T from the double angle 2θ of its major axis: T : T
[0098]
[0099] 4.13) Calculate the coordinates of another point on the major axis except the centroid according to the major axis θ of the test set image: T
[0100]
[0101] where 1 is the abscissa of this point, is the ordinate of the point;
[0102] 4.14) Calculate the Euclidean distance d from each point of the test set image after cutting to the major axis according to the two-point formula: at ;
[0103]
[0104] 4.15) Calculate the Euclidean distance d from each point of the test set image after cutting to the centroid according to the centroid coordinates: ct ;
[0105]
[0106] 4.16) Calculate the centroid principal axis mask M of the test set according to d at and d ct : PθT :
[0107]
[0108] where X and Y are the length and width of the training set image, and m″ xyt is the centroid principal axis mask M calculated according to d at and d ct ; the pixel value at the x-th row and y-th column is given by the following formula: PθT :
[0109] m″ xyt = 255 - d ct c p - d at a p
[0110] In the formula, a p is the principal axis attenuation coefficient, and c p is the centroid attenuation coefficient.
[0111] Step 5: Use the centroid principal axis semantic mask M of the training set PθV to construct a deformation registration network S B ′ trained with the semantic mask, and use S B ′ to obtain the deformation registration result from the centroid principal axis semantic mask M of the test set PθT .
[0112] After affine registration, fine-tuning of the edges and interior of the image is required, and deformation registration is used at this time. Deformation registration calculates a position mapping vector for each pixel point in the image, and the matrix composed of these vectors is called the deformation field. Then, the pixel points are moved and interpolated according to the vectors to complete the registration task.
[0113] This step uses the deformation registration network to calculate the deformation field of the image to be registered, and the specific implementation is as follows:
[0114] 5.1) Input the centroid principal axis mask M of the training set PθV into the deformation registration network S B to obtain the deformation field F of the training set V ;
[0115] 5.2) According to the deformation field F of the training set V , perform coordinate mapping on the centroid principal axis mask M of the training set PθV and the training set image I V ′ after affine registration;
[0116] 5.3) For the centroid principal axis mask M of the training set PθV and the training set image I after affine registration V ′, perform bilinear interpolation on the blank pixels after coordinate mapping, that is, obtain the centroid principal axis mask M of the training set by calculating the linear interpolation results in two directions PθV and the training set image I after affine registration V ′, and the pixel values of the blank pixels;
[0117] 5.4) Use the centroid principal axis mask M of the training set PθV and the training set image I after affine registration V ′ to fill the blank pixels in the centroid principal axis mask M of the training set PθV and the training set image I after affine registration V ′ in the coordinate mapping result, and obtain the centroid principal axis mask M of the training set after deformation registration PθV ′ and the image I V ″.
[0118] 5.5) According to the deformation registration network loss function, calculate the correlation coefficient loss value L of the deformation registration network Corr and the total variation loss value L TV :
[0119]
[0120]
[0121] where Ω represents the spatial voxel, I1 and I2 respectively represent the floating image and the reference image in the centroid principal axis mask M of the training set PθV ′, e i is 's natural basis, Cov[I1, I2] is the cosine similarity of I1 and I2, and the calculation formula is as follows:
[0122]
[0123] 5.6) Use the backpropagation algorithm to calculate the parameter gradient of the deformation registration network according to the loss value;
[0124] 5.7) Set the learning rate to 0.0001, and update the parameters of the deformation registration network using the adaptive moment estimation method according to the parameter gradient of the deformation registration network;
[0125] 5.8) After iterating steps 5.5 to 5.7 for 20,000 times, obtain the deformation registration network S B ′ trained with the semantic mask;
[0126] 5.9) The centroid principal axis mask M of the test setPθT Input into the deformation registration network S trained with semantic masks B ′, to obtain the deformation field F of the test set T ;
[0127] 5.10) According to the deformation field F of the test set T , perform coordinate mapping on the centroid principal axis mask M of the test set PθT and the test set image I′ after affine registration V ;
[0128] 5.11) Perform bilinear interpolation on the blank pixels after coordinate mapping of the centroid principal axis mask M of the test set PθT and the test set image I′ after affine registration V , that is, obtain the pixel values of the blank pixels of the centroid principal axis mask M of the test set PθT and the test set image I′ after affine registration V by calculating the linear interpolation results in two directions;
[0129] 5.12) Use the pixel values of the blank pixels obtained from the centroid principal axis mask M of the test set PθT and the test set image I′ after affine registration V to fill the blank pixels in the coordinate mapping results of the centroid principal axis mask M of the test set PθT and the test set image I′ after affine registration V , to obtain the centroid principal axis mask M PθT ′ of the test set after deformation registration and the image I T ″′.
[0130] The effects of the present invention can be further illustrated by the following simulations.
[0131] 1. Simulation conditions:
[0132] The simulation platform for this experiment is a desktop computer with an Intel Core i7-9700K CPU and 32GB of memory, the operating system is Windows10, a neural network model is constructed and trained using a mixed programming of python3.6, keras2.2.4 and tensorflow1.13.0, and acceleration is performed using an NVIDIA 1080Ti GPU and CUDA10.0.
[0133] The experimental data used in the simulation are 158 lower abdominal preoperative T2MRI and intraoperative CBCT patients from a certain hospital's radiology department and the radiology department. Each group includes more than 60 MRIs and more than 100 CBCTs, resampled to a spacing of (0.97, 0.97, 5) mm. Each group is manually aligned.
[0134] The batch size of the segmentation network is set to 2, the initial learning rate is set to 0.0001, and the optimizer used is Adam. There are more than 1200 pieces of labeled data segmented from three organs, and 150 pairs are separated as the test set.
[0135] The segmentation performance evaluation metrics used in the simulation include the Dice Similarity Coefficient (DSC), Mutual Information (MI), and Average Surface Distance (ASD). The specific calculation formulas are as follows:
[0136]
[0137] MI(A, B) = H(A) + H(B) - H(A, B)
[0138]
[0139] where A represents the ground truth label, B represents the prediction result, H(A) and H(B) represent the information entropy of A and B respectively, H(A, B) is the joint entropy of A and B, S(A) represents the surface pixels of the ground truth label, S(B) represents the surface pixels of the prediction result, d(s A , S(B)) represents the shortest distance from any pixel of the ground truth label to the surface pixels of the prediction result, and d(s B , S(A)) represents the shortest distance from any pixel of the prediction result to the surface pixels of the ground truth label.
[0140] Existing image registration methods used in the simulation include the integrated registration method elastix, the iterative registration method demons, the diffeomorphic version of demons Symmetric demons, the traditional method SyN, and the deep learning method VoxelMorph.
[0141] 2. Simulation Content
[0142] (2.1) Under the above simulation conditions, the present invention and the existing 5 registration methods are used to register the dataset, and the results are as Figure 2 shown. The first row shows the comparison of the registration results of different algorithms with the reference image, and the second row shows the stitching result of the reference image and the registration result. The upper left and lower right quarters are the reference images, and the upper right and lower left quarters are the registration results.
[0143] As Figure 2 can be seen from the first row, the existing registration methods of elastix, demons, Symmetric demons, SyN, and VoxelMorph cannot obtain accurate registration results, while the registration result of the present invention is relatively accurate.
[0144] As Figure 2As can be seen in the second row, the results of the existing SyN and the present invention can be smoothly stitched with the reference image, indicating that the registration results are relatively accurate. Compared with SyN, the results of the present invention can better align the contours in the stitching comparison, indicating that the present invention is more accurate in the registration of position and external contours.
[0145] (2.2) Calculate the quantitative indicators DSC, MI, and ASD of the registration tests of the existing elastix, demons, Symmetric demons, SyN, VoxelMorph registration methods and the present invention on the test set respectively, and the results are shown in Table 1.
[0146] Table 1 DSC, MI, and ASD results of image registration by different methods
[0147] Method Dice MI ASD elastix 0.315895 0.194039 15.876984 demons 0.280000 0.093330 16.075311 Symmetric demons 0.086978 0.012995 nan SyN 0.412759 0.128072 13.076204 VoxelMorph 0.337425 0.116464 16.062301 The present invention 0.971510 0.322582 0.561818
[0148] As can be seen from Table 1, in terms of the DSC, MI, and ASD indicators, the present invention has a significant improvement compared with other methods. Among them, the DSC indicator has increased by more than 0.6, and the ASD indicator has increased by more than 12. This is because other methods cannot obtain sufficient information in the cross-modal scenario and do not have sufficient deformation ability to accurately register large-deformation organs, while the present invention represents the information of different modalities through semantic masks and accurately registers large-deformation organs through multi-level registration.
[0149] The above comparison results show that the present invention can solve the problem of the inability to extract general semantic information during the registration of cross-modal images and improve the accuracy of large-deformation organ registration.
Claims
1. A cross-modal large deformation image registration method based on semantic masks, characterized in that, Including: (1) Using a cross-modal large deformation medical image dataset as the original data, it is divided into a training set and a test set according to a ratio of 4:
1. Both the training set and the test set include two parts: labels and images. Among them, the images in the training set are marked as I V , and the images in the test set are marked as I T ; (2) Input the images in the training set into the Mask-RCNN network, and use the labels in the training set and the output results of the Mask-RCNN network to iteratively update the network parameters until after 10,000 iterations, a trained Mask-RCNN network is obtained; (3) Input the images in the test set and the training set into the trained Mask-RCNN network to obtain the segmentation result D of the test set T and the segmentation result D of the training set V ; (4) Use the segmentation results to construct a multi-level registration framework, and use the multi-level registration framework to obtain multi-level registration results: (4a) Multiply the segmentation result D of the test set T and the segmentation result D of the training set V by a fixed coefficient L to obtain the mean semantic mask M of the test set T and the mean semantic mask M of the training set V ; (4b) Using the mean semantic mask M of the training set V , construct an affine registration network S trained with the semantic mask A ′, and use S A ′ to obtain the affine registration result from the mean semantic mask M of the test set T ; (4c) Generate the centroid principal axis semantic mask M of the training set using the affine registration result PθV and the centroid principal axis semantic mask M of the test set PθT ; (4d) Use the centroid principal axis semantic mask M of the training set PθV to construct a deformation registration network S trained with the semantic mask B ′, and use S B ′ to obtain the deformation registration result from the centroid principal axis semantic mask M of the test set PθT . The obtained deformation registration result is the multi-level registration result.
2. The method according to claim 1, characterized in that, In (4b), the mean semantic mask M of the training set is used V to construct an affine registration network S trained with the semantic mask A ′, and S A ′ is used to obtain the affine registration result from the mean semantic mask M of the test set T The implementation is as follows: (4b1) Input the mean semantic mask M of the training set V into the affine registration network S A to obtain the affine registration parameters H of the training set V ; (4b2) Use the affine registration parameter H of the training set V Perform affine registration on the segmentation result D of the training set respectively V and the image I V to obtain the affine registration result D V ' of D V and the affine registration result I' V of I V ; (4b3) Use the segmentation result D of the training set after affine registration V ′ and the affine registration parameter H of the training set V Iteratively update the network parameters until after 20,000 iterations, an affine registration network S trained with a semantic mask is obtained A ′; (4b4) Input the mean semantic mask M of the test set T into the affine registration network S A ' trained with the semantic mask to obtain the affine registration parameters H T ; Use the affine registration parameter H of the test set T , and using the same method as the training set, perform affine registration on the segmentation result D T of the test set and the image I T respectively to obtain the affine registration result D T ' of D T and the affine registration result I T ' of I. T '.
3. The method according to claim 1, characterized in that Generate the centroid principal axis semantic mask M of the training set and the test set using the affine registration results in (4c), and the implementation is as follows: PθV and the centroid principal axis semantic mask M of the test set PθT , as follows: (4c1) Multiplication and cutting: Multiply D V 's affine registration result D V ' and I V 's affine registration result I' V to obtain the cut training set image C V ; (4c2) Calculate the centroid P of the cut training set image C V and the principal axis θ V ; V ; (4c3) According to D V 's affine registration result D V ' and the centroid P of the cut training set image C V , the principal axis θ V , calculate the centroid principal axis semantic mask M of the training set V ; PθV ; (4c4) Multiplication and cutting: Multiply D T 's affine registration result D T ' and I T 's affine registration result I T ' to obtain the cut test set image C T ; (4c5) Calculate the centroid P of the cut test set image C using the same method as the training set T ; T the major axis θ T ; (4c6)According to D T ' of the affine registration result D T ', and the centroid P T of the cut test set image C T , major axis θ T , using the same method as the training set, calculate the centroid major axis semantic mask M PθT of the test set.
4. The method according to claim 1, wherein In (4d), using the centroid principal axis semantic mask M of the training set PθV , construct a deformation registration network S trained with the semantic mask B ′, and use S B ′ to obtain the deformation registration result from the centroid principal axis semantic mask M of the test set PθT , and the implementation is as follows: (4d1) Input the centroid principal axis semantic mask M of the training set PθV into the deformation registration network S B to obtain the deformation field F of the training set V ; (4d2) Use the deformation field F of the training set V Perform deformation registration on the centroid principal axis semantic mask M of the training set respectively PθV and the training set image I after affine registration V ′ to obtain the centroid principal axis semantic mask M of the training set after deformation registration PθV ′ and the image I V ″; (4d3) Use the centroid principal axis semantic mask M of the training set after deformation registration PθV ′ and the deformation field F of the training set V Iteratively update the network parameters until after 20,000 iterations, obtain the deformation registration network S trained with the semantic mask B ′; (4d4) Input the centroid principal axis semantic mask M of the test set PθT into the deformation registration network S B ' trained with the semantic mask to obtain the deformation field F T ; Use the deformation field F of the test set T , and using the same method as the training set, perform deformation registration on the centroid principal axis semantic mask M PθT of the test set and the image I T ' respectively, to obtain the deformed centroid principal axis semantic mask M PθT ' and the image I T ″ after deformation registration.
5. The method according to claim 1, characterized in that In the above (2), using the labels in the training set and the output results of the Mask-RCNN network to iteratively update the network parameters is achieved as follows: (2a) Calculate the loss value according to the Mask-RCNN network loss function: where y c is the pixel value in the label, p c is the pixel value in the network output result, i represents the sum of the calculation results for each point in the image, and L mask is the network loss value; (2b) Use the backpropagation algorithm to calculate the parameter gradients of the Mask-RCNN network according to the loss value; (2c) Set the learning rate to 0.0005, and update the parameters of the Mask-RCNN network using the stochastic gradient descent method according to the parameter gradients of the Mask-RCNN network.
6. The method according to claim 1, characterized in that The fixed coefficient L in (4a) is set according to the average number of pixels N of the organ to be registered l and the total number of pixels N of the image, that is:
7. The method according to claim 2, characterized in that, The affine registration parameter H of the training set used in (4b2) V Perform affine registration on the segmentation result D of the training set V and the image I V respectively, and the implementation is as follows: (4b2a)According to the affine registration parameter H of the training set V , perform coordinate mapping on the segmentation result D V of the training set and the image I V ; (4b2b) Perform bilinear interpolation on the blank pixels after coordinate mapping, that is, obtain the pixel values of the blank pixels by calculating the linear interpolation results in two directions; (4b2c) Fill the blank pixels in the coordinate mapping result with the pixel values of the blank pixels to obtain the segmented result D of the training set after affine registration V ′ and the training set image I V ′.
8. The method according to claim 2, wherein The segmentation result D V ′ of the training set after affine registration used in (4b3) and the affine registration parameters H V Iteratively update the network parameters to achieve the following: (4b3a)Calculate the determinant loss value \(L\) of the affine registration network according to the loss function of the affine registration network det and the orthogonal loss value \(L\) ortho : L det = (-1 + det(H V + I)) 2 where I is the identity matrix and λ is the singular value of H V +I, and det(H V +I) is the determinant operation on H V +I; (4b3b) Use the backpropagation algorithm to calculate the parameter gradients of the affine registration network according to the loss value; (4b3c) Set the learning rate to 0.0001, and update the parameters of the affine registration network using the adaptive moment estimation method according to the parameter gradients of the affine registration network.
9. The method according to claim 3, characterized in that, Calculate the centroid P V of the segmented training set image C V and the principal axis θ V as follows: (4c2a) Calculate the centroid P of the training set image C after cutting V V : where x and y are the coordinates of the pixels in the image, respectively, and b V (x, y) is the cropped training set image C V The pixel value at the coordinate (x, y), is the centroid P V of the image to be found; (4c2b) Calculate the double angle 2θ of the principal axis of the training set image C after cutting V V : Where a, b, and c are intermediate variables, and the calculation formulas are as follows: a = ∫∫x V ′ 2 b V (x V ′, y V ′)dx V ′dy V ′ b = 2∫∫x V ′y V ′b V (x V ′, y V ′)dx V ′dy V ′ c = ∫∫ y V ′ 2 b V (x V ′, y V ′)dx V ′dy V ′ wherein is the centroid coordinate of the image, (x, y) is the coordinate of the pixel in the image; (4c2c) Obtain the principal axis θ according to the half-angle formula V :[[]]END]] 10. The method according to claim 3, characterized in that, In the above (4c3), according to the segmentation result D of the training set after affine registration V ′ and the cut training set image C V 's centroid P V , principal axis θ V , calculate the centroid principal axis semantic mask M of the training set PθV, The implementation is as follows: (4c3a)According to the principal axis θ V Calculate the coordinates of another point on the principal axis except the centroid: where 1 is the abscissa of this point, and tanθ V is for θ V the tangent value of, and is the ordinate of the point; (4c3b) Calculate the Euclidean distance d from each point of the cut training set image to the main axis according to the two-point formula av ; (4c3c) Calculate the Euclidean distance d from each point of the segmented training set image to the centroid according to the centroid coordinates cv ; (4c3d)According to d av and d cv Calculate the centroid principal axis semantic mask M of the training set PθV : m″ xyv = 255 - d cv c p - d av a p where a p is the spindle attenuation coefficient, and c p is the centroid attenuation coefficient. The matrix composed of pixel points m″ xyv is the centroid spindle semantic mask M PθV of the training set.
11. The method according to claim 4, characterized in that The deformation field F of the training set is used in (4d2). V The centroid principal axis semantic mask M of the training set is separately PθV and the training set image I after affine registration V ′ are subjected to deformation registration to obtain the centroid principal axis semantic mask M of the training set after deformation registration PθV ′ and the image I V ″, which is achieved as follows: (4d2a)According to the deformation field F of the training set V , perform coordinate mapping on the centroid principal axis semantic mask M PθV of the training set and the training set image I V ' after affine registration; (4d2b) Perform bilinear interpolation on the blank pixels after coordinate mapping, that is, obtain the pixel values of the blank pixels by calculating the linear interpolation results in two directions; (4d2c) Fill the blank pixels in the coordinate mapping result with the pixel values of the blank pixels to obtain the semantic mask M of the centroid principal axis of the training set after deformation registration PθV ′ and the image I V ″.
12. The method according to claim 4, characterized in that, The centroid principal axis semantic mask M of the training set after deformation registration and the deformation field F of the training set are used in (4d3). PθV ' V Iteratively update the network parameters to achieve the following: (4d3a)Calculate the correlation coefficient loss value \(L\) of the deformation registration network according to the deformation registration network loss function Corr and the total variation loss value \(L\) TV : Among them, Ω represents the spatial voxel, and I1 and I2 respectively represent the floating image and the reference image in the semantic mask M PθV ′, and e i is the natural basis of, and Cov[I1, I2] is the cosine similarity of I1 and I2. The calculation formula is as follows: (4d3b) Use the backpropagation algorithm to calculate the parameter gradients of the deformation registration network according to the loss value; (4d3c) Set the learning rate to 0.0001, and update the parameters of the deformation registration network using the adaptive moment estimation method according to the parameter gradients of the deformation registration network.