Preoperative multi-modal image deep learning registration method for spinal surgery
By employing deep learning methods for deep affine registration and deformable registration networks, the slow speed and distortion issues of traditional tools in spinal image registration are resolved, achieving efficient and accurate multimodal image registration to meet surgical needs.
Patent Information
- Application Number
- CN202510855546.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-17
AI Technical Summary
Existing traditional medical image registration tools are slow, prone to distortion and size misalignment in multimodal image registration of the spine, and require re-registration for each new image pair, resulting in high time costs. Furthermore, deep learning methods cannot complete affine registration tasks.
We employ a deep learning approach based on CNN and Swin transformer, using a deep affine registration network and a deep deformable registration network to perform affine alignment and tissue detail alignment of images, respectively. We utilize deep learning to extract global and local features, achieving smooth deformation and efficient registration.
It achieves efficient and smooth registration of spinal CT and MRI images, reduces distortion, improves registration efficiency, and ensures that anatomical tissues are in the correct position to meet the needs of surgical diagnosis and planning.
Smart Images

Figure CN120807596A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of medical image processing, and particularly relates to a preoperative multi-modal image deep learning registration method for spine surgery. BACKGROUND
[0002] When imaging the spine, CT and nuclear magnetic resonance are the main imaging means. CT images have good imaging effect on bone tissue, while nuclear magnetic resonance has good imaging effect on soft tissues such as intervertebral discs and nerve roots. When doctors actually refer to the images, they need to look back and forth, which is easy to cause misalignment of the relative position. Therefore, multi-modal image automatic registration needs to be completed before spine-related surgery.
[0003] Currently, traditional tools that can complete the task of automatic registration of medical images include SimpleITK, Elastix, ANTs (SyN), deedsBCV and NiftyReg, and it is specified that one of the two images to be registered is a fixed image Fixed Image and the other is a moving image Moving Image. The above tools can align the anatomical tissues of the moving image to the corresponding positions in the fixed image coordinate space. In addition, existing deep learning-based medical image automatic registration algorithm models include VoxelMorph and Swin-VoxelMorph.
[0004] The existing process of automatic registration of medical images is as follows: a deep neural network is designed, which can be trained with CT as the fixed image and nuclear magnetic resonance as the moving image, or with nuclear magnetic resonance as the fixed image and CT as the moving image, to finally complete the automatic registration task. The existing similar neural network architecture is as shown in Figure 1 If the extraction convolutional layer and the fusion convolutional layer are based on CNN, it is VoxelMorph; if the extraction convolutional layer and the fusion convolutional layer are based on swin transformer, it is Swin-VoxelMorph.
[0005] However, the traditional medical image registration has the following defects: 1. The existing traditional medical image registration tools only perform iterative deformation on specific image pairs with similarity metrics such as MSE (mean square error), NCC (normalized cross correlation), MI (mutual information) as indicators, which is slow in registration speed. The human tissues in the image are prone to distortion in the deformation process, the deformation field is not smooth, and every time a new image pair needs to be registered, it must be re-registered for new data, which is time-consuming.
[0006] 2、Due to the size and modal differences of CT and nuclear magnetic images of the spine part are relatively large, especially the size space difference, the number of CT sagittal scanning layers is more, generally 300-600 layers, and the number of nuclear magnetic sagittal scanning layers is less, generally only 10-20 layers, the traditional registration tool has very poor registration effect for the multi-modal images of the spine part, and serious problems such as tissue confusion, position deviation and size misalignment are prone to occur.
[0007] 3、Many existing image registration algorithms based on deep learning require that the two images as input must be the same size, then the processing of these deep learning methods is only deformable registration, which cannot complete the affine registration task, for the spine image, this requires the image to undergo affine registration alignment size before inputting into the neural network, and the affine registration effect using the traditional tool is poor. SUMMARY
[0008] The application provides a preoperative multi-modal image deep learning registration method for spine surgery, which uses deep learning instead of traditional tools, is more smooth for tissue deformation and is not prone to distortion, considers the alignment of tissue area position and the connection between voxels, greatly reduces the error of affine alignment image size, and improves the registration efficiency.
[0009] To achieve the above object, the application provides the following technical scheme: A preoperative multi-modal image deep learning registration method for spine surgery, which comprises the following two stages: In the first stage, the affine registration aligns the size of the moving image and the fixed image: a deep affine registration network based on CNN and swin transformer is constructed; the moving image and the fixed image are input into the deep affine registration network, and image features are extracted through two different mechanism data flow paths, i.e., the data is encoded through an encoder Encoder, one data flow path passes through a CNN-based U-Net-like structure to learn the global combined affine registration logic of the image; the other data flow path is to first divide the image into several image blocks, then linearly project the image block data through a LinearProjection layer, and then process the projected linear image block data through several swin transformer layers composed of continuous swin transformer blocks to learn the local combined affine registration logic of the image; then the image feature data extracted by the encoder Encoder and the combined affine registration logic learned by the model are fused and decoded through a decoder Decoder, the decoder Decoder under the two data flow paths is composed of different numbers of CNN modules and up-sampling operation modules, the data under the two paths is fused through the decoder Decoder, and is decoded into a first weight file according to the function of the CNN module; the original moving image is warped using the first weight file to achieve affine registration and obtain an affine registration result image; the output result of the first stage is a medical image file in “.nii.gz” format, which is used as the input of the second stage; In the second stage, the tissue details in the images are aligned through deformable registration: a deep deformable registration network based on CNN and swin transformer is constructed; the medical image file output by the first stage is input into the deep deformable registration network; the fixed image and the affine registration result image are tensor spliced in the channel dimension; after splicing, image features are extracted through two different data flow paths, that is, data is encoded through an encoder, one data flow path passes through a CNN-based U-Net-like structure to learn the global deformable registration logic of the image; the other data flow path still divides the complete input data into blocks, then linearly projects, and then flows through several swin transformer layers composed of continuous swin transformer blocks to learn the local deformable registration logic of the image; then the extracted image feature data and the learned deformable registration logic of the model are fused and decoded, the decoders under the two data flow paths are respectively composed of a certain number of CNN modules and up-sampling operation modules, the data under the two paths are fused through the decoders, and are decoded into a second weight file according to the function of the CNN module; finally, the affine registration result image is deformed using the second weight file to obtain the final registration result.
[0010] Further, in the first stage, the combined affine registration logic is the multiplexing and combination logic of the translation, rotation, stretching, shearing and interpolation steps of the image.
[0011] Further, in the second stage, the deformable registration logic includes voxel correspondence logic, voxel movement logic and interpolation logic.
[0012] Further, in the second stage, the voxel movement logic and the interpolation logic are based on the velocity field theory or the optical flow field theory or the displacement vector field theory.
[0013] Further, the first weight file is an affine registration deformation field.
[0014] Further, the second weight file is a deformable registration deformation field.
[0015] Further, the input data, intermediate data and result data are medical image files in “.nii.gz” format.
[0016] Further, the mamba block with long-distance modeling capability is used to replace the swim transformer block.
[0017] Further, the KAN Linear layer with stronger fitting capability is used to replace the Linear Projection layer.
[0018] Compared with the prior art, the present application has the following beneficial effects: 1、The registration method of the present application is based on a deep affine registration network of deep learning, learns the global features of the entire data set for full data domain optimization, without iteration on each image pair, and the deformation of the tissue is more smooth and less prone to distortion.
[0019] 2、The registration method of the present application is based on a neural network of swin transformer, which can well model the global and local information of the image, and balance the alignment of the tissue area position and the connection between voxels.
[0020] 3、The affine registration method of the present application is based on a deep affine registration network of deep learning, which can not only realize the affine alignment of CT to nuclear magnetic size (registration of more layers to less layers) in the case of obvious difference between the number of layers of CT and nuclear magnetic image scanning, but also realize the affine alignment of nuclear magnetic to CT size (registration of less layers to more layers), and the registration effect of the present application is better than that of the traditional tool in terms of the commonly used quantitative indicators of medical image registration, and the registration efficiency is improved.
[0021] 4、The affine registration method of the present application is based on a deep deformable registration network of deep learning, which processes the two images after affine alignment, so that the anatomical tissues such as spinal bones, intervertebral discs and nerve roots in the result image are in the correct position in details (such as edges), thereby meeting the needs of assisting doctors in diagnosing diseases and planning operation paths. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is an architecture diagram of the prior similar neural network; Figure 2 is a flowchart of the preoperative multi-modal image deep learning registration method of the present application; Figure 3 is an architecture diagram of the deep affine registration network and the deep deformable registration network of the present application. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0024] The present application provides a preoperative multi-modal image deep learning registration method for spinal surgery, as shown in Figure 2 and Figure 3 The registration method includes the following two stages: In the first stage, the affine registration aligns the size of the moving image and the fixed image: Firstly, a deep affine registration network based on CNN and swin transformer is constructed; Secondly, the moving image and the fixed image are input into the deep affine registration network, and image features are extracted through two different data flow paths, i.e., the data is encoded by an encoder. One path of the encoder learns the reuse and combination logic of the translation, rotation, stretching, shearing and interpolation steps of the whole image through a U-Net-like structure based on CNN. The other path is to first divide the image into several image blocks, because swin transformer has a unique mechanism of dividing image windows, cutting image blocks and processing image blocks. Then the image block data is linearly projected through a Linear Projection layer. This data flow path processes the projected linear image block data through several swin transformer layers composed of consecutive swin transformer blocks to learn the reuse and combination logic of the translation, rotation, stretching, shearing and interpolation steps of the local image, which can be called combined affine registration logic. Then the image feature data extracted by the encoder and the combined affine registration logic learned by the model are fused and decoded by a decoder. The decoder in the two data flow paths is composed of different numbers of CNN modules and up-sampling operation modules. The data in the two paths is fused through the decoder, and the data is decoded into a first weight file according to the function of the CNN module, which can also be called an affine registration deformation field.
[0025] Then, the original moving image is warped using the first weight file (affine registration deformation field) to achieve good affine registration and obtain an affine registration result image.
[0026] The output result of the first stage is a medical image file in the “.nii.gz” format, which is also used as the input of the second stage. In the present application, the input data, intermediate data and result data of the model are all medical image files in the “.nii.gz” format.
[0027] The second stage is to align the tissue details in the images through deformable registration. The same network architecture as the first stage is still used. First, a deep deformable registration network based on CNN and swin transformer is constructed. Unlike the first stage, in the second stage, the affine moving image and the fixed image are first tensor concatenated in the channel dimension (the size is the same at this time, so tensor concatenation can be performed). After tensor concatenation, the image features are extracted through two different data flow paths, i.e., the data is encoded. One data path of the encoder is still based on CNN and similar to the U-Net structure. In another data path, the complete input data is first divided into blocks, then linearly projected, and then flows through several swin transformer layers composed of consecutive swin transformer blocks to learn the local image deformable registration logic. In this stage, the model learns the voxel correspondence logic, voxel movement logic, and interpolation logic based on the velocity field theory or the optical flow field theory or the displacement vector field theory (DVF). The voxel correspondence logic, voxel movement logic, and interpolation logic are included in the deformable registration logic. Then the extracted image feature data and the learned deformable registration logic are fused and decoded. The decoders in the two data flow paths are still composed of different numbers of CNN modules and up-sampling operation modules. The data in the two paths is fused through the decoder, and the CNN module is decoded into a second weight file, which can also be called a deformable registration deformation field. Finally, the second weight file (deformable registration deformation field) is used to deformably warp the affine-aligned moving image to obtain the final registration result.
[0028] The characteristics of the above registration method include the following aspects: 1. The working mechanism of the traditional registration tool is (taking the normalized cross correlation (NCC) between two registered images as an example of similarity measurement optimization): move the moving image to the fixed image space by one point each time, then interpolate, then calculate the NCC to measure the similarity between the two images at this time, then move one point, then interpolate, then calculate the NCC again, then move again, then interpolate again, and follow this iterative process until the NCC is less than a threshold. This is a method based on spatial information. This traditional method can only be used in one image pair, and it is easy to cause irregular deformation.
[0029] The deep learning method used in the registration method can extract features of all image pairs in the data set, model shallow spatial information and deep semantic information, and finally obtain a weight file (which can be regarded as a deformation field), which can be used to register all image pairs in the data set, greatly reducing the time cost and improving the registration efficiency.
[0030] 2. The interpolation method of the traditional registration tool in the registration process is only linear interpolation, nearest neighbor interpolation and other simple algorithms, and there is a significant layer gap between the spinal CT and the magnetic resonance image, which makes it difficult to accurately interpolate, especially when registering the magnetic resonance image to the CT space. Simple interpolation algorithms such as linear interpolation and nearest neighbor interpolation cannot obtain comprehensive local information, which results in a large amount of error when the traditional tool aligns the size in affine registration.
[0031] The feature extraction capability (spatial information, semantic information) of the deep learning method used in the registration method makes the model more comprehensively capture the local information around a voxel, and the use of local voxel average interpolation makes the interpolation result more accurate, thereby greatly reducing the error caused by the step of aligning the image size in affine alignment, and replacing the traditional registration tool for affine registration.
[0032] 3. VoxelMorph and Swin-VoxelMorph require the input of the neural network to be images of the same size, because the first step in the model is a tensor splicing process, and different sizes cannot be spliced, so the importance of affine is emphasized. The deep affine registration network used in the registration method of the present application is not available in other deep learning methods.
[0033] 4. For deformable registration, the present application uses an innovative neural network architecture that is different from most deep learning models. Compared with the existing VoxelMorph and Swin-VoxelMorph, the deep neural network used in the second stage of deformable registration in the present application has stronger feature extraction capability, stronger modeling and registration logic capability, and higher accuracy and better smoothness of the final registration result.
[0034] Obviously, those skilled in the art can make various modifications and variations to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A deep learning registration method for preoperative multimodal images in spinal surgery, characterized by: It includes the following two stages: In the first stage, affine registration aligns the sizes of the moving and fixed images: a deep affine registration network based on CNN and Swin transformer is constructed. The moving and fixed images are input into the deep affine registration network, and image features are extracted through two different data flow paths. The data is encoded through an encoder. One data flow path uses a CNN-based U-Net-like structure to learn the global combined affine registration logic of the image. In the other data flow path, the image is first divided into several image blocks, and then the image block data is linearly projected through a LinearProjection layer. The projected linear image block data is then processed through several Swin transformer layers composed of consecutive Swin transformer blocks to learn the combined affine registration logic of the local image. The image feature data extracted by the encoder and the combined affine registration logic learned by the model are then fused and decoded by the decoder. The decoders in the two data flow paths are composed of different numbers of CNN modules and upsampling operation modules. The data from the two paths are fused through the decoder and decoded into the first weight file according to the function of the CNN module. The original moving image is distorted using the first weight file to achieve affine registration and obtain the affine registration result image; the output of the first stage is a medical image file in the ".nii.gz" format, which serves as the input of the second stage; In the second stage, the tissue details in the image are aligned through deformable registration: a deep deformable registration network based on CNN and swin transformer is constructed; the medical image file output from the first stage is input into the deep deformable registration network; the fixed image and the affine registration result image are tensor-stitched in the channel dimension, and then the image features are extracted through two different data flow paths after splicing, that is, the data is encoded through the encoder. One of the data flow paths passes through a CNN-based U-Net structure to learn the global deformable registration logic of the image; the other data flow path still divides the complete input data into blocks, then linearly projects it, and then flows through several swin transformer blocks composed of consecutive swin transformer blocks. The transformer layer learns the deformable registration logic of the local image. The extracted image feature data and the deformable registration logic learned by the model are then fused and decoded. The decoders under the two data flow paths are composed of different numbers of CNN modules and upsampling operation modules. The decoder fuses the data under the two paths and decodes them into a second weight file based on the function of the CNN module. Finally, the second weight file is used to perform deformable distortion processing on the affine registration result image to obtain the final registration result.
2. The preoperative multimodal image deep learning registration method according to claim 1, characterized in that: In the first stage, the combined affine registration logic is the multiplexing and combination logic of image translation, rotation, stretching, shearing, and interpolation steps.
3. The preoperative multimodal image deep learning registration method according to claim 1, characterized in that: In the second stage, the deformable registration logic includes voxel correspondence logic, voxel movement logic, and interpolation logic.
4. The preoperative multimodal image deep learning registration method according to claim 3, characterized in that: In the second stage, the voxel movement logic and interpolation logic are based on velocity field theory or optical flow field theory or displacement vector field theory.
5. The preoperative multimodal image deep learning registration method according to claim 1, characterized in that: The first weight file is the affine registration deformation field.
6. The preoperative multimodal image deep learning registration method according to claim 1, characterized in that: The second weight file is the deformable registration deformation field.
7. The preoperative multimodal image deep learning registration method according to claim 1, characterized in that: The input data, intermediate data, and result data are all medical image files in the “.nii.gz” format.
8. The preoperative multimodal image deep learning registration method according to any one of claims 1 to 7, characterized in that: The swim transformer block is replaced by the mamba block, which has long-range modeling capabilities.
9. The preoperative multimodal image deep learning registration method according to any one of claims 1 to 7, characterized in that: The KAN Linear layer with stronger fitting ability is used to replace the Linear Projection layer.