A deep learning-based multi-modal image registration method
By preprocessing X-ray and neutron images and using pre-trained convolutional neural networks, the problems of high computational cost and noise impact in multimodal image registration are solved, achieving efficient and accurate image registration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies for multimodal image registration suffer from problems such as high computational cost, long registration time, and easy getting trapped in local optimization. This is especially true in industrial nondestructive testing, which makes the inspection operation time-consuming and labor-intensive. At the same time, the mixed noise in the neutron image affects the registration accuracy.
A pre-trained convolutional neural network is used to preprocess X-ray and neutron images to remove noise. After noise removal, the deformation field and transformation parameters are extracted by an encoder and a decoder. The spatial transformation network is used for image registration to reduce the iterative process.
It improves the accuracy and efficiency of image registration, reduces the impact of noise on deformation field and transformation parameters, and enhances the accuracy and efficiency of registration results.
Smart Images

Figure CN116523981B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of image registration, and in particular to a multimodal image registration method based on deep learning. Background Technology
[0002] Different imaging devices, imaging times, and imaging angles can cause the same object to exhibit different amounts of information. Since different rays attenuate the same element to varying degrees, the final imaging results will also differ. Therefore, the information between multimodal images is complementary, and multimodal registration technology can be used to register multimodal images to improve the accuracy of subsequent image fusion.
[0003] In the field of industrial non-destructive testing (NDT), X-rays are commonly used to irradiate industrial devices. However, X-rays are not sensitive enough to some materials, resulting in their inability to detect residual materials inside devices at higher precision levels. Neutrons, on the other hand, are highly sensitive to elements with light atomic numbers and can detect core residues that X-rays fail to detect at higher precision levels. The core residue information from X-ray and neutron images can be complementary, enabling comprehensive and accurate NDT through multimodal image fusion.
[0004] However, current methods for registering X-ray and neutron images typically employ iterative maximization of similarity measures to find the optimal transformation parameters for the image pair. While this achieves image pair registration optimization, iterative optimization suffers from massive computational demands, long registration times, and a tendency to get trapped in local optima. Especially in industrial nondestructive testing, where massive amounts of data are collected, these issues lead to time-consuming and labor-intensive inspection operations, significantly increasing time costs.
[0005] To address the aforementioned issues, a convolutional neural network can be used to obtain the deformation field between the X-ray and neutron images. This deformation field can then be used to spatially transform the neutron image, enabling registration between the two images. However, neutron images contain mixed noise. If the neutron image is directly input into the convolutional neural network, the obtained deformation field will have significant errors due to the mixed noise, leading to poor accuracy in the registration results between the X-ray and neutron images. Summary of the Invention
[0006] In view of the above problems, the present invention is proposed to provide a deep learning-based multimodal image registration method that overcomes or at least partially solves the above problems, and can solve the problem of poor image registration accuracy in the prior art.
[0007] Specifically, in order to at least solve the above-mentioned technical problems, the present invention provides a deep learning-based multimodal image registration method, comprising:
[0008] Acquire two sets of initial images with different modalities, and obtain a pre-trained convolutional neural network model;
[0009] The initial image is preprocessed, and the two preprocessed initial images are respectively used as a fixed image and a floating image, wherein the preprocessing includes noise reduction processing;
[0010] The fixed image and the floating image are input into the pre-trained convolutional neural network model to obtain the deformation field and transformation parameters between the fixed image and the floating image;
[0011] Based on the deformation field and the transformation parameters, the floating image is spatially transformed to obtain the registration result.
[0012] According to one embodiment of the present invention, the pre-trained convolutional neural network model includes an encoder for downsampling the fixed image and the floating image to extract features of the fixed image and the floating image, wherein the number of downsampling operations on the floating image is a predetermined number.
[0013] According to one embodiment of the present invention, the set number of times is not greater than a set threshold, and the encoder downsampling the floating image includes: downsampling the floating image using strided convolution; or
[0014] The set number of times is greater than the set threshold, and the encoder downsampling the floating image includes:
[0015] In response to the fact that the number of downsampling operations on the floating image is not greater than the set threshold, strided convolution is used to downsample the floating image;
[0016] In response to the number of downsampling operations on the floating image exceeding the set threshold, dilated convolution is used to downsample the floating image.
[0017] According to one embodiment of the present invention, the encoder downsampling the floating image includes:
[0018] Before each downsampling operation, the image to be downsampled is filtered.
[0019] According to one embodiment of the present invention, the encoder is a multi-scale encoder composed of a pooling layer, a convolutional layer and a linear activation layer.
[0020] According to one embodiment of the present invention, the pre-trained convolutional neural network model further includes a decoder, which is used to perform multiple upsampling operations on the features extracted by the encoder, and after each upsampling operation, a skip connection is used to pass the corresponding feature map in the encoder to the decoder and fuse it with the corresponding feature map in the decoder.
[0021] According to an embodiment of the present invention, inputting the fixed image and the floating image into the pre-trained convolutional neural network model includes:
[0022] The floating image and the fixed image are superimposed on their corresponding channels; and
[0023] The tensor of the superimposed channels is input into the pre-trained convolutional neural network.
[0024] According to an embodiment of the present invention, the spatial transformation of the floating image based on the deformation field and the transformation parameters includes:
[0025] The deformation field and the floating image are input into a spatial transformation network, and the transformation parameters are applied to the floating image using an interpolation function.
[0026] According to one embodiment of the present invention, the preprocessing further includes:
[0027] Perform affine transformation processing on the two sets of initial images;
[0028] The initial image after the affine transformation is normalized.
[0029] According to an embodiment of the present invention, obtaining the pre-trained convolutional neural network model includes:
[0030] Obtain a training dataset, which contains multiple image pairs, each image pair including two sets of preprocessed training images;
[0031] A convolutional neural network model is constructed, and the image similarity and deformation field regularization values are used as loss functions. The convolutional neural network model is trained using the training dataset to obtain the pre-trained convolutional neural network model.
[0032] The technical solution provided by this invention first preprocesses two sets of initial images with different modalities, including denoising, to obtain a fixed image and a floating image. Then, a pre-trained convolutional neural network is used to obtain the deformation field and transformation parameters between the fixed and floating images. Finally, spatial transformation is performed on the floating image based on the deformation field and transformation parameters to obtain the registration result. In this technical solution, since image noise is removed during the preprocessing of the initial images, the influence of noise in the images can be reduced when the pre-trained convolutional neural network obtains the deformation field and transformation parameters between the fixed and floating images, thereby improving the accuracy of the obtained deformation field and transformation parameters, and thus improving the accuracy of the registration result.
[0033] The above and other objects, advantages and features of the present invention will become more apparent to those skilled in the art from the following detailed description of specific embodiments of the invention in conjunction with the accompanying drawings. Attached Figure Description
[0034] The following sections will describe some specific embodiments of the invention in detail by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0035] Figure 1 This is a flowchart of a deep learning-based multimodal image registration method according to an embodiment of the present invention;
[0036] Figure 2 This is a flowchart illustrating how an encoder downsamples a floating image when the set number of iterations exceeds a preset value, according to an embodiment of the present invention.
[0037] Figure 3 This is a flowchart illustrating the input of fixed and floating images into a pre-trained convolutional neural network model according to an embodiment of the present invention;
[0038] Figure 4 This is a flowchart of preprocessing an initial image according to an embodiment of the present invention;
[0039] Figure 5 This is a flowchart of obtaining a pre-trained convolutional neural network model according to an embodiment of the present invention. Detailed Implementation
[0040] In the description of this embodiment, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0041] This application provides a deep learning-based multimodal image registration method for matching and overlaying two or more images. The two or more images can be images of the same object acquired by different image acquisition devices, or images of the same object acquired at different times, from different angles, or under different environments. This registration method can be applied to fields such as remote sensing data analysis, computer vision, or image processing.
[0042] The deep learning-based multimodal image registration method provided in this application has the following process: Figure 1 As shown, this registration method involves obtaining the corresponding deformation field and transformation parameters through a pre-trained convolutional neural network. This avoids the problem of long registration times caused by iterations in existing image registration methods, thereby improving the efficiency of image registration. Furthermore, before inputting the initial image for registration into the pre-trained convolutional neural network, preprocessing such as denoising is performed on the initial image to eliminate noise, thereby improving the accuracy of the obtained deformation field and transformation parameters, and ultimately improving the accuracy of image registration. The following section combines... Figure 1 The flowchart shown provides a detailed introduction to the multimodal image registration method of this application.
[0043] like Figure 1 As shown, the deep learning-based multimodal image registration method of this application includes the following steps:
[0044] Step S1: Obtain two sets of initial images with different modalities and obtain a pre-trained convolutional neural network model.
[0045] In this embodiment, the two sets of initial images of different modes can be X-ray images and neutron images, or CT (Computed Tomography) images and MRI (nuclear magnetic resonance) images, or infrared images and RGB images, respectively.
[0046] In step S1, images of the target object can be acquired using different image acquisition devices to obtain initial images in different modalities. For example, an X-ray camera and a neutron camera can be used to acquire X-ray and neutron images of the target object, respectively, and these images can be used as the initial images. In other embodiments, the initial images can be stored in storage devices such as USB flash drives or computer hard drives. When performing step S1, the initial images are read from the storage device based on their storage address and storage identifier.
[0047] The aforementioned pre-trained convolutional neural network (CNN) model refers to a CNN model trained on data, where the weight parameters are fixed. When obtaining a pre-trained CNN model, a CNN can be constructed first to determine its structure. Then, the constructed CNN can be trained using a dataset and a relevant loss function to solidify the weight parameters, resulting in the pre-trained CNN model. Alternatively, the pre-trained CNN model can be stored in a database. When executing step S1, the pre-trained CNN model can be retrieved from the database via data reading.
[0048] Step S2: Preprocess the two sets of initial images of different modalities, and use the preprocessed initial images as fixed images and floating images respectively.
[0049] In this embodiment, preprocessing includes denoising, which refers to performing denoising on both sets of initial images separately. When denoising the initial images, existing image filtering methods can be used. For example, an image filter can be used to first degrade the initial images using a degradation function to obtain a corresponding degraded initial image. Then, noise is removed from the degraded initial image. For instance, noise can be identified and denoised based on its cause, or noise with fixed attributes during the degradation process can be identified and denoised, thus obtaining a denoised degraded initial image. Finally, the denoised degraded image is restored, which is the reverse of the degradation process, to obtain the denoised initial image, i.e., the preprocessed initial image.
[0050] Step S3: Input the above fixed image and floating image into the pre-trained convolutional neural network model to obtain the deformation field and transformation parameters between the above fixed image and floating image.
[0051] The pre-trained convolutional neural network model mentioned above includes an encoder and a decoder. The encoder is used to extract features from the fixed image and the floating image, and the decoder is used to upsample the features extracted by the encoder to obtain the deformation field and transformation parameters between the fixed image and the floating image.
[0052] In this embodiment, the setup only requires inputting the fixed image and the floating image into the pre-trained convolutional neural network to obtain the corresponding deformation field and transformation parameters. This process does not require repeated iterations, thus reducing time costs compared to existing image registration methods, such as B-spline-based image registration methods.
[0053] Step S4: Perform spatial transformation on the floating image based on the deformation field and transformation parameters between the fixed image and the floating image to obtain the registration result of the fixed image and the floating image.
[0054] In summary, the technical solution of this application first preprocesses the initial images of different modalities to eliminate noise, and then uses the preprocessed initial images as fixed and floating images respectively. The fixed and floating images are then input into a pre-trained convolutional neural network model to obtain the corresponding deformation field and transformation parameters. Finally, spatial transformation is performed on the floating image based on the obtained deformation field and transformation parameters to obtain the registration result of the fixed and floating images. Because the technical solution of this application performs denoising processing on the initial images before inputting them into the pre-trained convolutional neural network model to eliminate image noise, the influence of image noise in the initial images can be reduced when obtaining the corresponding deformation field and transformation parameters through the pre-trained convolutional neural network model, thereby improving the accuracy of the obtained deformation field and transformation parameters, and ultimately improving the accuracy of the registration result.
[0055] The above text provides a detailed description of the implementation process and technical effects of the registration method of this application. The following text, in conjunction with specific implementation methods, further describes the structure and working method of the above-mentioned pre-trained convolutional neural network model. It should be understood that the implementation methods shown below are exemplary and not restrictive.
[0056] In one embodiment, the encoder in the pre-trained convolutional neural network model is used to extract features of the fixed and floating images by downsampling the fixed and floating images, and the number of downsampling operations on the floating image is a set number.
[0057] When the floating image is of low quality, excessive downsampling can lead to the loss of key features. This embodiment limits the number of downsampling operations on the floating image to a predetermined number. This reduces feature loss during encoder downsampling, improving the accuracy of the deformation field and transformation parameters between the fixed and floating images, thereby enhancing the accuracy of the registration results.
[0058] In one embodiment, the number of iterations is not greater than a set threshold. When the encoder in the pre-trained convolutional neural network model downsamples the floating image, the downsampling method used is strided convolution, that is, strided convolution is performed on the floating image a set number of times to extract the features of the floating image. For example, when the fixed image is an X-ray image and the floating image is a neutron image, the set threshold can be set to 3. When the number of times the encoder needs to downsample the floating image is no greater than 3, strided convolution is used to downsample the floating image.
[0059] In this embodiment, strided convolution is used to downsample the floating image. Compared with the pooling operation and other downsampling methods in the prior art, this method can reduce feature loss during the downsampling of the floating image and improve the accuracy of the obtained deformation field and transformation parameters.
[0060] In another embodiment, if the number of preset attempts exceeds a preset threshold, the encoder in the preset convolutional neural network model uses the following downsampling method to downsample the floating image: Figure 2 As shown, it includes the following steps:
[0061] Step S101: When the number of downsampling operations on the floating image is not greater than the above-mentioned set threshold, strided convolution is used to downsample the floating image;
[0062] Step S102: When the number of downsampling operations on the floating image exceeds the above-mentioned set threshold, dilated convolution is used to downsample the floating image.
[0063] By using the configuration method in this embodiment, when the encoder downsamples the floating image more than a preset value, dilated convolution can be used for downsampling to reduce the loss of features in the floating image during the downsampling process, thereby improving the accuracy of the obtained deformation field and transformation parameters.
[0064] In one embodiment, the encoder in the pre-trained convolutional neural network model filters the image to be downsampled before downsampling the floating image, so as to avoid losing the main features of the floating image during the downsampling process.
[0065] In this embodiment, an image filter can be used to filter the image. For example, when downsampling the floating image for the first time, an image filter can be used to filter the floating image, and then the filtered floating image can be downsampled. When downsampling the floating image for a subsequent time, an image filter can be used to filter the feature map of the floating image obtained from the previous downsampling, and then the filtered feature map can be downsampled.
[0066] By using the configuration method in this embodiment, the image to be downsampled can be filtered before each downsampling process, thereby avoiding the loss of the main features of the floating image and improving the accuracy of the obtained deformation field and transformation parameters.
[0067] In one embodiment, the encoder in the aforementioned pre-trained convolutional neural network model is a multi-scale encoder composed of pooling layers, convolutional layers, and linear activation layers. The pooling layers can be used to downsample a fixed image through pooling operations, the convolutional layers can be used to downsample a floating image and / or a fixed image through convolutional operations (such as dilated convolution, strided convolution, etc.), and the linear activation layers are used to activate the resulting feature map after each convolutional or pooling downsampling operation. Furthermore, the multi-scale encoder can extract features from both fixed and floating images at various scales, thereby improving the accuracy of the obtained deformation field and transformation parameters.
[0068] The above text provides a detailed introduction to the implementation and working principle of the encoder in a pre-trained convolutional neural network model, combined with specific application scenarios. The following text will provide a detailed explanation of the decoder in a pre-trained convolutional neural network model, combined with specific embodiments. It should be understood that the decoder shown below is exemplary and not restrictive.
[0069] In one embodiment, the decoder in the pre-trained convolutional neural network model is used to upsample the features extracted by the encoder multiple times to obtain the deformation field and transformation parameters between the fixed image and the floating image. After each upsampling, a skip connection is used to pass the corresponding feature map in the encoder to the decoder and fuse it with the corresponding feature map in the decoder.
[0070] For example, assuming the encoder performs M downsampling operations on both the fixed and floating images, the decoder needs to perform M upsampling operations on the features extracted by the encoder to obtain a deformation field with the same size as the floating image. When the decoder performs the m-th upsampling operation on the features extracted by the encoder, let the resulting feature map be the first feature map. During skip connections, the corresponding feature maps in the encoder include the fixed image feature map obtained from the Mm-th downsampling operation on the fixed image and the floating image feature map obtained from the Mm-th downsampling operation on the floating image. Let the fixed image feature map be the second feature map and the floating image feature map be the third feature map.
[0071] After passing the second and third feature maps to the decoder via skip connections, the second and third feature maps are fused with the first feature map. After fusion, the feature map resulting from the fusion of the first and second feature maps is subjected to two strided convolutions, and after each strided convolution, it is activated by a linear activation function to decouple the features. The feature map resulting from the fusion of the first and third feature maps is subjected to two dilated convolutions, and after each dilated convolution, it is activated by a linear activation function to decouple the features.
[0072] With the configuration method of this embodiment, after each upsampling is completed, the information in the upsampled feature map can be supplemented by skip connections, thereby improving the accuracy of the obtained deformation field and transformation parameters, and thus improving the accuracy of the image registration result.
[0073] In one embodiment, the process of inputting the fixed image and the floating image into the pre-trained convolutional neural network model in step S3 above is as follows: Figure 3 As shown, it includes the following steps:
[0074] Step S201: Through a concatenation operation, the fixed image and the floating image are respectively superimposed on their respective channels. In this embodiment, the pre-trained convolutional neural network model has two channels for the input image, and the fixed image and the floating image are superimposed on their respective channels.
[0075] Step S202: Input the tensors of the channels of the superimposed fixed image and the channels of the superimposed floating image into the pre-trained convolutional neural network model.
[0076] In this embodiment, after the fixed image and the floating image are superimposed on their respective channels, the Tensor() function can be used to convert each channel into data that can be driven by the convolutional neural network model, i.e., channel tensors. These tensors are then input into the pre-trained convolutional neural network model. Through this configuration, the fixed image and the floating image can be quickly input into the pre-trained convolutional neural network.
[0077] In one embodiment, the method for spatially transforming the floating image based on the obtained deformation field and transformation parameters in step S4 above includes:
[0078] The deformation field between the fixed image and the floating image is input into the spatial transformation network. The spatial transformation network applies the transformation parameters to the floating image through the interpolation function to perform spatial transformation on the floating image to obtain the registration result.
[0079] The configuration method of this embodiment allows for the spatial transformation network to quickly and accurately apply transformation parameters to the floating image through an interpolation function, thereby performing spatial transformation on the floating image and obtaining the registration result. Therefore, the configuration method of this embodiment can improve the efficiency of image registration.
[0080] In one embodiment, in step S2 above, the preprocessing of the initial image includes not only noise reduction but also, as well as... Figure 4 The steps shown are as follows:
[0081] Step S301: Perform affine transformation on the two sets of initial images with different modalities to align the two sets of initial images.
[0082] Step S302: Normalize the initial image after affine transformation.
[0083] The above-mentioned methods for radiometric transformation and normalization can both employ existing image radiometric transformation and image normalization methods.
[0084] By using the configuration method of this embodiment, the two sets of initial images can be aligned and their pixels normalized during the preprocessing of the initial image, thereby reducing the difference between the two sets of initial images, that is, reducing the difference between the fixed image and the floating image. This allows the pre-trained convolutional neural network to quickly obtain the deformation field and transformation parameters between the fixed image and the amplitude image, thereby improving the efficiency of image registration.
[0085] In one embodiment, the multimodal image registration method of this application further includes a method for obtaining a preset convolutional neural network, the process of which is as follows: Figure 5 As shown, it includes the following steps:
[0086] Step S401: Obtain the training dataset, which contains multiple image pairs, each image pair including two sets of preprocessed training images.
[0087] Step S402: Construct a convolutional neural network model, and use the image similarity and deformation field regularization values as loss functions. Train the constructed convolutional neural network model using the above training dataset to obtain a pre-trained convolutional neural network model.
[0088] In this embodiment, the method for training the constructed convolutional neural network model using the above-mentioned training dataset includes:
[0089] The convolutional neural network model is constructed by inputting two sets of images from each image pair in the training dataset as fixed and floating images, respectively, to obtain the corresponding deformation field and transformation parameters. Then, the two sets of training images in the image pair are registered according to the obtained deformation field and transformation parameters, and the registered training images are calculated according to the above loss function to determine whether the convolutional neural network model has converged, that is, whether the calculation result is stable within the preset low value range. If it has not converged, the weight parameters in the convolutional neural network model are updated and backpropagation is performed, and the convolutional neural network model continues to be trained. If it has converged, the training of the convolutional neural network model is completed, and the pre-trained convolutional neural network model is obtained.
[0090] In this embodiment, the two sets of training images in each image pair of the training dataset can be designated as image X and image R, respectively. During training, the mutual information of the images is used as the image similarity. The resulting loss function is:
[0091] L loss =L sim (X, R) + L smooth (flow)
[0092] Where L sim (X, R) represents the similarity between image X and image R, and L... smooth (flow) represents the gradient smoothing L2 regularization loss of the deformation field, and
[0093]
[0094]
[0095] Where p represents the deformation displacement along each coordinate axis, Ω represents the entire set, u(p) represents the spatial displacement of an image pixel, H(X) represents the information entropy of image X, H(R) represents the information entropy of image R, and I(X,R) is the mutual information between image X and image R, which can be calculated using the following formula.
[0096] I(X,R)=H(X)+H(R)-H(X,R)
[0097] In the above formula, H(X, R) represents the joint information entropy of image X and image R, and H(X), H(R), and H(X, R) can be calculated by the following formulas:
[0098]
[0099]
[0100]
[0101] In the above calculation formulas, N represents the gray levels of the image, and a j Let b be the total number of pixels with gray value j in image X. j Let ab be the total number of pixels with gray value j in image R. j This represents the total number of pixels with grayscale value j at corresponding positions in images X and R.
[0102] The configuration method of this embodiment allows the pre-processed training images to be used to train the convolutional neural network model, resulting in a pre-trained convolutional neural network model. Since the pre-processed training images are free from image noise interference, the obtained pre-trained convolutional neural network model has high accuracy, and the accuracy of the convolution field and transformation parameters obtained using this pre-trained convolutional neural network model is also higher.
[0103] Therefore, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the present invention can be directly determined or derived from the disclosure of the present invention without departing from the spirit and scope of the invention. Thus, the scope of the present invention should be understood and construed as covering all such other variations or modifications.
Claims
1. A multimodal image registration method based on deep learning, characterized in that, include: Acquire two sets of initial images with different modalities, and obtain a pre-trained convolutional neural network model; The initial image is preprocessed, and the two preprocessed initial images are respectively used as a fixed image and a floating image, wherein the preprocessing includes noise reduction processing; The fixed image and the floating image are input into the pre-trained convolutional neural network model to obtain the deformation field and transformation parameters between the fixed image and the floating image; Based on the deformation field and the transformation parameters, the floating image is spatially transformed to obtain the registration result; The pre-trained convolutional neural network model includes an encoder, which is used to downsample the fixed image and the floating image to extract features of the fixed image and the floating image, wherein the number of downsampling operations on the floating image is a set number. The set number of iterations is not greater than a set threshold, and the encoder's downsampling of the floating image includes: downsampling the floating image using strided convolution; or The set number of times is greater than the set threshold, and the encoder downsampling the floating image includes: in response to the number of times the floating image is downsampled not being greater than the set threshold, using strided convolution to downsample the floating image; and in response to the number of times the floating image is downsampled being greater than the set threshold, using dilated convolution to downsample the floating image.
2. The multimodal image registration method according to claim 1, characterized in that, The encoder downsamples the floating image by: Before each downsampling operation, the image to be downsampled is filtered.
3. The multimodal image registration method according to claim 1, characterized in that, The encoder is a multi-scale encoder composed of pooling layers, convolutional layers, and linear activation layers.
4. The multimodal image registration method according to claim 1, characterized in that, The pre-trained convolutional neural network model also includes a decoder, which is used to perform multiple upsampling operations on the features extracted by the encoder, and after each upsampling operation, a skip connection is used to pass the corresponding feature map in the encoder to the decoder and fuse it with the corresponding feature map in the decoder.
5. The multimodal image registration method according to claim 1, characterized in that, The step of inputting the fixed image and the floating image into the pre-trained convolutional neural network model includes: The floating image and the fixed image are superimposed on their corresponding channels; and The tensor of the superimposed channels is input into the pre-trained convolutional neural network.
6. The multimodal image registration method according to claim 1, characterized in that, The method of performing spatial transformation on the floating image based on the deformation field and the transformation parameters includes: The deformation field and the floating image are input into a spatial transformation network, and the transformation parameters are applied to the floating image using an interpolation function.
7. The multimodal image registration method according to claim 1, characterized in that, The preprocessing also includes: Perform affine transformation processing on the two sets of initial images; The initial image after the affine transformation is normalized.
8. The multimodal image registration method according to claim 1, characterized in that, The process of obtaining a pre-trained convolutional neural network model includes: Obtain a training dataset, which contains multiple image pairs, each image pair including two sets of preprocessed training images; A convolutional neural network model is constructed, and the image similarity and deformation field regularization values are used as loss functions. The convolutional neural network model is trained using the training dataset to obtain the pre-trained convolutional neural network model.
Citation Information
Patent Citations
Cross-modal medical image registration method and device
CN111862174A
Image registration method and image processing method for medical images, and medium
CN112686932A