Image Registration Method, Apparatus, Storage Medium and Computer Device
The unsupervised convolutional neural network model processes image data of multiple image layers, and solves the problems of inaccurate image coarse registration and low efficiency, achieving more efficient image registration.
Patent Information
- Application Number
- CN202111353713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-16
AI Technical Summary
The problems of inaccurate image coarse registration positioning and low registration efficiency.
By acquiring the registration reference image and the registration moving image with multiple image layers and inputting it into the trained unsupervised convolutional neural network model, the confidence scores of each image layer are obtained, and the coarse positioning results of the area to be registered are determined.
It effectively reduces the training difficulty and training time of the rough registration model, avoids registration errors caused by excessive search range, and improves image registration efficiency and results.
Smart Images

Figure CN114241017B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular, to an image registration method, apparatus, storage medium, and computer device. Background Art
[0002] Image registration is a necessary processing step required for clinical applications such as medical image comparison, data fusion, and change analysis. During the image registration process, it is often necessary to perform registration processing on multiple groups of images with different phases during the scanning process. As a common preprocessing link in medical image analysis, image registration is of great significance for applications such as medical diagnosis and surgical planning.
[0003] Generally speaking, the image registration process is mainly implemented in two stages. The first stage is called rough registration, that is, roughly aligning the parts, which is used to solve large offset corrections, determine the overlapping area, and provide initial parameter estimates for the second-stage processing. The second stage is called fine registration, which optimizes the registration parameters to make the registration effect reach the best, and is used to determine rotation, stretching, and non-rigid deformation. The quality of the fine registration processing depends to a large extent on the accuracy of the rough registration. If the rough registration fails to provide a good initial estimate, the fine registration will fall into a local optimal solution and cannot generate the expected result, resulting in registration failure.
[0004] Currently, the more commonly used rough registration schemes mainly include methods such as sliding window search, specific region extraction based on segmentation methods, and model training with manually designed feature points. However, directly using the sliding window search technology to search for matching positions on three-dimensional volume data has a large amount of calculation and is very time-consuming. The search range is too large, and errors may occasionally occur, so manual assistance is sometimes required. Secondly, for the rough registration method based on the segmentation method, it is necessary to train the model after voxel-by-voxel annotation of data for a specific organ or part. The annotation is difficult and time-consuming. Moreover, the segmentation method is a voxel-level prediction task, and the edge prediction difficulty and calculation amount are both large. Finally, compared with deep learning abstract features, using traditional features or manually designed features has certain limitations, resulting in inaccurate image feature expression, which will in turn affect the implementation of subsequent similarity measurement and other links. Summary of the Invention
[0005] In view of this, the present application provides an image registration method, apparatus, storage medium, and computer device, mainly aiming to solve the technical problems of inaccurate rough registration positioning and low registration efficiency of images.
[0006] According to the first aspect of the present invention, an image registration method is provided, and the method includes:
[0007] Obtain a registration reference image and a registration moving image, where the registration reference image includes a plurality of reference image layers, and the registration moving image includes a plurality of moving image layers;
[0008] The registered reference image and the registered moving image are respectively input into the trained unsupervised convolutional neural network model to obtain the confidence scores of each reference image layer of the registered reference image and the confidence scores of each moving image layer of the registered moving image;
[0009] The confidence scores of each reference image layer and the confidence scores of each moving image layer are respectively compared with the confidence score range of the preset area to be registered, to obtain a plurality of roughly registered reference image layers of the registered reference image and a plurality of roughly registered moving image layers of the registered moving image.
[0010] According to the second aspect of the present invention, there is provided an image registration device, which includes:
[0011] An image acquisition module, configured to acquire a registered reference image and a registered moving image, wherein the registered reference image includes a plurality of reference image layers, and the registered moving image includes a plurality of moving image layers;
[0012] An image processing module, configured to respectively input the registered reference image and the registered moving image into the trained unsupervised convolutional neural network model to obtain the confidence scores of each reference image layer of the registered reference image and the confidence scores of each moving image layer of the registered moving image;
[0013] An image comparison module, configured to respectively compare the confidence scores of each reference image layer and the confidence scores of each moving image layer with the confidence score range of the preset area to be registered, to obtain a plurality of roughly registered reference image layers of the registered reference image and a plurality of roughly registered moving image layers of the registered moving image.
[0014] According to the third aspect of the present invention, there is provided a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above image registration method is implemented.
[0015] According to the fourth aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above image registration method is implemented.
[0016] An image registration method, device, storage medium, and computer device provided by the present invention first obtain a registration reference image and a registration moving image with multiple image layers, then obtain the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image through a trained unsupervised convolutional neural network model, and finally determine the rough positioning result of the area to be registered through the confidence scores of each image layer. In the above method, only an unsupervised learning method is adopted, and the convolutional neural network model can be trained without data annotation. By using the trained convolutional neural network model, the confidence scores of each image layer can be obtained, and the rough positioning result of the area to be registered can be determined, thereby effectively reducing the training difficulty and training time of the rough registration model, and avoiding the problem of registration errors caused by too large a search range in the image rough registration stage. In addition, by using an unsupervised convolutional neural network model to obtain partial image layers of the area to be registered and then performing local registration, the image registration efficiency can be greatly improved, and the registration result can be improved.
[0017] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Brief Description of the Drawings
[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0019] Figure 1 The flowchart of an image registration method provided by an embodiment of the present invention is shown;
[0020] Figure 2 The flowchart of an image registration method in a medical scenario provided by an embodiment of the present invention is shown;
[0021] Figure 3 The flowchart of another image registration method provided by an embodiment of the present invention is shown;
[0022] Figure 4 The schematic diagram of the confidence score range of the area to be registered in a medical scenario provided by an embodiment of the present invention is shown;
[0023] Figure 5 The schematic diagram of the relationship between the number of image layers and the confidence scores of image layers of a three-dimensional image provided by an embodiment of the present invention is shown;
[0024] Figure 6Shows a schematic diagram of the effect of converting an image layer into an image pyramid provided by an embodiment of the present invention;
[0025] Figure 7 Shows a schematic diagram of the scenario of sliding and extracting image features of an image pyramid by a fully convolutional network provided by an embodiment of the present invention;
[0026] Figure 8 Shows a schematic flowchart of feature extraction, feature fusion, and feature classification for an image layer provided by an embodiment of the present invention;
[0027] Figure 9 Shows a schematic diagram of the effect of candidate regions of an image layer provided by an embodiment of the present invention;
[0028] Figure 10 Shows a schematic diagram of the center positioning result of a three-dimensional image provided by an embodiment of the present invention;
[0029] Figure 11 Shows a comparison diagram of the effect of an image registration method provided by an embodiment of the present invention;
[0030] Figure 12 Shows a schematic structural diagram of an image registration device provided by an embodiment of the present invention;
[0031] Figure 13 Shows a schematic structural diagram of another image registration device provided by an embodiment of the present invention. Detailed implementation manners
[0032] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0033] Before specifically describing each embodiment, a brief introduction to the image registration process is first made. Among them, image registration refers to performing translation, rotation, stretching, and other non-linear spatial transformation operations on an image to achieve position matching and alignment between different images. For a set of registered images (mainly referring to the registration reference image and the registration moving image), the information at the same image position (m, n) (m represents the row label, n represents the column label) corresponds to the same image region. Therefore, the registered images can be used for applications such as lesion change analysis. On this basis, multi-phase image registration refers to registering images of the same individual collected at different times. However, the body positions of the images in different phases will change, which will undoubtedly increase the processing difficulty of image registration.
[0034] Generally speaking, image registration goes through a rough registration stage and a fine registration stage successively. Among them, the quality of the fine registration processing depends to a large extent on the accuracy of the rough registration. If the rough registration fails to provide a good initial estimate, the fine registration will fall into a local optimal solution and cannot generate the expected result, ultimately resulting in registration failure. Therefore, how to improve the initial positioning accuracy and positioning efficiency in the rough registration stage has become a very important issue.
[0035] Based on this, in one embodiment, as Figure 1 shown, a method for image registration is provided. Taking the application of this method to a computer device as an example, it includes the following steps:
[0036] 101. Obtain a registration reference image and a registration moving image, where the registration reference image includes multiple reference image layers, and the registration moving image includes multiple moving image layers.
[0037] Specifically, for image registration, one image needs to be selected as the reference image and kept stationary, and then another image is selected as the moving image. Then, translation, rotation, stretching, and non-rigid spatial transformation are performed on the moving image so that the transformed moving image can match the reference image. In this embodiment, the image that remains stationary is called the registration reference image, and the image that needs to be moved is called the registration moving image. Both the registration reference image and the registration moving image are three-dimensional images and both include multiple two-dimensional image layers. That is, for the registration reference image, it includes multiple two-dimensional reference image layers, and for the registration moving image, it includes multiple two-dimensional moving image layers. Specifically in the medical image registration scenario, the registration reference image and the registration moving image can be three-dimensional CT images. Through image registration, the information at the same position on two three-dimensional CT images can be corresponding to the same human body part or organ structure.
[0038] 102. Input the registration reference image and the registration moving image into the trained unsupervised convolutional neural network model respectively to obtain the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image.
[0039] Specifically, for a three-dimensional image, the multiple two-dimensional images it contains have certain internal structural relationships and / or upper and lower layer sequence information. For example, a three-dimensional image can be segmented into several cross-sectional image layers arranged one by one. Among these image layers, the order of each image layer relative to the previous arranged image layer is increasing, and the distance between two adjacent image layers is equal. By using this internal structural information and sequence information of the image, a convolutional neural network model can be trained. Furthermore, through the trained convolutional neural network model, the positions of each image layer in the three-dimensional image can be predicted, and this position is the output of the model. In this embodiment, the convolutional neural network model can be a convolutional neural network regression model or a convolutional neural network generation model. The output of the convolutional neural network model can be the confidence scores of each image layer of the three-dimensional image, where the confidence score can be used to represent the position of each image layer in the three-dimensional image. Specifically in the medical image registration scenario, the confidence score can be used to represent the body part or organ structure corresponding to each image layer in the CT medical image. Further, by processing a large number of CT medical images, the confidence score range corresponding to each body part or organ structure can be determined.
[0040] In this embodiment, through the internal structural information and / or sequence information of each image layer of the three-dimensional image, a deep convolutional neural network model that can predict the confidence scores of each image layer can be trained. Since this model is trained through an unsupervised learning algorithm, it is also called an unsupervised convolutional neural network model. Further, by inputting the registration reference image and the registration moving image into the trained unsupervised convolutional neural network model respectively, the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image can be obtained.
[0041] In this embodiment, compared with traditional manually designed features, the unsupervised convolutional neural network model can automatically learn more useful features by constructing a machine learning model with many hidden layers and a large amount of training data, and finally can improve the accuracy of image classification or prediction.
[0042] 103. Compare the confidence scores of each reference image layer and the confidence scores of each moving image layer with the confidence score range of the preset area to be registered respectively, and obtain multiple roughly registered reference image layers of the registration reference image and multiple roughly registered moving image layers of the registration moving image.
[0043] Among them, the area to be registered refers to the area where the registration reference image and the registration moving image need to be registered. By inputting the three-dimensional image into the convolutional neural network model, the confidence scores of each image layer of the three-dimensional image can be obtained. By synthesizing the output data of multiple groups of three-dimensional image data and statistically analyzing the confidence scores of the area to be registered, the confidence score range of the area to be registered can be obtained.
[0044] Specifically, after obtaining the confidence scores of each reference image layer and each moving image layer, the confidence score of each reference image layer and the confidence score of each moving image layer can be compared with the preset confidence score range of the area to be registered respectively, so as to screen out the image layers whose confidence scores are within the confidence score range of the area to be registered, and obtain multiple roughly registered reference image layers of the registration reference image and multiple roughly registered moving image layers of the registration moving image. Among them, the obtained multiple roughly registered reference image layers and multiple roughly registered moving image layers are the rough positioning results of the image rough configuration, and this rough positioning result can effectively narrow the search range in the image fine registration stage.
[0045] In this embodiment, by using the confidence scores of the image layers of the three-dimensional image for rough image registration, the roughly registered image layers within the same registration range of different images can be obtained, so as to effectively narrow the search range of the matching position of the registration moving image relative to the registration reference image, and effectively avoid the obvious error of matching different registration areas together.
[0046] The image registration method provided in this embodiment first obtains a registration reference image and a registration moving image with multiple image layers, then obtains the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image through a trained unsupervised convolutional neural network model, and finally determines the rough positioning result of the area to be registered through the confidence scores of each image layer. In the above method, only the unsupervised learning method is adopted, and the convolutional neural network model can be trained without data annotation. By using the trained convolutional neural network model, the confidence scores of each image layer can be obtained, and the rough positioning result of the area to be registered can be determined, thus effectively reducing the training difficulty and training time of the rough registration model, and avoiding the problem of registration errors caused by too large a search range in the rough image registration stage. In addition, by using the unsupervised convolutional neural network model to obtain some image layers of the area to be registered and then performing local registration, the image registration efficiency can be greatly improved and the registration result can be enhanced.
[0047] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the implementation process of this embodiment, an image registration method is provided, as Figure 2As shown in the figure, taking the liver region registration in the medical scenario as an example, this registration method mainly covers three stages of image registration, including: in the first stage (corresponding to steps 201 to 203), a deep convolutional neural network is used to obtain the confidence scores of each image layer to determine the rough positioning result of the region to be registered; in the second stage (corresponding to steps 204 to 206), a deep convolutional neural network feature fusion network is used to achieve the fine positioning of the region to be registered; in the third stage (corresponding to steps 207 and 208), the multi-phase image local registration is performed using the range of the region to be registered that has been finely positioned, thereby shortening the registration time and improving the registration effect.
[0048] Specifically, as Figure 3 shown, this method specifically includes the following steps:
[0049] 201. Obtain a registration reference image and a registration moving image, where the registration reference image includes multiple reference image layers, and the registration moving image includes multiple moving image layers.
[0050] Specifically, when performing image registration, it is first necessary to obtain a registration reference image and a registration moving image. In this embodiment, both the registration reference image and the registration moving image are three-dimensional images. Among them, the registration reference image includes multiple reference image layers, and the registration moving image includes multiple moving image layers. Specifically in the medical image registration scenario, the registration reference image and the registration moving image can be three-dimensional CT images. Through image registration, the information at the same position on the two three-dimensional CT images can be corresponding to the same human body part or organ structure.
[0051] 202. Input the registration reference image and the registration moving image into the trained unsupervised convolutional neural network model respectively to obtain the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image.
[0052] Specifically, through the internal structure information and / or sequential information of each image layer of the three-dimensional image, a deep convolutional neural network model that can predict the confidence scores of each image layer can be trained. Further, by inputting the registration reference image and the registration moving image into the trained unsupervised convolutional neural network model respectively, the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image can be obtained.
[0053] In one embodiment, the above-mentioned unsupervised convolutional neural network model can be a convolutional neural network regression model or a convolutional neural network generation model, and this model can be trained through the following method: First, obtain multiple groups of image samples, where each group of image samples contains multiple image layers. Then, according to each image layer of the multiple groups of image samples, use a preset sequential error loss function and distance error loss function to iteratively train the parameters of the initialized deep convolutional neural network model. When the sum of the sequential error loss function and the distance error loss function reaches a set loss value, the trained unsupervised convolutional neural network model can be obtained. The training method of the above-mentioned deep convolutional neural network model makes full use of the internal structure information and sequential information of each image layer of the three-dimensional image, adopts an unsupervised learning method, and does not require data annotation, effectively reducing the training difficulty and training time of the model, and thus improving the efficiency of image rough registration.
[0054] Next, the training process of the deep convolutional neural network model in the above embodiment will be described in combination with specific examples. Here, taking the training of a deep convolutional neural regression model as an example, the network structure of the deep convolutional neural network regression model can be based on the VGGNet-D network. The convolutional layers 1 to 5 of the network structure can adopt the pre-trained VGGNet-D model parameters. The activation layers 1 to 6 of the network structure can be used to learn different abstract features of each organ of the body. Then, a global average pooling layer can be used to map each activation map into a value to obtain a 512-dimensional feature vector. Finally, a fully connected layer can be used to obtain the data layer confidence score of each feature vector.
[0055] For the sequential error loss function of the regression model, it can be expressed by the following formula:
[0056]
[0057] For the distance error loss function of the regression model, it can be expressed by the following formula:
[0058]
[0059] Δ i,j = S(i,j) - S(i,j - 1)
[0060] In the above two loss functions, g is the number of image layers processed in a batch during the training process, m is the number of image layers, S(i,j) represents the confidence score of the j-th layer of image i. As the number of image layers increases, the confidence score will be higher and higher, and the spatial distance between different layers of the image is approximately proportional.
[0061] In the training stage of the regression model, 219 groups of data of the collected whole-body CT images can be utilized. Each group of data contains different phases, resulting in a total of 1324 groups of data as training samples, including 269,248 axial image slices. The VGG Net-D network model can be fine-tuned using this training sample. In the testing stage of the regression model, a CT image with 1601 image slices can be used as the test data. By synthesizing the confidence score prediction results of multiple layers of data, the confidence score values of different regions of the human body in the regression can be presented. As Figure 4 can be seen from the confidence score range of the to-be-registered region shown, the confidence score range of the head region of the CT image is from -14.5 to -12.0, the shoulder range is from -13.5 to -11.5, the heart range is from -9.0 to -7.0, the liver range is from -6.0 to -5.0, the kidney range is from -2.0 to -1.0, the hip range is from 3.0 to 4.0, and the leg range is from 6.0 to 8.0. Further, from Figure 5 the relationship between the number of image slices of the shown CT image and the confidence score of the image slice, it can be seen that the confidence score increases with the increase in the number of image slices and shows an approximately linear growth. This prediction result also conforms to the characteristic that medical images have internal structure or upper and lower layer sequence information.
[0062] 203. Compare the confidence scores of each reference image slice and each moving image slice with the preset confidence score range of the to-be-registered region respectively to obtain multiple roughly registered reference image slices of the registered reference image and multiple roughly registered moving image slices of the registered moving image.
[0063] Specifically, after obtaining the confidence scores of each reference image slice and each moving image slice, the confidence score of each reference image slice and the confidence score of each moving image slice can be compared with the preset confidence score range of the to-be-registered region in sequence, so as to screen out all the image slices whose confidence scores are within the confidence score range of the to-be-registered region, and obtain multiple roughly registered reference image slices of the registered reference image and multiple roughly registered moving image slices of the registered moving image.
[0064] In one embodiment, the following method can be used to determine whether each image slice is a roughly registered image slice: First, determine whether the confidence scores of each reference image slice and each moving image slice are within the preset confidence score range of the to-be-registered region respectively. If the confidence score of any reference image slice is within the preset confidence score range of the to-be-registered region, then determine that this reference image slice is a roughly registered reference image slice. If the confidence score of any moving image slice is within the preset confidence score range of the to-be-registered region, then determine that this moving image slice is a roughly registered moving image slice.
[0065] In one embodiment, the confidence score range of the area to be registered can be obtained by the following method: First, obtain multiple groups of image samples, where each group of image samples includes multiple image layers. Then, input the multiple groups of image samples into the trained unsupervised convolutional neural network model respectively to obtain the confidence scores of each image layer of the multiple groups of image samples. Finally, determine the confidence score range of the area to be registered according to the confidence scores of the image layers where the areas to be registered of the multiple groups of image samples are located.
[0066] In this embodiment, rough image registration is performed through the confidence scores of the image layers of the three-dimensional image, and rough registration image layers within the same registration range of different images can be obtained, so that the search range of the matching position of the registration moving image relative to the registration reference image can be effectively reduced, and obvious errors of matching different registration areas together can be effectively avoided.
[0067] 204. Input the multiple rough registration reference image layers and the multiple rough registration moving image layers into the trained first feature extraction model and second feature extraction model respectively to obtain the first image features and second image features of each rough registration reference image layer, and the first image features and second image features of each rough registration moving image layer.
[0068] Specifically, after obtaining the multiple rough registration reference image layers and the multiple rough registration moving image layers, each rough registration reference image layer and each rough registration moving image layer can be input into the trained first feature extraction model and second feature extraction model in sequence to obtain the first image features and second image features of each rough registration reference image layer, and the first image features and second image features of each rough registration moving image layer. In this embodiment, by adopting different deep network structures to extract the feature information in each image layer respectively, more and richer feature information of each image layer can be obtained, and thus the classification performance of the image can be effectively improved.
[0069] In one embodiment, step 204 can be specifically implemented by the following method: First, scale the multiple rough registration reference image layers and the multiple rough registration moving image layers, and construct a reference image pyramid for each rough registration reference image layer and a moving image pyramid for each rough registration moving image layer according to the scaled images. Then, slide and extract the image features of each reference image pyramid and each moving image pyramid through the fully convolutional network of the first feature extraction model to obtain the first image features of each rough registration reference image layer and the first image features of each rough registration moving image layer. Finally, slide and extract the image features of each reference image pyramid and each moving image pyramid through the fully convolutional network of the second feature extraction model to obtain the second image features of each rough registration reference image layer and the second image features of each rough registration moving image layer.
[0070] Next, the feature extraction process in the above embodiments will be described in conjunction with specific examples. Here, taking the detection of the liver organ in medical images as an example, the image data layer with an output range of -6.0 to -5.0 is used as the input data for two feature extraction models. First, a sliding window strategy can be adopted to detect livers of different sizes in each image. Among them, Figure 6 FIG. shows a schematic diagram of the effect of converting an image layer into an image pyramid. Here, a scaling scale of 3 and a scaling factor of 0.9057 can be selected to construct the image pyramid. Since the size of the network input (i.e., the detection window) is 224×224, a liver region with a minimum size of 224 / 3 = 75 pixels can be detected. Then, the fully connected layer of the network structure of the two feature extraction models can be converted into a fully convolutional layer and the layer parameters can be redefined. The entire detection process is essentially a sliding window, but the original sliding window method has a large computational amount. To reduce the computational amount, a fully convolutional network is adopted, which can process input images of any size. For the 6×6 candidate window obtained from the fc6-conv layer in the fully convolutional network, it corresponds to a 224×224 detection window on the original image. As Figure 7 shown, through forward calculation, a feature vector with a fixed dimension in the fc6-conv layer can be obtained for each candidate region, thereby realizing multi-scale detection.
[0071] In one embodiment, the above first feature extraction model and second feature extraction model can be trained by the following method: First, multiple groups of image samples are obtained, where each group of image samples contains multiple coarsely registered image layers with target regions marked. Then, multiple sample regions are marked on each coarsely registered image layer, and positive and negative samples for each coarsely registered image layer are constructed according to the overlap rate between the sample region and the target region. Finally, according to the positive and negative samples of each coarsely registered image layer, the parameters of the initialized first deep convolutional neural network model and the parameters of the initialized second deep convolutional neural network model are iteratively trained to obtain the trained first feature extraction model and second feature extraction model.
[0072] Next, the feature extraction model training process in the above embodiments will be described with specific examples. Taking the medical image feature extraction model as an example, when training the model, the parameters of the deep network structures Clarifai and VGG Net-D can be initialized using the already trained model, and then the model can be fine-tuned using medical image samples. Among them, the first convolutional layer of the Clarifai network structure can use a relatively small 7×7 receptive field for filtering to perform dense filtering on the image. This network structure can ensure that more feature information is included in the first and second layers, improving the classification performance. Further, the VGG Net-D network structure can use a small 3×3 convolutional structure, which indirectly increases the network depth. At the same time, the convolution stride is set to 1 to ensure that the feature information is not lost. In addition to the above Clarifai and VGGNet-D network structures, other network structures can also be used to train the first feature extraction model and the second feature extraction model, such as UNet, ResNet, etc. This embodiment does not make specific limitations here.
[0073] Further, when training the model, the models trained by the Clarifai and VGG Net-D networks on the sample dataset can be used to initialize the network parameters, and then the processed training data can be used to fine-tune the two networks to obtain two feature models. Among them, when fine-tuning the Clarifai network, the dataset collected by multiple devices with annotations can be used. The image regions are cropped as positive samples according to the standard that the Intersection over Union (IOU, the overlap rate between the detection region and the target region) with the true annotation is greater than the threshold of 0.7, and the data is cropped according to the standard that the IOU with the true annotation is less than the threshold of 0.3 and a part is randomly selected as negative samples. Finally, the cropped grayscale images are fixed to a size of 224×224, and the training data is augmented through mirror flipping and then used for network training to finally obtain the first feature extraction model. For the VGG Net-D network, the VGG Net-D network model can be fine-tuned using the same training samples to finally obtain the second feature extraction model.
[0074] 205. Feature fusion, feature dimensionality reduction, and feature classification are respectively performed on the first image features and the second image features of each coarsely registered reference image layer, and the first image features and the second image features of each coarsely registered moving image layer to obtain the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer.
[0075] Specifically, after obtaining the first image features and second image features of each coarsely registered reference image layer, as well as the first image features and second image features of each coarsely registered moving image layer, the first image features and second image features of each image layer can be successively subjected to feature fusion, feature dimensionality reduction, and feature classification, and finally the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer are obtained. In this embodiment, the fused features can make up for the insufficient information extracted by a single network and depict richer data information. However, there is often a certain correlation between the fused features, resulting in information redundancy, and the dimensionality of the fused features is relatively high, increasing the computational complexity. Therefore, the method of feature dimensionality reduction can be used to obtain comprehensive features that can reflect the essence of classification. Feature fusion is beneficial to fully learning image features, depicting rich internal information of the data, and showing strong robustness to morphological changes, small scales, and illumination. The dimensionality-reduced features comprehensively express each part of the information of the region to be registered, solve problems such as the sparsity of feature vectors, and greatly improve the object detection in an unconstrained environment. Finally, through feature classification, the accuracy of target region detection can be further improved. In this embodiment, the fine positioning of the region to be registered is realized through feature fusion, feature dimensionality reduction, and feature classification. Compared with the voxel-by-voxel prediction in the existing segmentation method, the positioning efficiency of the region to be registered can be greatly improved.
[0076] In one embodiment, step 205 can be specifically implemented by the following method: First, the first image features and second image features of each coarsely registered reference image layer are respectively subjected to feature concatenation fusion with the first image features and second image features of each coarsely registered moving image layer to obtain the high-dimensional fusion features of each coarsely registered reference image layer and the high-dimensional fusion features of each coarsely registered moving image layer. Then, through the principal component analysis algorithm, feature dimensionality reduction operations are performed on the high-dimensional fusion features of each coarsely registered reference image layer and the high-dimensional fusion features of each coarsely registered moving image layer to obtain the fusion features of each coarsely registered reference image layer and the fusion features of each coarsely registered moving image layer. Finally, through the trained feature classification model, feature binary classification is performed on the fusion features of each coarsely registered reference image layer and the fusion features of each coarsely registered moving image layer to obtain the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer.
[0077] Next, the processes of feature fusion, feature dimensionality reduction, and feature classification in the above embodiment will be described in combination with specific examples. As Figure 8As shown in the figure, first, for each candidate region (i.e., target detection region) in the image layer, feature extraction is performed using the first feature extraction model and the second feature extraction model respectively, and the feature vectors of the same dimension (4096 dimensions) extracted are concatenated into a high-dimensional (8192 dimensions) fusion feature. Subsequently, the PCA (Principal Component Analysis) algorithm is used for feature dimensionality reduction to obtain a comprehensive feature that can reflect the essence of the classified image. The basic idea of the PCA algorithm is to find a projection direction such that the variance is maximized after the data is projected onto this projection direction. Calculating PCA mainly can be divided into two steps:
[0078] The first step is to zero-mean the samples, that is, to obtain the mean of the samples, and at the same time subtract this mean from each sample. After this step of processing, the mean of the samples is zero.
[0079] The second step is to calculate the projection direction with the largest sample variance. First, perform singular value decomposition on the covariance matrix of the samples, then select eigenvectors according to the size of the eigenvalues to construct a projection matrix, and select eigenvectors corresponding to a certain proportion of the eigenvalues (such as the first 50% of the eigenvalues) to construct a projection direction matrix. The resulting feature matrix after dimensionality reduction is m×4096 dimensions, where m is the number of pictures.
[0080] Furthermore, the SVM (Support Vector Machine) classification model can be used to classify each dimension-reduced feature. Specifically, when classifying the fused and dimension-reduced features, the trained SVM model can be used to classify the extracted features to obtain a probabilistic representation (i.e., confidence) of each candidate region, which corresponds to the confidence of a 224×224 detection window in the input image. Compare the confidence of the obtained candidate window with a pre-set probability threshold, and judge the region higher than the threshold as a candidate region (such as the liver region), and the region lower than the probability threshold as a non-candidate region. Although the detection speed of SVM is slow, it has a small risk of misclassification.
[0081] In one embodiment, the above feature classification model can be trained through the following method: First, obtain multiple groups of image samples, where each group of image samples contains multiple coarsely registered image layers with target regions marked. Then, mark multiple sample regions on each coarsely registered image layer respectively, and construct positive and negative samples for each coarsely registered image layer according to the overlap rate between the sample regions and the target regions. Finally, according to the positive and negative samples of each coarsely registered image layer, iteratively train the parameters of the initialized support vector machine classification model to obtain the trained feature classification model.
[0082] Next, the model training process in the above embodiments will be described with specific examples. Specifically, when training the model, the labeled dataset can be cropped using the true annotations, and a part with an IOU greater than the threshold of 0.7 is selected as the positive samples, and the standard cropped dataset with an IOU of the true annotation less than the threshold of 0.3 is selected as the negative samples. The positive and negative samples are randomly selected in a 1:1 ratio to train the SVM classifier. The steps are as follows:
[0083] First, select an appropriate kernel function (parameters of the SVM classifier). After comparing with other linear kernel functions and RBF kernel functions, the polynomial kernel function is selected as the final kernel function, and the classification function is:
[0084]
[0085] where w and b are the normal vector and intercept of the classification hyperplane, and k(x i , x) is the polynomial kernel function.
[0086] Then, introduce a relaxation factor (which has better adaptability for non-linear functions), and use the Lagrange multiplier method to obtain the optimization objective as:
[0087]
[0088] where C is a determined constant used to control the weights of each term in the objective function. ε i (i = 1, 2,..., n) is the relaxation variable, corresponding to the data point x i allowing the value to deviate from the function margin. α i is the Lagrange multiplier, and converting it into a dual problem gives:
[0089]
[0090] 0 ≤ α i ≤ C, i = 1, 2,..., n
[0091]
[0092] Finally, use the SMO algorithm (the intermediate theoretical link) to solve the parameters to obtain the final classification model.
[0093] 206. Post-process the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer to obtain the center localization results of the registered reference image and the center localization results of the registered moving image.
[0094] Specifically, after obtaining the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer, the candidate regions of each image layer can be post-processed in sequence, such as performing central position localization processing on the candidate regions of each image layer and performing clustering processing on the central positions of each image layer of the registered reference image and the registered moving image, etc., to obtain the central localization results of the registered reference image and the central localization results of the registered moving image.
[0095] In one embodiment, step 206 can be specifically implemented by the following method: First, through the non-maximum suppression algorithm, perform a merging operation on the overlapping regions of the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer to obtain the localization regions of each coarsely registered reference image layer and the localization regions of each coarsely registered moving image layer. Then, perform localization operations on the central positions of the localization regions of each coarsely registered reference image layer and the localization regions of each coarsely registered moving image layer respectively to obtain the central localization results of each coarsely registered reference image layer and the central localization results of each coarsely registered moving image layer. Finally, perform clustering calculations on the central localization results of each coarsely registered reference image layer and the central localization results of each coarsely registered moving image layer respectively to obtain the central localization results of the registered reference image and the central localization results of the registered moving image.
[0096] Next, the post-processing process in the above embodiment will be described in combination with specific examples. Specifically, although the features obtained after fusion, dimensionality reduction, and classification have obtained more detection information of different sizes, the detection boxes output by the feature classification model will have a high overlap. By using the NMS (Non-Maximum Suppression) method, candidate windows with a high overlap region can be merged, thereby realizing the post-processing operation and obtaining the final candidate region. In this implementation, the detection window with the highest confidence score can be selected, the detection windows with an IOU rate higher than a preset threshold can be removed, and then some detection windows that meet the overlap threshold can be aggregated to obtain the final detection box, that is, the candidate region.
[0097] Further, as Figure 9 shown, the white rectangular box shows the 2D localization result of the liver in the image data layer. By using the 2D localization results of each image layer of the three-dimensional image for central localization and clustering, the 3D fine localization result of the liver can be obtained. As Figure 10 shown, the multiple black dots on the left are the central localization results of the liver in multiple image layers. Through the clustering algorithm, the black dot in the image on the right can be obtained, and this black dot is the final fine central localization result of the liver.
[0098] 207. Obtain multiple registration reference image layers of the registration reference image and multiple registration moving image layers of the registration moving image according to the center positioning results of the registration reference image and the center positioning results of the registration moving image.
[0099] 208. Perform local registration on the multiple registration reference image layers of the registration reference image and the multiple registration moving image layers of the registration moving image to obtain the local registration results of the registration reference image and the registration moving image.
[0100] Specifically, after obtaining the center positioning results of the registration reference image and the center positioning results of the registration moving image, a predetermined range of the center positioning results of the two images (i.e., the upper and lower data layer ranges of the center positioning points) can be obtained as the registration area. For example, a range including all the areas to be registered can be selected for local registration. The registration results are as Figure 11 shown. The left side shows the registration effect and registration time of the global registration method in the prior art, and the right side shows the registration effect and registration time of the local registration method proposed in this embodiment. As Figure 11 shown, the registration effect of the image registration method proposed in this embodiment is more stable, and the registration time can be shortened by about 3 times, which can greatly improve the registration efficiency.
[0101] In this embodiment, through the registration area positioning in the rough registration stage and the fine registration stage, the registration reference image layers and the registration moving image layers within the same registration range can be obtained. Therefore, the search range for the matching position of the moving image relative to the reference image can be effectively reduced, and at the same time, obvious errors (such as the situation of matching different parts together) can be avoided. Compared with global registration, in this embodiment, by setting the registration ranges of the registration reference image and the registration moving image by itself, the local registration time can be controlled, so that the image registration efficiency can be greatly improved, and at the same time, the registration effect can be improved.
[0102] The image registration method provided in this embodiment first, in the coarse registration stage, adopts an unsupervised learning method. Without data annotation, by using the spatial and distance relationships between adjacent data layers of the image data, the confidence scores of each image layer can be obtained. Further, in the fine registration stage, based on the coarsely registered image layers obtained in the coarse registration stage, different types of deep convolutional neural networks are used to automatically learn the feature representations of the images from a large amount of data, and the features extracted by different networks are fused to obtain a rich feature description of the images. Then, the principal component analysis method is used to reduce the dimension of the fused features to filter out sparse features and obtain the comprehensive feature representation of the images. Finally, a feature classification model is used for binary classification to achieve the fine positioning of the area to be registered. Finally, in the local registration stage, first, the central positioning point of the area to be registered is obtained, and then the image layers within a certain range of the central positioning point are used for local registration of the image. The time for this local registration can be custom-controlled according to the registration range, making the time of the application preprocessing link controllable, thereby shortening the registration time and improving the registration effect.
[0103] Further, as Figures 1 to 11 a specific implementation of the method shown, this embodiment provides an image registration device, as Figure 12 shown. The device includes: an image acquisition module 31, an image processing module 32, and an image comparison module 33.
[0104] The image acquisition module 31 is configured to acquire a registration reference image and a registration moving image. Among them, the registration reference image includes multiple reference image layers, and the registration moving image includes multiple moving image layers.
[0105] The image processing module 32 is configured to input the registration reference image and the registration moving image into a trained unsupervised convolutional neural network model respectively, to obtain the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image.
[0106] The image comparison module 33 is configured to compare the confidence scores of each reference image layer and the confidence scores of each moving image layer with the preset confidence score range of the area to be registered respectively, to obtain multiple coarsely registered reference image layers of the registration reference image and multiple coarsely registered moving image layers of the registration moving image.
[0107] In a specific application scenario, the image comparison module 33 can specifically be used to respectively determine whether the confidence scores of each reference image layer and each moving image layer are within the confidence score range of a preset area to be registered; if the confidence score of any reference image layer is within the confidence score range of the preset area to be registered, then determine the reference image layer as a roughly registered reference image layer; if the confidence score of any moving image layer is within the confidence score range of the preset area to be registered, then determine the moving image layer as a roughly registered moving image layer.
[0108] In a specific application scenario, such as Figure 13 shown, the device further includes a comparison range determination module 34. The comparison range determination module 34 can specifically be used to obtain multiple groups of image samples, where each group of image samples includes multiple image layers; input the multiple groups of image samples into a trained unsupervised convolutional neural network model respectively to obtain the confidence scores of each image layer of the multiple groups of image samples; determine the confidence score range of the area to be registered according to the confidence scores of the image layers where the areas to be registered of the multiple groups of image samples are located.
[0109] In a specific application scenario, such as Figure 13 shown, the device further includes a model training module 35. The model training module 35 can specifically be used to obtain multiple groups of image samples, where each group of image samples includes multiple image layers; perform iterative training on the parameters of the initialized unsupervised convolutional neural network model according to each image layer of the multiple groups of image samples through a preset sequential error loss function and distance error loss function; when the sum of the sequential error loss function and the distance error loss function reaches a set loss value, obtain a trained unsupervised convolutional neural network model.
[0110] In a specific application scenario, such as Figure 13As shown in the figure, the present device further includes a feature extraction module 36, a feature fusion module 37, and a center positioning module 38. Among them, the feature extraction module 36 can be used to input multiple coarsely registered reference image layers and multiple coarsely registered moving image layers into a trained first feature extraction model and a second feature extraction model respectively, to obtain the first image features and the second image features of each coarsely registered reference image layer, as well as the first image features and the second image features of each coarsely registered moving image layer; the feature fusion module 37 can be used to perform feature fusion, feature dimensionality reduction, and feature classification on the first image features and the second image features of each coarsely registered reference image layer, as well as the first image features and the second image features of each coarsely registered moving image layer, to obtain the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer; the center positioning module 38 can be used to post-process the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer, to obtain the center positioning results of the registered reference image and the center positioning results of the registered moving image.
[0111] In a specific application scenario, the feature extraction module 36 can specifically be used to perform scale scaling on multiple coarsely registered reference image layers and multiple coarsely registered moving image layers, and construct a reference image pyramid for each coarsely registered reference image layer and a moving image pyramid for each coarsely registered moving image layer according to the scaled images; slide and extract the image features of each reference image pyramid and the image features of each moving image pyramid through the fully convolutional network of the first feature extraction model, to obtain the first image features of each coarsely registered reference image layer and the first image features of each coarsely registered moving image layer; slide and extract the image features of each reference image pyramid and the image features of each moving image pyramid through the fully convolutional network of the second feature extraction model, to obtain the second image features of each coarsely registered reference image layer and the second image features of each coarsely registered moving image layer.
[0112] In a specific application scenario, the feature fusion module 37 can specifically be used to perform feature concatenation fusion on the first image features and the second image features of each coarsely registered reference image layer, as well as the first image features and the second image features of each coarsely registered moving image layer, to obtain the high-dimensional fusion features of each coarsely registered reference image layer and the high-dimensional fusion features of each coarsely registered moving image layer; perform feature dimensionality reduction operations on the high-dimensional fusion features of each coarsely registered reference image layer and the high-dimensional fusion features of each coarsely registered moving image layer through the principal component analysis algorithm, to obtain the fusion features of each coarsely registered reference image layer and the fusion features of each coarsely registered moving image layer; perform feature binary classification on the fusion features of each coarsely registered reference image layer and the fusion features of each coarsely registered moving image layer through a trained feature classification model, to obtain the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer.
[0113] In a specific application scenario, the model training module 35 may specifically be further configured to obtain multiple groups of image samples, where each group of image samples includes multiple coarsely registered image layers with target regions marked thereon; mark multiple sample regions on each coarsely registered image layer respectively, and construct positive and negative samples for each coarsely registered image layer according to the overlap rate between the sample regions and the target regions; and iteratively train the parameters of the initialized support vector machine classification model according to the positive and negative samples of each coarsely registered image layer to obtain a trained feature classification model.
[0114] In a specific application scenario, the center positioning module 38 may specifically be configured to perform a merging operation on the overlapping regions of the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer through a non-maximum suppression algorithm to obtain the positioning regions of each coarsely registered reference image layer and the positioning regions of each coarsely registered moving image layer; perform a positioning operation on the center positions of the positioning regions of each coarsely registered reference image layer and the positioning regions of each coarsely registered moving image layer respectively to obtain the center positioning results of each coarsely registered reference image layer and the center positioning results of each coarsely registered moving image layer; and perform a clustering calculation on the center positioning results of each coarsely registered reference image layer and the center positioning results of each coarsely registered moving image layer respectively to obtain the center positioning result of the registered reference image and the center positioning result of the registered moving image.
[0115] In a specific application scenario, the model training module 35 may specifically be further configured to obtain multiple groups of image samples, where each group of image samples includes multiple coarsely registered image layers with target regions marked thereon; mark multiple sample regions on each coarsely registered image layer respectively, and construct positive and negative samples for each coarsely registered image layer according to the overlap rate between the sample regions and the target regions; and iteratively train the parameters of the initialized first deep convolutional neural network model and the parameters of the initialized second deep convolutional neural network model respectively according to the positive and negative samples of each coarsely registered image layer to obtain a trained first feature extraction model and a trained second feature extraction model.
[0116] In a specific application scenario, as Figure 13 shown, the apparatus further includes a local registration module 39, and the local registration module 39 may specifically be configured to obtain multiple coarsely registered reference image layers of the registered reference image and multiple coarsely registered moving image layers of the registered moving image according to the center positioning result of the registered reference image and the center positioning result of the registered moving image; perform local registration on the multiple coarsely registered reference image layers of the registered reference image and the multiple coarsely registered moving image layers of the registered moving image to obtain the local registration result of the registered reference image and the registered moving image.
[0117] It should be noted that for other corresponding descriptions of each functional unit involved in the image registration device provided in this embodiment, reference can be made to Figures 1 to 11 the corresponding description in
[0118] Based on the method as described above Figures 1 to 11 shown, correspondingly, this embodiment further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image registration method as described above Figures 1 to 11 shown.
[0119] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. The software product to be recognized can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of this application.
[0120] Based on the method as described above Figures 1 to 11 shown, and Figure 12 and Figure 13 the image registration device embodiment shown, in order to achieve the above object, this embodiment further provides an entity device for image registration, which can specifically be a personal computer, a server, a smart phone, a tablet computer, a smart watch, or other network devices, etc. The entity device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as described above Figures 1 to 11 shown.
[0121] Optionally, the entity device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0122] Those skilled in the art can understand that the structure of the entity device for image registration provided in this embodiment does not constitute a limitation on the entity device, and it may include more or fewer components, or combine some components, or have different component arrangements.
[0123] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware of the above-mentioned physical device and the software resources to be recognized, and supports the operation of the information processing program and other software and / or programs to be recognized. The network communication module is used to implement communication between components inside the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can be implemented by hardware. By applying the technical solution of the present application, first, a registration reference image and a registration moving image with multiple image layers are obtained, then the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image are obtained through a trained unsupervised convolutional neural network model, and finally, the rough positioning result of the area to be registered is determined through the confidence scores of each image layer. Compared with the prior art, the above method can effectively reduce the training difficulty and training time-consuming of the rough registration model, and avoid the problem of registration errors caused by too large a search range in the image rough registration stage. In addition, by using an unsupervised convolutional neural network model to obtain partial image layers of the area to be registered and then performing local registration, the image registration efficiency can be greatly improved and the registration result can be enhanced.
[0125] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.
[0126] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure is only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.
Claims
1. An image registration method, characterized in that, The method includes: Obtaining a registration reference image and a registration moving image, where the registration reference image includes a plurality of reference image layers, and the registration moving image includes a plurality of moving image layers; Inputting the registration reference image and the registration moving image into a trained unsupervised convolutional neural network model respectively, to obtain the confidence scores of each reference image layer of the registration reference image and the confidence scores of each moving image layer of the registration moving image; Comparing the confidence scores of each reference image layer and the confidence scores of each moving image layer with a preset confidence score range of the area to be registered respectively, to obtain a plurality of roughly registered reference image layers of the registration reference image and a plurality of roughly registered moving image layers of the registration moving image; Inputting the plurality of roughly registered reference image layers and the plurality of roughly registered moving image layers into a trained first feature extraction model and a second feature extraction model respectively, to obtain the first image features and the second image features of each roughly registered reference image layer, and the first image features and the second image features of each roughly registered moving image layer; Performing feature fusion, feature dimensionality reduction and feature classification on the first image features and the second image features of each roughly registered reference image layer respectively, and performing feature fusion, feature dimensionality reduction and feature classification on the first image features and the second image features of each roughly registered moving image layer respectively, to obtain the candidate regions of each roughly registered reference image layer and the candidate regions of each roughly registered moving image layer; Performing post-processing on the candidate regions of each roughly registered reference image layer and the candidate regions of each roughly registered moving image layer, to obtain the center positioning result of the registration reference image and the center positioning result of the registration moving image; According to the center positioning result of the registration reference image and the center positioning result of the registration moving image, obtaining a plurality of registered reference image layers of the registration reference image and a plurality of registered moving image layers of the registration moving image; Performing local registration on the plurality of registered reference image layers of the registration reference image and the plurality of registered moving image layers of the registration moving image, to obtain the local registration result of the registration reference image and the registration moving image.
2. The method according to claim 1, wherein The step of comparing the confidence scores of each reference image layer and the confidence scores of each moving image layer with a preset confidence score range of the area to be registered respectively, to obtain a plurality of roughly registered reference image layers of the registration reference image and a plurality of roughly registered moving image layers of the registration moving image, includes: Judging whether the confidence scores of each reference image layer and the confidence scores of each moving image layer are within the preset confidence score range of the area to be registered respectively; If the confidence score of any reference image layer is within the preset confidence score range of the area to be registered, determining the reference image layer as a roughly registered reference image layer; If the confidence score of any moving image layer is within the preset confidence score range of the area to be registered, determining the moving image layer as a roughly registered moving image layer.
3. The method according to claim 2, wherein Before comparing the confidence scores of each reference image layer and the confidence scores of each moving image layer with the confidence score range of the preset region to be registered respectively, the method further includes: Obtain multiple groups of image samples, where each group of image samples includes multiple image layers; Input the multiple groups of image samples into the trained unsupervised convolutional neural network model respectively, and obtain the confidence scores of each image layer of the multiple groups of image samples; Determine the confidence score range of the region to be registered according to the confidence scores of the image layers where the regions to be registered in the multiple groups of image samples are located.
4. The method according to any one of claims 1 to 3, characterized in that, The training method of the unsupervised convolutional neural network model includes: Obtain multiple groups of image samples, where each group of image samples includes multiple image layers; According to each image layer of the multiple groups of image samples, iteratively train the parameters of the initialized unsupervised convolutional neural network model through a preset sequential error loss function and distance error loss function; When the sum of the sequential error loss function and the distance error loss function reaches a set loss value, obtain the trained unsupervised convolutional neural network model.
5. The method according to claim 1, wherein The step of inputting the multiple coarsely registered reference image layers and the multiple coarsely registered moving image layers into the trained first feature extraction model and second feature extraction model respectively to obtain the first image features and second image features of each coarsely registered reference image layer, and the first image features and second image features of each coarsely registered moving image layer includes: Perform scale scaling on the multiple coarsely registered reference image layers and the multiple coarsely registered moving image layers, and construct a reference image pyramid for each coarsely registered reference image layer and a moving image pyramid for each coarsely registered moving image layer according to the scaled images; Slide and extract the image features of each reference image pyramid and the image features of each moving image pyramid through the fully convolutional network of the first feature extraction model to obtain the first image features of each coarsely registered reference image layer and the first image features of each coarsely registered moving image layer; Slide and extract the image features of each reference image pyramid and the image features of each moving image pyramid through the fully convolutional network of the second feature extraction model to obtain the second image features of each coarsely registered reference image layer and the second image features of each coarsely registered moving image layer.
6. The method according to claim 1, characterized in that, The step of respectively performing feature fusion, feature dimensionality reduction and feature classification on the first image features and second image features of each coarsely registered reference image layer, and respectively performing feature fusion, feature dimensionality reduction and feature classification on the first image features and second image features of each coarsely registered moving image layer to obtain candidate regions of each coarsely registered reference image layer and candidate regions of each coarsely registered moving image layer includes: Respectively perform feature concatenation fusion on the first image features and second image features of each coarsely registered reference image layer, and respectively perform feature concatenation fusion on the first image features and second image features of each coarsely registered moving image layer to obtain high-dimensional fusion features of each coarsely registered reference image layer and high-dimensional fusion features of each coarsely registered moving image layer; By using the principal component analysis algorithm, perform feature dimension reduction operations on the high-dimensional fusion features of each coarsely registered reference image layer and the high-dimensional fusion features of each coarsely registered moving image layer to obtain the fusion features of each coarsely registered reference image layer and the fusion features of each coarsely registered moving image layer; By using the trained feature classification model, perform feature binary classification on the fusion features of each coarsely registered reference image layer and the fusion features of each coarsely registered moving image layer to obtain the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer.
7. The method according to claim 6, characterized in that, The training method of the feature classification model includes: Obtain multiple groups of image samples, where each group of image samples contains multiple coarsely registered image layers marked with target regions; Mark multiple sample regions on each of the coarsely registered image layers, and construct positive samples and negative samples for each coarsely registered image layer according to the overlap rate between the sample regions and the target regions; According to the positive samples and negative samples of each coarsely registered image layer, perform iterative training on the parameters of the initialized support vector machine classification model to obtain the trained feature classification model.
8. The method according to claim 1, wherein The post-processing of the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer to obtain the center positioning result of the registered reference image and the center positioning result of the registered moving image includes: By using the non-maximum suppression algorithm, perform a merging operation on the overlapping regions of the candidate regions of each coarsely registered reference image layer and the candidate regions of each coarsely registered moving image layer to obtain the positioning regions of each coarsely registered reference image layer and the positioning regions of each coarsely registered moving image layer; Perform positioning operations on the center positions of the positioning regions of each coarsely registered reference image layer and the positioning regions of each coarsely registered moving image layer respectively to obtain the center positioning results of each coarsely registered reference image layer and the center positioning results of each coarsely registered moving image layer; Perform clustering calculations on the center positioning results of each coarsely registered reference image layer and the center positioning results of each coarsely registered moving image layer respectively to obtain the center positioning result of the registered reference image and the center positioning result of the registered moving image.
9. The method according to any one of claims 1, 5, 6, 7, and 8, characterized in that, The training methods of the first feature extraction model and the second feature extraction model include: Obtain multiple groups of image samples, where each group of image samples contains multiple coarsely registered image layers marked with target regions; Mark multiple sample regions on each of the coarsely registered image layers, and construct positive samples and negative samples for each coarsely registered image layer according to the overlap rate between the sample regions and the target regions; According to the positive samples and negative samples of each coarsely registered image layer, perform iterative training on the parameters of the initialized first deep convolutional neural network model and the parameters of the initialized second deep convolutional neural network model respectively to obtain the trained first feature extraction model and second feature extraction model.
10. An image registration device, characterized in that, The device includes: An image acquisition module, configured to acquire a registration reference image and a registration moving image, wherein the registration reference image includes a plurality of reference image layers, and the registration moving image includes a plurality of moving image layers; An image processing module, configured to respectively input the registration reference image and the registration moving image into a trained unsupervised convolutional neural network model to obtain confidence scores of each reference image layer of the registration reference image and confidence scores of each moving image layer of the registration moving image; An image comparison module, configured to respectively compare the confidence scores of each reference image layer and the confidence scores of each moving image layer with a preset confidence score range of a region to be registered, to obtain a plurality of roughly registered reference image layers of the registration reference image and a plurality of roughly registered moving image layers of the registration moving image; A feature extraction module, configured to respectively input the plurality of roughly registered reference image layers and the plurality of roughly registered moving image layers into a trained first feature extraction model and a second feature extraction model to obtain first image features and second image features of each roughly registered reference image layer, and first image features and second image features of each roughly registered moving image layer; A feature fusion module, configured to respectively perform feature fusion, feature dimension reduction and feature classification on the first image features and the second image features of each roughly registered reference image layer, and respectively perform feature fusion, feature dimension reduction and feature classification on the first image features and the second image features of each roughly registered moving image layer, to obtain candidate regions of each roughly registered reference image layer and candidate regions of each roughly registered moving image layer; A center positioning module, configured to perform post-processing on the candidate regions of each roughly registered reference image layer and the candidate regions of each roughly registered moving image layer to obtain a center positioning result of the registration reference image and a center positioning result of the registration moving image; A local registration module, configured to obtain a plurality of registered reference image layers of the registration reference image and a plurality of registered moving image layers of the registration moving image according to the center positioning result of the registration reference image and the center positioning result of the registration moving image; perform local registration on the plurality of registered reference image layers of the registration reference image and the plurality of registered moving image layers of the registration moving image to obtain a local registration result of the registration reference image and the registration moving image.
11. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-modal image registration method and device, electronic equipment and storage medium
CN111311655A
Unsupervised Deep Representation Learning for Fine-grained Body Part Recognition
US20180060652A1