Feature point detection
By training the model for feature point detection, selecting and matching the points of interest in the image, and updating the model, the problem of difficult processing of low-quality image in medical image registration is solved in the prior art, and higher registration accuracy and stability are achieved.
Patent Information
- Application Number
- CN202080031637.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-30
- Filing Date
- 2020-03-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-03-12
AI Technical Summary
The prior art is difficult to effectively process low-quality images in medical image registration, especially in the case of light changes and small overlapping areas, resulting in low registration quality.
A method of training a model for feature point detection is adopted. By acquiring two images, a score map is generated and the points of interest are selected, paired matches are performed, and the correctness of the match is checked based on the basic truth value transformation, a reward map is generated, and the model is finally updated.
Improve the accuracy and stability of image registration, especially under low-quality images and challenging conditions, and enhances the directness and targeting of the model training.
Smart Images

Figure CN113826143B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to feature detection. The present invention more particularly relates to feature detection for image registration. The present invention also relates to image registration using detected feature points. Background Art
[0002] Image registration (the process of aligning two or more images to the same global spatial reference) is a key element in the fields of computer vision, pattern recognition, and medical image analysis. Imaging in the medical field is a challenging task. This can sometimes result in inferior image quality compared to regular photography, for example. The presence of noise, blur, and other imaging artifacts, combined with the nature of the imaged tissue, do not lend themselves to classical, state-of-the-art feature detectors optimized for natural images. Furthermore, domain-specific handcrafting of new feature detectors is a time-consuming task with no guarantee of success.
[0003] Existing registration algorithms can be divided into region-based and feature-based methods. Region-based methods typically rely on similarity measures such as cross-correlation
[15] , mutual information [33, 4], or phase correlation
[24] to compare the intensity patterns of an image pair and estimate the transformation. However, in the presence of illumination variations or small overlapping regions, region-based methods become challenging or infeasible. In contrast, feature-based methods extract corresponding points on an image pair and search for a transformation that minimizes the distance between the detected feature points. Compared to region-based registration techniques, they are more robust to variations in intensity, scale, and rotation and are therefore considered more applicable to problems such as medical image registration. Typically, feature extraction and matching of two images involves four steps: detection of interest points, description of each feature, matching of corresponding features, and estimation of the transformation between the images using the matching. As can be seen, the detection step affects each step and is therefore crucial for successful registration. It requires high image coverage and stable keypoints in low-contrast images.
[0004] In the literature, local interest point detectors have been evaluated in detail. SIFT
[29] is probably the most famous detector / descriptor in computer vision. It computes corners and blobs at different scales to add scale invariance and uses local gradients to extract descriptors. Root SIFT [9] has been shown to give enhanced results compared to SIFT. Speeded Up Robust Features (SURF)
[12] is a faster alternative that uses Haar filters and integral images, while KAZE [7] exploits non-linear scale space for more accurate keypoint detection.
[0005] In the field of fundus imaging, widely used techniques rely on vessel tree and branch point analysis [28, 21]. However, accurate segmentation of the vessel tree is challenging and registration often fails on images with few vessels. Alternative registration techniques are based on matching repeatable local features: Chen et al. detected Harris corners on low-quality multimodal retinal images
[22] and assigned them partial intensity invariant feature (Harris-PIIFD) descriptors
[14] . They achieved good results on low-quality images with an overlap area of more than 30%, but they are characterized by low repeatability. Wang et al.
[41] used SURF features to increase repeatability and introduced a new point matching method to reject a large number of outliers, but the success rate dropped significantly when the overlap area was reduced to below 50%.
[0006] Cattin et al.
[13] also demonstrated that the SURF method can be effectively used to create mosaics of retinal images, even in the absence of significant angiogenesis. However, this technique was only successful in the case of highly self-similar images. The D-saddle detector / descriptor
[34] outperformed previous methods in terms of successful registration rate on the Fundus Image Registration (FIRE) dataset
[23] and was able to detect points of interest on low-quality regions.
[0007] Recently, with the advent of deep learning, learned detectors based on CNN architectures have been shown to outperform state-of-the-art computer vision detectors [19, 17, 44, 32, 10]. The Learned Invariant Feature Transform (LIFT)
[32] uses patches to train a fully differentiable deep CNN for interest point detection, orientation estimation, and descriptor computation supervised by the classic Structure from Motion (SfM) system. SuperPoint
[17] introduces a self-supervised framework to train interest point detectors and descriptors. It achieves state-of-the-art homography matrix estimation results on HPatches
[11] compared to LIFT, SIFT, and Oriented Rapid Rotation Briefs (ORB). However, the training process is complex and their self-supervision means that the network can only find points on corners. The Local Feature Network (LF-NET)
[32] is closest to our approach: Ono et al. trained a keypoint detector and descriptor end-to-end in a two-branch setting, where one branch is differentiable and the output of the other branch is non-differentiable. They optimized the detector for repeatability between image pairs.
[0008] Truong et al.
[39] evaluated Root-SIFT, SURF, KAZE, ORB
[36] , Binary Robust Invariant Scalable Keypoints (BRISK)
[27] , Fast Retinal Keypoints (FREAK) [6], LIFT, SuperPoint
[17] , and LF-NET
[32] in terms of image matching and registration quality for retinal fundus images. They found that while SuperPoint outperformed all others in terms of matching performance, LIFT showed the best results in terms of registration quality, followed closely by KAZE and SIFT. They emphasized that the problem with these detectors is that they detect feature points that are densely located with each other and may have similar descriptors. This can lead to false matches, resulting in inaccurate or failed registration. Summary of the invention
[0009] An aspect of the present invention is to solve at least one of the problems outlined above, or to provide at least one of the advantages described herein.
[0010] According to a first aspect of the present invention, a method for training a model for feature point detection comprises:
[0011] Acquire a first image and a second image;
[0012] generating a first score map for the first image and a second score map for the second image using the model;
[0013] selecting a first plurality of interest points in the first image based on the first score map;
[0014] selecting a second plurality of interest points in the second image based on the second score map;
[0015] pair-matching a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest;
[0016] checking correctness of the pairwise matches based on a ground truth transformation between the first image and the second image to generate a reward map;
[0017] combining or comparing the score map and the reward map; and
[0018] The model is updated based on a result of the combining or comparing.
[0019] Based on the result of combining or comparing the score map and the reward map based on the ground truth transformation between the first image and the second image, updating the model provides a highly direct reward, more targeted to align the two images. Therefore, the training of the model is improved. In any embodiments disclosed herein, the model can be, for example, a learning function, an artificial neural network, or a classifier.
[0020] Selecting multiple interest points may include imposing a maximum limit on the distance from any point in the image to the nearest one of the interest points. This helps avoid situations where most interest points are clustered in a small area of the image. This feature helps obtain matching points, thereby providing an overall better image registration.
[0021] The pairwise matching may be performed based on similarities between features detected at the first interest point in the first image and features detected at the second interest point in the second image. This allows matching an interest point with an interest point having similar feature descriptors in another image.
[0022] The matching may be performed in the first direction by matching the first interest point with a second interest point of the plurality of second interest points that has features most similar to the features of the first interest point. This finds a suitable selection of matching interest points.
[0023] Matching can be further performed in the second direction by matching the second interest point with a first interest point among the plurality of first interest points that has features most similar to the features of the second interest point, which helps improve the selection of matching interest points.
[0024] The reward map may indicate rewards for successfully matched interest points according to the ground truth data, while indicating no rewards for unsuccessfully matched interest points according to the ground truth data. This helps improve the training process by providing a more targeted reward map.
[0025] The combining or comparing may include combining or comparing the score maps and reward maps of only the points of interest. The other points (non-points of interest) are numerous and may not add enough information to help the training process.
[0026] Combining or comparing may include balancing some true positive matches with some false positive matches by (potentially randomly) selecting the false positive matches and combining or comparing the score map and reward map only for the selection of the false positive matches and the true positive matches, where the true positive matches are points of interest that pass the correctness check and the false positive matches are points of interest that do not pass the correctness check. This helps to further reduce any bias towards "mismatched" training.
[0027] The combining or comparing may comprise calculating the sum of squared differences between the score map and the reward map. This may comprise a suitable component of a training procedure.
[0028] According to another aspect of the present invention, a device for training a feature point detection model includes a control unit; and a memory, which includes instructions for causing the control unit to perform the following steps: acquiring a first image and a second image; generating a first score map for the first image and a second score map for the second image using the model; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; pair-matching a first interest point in the first plurality of interest points with a second interest point in the second plurality of interest points; checking the correctness of the pair-wise matching based on a ground truth transformation between the first image and the second image to generate a reward map; combining or comparing the score map and the reward map; and updating the model based on the result of the combination or comparison.
[0029] According to another aspect of the present invention, a method for registering a first image to a second image is provided, the method comprising acquiring a first image and a second image; generating a first score map for the first image and a second score map for the second image using a model trained by any of the methods or devices of the preceding items; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; and pair-matching a first interest point in the first plurality of interest points with a second interest point in the second plurality of interest points.
[0030] Selecting a plurality of points of interest may comprise imposing a maximum limit on the distance from any point in the image to the nearest one of the points of interest.
[0031] The pairwise matching may be performed based on similarities between features detected at the first interest point in the first image and features detected at the second interest point in the second image.
[0032] The matching may be performed in the first direction by matching the first interest point with a second interest point having a feature most similar to that of the first interest point among a plurality of second interest points.
[0033] Matching may be further performed in the second direction by matching the second interest point with a first interest point having a feature most similar to that of the second interest point among the plurality of first interest points.
[0034] Selecting a plurality of points of interest may comprise imposing a maximum limit on the distance from any point in the image to the nearest one of the points of interest.
[0035] The pairwise matching may be performed based on similarities between features detected at the first interest point in the first image and features detected at the second interest point in the second image.
[0036] The matching may be performed in the first direction by matching the first interest point with a second interest point having a feature most similar to that of the first interest point among a plurality of second interest points.
[0037] Matching may be further performed in the second direction by matching the second interest point with a first interest point having a feature most similar to that of the second interest point among the plurality of first interest points.
[0038] According to another aspect of the present invention, there is provided an apparatus for registering a first image to a second image, comprising a control unit, for example, at least one computer processor, and a memory, comprising instructions for causing the control unit to perform the following steps: acquiring a first image and a second image; generating a first score map for the first image and a second score map for the second image using a model generated by the aforementioned method or apparatus; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; and pair-matching a first interest point in the first plurality of interest points with a second interest point in the second plurality of interest points.
[0039] According to another aspect of the present invention, a model generated by the aforementioned method or device is provided.
[0040] One aspect of the invention is a semi-supervised learning method for keypoint detection. Detectors are typically optimized for repeatability (e.g., LF-NET) rather than for the quality of relevant matches between pairs of images. One aspect of the invention is a training process that uses reinforcement learning to extract repeatable, stable points of interest with dense coverage and is specifically designed to maximize correct matches on a specific domain. An example of such a specific domain is challenging retinal slit-lamp images.
[0041] Those skilled in the art will appreciate that the above features may be combined in any manner deemed useful. In addition, modifications and variations described with respect to the system and apparatus may also be applied to the method and computer program product, and modifications and variations described with respect to the method may also be applied to the system and apparatus. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In the following, aspects of the invention will be explained by way of examples with reference to the accompanying drawings. The drawings are schematic and may not be drawn to scale.
[0043] Figure 1 Illustrated how to match points in a real-world example.
[0044] Figure 2A An example of the steps of training image pairs is shown.
[0045] Figure 2B An example of loss calculation is shown.
[0046] Figure 2C An example schematic diagram showing Unet-4.
[0047] Figure 3 Examples of images from the slit-lamp dataset are shown.
[0048] Figure 4A A summary of detector / descriptor performance metrics evaluated on 206 pairs of slit-lamp datasets using non-preprocessed data is shown.
[0049] Figure 4B A summary of detector / descriptor performance metrics evaluated on 206 pairs of slit-lamp datasets using pre-processed data is shown.
[0050] Figure 5 It is shown how to match points in another practical example.
[0051] Figure 6 A mosaic formed by the registration of consecutive images is shown.
[0052] Figure 7 A block diagram of a system for training a feature point detection model is shown.
[0053] Figure 8 A flow chart showing a method for training a feature point detection model.
[0054] Fig. 9 A block diagram of a system for registering a first image to a second image is shown.
[0055] Fig.10 A flow chart illustrating a method of registering a first image to a second image. DETAILED DESCRIPTION
[0056] Certain exemplary embodiments will be described in more detail with reference to the accompanying drawings and text.
[0057] Matters disclosed in the description, such as detailed configurations and elements, are provided to help a comprehensive understanding of the exemplary embodiments. Therefore, it is apparent that the exemplary embodiments can be carried out without those specifically defined matters. In addition, well-known operations or structures are not described in detail because they would obscure the description with unnecessary details.
[0058] The techniques disclosed herein can be applied to image registration in any application field.
[0059] Known fully supervised machine learning solutions to this problem require manually annotated ground truth that relates locations in images from two independent viewpoints. While in natural images ground truth can be created in a static setting, medical data is highly dynamic and involves patients. This makes obtaining ground truth very difficult to infeasible.
[0060] The features detected by most feature detectors or keypoint detectors are concentrated on edges and corners. In the medical field, large regions are usually featureless, which leads to the aggregation of feature points and thus inaccurately transformed matching.
[0061] The following aspects are presented as examples.
[0062] 1. In contrast to prior art that uses indirect metrics such as repeatability, training is performed on the final matching success on the target domain.
[0063] 2. The algorithm is trained only on synthetic augmentations, solving the problem of ground truth data. It allows the detector to be trained using only target data.
[0064] 3. The feature points are evenly distributed over the entire image.
[0065] The following advantages can be achieved partially or completely:
[0066] 1. By using certain embodiments of the present invention, feature point detectors can be learned and optimized for specific imaging domains that would otherwise be infeasible. In addition, good feature descriptors can be reused.
[0067] 2. No ground truth is required, only samples from the target domain.
[0068] 3. No pre-processing or manual feature creation is required to achieve better matching rates.
[0069] 4. With the uniform distribution of feature points (and matches), the accuracy of estimating the transformation between two images is greatly improved.
[0070] The following further advantages may be achieved partially or completely:
[0071] 1. If more data is available from the target domain, it can be used to further improve the detector.
[0072] 2. The detector can best fit the descriptor algorithm. If a better descriptor is found, a new detector can be trained with no additional implementation cost or data.
[0073] One aspect of the present invention is a training process for training a model to detect feature points in an image. This training process can be applied to any type of image. For example, in the medical field, this can be used for satellite images or outdoor images. The images can also be formed by a mobile phone or handheld camera, an ultrasound device, or an ophthalmic slit lamp imaging device.
[0074] For example, the image may be a 2D image or a 3D image. In the case of a 2D image, 2D (or 1D) features may be detected and the coordinates of the points may be two-dimensional. In the case of a 3D image, 1D, 2D or 3D features may be detected and the coordinates of the points may be three-dimensional.
[0075] For example, the image may be a 1D image. In the case of a 1D image, 1D features may be detected and the coordinates of the points may be one-dimensional. In general, for any positive integer N, the image may be N-dimensional.
[0076] For example, the image is a photograph or an X-ray image.
[0077] For example, feature points detected by the model can be used for image registration.
[0078] One aspect of the invention is a method or apparatus for finding the best feature point detection for a particular feature point descriptor. In this sense, the algorithm finds points of interest in the image that optimize the matching capabilities of an object descriptor algorithm. The object descriptor algorithm can be root SIFT, for example, as described in the accompanying article, but can be any other descriptor, such as ORB, BRISK, BRIEF, etc.
[0079] In the following description, the terms interest point and key point are used interchangeably.
[0080] The following steps can be used to train the model.
[0081] 1. Given a pair of images I∈R H×W and I'∈R H×W With the basic truth homography matrix H = H I,I’ Related, the model can provide a score map for each image:
[0082] S=f θ (I) and S' = f θ (I 0 ).
[0083] In this step, two homography matrices can be used to transform HI and H respectively. I' Two images I and I' are generated from the original image. The homography matrix between these two images is H = H I *H I'For example, the homography matrix transform H I and H I' You can use the random generator to generate random numbers. θ Generate two key point probability maps S and S' as follows: S = f θ (I) and S' = f θ (I').
[0084] 2. The locations of interest points can be extracted on the two score maps using a window of size w using the standard non-differentiable NonMax-Supression (NMS). In this step, the locations of interest points can be extracted on the two score maps S and S' using non-maximum suppression with a window size of w. This means that only the maximum value is locally retained in all squares w×w, and all other values are set to 0. This results in noticeably sparse points in the image. In other words, the number of key points is reduced because only the maximum value key points are retained in each square w×w of the key point probability maps S and S'.
[0085] The window size w can be chosen by trial and error. This width w can be predetermined (fixed algorithm parameter) or dynamically determined based on the image I at hand.
[0086] It should be understood that this is an optional step in the process. Furthermore, alternative algorithms may be used instead of the standard non-differentiable NMS.
[0087] 3. 128 root SIFT feature descriptors may be calculated for each detected keypoint. For example, in this step, we assign a feature descriptor vector to each keypoint found in step 2 using the root SIFT descriptor algorithm. The feature descriptor vector may have a length of, for example, 128. Other lengths may be used instead. As mentioned elsewhere in this document, another type of feature descriptor may be used instead.
[0088] For example, the feature descriptor is calculated based on image information present in the image I or I'. For example, the SIFT feature descriptor describes the image gradient around a key point in the image I or I'.
[0089] 4. A brute force matcher may be used to match key points in image I with key points in image I' and vice versa, such as [1]. For example, only matches found in both directions are retained. In this step, the key points of image I may be matched with the key points of image I' using the closest descriptor of the other image I' calculated in step 3). To do this, the descriptor output calculated for the key points of image I may be compared with the descriptor output calculated for the key points of image I'. Here, "closest descriptor" refers to the most similar descriptor output according to a predefined similarity measure. In some implementations, the match is discarded if the second closest point has a very similar descriptor.
[0090] Matching can be done in both directions (from image I to image I' and from image I' to image I). Only matches found in both directions are retained. In alternative embodiments, some or all of the matches found in only one direction may be retained in the set of matches.
[0091] Any suitable kind of matcher may be used. This step is not limited to brute force matchers.
[0092] In some implementations, the matching in step 4 is performed based on the output of the descriptors in images I and I' without considering the ground truth homography matrix H (or H I or H I' ).
[0093] 5. Check the match against the ground truth homography matrix H. If the corresponding keypoint x in image I falls into I after applying H 0 Point x in 0 The match is defined as a true positive if the neighborhood of is greater than . This can be expressed as:
[0094] ||H*xx′||≤ε,
[0095] where ε can be chosen to be, for example, 3 pixels. Let T denote the set of ground truth matches.
[0096] In this step, for all matching points found in step 4), the matching points of image I are transformed into points of image I' using the homography matrix H. For example, the coordinates of the matching points in image I are transformed into the coordinates of the corresponding points in image I' using the homography matrix H. In some embodiments, only a subset of the matching points found in step 4) are transformed.
[0097] A match can be defined as a true positive if a keypoint x in image I is close enough to a matching point x' found in step 4 after being transformed using the homography matrix H to obtain the ground truth corresponding point in image I'.
[0098] For example, "close enough" can be defined as the distance between the ground truth corresponding point of key point x and the matching point x' is less than a certain threshold ε. The distance can be, for example, the Euclidean distance. In some implementations, ε can be 3 pixels. However, this is not a limitation.
[0099] Model f θ It may be a neural network, such as a convolutional neural network. However, this is not a limitation. In alternative embodiments, other types of models may be used, such as statistical models.
[0100] According to one aspect of the invention, the model may be trained using steps 2, 3, 4 and / or 5. Such training may involve back propagation of the model or neural network.
[0101] For example, the function to be optimized by the training process (ie, the cost function or loss) may be based on the matches found in step 5.
[0102] Therefore, the matching ability of key points can be optimized.
[0103] The loss function can be provided as follows:
[0104] L simple =(f θ (I)-R) 2 ,
[0105] Among them, the reward matrix R can be defined as follows:
[0106]
[0107] However, the above L simple One drawback of the formula may be that in some cases there is a relatively large class imbalance between positive reward points and empty reward points, with the latter dominating, especially in the first stage of training. Given a reward R of almost zero value, model f θ It may converge to a zero output rather than the desired indication of keypoints corresponding to image features used for registration.
[0108] Preferably, to counteract the imbalance, we use sample mining: we select all n true positive points and randomly sample an additional number of n points from the set of false positives instead of all false positives. We backpropagate through only the 2n true positive feature points and mine out the false positive keypoints.
[0109] Alternatively, the number of true positive points and false positive points used for back propagation need not be the same. It is sufficient if the number of false positive points is reduced to a certain extent relative to the number of true positive points.
[0110] This way, the number of true positive keypoints and false positive keypoints is (approximately) the same. This may help avoid an imbalance in the number of true positive and false positive samples.
[0111] If there are more true positives than false positives, gradients may be backpropagated through all matches found.
[0112] This mining can be mathematically formulated as a mask M, equal to 1 at the locations of true positive keypoints and at the locations of a (random) subset of mined (false positive) feature points, and equal to 0 otherwise.
[0113] Therefore, the loss can be formulated, for example, as follows:
[0114] L=∑(f e (I)-R) 2 M. Preferably, according to the mask M, the loss can be formulated as the average of the losses of each participating feature point:
[0115]
[0116] The dot represents element-by-element multiplication, the superscript 2 represents element-by-element squaring, and the minus sign represents element-by-element subtraction, and the sum is the sum of all elements of the matrix. Model f θ The output of can indicate the probability or likelihood that the point is a keypoint.
[0117] Preferably, the loss function is calculated only based on points that are found to be matches according to the feature descriptors.In certain applications, better and / or more stable results may be obtained by not training on true negatives.
[0118] Step 2 may help to distribute the keypoints more regularly over the image. For example, there is a fixed non-maximum suppression window of width w and height h on the keypoint probability map, ensuring that the maximum distance between two keypoints is 2*w in the x-direction and 2*h in the y-direction. For example, the width w of the window can be equal to the height h of the window. Alternatively, the window size can be made image-dependent, for example, depending on the amount of information in the image.
[0119] The descriptors generated in step 3 can be used to match the keypoints of image I with the keypoints of image I' based on the similarity between the descriptors of the features present at the keypoints in the respective images.
[0120] Different types of descriptors can be used. For example, SIFT might be a suitable descriptor (it has a detector and a descriptor component).
[0121] Generally speaking, matching can be done in three steps:
[0122] 1) Detect the feature points of interest in the two images;
[0123] 2) Generate a unique description (feature vector) for each detected feature point; and
[0124] 3) Matching feature points of one image with those of another image using a similarity or distance metric (e.g., Euclidean distance) between feature vectors.
[0125] The advantage of this method may be that the image may not need to be preprocessed before generating the keypoint probability map. When a model or neural network is used to generate the keypoint probability map in step 1, the model or neural network may learn to implicitly perform preprocessing to obtain better keypoints during the optimization process.
[0126] In the described examples, images were synthetically transformed (using random homography matrices) to create the training set. However, the method can also be used to train on real data pairs. For example, the ground truth homography matrix H can then be determined by a human observer or in other ways.
[0127] A network trained using the techniques disclosed in this article can predict the locations of stable points of interest called GLAMpoints on full-size grayscale images. In the following, examples of the generation of the training set used and the training process are disclosed. Since standard convolutional network architectures can be used, we only briefly discuss these architectures at the end.
[0128] The inventors have conducted a study that focused on digital fundus images of the human retina, which are widely used to diagnose various ocular diseases such as diabetic retinopathy (DR), glaucoma, and age-related macular degeneration (AMD) [37,47]. For retinal images acquired during the same session and presented with small overlaps, registration can be used to create mosaics depicting larger areas of the retina. Through image stitching, ophthalmologists can display the retina in a large picture, which helps them make diagnosis and treatment plans. In addition, stitching of retinal images taken at different times has been shown to be important for monitoring the progression or identification of ocular diseases. More importantly, registration applications have been explored in ocular laser treatments for DR. They allow real-time tracking of blood vessels during surgery to ensure accurate application of lasers on the retina and minimize damage to healthy tissue.
[0129] Stitching usually relies on extracting repeatable points of interest from the images and searching for transformations associated with them. Therefore, keypoint detection is the most critical stage of this pipeline as it determines all further steps and thus the success of the registration.
[0130] At the same time, classical feature detectors are generic and hand-optimized for outdoor, in-focus, low-noise images with sharp edges and corners. They often fail to handle medical images, which may be upscaled and distorted, noisy, not guaranteed to be in focus, and depict soft tissue without sharp edges (see Figure 3 ).
[0131] Figure 3 Examples of images from the slit lamp dataset are shown, demonstrating challenging conditions for registration: A) Hypovascularization and overexposure leading to weak contrast and corners, B) Motion blur, C) Focus blur, D) Acquisition artifacts and reflections.
[0132] On these images, traditional methods perform suboptimally, requiring more complex optimizations in subsequent steps of registration, such as Random Sample Consensus (RanSaC)
[20] , bundle adjustment
[38] , and Simultaneous Localization and Mapping (SLAM)
[18] techniques. In these cases, supervised learning methods fail or are not applicable due to the lack of ground truth for feature points.
[0133] In the present disclosure, a method for learning feature points in a semi-supervised manner is disclosed. Learned feature detectors outperform heuristic-based methods, but they are usually optimized for repeatability, and therefore they may perform poorly in the final matching process. In contrast, according to the present disclosure, key points called GLAMpoints can be trained for final matching capabilities, and when associated with, for example, scale-invariant feature transform (SIFT) descriptors, they outperform the state-of-the-art in terms of matching performance and registration quality.
[0134] Figure 1 An example of detected keypoints and matching of keypoints between a pair of slit lamp images is shown. The first column shows the original image, while the second column describes the pre-processed data. The detected points are shown in white. White lines represent true positive matches, while black lines represent false positive matches.
[0135] like Figure 1 As shown, GLAMpoints (shown in row A) produces more correct matches than SIFT (shown in row B). Feature point based registration is inherently non-differentiable due to point matching and transformation estimation. Using reinforcement learning (RL) can avoid this problem by assuming that the detected points are decisions in the classical RL sense. It makes it possible to directly use the key performance metric (i.e. matching ability) to train a convolutional neural network (CNN) specialized for a specific image modality.
[0136] The trained network can predict the locations of stable interest points on full-size grayscale images. These interest points are referred to as "GLAMpoints" in this article. The generation method of the training set and the training process are disclosed below. Since the standard convolutional network architecture can be used, only this architecture is briefly discussed at the end.
[0137] For example, a training set from the field of ophthalmology for laser treatment, namely slit lamp fundus videos, was selected. In this application, real-time registration is used to accurately ablate retinal tissue. The exemplary training dataset contains images from a base set of 1336 images with different resolutions, ranging from 300 pixels to 700 pixels by 150 pixels to 400 pixels. These images were taken from multiple cameras and devices to cover the huge variation of imaging modalities. They come from ophthalmological examinations of 10 different patients, who are either healthy or suffering from diabetic retinopathy. The full-resolution samples are scaled to 256×256 pixels by padding zeros or random cropping on larger images. This size reduction is performed to speed up the training process and increase the number of images. However, it should be noted that the above dimensions and the content of the training set are provided here only as non-limiting examples.
[0138] Let B be a base image set of size H × W. At each step i, the image is transformed by applying two separate, randomly sampled homography matrices g i , g′ i From the original image B i Generate image pair I i , I′ i Image I i and I' i Therefore, according to the homography matrix Related. The homography matrix generation method is elaborated in detail elsewhere in this specification. On top of the geometric transformations, standard data augmentation can be used: Gaussian noise, changes in contrast, illumination, gamma, motion blur, and inverse of the image. A subset of these appearance transformations can be randomly selected for each image I and I'. In some embodiments, different geometric and appearance transformations are applied to the base image at each step so that the network does not see any identical image pair twice.
[0139] To train a feature point detector, some features from classical reinforcement learning (RL) can be used. RL focuses on estimating the probability of actions in an environment to maximize the reward over multiple steps. Feature point detection can be viewed as taking a single action at each location in the image, i.e. selecting it as a feature point or background. The learning function can be defined as
[0140] f θ (I)→-S,
[0141] Where S represents the pixel-level feature point probability map. In the absence of direct ground truth for keypoint locations, a reward can be calculated instead. This reward can be based on matching success. This matching success can be calculated after the classic matching step in computer vision.
[0142] Training can be done as follows:
[0143] 1. Given a pair of images I∈R H×W and I'∈R H×W With the basic truth homography matrix H = H I,I’ Related, the model can provide a score map for each image:
[0144] S=f θ (I) and S' = f θ (I 0 ).
[0145] 2. The standard non-differentiable NonMax-Supression (NMS) can be used to extract the locations of the points of interest on the two score maps using a window of size w.
[0146] 3. 128 root SIFT feature descriptors can be calculated for each detected keypoint.
[0147] 4. A brute force matcher can be used to match keypoints in image I with keypoints in image I' and vice versa, such as [1]. For example, only matches found in both directions are retained.
[0148] 5. Check the match against the ground truth homography matrix H. If the corresponding keypoint x in image I falls into I after applying H 0 Point x in 0 The match is defined as a true positive if the neighborhood of is greater than . This is expressed as:
[0149] ||H*xx′||≤ε
[0150] where ε can be chosen to be, for example, 3 pixels. Let T denote the set of ground truth matches.
[0151] In the classic RL framework, if a given action is taken — i.e. a feature point is selected — and ends up in the set of true positive points, it receives a positive reward. All other points / pixels receive a reward of 0. The reward matrix R can then be defined, for example, as follows:
[0152]
[0153] This yields the following loss function:
[0154] L Simple =(fθ (I)-R) 2 .
[0155] However, a major drawback of this formulation is that there is a large class imbalance between positive reward points and empty reward points, where the latter may dominate, especially in the first stage of training. Given a reward R with an almost zero value, f θ may converge to zero output. Hard mining has been shown to improve the training of descriptors
[37] . Negative hard mining of false positive matches may also improve the performance of our method, but has not been studied in this work. Instead, to counteract the imbalance, sample mining can be used: all n true positive points and an additional n can be randomly sampled from the set of false positives. Backpropagation can be performed through the 2n true positive feature points and the mined false positive key points. In some embodiments, backpropagation is performed only through the 2n true positive feature points and the mined false positive key points. If there are more true positives than false positives, the gradients may be backpropagated through all found matches. This mining can be mathematically formulated as a mask M that is equal to 1 at the locations of the true positive key points and the locations of a subset of the mined feature points, and equal to 0 otherwise. Therefore, the loss can be formulated as follows:
[0156]
[0157] Here, the symbol · represents element-by-element multiplication.
[0158] Figure 2A The steps of training a pair of images I and I' at epoch i for a particular base image B are shown. Figure 2B An example of loss calculation is shown. Figure 2C A schematic diagram of Unet-4 is shown.
[0159] Figure 2A An overview of the training steps is given in . Importantly, it can be observed that only step 1 is differentiable with respect to the loss. Learning is done directly on the reward, which is the result of a non-differentiable action, without supervision. It can be noted that the descriptor used is a version of SIFT without rotation invariance. The reason is that on slit lamp images, rotation dependent SIFT detectors / descriptors outperform SIFT detectors / descriptors with rotation invariance. The purpose of the evaluation is to study only the detector and therefore rotation dependent SIFT descriptors are used for consistency.
[0160] A standard 4-level deep Unet
[34] with a final sigmoid activation is used to learn f θ It contains 3×3 convolutional blocks with batch normalization and rectified linear unit (ReLU) activation (see Figure 2C). Since the task of keypoint detection is similar to pixel-wise binary segmentation (interest point class or not), the Unet model looks promising due to its past success in binary and semantic segmentation tasks.
[0161] In the following, a description of the test datasets and the evaluation protocol is provided. Prior art detectors are quantitatively and qualitatively compared with the disclosed technique (eg, GLAMpoints).
[0162] In this study, the trained models were tested on multiple fundus image datasets and natural images. For medical images, two datasets were used:
[0163] First, the "slit lamp" dataset: a set of 206 frame pairs are randomly selected from retinal videos of 3 different patients as test samples, with sizes ranging from 338 pixels (px) to 660 pixels × 190 pixels to 350 pixels. The examples are shown in Figure 3 , and it can be seen that they present multiple artifacts, making it a particularly challenging dataset. The pairs were selected to have overlaps ranging from 20% to 100%. They are related by affine transformations and rotations of up to 15 degrees. All image pairs were manually annotated with at least 5 corresponding points using a dedicated software tool (Omniviewer1). These annotations were used to estimate the ground truth homography matrix associated with these pairs. To determine the correct matches, experts have verified each estimated homography matrix and corrected incorrect matches.
[0164] Secondly, the FIRE dataset
[22] . This is a publicly available retinal image registration dataset with ground truth annotations. It consists of 129 retinal images, forming 134 image pairs. The original images of 2912 × 2912 pixels are scaled down to 15% of the original size to have a resolution similar to the training set. Examples of such images are Figure 5 Presented in.
[0165] For the fundus image test, as preprocessing, we isolated the green channel and applied adaptive histogram equalization and bilateral filter to reduce noise and enhance the appearance of edges. The effect of preprocessing can be seen in Figure 1 . According to DeZanet et al.
[46] , this procedure leads to improved results during detection and description. However, GLAMpoints performs well on both original and preprocessed images. Therefore, to compare the performance, we present the evaluation results for both cases.
[0166] In addition, keypoint detection and matching is performed on natural images. For this purpose, the following are used: Oxford dataset
[30] , EF dataset
[48] , webcam dataset [40, 24] and viewpoint dataset
[45] , resulting in a total of 195 pairs. These datasets may result in a total of 195 pairs. More details are given elsewhere in this specification.
[0167] The evaluation criteria considered include the following: repeatability, average number of keypoints detected, and success rate. These are described in detail below.
[0168] 1. Repeatability describes the percentage of corresponding points that appear in the two images:
[0169]
[0170] For images I and I', the detected point sets are denoted as P and P', respectively, with corresponding key points x and x'. I,I' is the ground truth homography matrix relating the reference image to the transformed image. ε is the distance cutoff between two points (set to 3 pixels).
[0171] 2. Average number of detected keypoints per image: As proposed in
[28] , matches are found using the nearest neighbor distance ratio (NNDR) strategy: two keypoints are matched if the descriptor distance ratio between the first and second nearest neighbors is below a certain threshold t. In terms of matching performance, the following metrics are evaluated:
[0172] (a) AUC, which is the area under the receiver operating characteristic (ROC) curve (created by varying the value of t). This allows the discriminatory power of each method to be assessed, in line with [15, 43, 42].
[0173] (b) Matching score, defined as the ratio of correct matches in the shared viewpoint region to the total number of features extracted by the detector
[29] . This metric allows evaluating the performance of the entire feature pipeline.
[0174] (c) Coverage, which measures the coverage of an image by correctly matched keypoints. To compute it, the technique proposed in [7] is used: a coverage mask is generated from the correctly matched keypoints, each of which is added with a disk of fixed radius (25px).
[0175] The homography matrix relating the reference to the transformed image is computed by applying the RanSaC algorithm to remove outliers from the detected matches. All the above indicators refer to matching performance.
[0176] 3. Success rate: We evaluate the registration quality and accuracy achieved by each detector, as in [13, 41]. To this end, we compare the reprojection error of six fixed points of the reference image, denoted as
[0177] c i , i = {1, ..., 6},
[0178] On the other hand, for each image pair for which the homography matrix was found, the quality of the registration was evaluated using the median error MEE, the maximum error MAE, and the root mean square error RMSE of the correspondence. Using these metrics, we defined different thresholds on MEE and MAE that define “acceptable” and “inaccurate” registrations. For the slit lamp dataset, we classified an image pair as “acceptable” registration when (MEE < 10 and MAE < 30) and “inaccurate” registration otherwise. On the other hand, for the full retinal images from the FIRE dataset
[22] , “acceptable” registration corresponds to (MEE < 1:50 and MAE < 10). The values of the thresholds were found empirically after looking at the results.
[0179] Finally, the success rate for each class is calculated, which is equal to the percentage of image pairs whose registrations fall into each class. These metrics may be considered the most important quantitative evaluation criteria for overall performance in real-world settings.
[0180] To evaluate the performance of detectors relative to the root SIFT descriptor, the matching ability and registration quality are compared with well-known detectors and descriptors. Among them, Truong et al.
[39] demonstrated that SIFT [2], root SIFT [28, 8], KAZE [6], and LIFT
[44] perform well on fundus images. In addition, the method is compared with other CNN-based detector descriptors: LF-NET
[31] and SuperPoint
[16] . The inventors used their implementations of LIFT (pre-trained on Picadilly), SuperPoint, and LF-NET (pre-trained on indoor data, which may produce better results on fundus images compared to the versions pre-trained on outdoor data) and OpenCV implementations for SIFT and KAZE. The rotation-dependent version of the root SIFT can be used because it has better performance on our test set compared to the rotation-invariant version.
[0181] Training of GLAMpoints was performed using Tensorflow[4] with a mini-batch size of 5 and the Adam optimizer
[25] with default parameters, learning rate = 0.001 and β = (0.9, 0.999). The model was cross-validated with 4-folds and showed similar results with a standard deviation of 1% in the success rate. It was then retrained on the entire dataset consisting of 8936 base images for 35 epochs. GLAMpoints(NMS10) was trained and tested using an NMS window equal to 10px. It must be noted that other NMS windows can be applied to obtain similar performance.
[0182] Table 1: Success rate (%) for each registration class for each detector on 206 images of the Slit Lamp dataset. Acceptable registration is defined as (MEE < 10 and MAE < 30). The best results are in bold.
[0183] a) Non-preprocessed data
[0184]
[0185] b) Preprocessing data
[0186]
[0187]
[0188] Table 1 shows the registration success rates evaluated on the Slit Lamp dataset. Without preprocessing, most of the detectors used for comparison perform worse than on preprocessed images. In contrast, the proposed model performs well on unprocessed images. This is highlighted in Table 1, where the success rates of acceptable registration (A) for SIFT, KAZE, and SuperPoint drop by 20% to 30% between preprocessed and non-preprocessed images, while GLAMpoints as well as LIFT and LF-NET only show a drop of 3% to 6%. Moreover, LF-NET, LIFT, and GLAMpoints detect a stable average number of keypoints independent of preprocessing (around 485 for the first and around 350 for the latter), while the other detectors decrease by a factor of 2.
[0189] In the tested embodiments, GLAMpoints may outperform KAZE, SIFT, and SuperPoint by at least 18% in terms of acceptable registration success rate. In the same category, it outperforms LF-NET and LIFT by 3% and 5% on raw data and 7% and 8% on pre-processed data, respectively. It is also important to note that if LF-NET is trained on fundus images, while LF-NET may achieve similar results for specific metrics and datasets, its training process utilizes image pairs with their relative poses and corresponding depth maps, which is extremely difficult (if not impossible) for fundus images.
[0190] Furthermore, independent of preprocessing, the GLAMpoints model has the smallest MEE and RMSE for inaccurate registrations (I) as well as for globally successful registrations. The MEE and RMSE for acceptable registrations are similar to within 1 pixel for all detectors. Details on the MEE and RMSE for each class can be found elsewhere in this note. Robust results for GLAMpoints independent of preprocessing show that while the detector performs as well or better than SIFT on high-quality images, its performance does not degrade with low-quality, poorly contrasted images.
[0191] Although SIFT extracts a large number of keypoints (205.69 on average for unprocessed images and 431.03 on average for preprocessed images), they appear in clusters. As a result, the close localization of interest points leads to a large number of rejected matches and few true positive matches even though the reproducibility is relatively high (many possible valid matches) due to the nearest neighbor distance ratio (NNDR). Figure 4A and Figure 4B It leads to small M: scores and AUC. Figure 4A and Figure 4B As shown, with similar repeatability values, our method extracts widely distributed points of interest and is trained on their matching ability (highest coverage), resulting in more true positive matches (second highest M: score and AUC).
[0192] Figure 4A and Figure 4B A summary of the detector / descriptor performance metrics evaluated on the 206-pair slit-lamp dataset is provided. Figure 4A shows the results for non-preprocessed data, while Figure 4B The results of preprocessing the data are shown.
[0193] It can also be noted that even with a relatively small coverage, SuperPoint scores with the highest M: score and AUC (see Figure 4A and Figure 4B). However, in this case, the M: score and AUC are artificially inflated because Super-Point detects very few keypoints (35, 88 and 59, 21 on average for non-preprocessed and preprocessed images, respectively) and has one of the lowest repeatability, resulting in almost no possible true positive matches. Its matching performance then looks high, even though it does not find many true positive matches. This is evidenced by the large number of inaccurate and failed registrations (48.54 and 17.48% for the unprocessed data, compared to 51.46 and 7.77% for the preprocessed images, Table 1).
[0194] Finally, it is worth noting that LF-NET has very high reproducibility scores (highest for raw data, second highest for preprocessed images), but its M.score and AUC are at the bottom of the ranking ( Figure 4A and Figure 4B ). This can be explained by the training of the LF-NET detector, which prioritizes the repetitiveness of matching targets.
[0195] Table 2: Success rate (%) of each detector on non-preprocessed images of the FIRE dataset. Acceptable registration is defined as having (MEE < 1:5 and MAE < 10). The best results are in bold.
[0196]
[0197]
[0198] GLAMpoints was also evaluated on the FIRE dataset. Since all images presented good quality with high contrast vascularization, no preprocessing was applied. Table 2 shows the success rate of the registration. The mean and standard deviation of MEE and RMSE for each class can be found elsewhere in this note. The proposed method performs well in terms of success rate and global accuracy of non-failed registrations. It is interesting to notice that there is a 41.04% gap in the acceptable registration success rate between GLAMpoints and SIFT. Since both use the same descriptor (SIFT), this difference is explained solely by the detector. In fact, as Figure 5 As shown, while SIFT detects a limited number of keypoints that are densely localized only on the vessel tree and image boundaries, GLAMpoints (NMS10) extracts points of interest over the entire retina, including challenging regions such as the fovea and avascular areas, leading to a substantial increase in the number of correct matches.
[0199] Figure 5Detected interest points and corresponding matches for a pair of images from the FIRE dataset without preprocessing are shown. Black dots represent detected interest points. White lines represent true positive matches, while black lines represent false positive matches. Row A) shows interest points and matches achieved using GLAMpoints, while row B) shows interest points and matches achieved using SIFT.
[0200] Although GLAMpoints (NMS10) outperforms all other detectors, LIFT and SuperPoint also perform well on the FIRE dataset. In fact, this dataset presents well-defined corners on contrasting vascular trees. LIFT manages to extract keypoints distributed over the entire image, and SuperPoint is trained to detect corners on synthetic primitive shapes. However, as demonstrated on the Slit Lamp dataset, SuperPoint's performance deteriorates severely on images where features are less clear.
[0201] For the proposed GLAMpoints (NMS10) and SIFT methods, Figure 5 Matches between pairs of images from the FIRE dataset are shown. More examples of matches for all detectors can be found elsewhere in this note.
[0202] In certain embodiments, the feature detectors disclosed herein can be used in a system or method that can create a mosaic from a fundus slit lamp video. To this end, keypoints and descriptors are extracted on each frame and the homography matrix between consecutive images is estimated using RanSaC. The image is then warped according to the calculated homography matrix. Using 10 videos containing 25 to 558 images, a mosaic is generated by registering consecutive frames. The average number of frames before registration fails (due to missing extracted keypoints or matches between a pair of images) is calculated. In these ten videos, the average number of registered frames before failure is 9.98 for GLAMpoints (NMS15) and 1.04 for SIFT. An example of such a mosaic is shown in Figure 6 Presented in.
[0203] Figure 6 Mosaics formed from registration of consecutive images until failure are shown. A) GLAMpoints, non-preprocessed data, 53 frames, B) SIFT, preprocessed data, 34 frames, C) SIFT, non-preprocessed images, 11 frames.
[0204] From the same video, SIFT fails after 34 frames of preprocessing and only after 11 frames on the original data, while GLAMpoints successfully registers 53 consecutive images with no visual errors. It may be noted that the mosaics are created by frame-to-frame matching without bundle adjustment. The same hybrid approach as described in
[46] is used.
[0205] The runtime of detection is calculated on 84 pairs of images with a resolution of 660 pixels × 350 pixels. The GLAMpoints architecture runs on a GeForce GTX GPU, while NMS and SIFT use the CPU. The mean and standard deviation of the runtime of GLAMpoints (NMS10) and SIFT are shown in Table 3.
[0206] Table 3: Average running time [ms] for detecting an image using our detector and SIFT detector
[0207]
[0208] The results for natural images are computed using GLAMpoints trained on slit-lamp images. On global natural images, GLAMpoints achieves 75.38% success rate for acceptable registration, compared to 85.13% and 83.59% for the best performing detector (SIFT with rotation invariance) and SuperPoint, respectively. It also achieves state-of-the-art results in terms of AUC, M:score and coverage, obtaining second, second and first place, respectively. In terms of repeatability, GLAMpoints obtains the second-to-last position after SIFT, KAZE and LF-NET, although it successfully registers more images than the latter, again indicating that repeatability is not the most appropriate metric for measuring detector performance. The details of the metrics are given elsewhere in this description. Finally, it may be noted that the outdoor images of this dataset are quite different from the medical fundus images on which GLAMpoints was trained, suggesting a strong generalization property.
[0209] The proposed method uses deep RL to train a learnable detector, referred to as GLAMpoints in this description. For example, the detector can outperform the state of the art in image matching and registration of medical fundus images. Experiments show that (1) the detector is directly trained for matching capabilities associated with specific descriptors, and only a part of the pathway is differentiable. Most other detectors are designed for reproducibility, which can be misleading as keypoints may be repeated but not suitable for matching purposes. (2) It can be trained using only synthetic data. This eliminates the need for time-consuming manual annotation and provides flexibility in the amount of training data used. (3) The training method is domain-flexible and while optimized for success on medical fundus images, it can also be applied to other types of images. (4) Compared with other state-of-the-art detectors, the trained CNN is able to detect more keypoints and thus achieve correct matching in low-texture images, which do not present many corners / features. As a result, it is found that no explicit pre-processing of the images is required. (5) Any existing feature descriptor can potentially be improved by the corresponding trained detector.
[0210] In an alternative embodiment, rotation invariant descriptors can be computed along with the keypoint locations. Both detection and description can be trained end-to-end using similar methods. In addition, while the current experiments were conducted using the U-Net CNN architecture, other CNN architectures can also be applied, which may provide better performance than U-Net (UNet) in some cases.
[0211] In the following, additional details about the training method are disclosed. It should be understood that these details are to be regarded as non-limiting examples.
[0212] A performance comparison between SIFT descriptors with / without rotation invariance was performed. The Greedy Learned Accurate Matching Points (GLAMpoints) detector was trained and tested in association with the Scale Invariant Feature Transform (SIFT) descriptor rotation-dependently, since the SIFT descriptor without rotation invariance performs better than the rotation invariant version on fundus images. Details of the metrics evaluated on the pre-processed slit-lamp dataset for both versions of SIFT descriptors are disclosed in Table 4.
[0213] Table 4: Metrics calculated for the 206 pairs of preprocessed slit lamp dataset. The best result in each category is in bold
[0214]
[0215] An exemplary method for homography generation is outlined below. This homography generation can be performed to generate training data comprising image pairs from data comprising a single image. Let B be a base image set of size H×W for training. At each step i, the homography transformation g is transformed by applying two separate, randomly sampled homography transformations g i , g i 'From the original image B i Generate image pair I i ,I i '. Each of these homography matrix transformations is a combination of rotation, shear, perspective, scaling, and translation elements. Other combinations of transformation types may be used instead. Exemplary minimum and maximum values for the parameters are given in Table 5.
[0216] Table 5: Example parameters for random homography matrix generation during training
[0217]
[0218] Figure 7 A system 701 for training a model for feature point detection is shown. The system 701 may include a control unit 705, a communication unit 704, and a memory 706. The control unit 705 may include any processor or multiple cooperating processors. The control unit 705 may alternatively be implemented by a dedicated electronic circuit. The communication unit 704 may include any type of interface to connect peripheral devices, such as a camera 702 or a display 703, and may include a network connection, such as for data exchange and / or control of external devices. In an alternative embodiment, the camera 702 and / or the display may be incorporated into the system 701 as a single device. In an alternative embodiment, the image captured by the camera 702 may be stored in an external database (not shown) and subsequently transmitted to the communication unit 704. Similarly, the data generated by the system 701 may be stored in an external database before being displayed on the display. To this end, the communication unit 704 may be connected to a data server, for example, via a network.
[0219] The control unit 705 controls the operation of the system 701. For example, the control unit 705 executes the code stored in the memory 706. The memory 706 may include any storage device, such as RAM, ROM, flash memory, disk, or any other volatile or non-volatile computer-readable medium, or a combination thereof. For example, computer instructions may be stored in a non-volatile computer-readable medium. The memory 706 may also include data 707, such as an image 709, a model 708, and any other data. The program code may be divided into functional units or modules. However, this is not a limitation.
[0220] In operation, the control unit 705 may be configured to retrieve a plurality of images from the camera 702 or retrieve images captured by the camera 702 from an external storage medium and store the images 709 in the memory 706 .
[0221] The system may include a training module 711 for training the model 708. The model may include a neural network, such as a convolutional neural network, or other models, such as statistical models, whose model parameters may be adjusted by a training process performed by the training module 711.
[0222] The training module 711 may be configured to perform a training process including feeding input values to the model 708 , evaluating output values output by the model 708 in response to the input values, and adjusting model parameters of the model 708 based on the evaluation results.
[0223] The communication unit 704 may be configured to receive a plurality of images under the control of the control unit 705. These images 709 may be stored in the memory 706.
[0224] Optionally, the control unit 705 controls the camera 702 (internal or external camera) to generate images and send them to the communication unit 704 and store them in the memory 706 .
[0225] The system 701 may include a pre-processing module 717 that generates a pair of images of a single image. For example, the pre-processing module 717 is configured to generate a random transformation and generate a second image by applying the transformation to the first image. Alternatively, the pre-processing module 717 may be configured to generate two random transformations and generate a first image from a specific image by applying a first random transformation, and generate a second image from the same specific image by applying a second random transformation to the specific image. The type of random transformation to be generated may be carefully configured to correspond to typical motions that occur in a specific application domain. In an alternative embodiment, the two images generated by the camera are manually registered so that a transformation between the two images becomes feasible.
[0226] The processor may obtain such a pair of images including a first image and a second image from the memory 706 .
[0227] The system 701 may include a score map generator 713 configured to generate a first score map for the first image and a second score map for the second image using the model. That is, the score map generator 713 may be configured to perform optional preprocessing operations (normalization, etc., or other kinds of preprocessing). However, it is observed that good results are obtained without preprocessing. The resulting image may be provided as an input to the model 708. The corresponding output generated by the model 708 in response to the input may include another image (score map) in which each pixel is associated with a probability that the point is a suitable point of interest, so as to register the image to the other image. It can be seen that the score map generator 713 may be configured to process the first image and the second image separately (one is independent of the other), i.e., without using any knowledge about the content of the other image.
[0228] The system 701 may also include an interest point selector 712. The interest point selector 712 may be configured to select a first plurality of interest points in the first image based on the first score map, and to select a second plurality of interest points in the second image based on the second score map. Likewise, the processing of the two images may be separate independent processes. For example, the point with the largest score on the score map may be selected as the interest point. In some embodiments, the maximum and / or minimum distance between adjacent interest points may be imposed algorithmically. For example, in each N×M pixel block, only the pixel with the highest score is selected. Other algorithms for influencing the maximum and / or minimum distance between adjacent points may be envisioned.
[0229] The system may include a matching module 714. The matching module 714 may be configured to process images in pairs. Specifically, the matching module 714 matches a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest in pairs. In other words, the points of interest in the first image are matched with the points of interest in the second image. For example, feature descriptors are calculated at the points of interest in both images, and a similarity measure is calculated between the feature description of the point of interest in the first image and the feature description of the point of interest in the second image. The pair with the highest similarity measure may be selected as a matching pair. Other matching methods are contemplated.
[0230] The system may include a verification module 715 configured to check the correctness of the pairwise matches generated by the matching module 714. To this end, the verification module 715 may access ground truth information. For example, when a pair of images is artificially generated using different (affine) transformations of the same image, the transformation contains ground truth matches for the pair of image points. Therefore, applying a transformation to a point in the first image should produce a corresponding matching point in the second image. The distance (e.g., Euclidean distance) between a matching point in the second image and the ground truth transformation of the point in the first image can be regarded as the error of the matching point. The reward may be based on such an error metric: the lower the error, the higher the reward, and vice versa. In this way, for each matching point, a reward can be calculated. This may produce a reward map or reward matrix. Therefore, as found by the matching module 714, the reward for the point of interest in the first image is related to the success of the match of the matching point in the second image.
[0231] The system may include a combining module 716 for combining or comparing the score map and the reward map. That is, if the score of the point generated by the score map generator 713 is high and the reward for the point generated by the verification module 715 is also high (a "true positive"), the combining module 716 may determine a value to strengthen the model 708 to identify similar points of interest in the future. On the other hand, if the score of the point generated by the score map generator 713 is high, but the reward for the point generated by the verification module 715 is low (a "false positive"), the combining module 716 may determine a value to strengthen the model 708 to avoid identifying similar points of interest in the future. In some embodiments, the combining module 716 is configured to determine rewards only for a subset of false positives. For example, if the number of true positives in the image is M, at most M false positives are considered. For example, all values may be added together to calculate the total reward function.
[0232] The training module 711 can be configured to update the model based on the results of the combination or comparison. This is the training step of the model, which is known in the art itself. The exact parameters to be updated depend on the type of model used, such as nearest neighbor, neural network, convolutional neural network, U-net or other types of models.
[0233] The point of interest selector 712 may be configured to impose a maximum limit on the distance from any point in the image to the nearest one of the points of interest.
[0234] The matching module 714 may be configured to perform pairwise matching based on similarities between features detected at a first point of interest in the first image and features detected at a second point of interest in the second image.
[0235] The matching module 714 may be configured to perform matching in a first direction by matching the first interest point with a second interest point of the second interest point having features most similar to the features at the first interest point. The matching module 714 may be configured to perform further matching in a second direction by matching the second interest point with the first interest point of the first interest point having features most similar to the features at the second interest point. The matching module 714 may be configured to discard any and all matches that do not match in both directions.
[0236] The verification module 715 may be configured to reward the interest points that are successfully matched according to the basic truth data through a reward graph, and not reward the interest points that are not successfully matched according to the basic truth data.
[0237] The combining module 716 may be configured to combine or compare only the score map and the reward map of the point of interest.
[0238] The combining module 716 can be configured to balance some true positive matches with some false positive matches by (possibly randomly) selecting false positive matches and combining or comparing the score maps and reward maps for the selections only for the false positive matches and the true positive matches, where the true positive matches are points of interest that pass the correctness check and the false positive matches are points of interest that do not pass the correctness check. Instead of making a random selection, the combining module 716 can be configured to, for example, select the false positive match with the lowest reward map value.
[0239] The combining module 716 may be configured to calculate the sum of squared differences between the score map and the reward map.
[0240] The communication unit 704 may be used to output interesting information such as progress information or the value most recently output by the combination module through the display 703 .
[0241] Figure 8 A method of training a model for feature point detection is shown. In step 801, the preprocessing step may include generating image pairs from original images using, for example, random transformations, and storing the random transformations as ground truth information for the image pairs thus created.
[0242] Step 802 may include obtaining a first image and a second image of a pair of images. Step 803 may include generating a first score map of the first image and a second score map of the second image using the model. Step 804 may include selecting a first plurality of interest points in the first image based on the first score map, and selecting a second plurality of interest points in the second image based on the second score map.
[0243] Step 805 may include pair-matching a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest. Step 806 may include checking the correctness of the pair-wise matches based on a ground truth transformation between the first image and the second image to generate a reward map. Step 807 may include combining or comparing the score map and the reward map. Step 808 may include updating the model based on the result of the combination or comparison. At step 809, the process may check whether more training is needed. If more training is needed, the process may start from step 802 by obtaining the next pair of images.
[0244] Fig. 9 A system 901 for registering a first image to a second image is shown. The device includes a control unit 905, a communication unit 904, and a memory 906 to store data 907 and program code 910. The considerations of hardware and alternative implementation options related to the system 701 for training a model for feature point detection above also apply to the system 901 for registering a first image to a second image. For example, the program code 910 can be implemented by a dedicated electronic circuit instead. In addition, the camera 902 and the display 903 can be optional external devices, or can be integrated into the system 901 to form an integrated device. The system 901 includes a control unit 905 (e.g., at least one computer processor), a communication unit 904 for communicating with, for example, the camera 902 and / or the display 903, and a memory 906 (any kind of storage medium) including program code 910 or instructions for causing the control unit to perform certain steps. The memory 906 is configured to be able to store a trained model 908 and an image 909. For example, an image received from the camera 902 can be stored in the memory 906.
[0245] The system may include an acquisition module 917 configured to acquire a first image and a second image, such as two images captured and received from the camera 902, such as from the stored images 909. The system may also include a score map generator 913 configured to generate a first score map for the first image and a second score map for the second image using the trained model 908. For example, after optional preprocessing (e.g., normalization), the two images may be input to the model 908, and the output generated in response to each image is a score map.
[0246] The system may include an interest point selector configured to select a first plurality of interest points in the first image based on the first score map, and to select a second plurality of interest points in the second image based on the second score map. The selection may be performed similarly to the selection performed by the interest point selector 712, using feature descriptors and similarities between feature descriptions of the interest points in the two images. The system may include a matching module 914 configured to pairwise match a first interest point in the first plurality of interest points with a second interest point in the second plurality of interest points.
[0247] The system may also include a registration module 918. The registration module 918 may be configured to determine a morphological transformation based on the matched points of interest. The registration module 918 may be configured to map each point in the first image to a corresponding point in the second image based on the matched points of interest. For example, an affine or non-affine transformation may be determined based on the matching points. This may involve another parameter fitting procedure. However, the manner in which such a transformation is generated from a set of matching points is known in the art and is not described in detail herein. For example, the transformation may be applied to the first image or the second image. The communication unit 904 may be configured to output the transformed image through the display 903.
[0248] The point of interest selector 912 may be configured to impose a maximum limit on the distance from any point in the image to the nearest one of the points of interest.
[0249] The matching module 914 may be configured to perform pairwise matching based on similarities between features detected at a first point of interest in a first image and features detected at a second point of interest in a second image.
[0250] The matching module 914 may be configured to perform matching in a first direction by matching the first interest point with the second interest point of the second interest point having the most similar features to the features at the first interest point. The matching module 914 may also be configured to perform matching in a second direction by matching the second interest point with the first interest point of the first interest point having the most similar features to the features at the second interest point. For example, the matching module 914 may be configured to discard matches that do not match in both directions.
[0251] Fig.10An exemplary method of registering a first image to a second image is shown. Step 1002 may include obtaining a first image and a second image, such as two images captured from a camera. Step 1003 may include generating a first score map for the first image and a second score map for the second image using an appropriately trained model (e.g., a model trained by the system or method disclosed herein). Step 1004 may include selecting a first plurality of interest points in the first image based on the first score map, and selecting a second plurality of interest points in the second image based on the second score map. Step 1005 may include pair-matching a first interest point in the first plurality of interest points with a second interest point in the second plurality of interest points, for example, based on feature descriptors associated with the interest points. Optionally, step 1006 may include registering the pair of images by generating a morphological transformation based on the matched interest points. Optionally, step 1006 includes applying a transformation to the images and outputting the transformed image using a display or storing the transformed image. Alternatively, step 1006 includes stitching the pair of images together to use the matched interest points and storing or displaying the stitched image.
[0252] Some or all aspects of the present invention may be suitable for implementation in the form of software, specifically a computer program product. A computer program product may include a computer program stored on a non-transitory computer readable medium. In addition, the computer program may be represented by a signal (such as an optical signal or an electromagnetic signal) carried by a transmission medium such as an optical cable or air. The computer program may be in part or in whole in the form of source code, object code, or pseudocode suitable for execution by a computer system. For example, the code may be executed by one or more processors.
[0253] The examples and embodiments described herein are intended to illustrate rather than limit the present invention. As defined by the appended claims and their equivalents, those skilled in the art will be able to design alternative embodiments without departing from the spirit and scope of the present disclosure. Reference symbols in brackets in the claims should not be interpreted as limiting the scope of the claims. Items described as separate entities in the claims or specification may be implemented as a single hardware or software item that combines the features of the described items.
[0254] The following topics are exposed as items.
[0255] 1. A method for training a classifier for feature point detection, the method comprising:
[0256] Acquire a first image and a second image;
[0257] generating a first score map for the first image and a second score map for the second image using the classifier;
[0258] selecting a first plurality of interest points in the first image based on the first score map;
[0259] selecting a second plurality of interest points in the second image based on the second score map;
[0260] pair-matching a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest;
[0261] checking correctness of pairwise matches based on a ground truth transformation between the first image and the second image to generate a reward map;
[0262] combining or comparing score maps and reward maps; and
[0263] The classifier is updated based on the result of the combination or comparison.
[0264] 2. A method according to item 1, wherein selecting the plurality of interest points comprises imposing a maximum limit on the distance from any point in the image to the nearest one of the interest points.
[0265] 3. A method according to any preceding item, wherein pairwise matching is performed based on similarities between features detected at a first point of interest in the first image and features detected at a second point of interest in the second image.
[0266] 4. A method according to any preceding item, wherein matching is performed in the first direction by matching the first interest point with a second interest point of the second interest point having a feature most similar to that of the first interest point.
[0267] 5. The method according to item 4, wherein the matching is further performed in the second direction by matching the second interest point with a first interest point among the first interest points having a feature most similar to that of the second interest point.
[0268] 6. A method according to any preceding clause, wherein the reward map indicates a reward for successfully matched points of interest and no reward for unsuccessfully matched points of interest.
[0269] 7. A method according to any preceding clause, wherein combining or comparing comprises combining or comparing the score map and the reward map only for points of interest.
[0270] 8. A method according to any preceding clause, wherein combining or comparing comprises balancing some true positive matches with some false positive matches by selecting false positive matches (possibly randomly) and combining or comparing the score map and the reward map only for the selection of false positive matches and the true positive matches,
[0271] Among them, a true positive match is a point of interest that passes the correctness check, and a false positive match is a point of interest that fails the correctness check.
[0272] 8. A method according to any preceding clause, wherein combining or comparing comprises calculating the sum of squared differences between the score map and the reward map.
[0273] 9. A device for training a feature point detection classifier, the device comprising
[0274] a control unit, for example at least one computer processor, and
[0275] A memory comprising instructions for causing the control unit to perform the following steps:
[0276] Acquire a first image and a second image;
[0277] generating a first score map for the first image and a second score map for the second image using the classifier;
[0278] selecting a first plurality of interest points in the first image based on the first score map;
[0279] selecting a second plurality of interest points in the second image based on the second score map;
[0280] pair-matching a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest;
[0281] checking correctness of pairwise matches based on a ground truth transformation between the first image and the second image to generate a reward map;
[0282] combining or comparing score maps and reward maps; and
[0283] The classifier is updated based on the result of the combination or comparison.
[0284] 10. A method for registering a first image to a second image, the method comprising:
[0285] Acquire a first image and a second image;
[0286] Generate a first score map for the first image and a second score map for the second image using a classifier generated by any of the methods or apparatus of the preceding items;
[0287] selecting a first plurality of interest points in the first image based on the first score map;
[0288] selecting a second plurality of interest points in the second image based on the second score map;
[0289] A first point of interest in the first plurality of points of interest is matched in pairs with a second point of interest in the second plurality of points of interest.
[0290] 11. A method according to item 10, wherein selecting the plurality of interest points comprises imposing a maximum limit on the distance from any point in the image to the nearest one of the interest points.
[0291] 12. A method according to item 10 or 11, wherein pairwise matching is performed based on similarity between features detected at a first point of interest in the first image and features detected at a second point of interest in the second image.
[0292] 13. A method according to any one of items 10 to 12, wherein the matching is performed in the first direction by matching the first interest point with a second interest point of the second interest point having a feature most similar to that of the first interest point.
[0293] 14. The method according to item 13, wherein the matching is further performed in the second direction by matching the second interest point with a first interest point among the first interest points having a feature most similar to the feature of the second interest point.
[0294] 15. A device for registering a first image to a second image, the device comprising
[0295] a control unit, for example at least one computer processor, and
[0296] A memory comprising instructions for causing the control unit to perform the following steps:
[0297] Acquire a first image and a second image;
[0298] A classifier generated using any of the methods or apparatuses of the preceding items generates a first score map for the first image and a second score map for the second image;
[0299] selecting a first plurality of interest points in the first image based on the first score map;
[0300] selecting a second plurality of interest points in the second image based on the second score map;
[0301] A first point of interest in the first plurality of points of interest is matched in pairs with a second point of interest in the second plurality of points of interest.
[0302] Reference List
[0303] [1]OpenCV: cv::BFMatcher Class Reference.
[0304] [2]OpenCV: cv::xfeatures2d::SIFT Class Reference.
[0305] [3]Improving Accuracy and Efficiency of Mutual Information for Multi-modal Retinal Image Registration using Adaptive Probability DensityEstimation,Computerized Medical Imaging and Graphics,2013,37(7-8):597–606.
[0306] [4]M.Abadi,A.Agarwal,P.Barham,E.Brevdo,Z.Chen,C.Citro,GSCorrado,A.Davis,J.Dean,M.Devin,S.Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Man′e, and R. Mon ga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke V. Vasudevan, F. Vi′egas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, Tensor-Flow:Large-scale Machine learning on heterogeneoussystems, 2015, tensorflow.org's newsletter.
[0307] [5]A.Alahi、R.Ortiz&P.Vandergheynst,FREAK:Fast Retina Keypoint,Proceedings of the Conference on Computer Vision and Pattern Recognition,2012,Pages 510-517.
[0308] [6] P.F. Alcantarilla, A. Bartoli, and A.J. Davison, KAZE features. Lecture Notes in Computer Science, 2012, 7577 LNCS(6): 214–227.
[0309] [7] J. Aldana-Iuit, D. Mishkin, O. Chum, and J. Matas, In the saddle: Chasing fast and repeatable features, 23rd International Conference on Pattern Recognition, 2016, pp. 675-680.
[0310] [8] R. Arandjelovic and A. Zisserman, Three things everyone should know to improve object retrieval c.
[0311] [9] V. Balntas, E. Johns, L. Tang, and K. Mikolajczyk, PN-Net: Conjoined Triple Deep Network for Learning Local Image Descriptors, 2016, CoRR, abs / 1601.05030.
[0312]
[10] V. Balntas, K. Lenc, A. Vedaldi, and K. Mikolajczyk. HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors, Conference on Computer Vision and Pattern Recognition, 2017, pp. 3852-3861.
[0313]
[11] H. Bay, T. Tuytelaars, and L. Van Gool, SURF: Speeded up Robust Features, Lecture Notes in Computer Science, 2006, 3951 LNCS: 404-417.
[0314]
[12] P.C. Cattin, H. Bay, L.J.V. Gool, and G. Székely, Retina Mosaicing Using Local Features, Medical Image Computing and Computer-Assisted Intervention, 2006, pp. 185 - 192.
[0315]
[13] J. Chen, J. Tian, N. Lee, J. Zheng, T.R. Smith, and A.F. Laine, A Partial Intensity Invariant Feature Descriptor for Multimodal Retinal Image Registration, IEEE Transactions on Biomedical Engineering, 2010, 57(7): 1707–1718.
[0316]
[14] A.V. Cideciyan, Registration of Ocular Fundus Images, IEEE Engineering in Medicine and Biology Magazine, 1995, 14(1): 52–58.
[0317]
[15] A.L. Dahl, H. and K.S. Pedersen, Finding the Best Feature Detector-Descriptor Combination, International Conference on 3D Imaging, Modeling, Processing, Visualization and Transmission, 2011, pp. 318 - 325.
[0318]
[16] D. DeTone, T. Malisiewicz, and A. Rabinovich. SuperPoint: Self-Supervised Interest Point Detection and Description, IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 224 - 236.
[0319]
[17] H. Durrant-Whyte and T. Bailey, Simultaneous Localisation and Mapping (SLAM): Part I The Essential Algorithms, Technical report.
[0320]
[18] P. Fischer and T. Brox, Descriptor Matching with Convolutional Neural Networks: a Comparison to SIFT, pages 1 - 10.
[0321]
[19] M. A. Fischler and R. C. Bolles, Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography, Communications of the ACM, June 1981, 24(6): 381–395.
[0322]
[20] Y. Hang, X. Zhang, Y. Shao, H. Wu and W. Sun, Retinal Image Registration Based on the Feature of Bifurcation Point, 10th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics, CISPBMEI, 2017, pages 1 - 6.
[0323]
[21] C. G. Harris and M. Stephens, A Combined Corner and Edge Detector, Proceedings of the Alvey Vision Conference, 1988, pages 1 - 6.
[0324]
[22] C. Hernandez-Matas, X. Zabulis, A. Triantafyllou, and P. Anyfanti, FIRE: Fundus Image Registration dataset, Journal for Modeling in Ophthalmology, 2017, 4: 16 - 28.
[0325]
[23] J. Z. Huang, T. N. Tan, L. Ma, and Y. H. Wang, Phase Correlation-based Iris Image Registration Kodel, Journal of Computer Science and Technology, 2005, 20(3): 419–425.
[0326]
[24] N. Jacobs, N. Roman, and R. Pless, Consistent Temporal Variations in Many Outdoor Scenes, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2007.
[0327]
[25] D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization, 2014, CoRR, abs / 1412.6980.
[0328]
[26] S. Leutenegger, M. Chli, and R. Y. Siegwart, BRISK: Binary Robust Invariant Scalable Keypoints, Proceedings of the IEEE International Conference on Computer Vision, 2011, pp. 2548 - 2555.
[0329]
[27] P. Li, Q. Chen, W. Fan, and S. Yuan, Registration of OCT Fundus Images with Color Fundus Images Based on Invariant Features, Cloud Computing and Security - Third International Conference, 2017, pp. 471 - 482.
[0330]
[28] D. G. Lowe, Distinctive Image Features from Scale Invariant keypoints, International Journal of Computer Vision, Vol. 60, 2004.
[0331]
[29] K. Mikolajczyk and C. Schmid, A Performance Evaluation of Local Descriptors, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2005, 27(10): 1615–1630.
[0332]
[30] K. Mikolajczyk, T. Tuytelaars, C. Schmid, A. Zisserman, J. Matas, F. Schaffalitzky, T. Kadir, and L. Van Gool, A Comparison of Affine Region Detectors, International Journal of Computer Vision.
[0333]
[31] Y. Ono, E. Trulls, P. Fua, and K. M. Yi. LF - Net: Learning Local Features from Images, In Advances in Neural Information Processing Systems, 2018, pp. 6237 - 6247.
[0334]
[32] J.P.W. Pluim, J.B.A. Maintz, and M.A. Viergever, Mutual Information Based Registration of Medical Images: A Survey, IEEE Transactions on Medical Imaging, 2003, 22(8): 986–1004.
[0335]
[33] R. Ramli, M. Yamani, I. Idris, K. Hasikin, N.K.A. Karim, A. Wahid, A. Wahab, I. Ahmedy, F. Ahmedy, N.A. Kadri, and H. Arof, Feature-Based Retinal Image Registration Using D-Saddle Feature, 2017.
[0336]
[34] O. Ronneberger, P. Fischer, and T. Brox, U-Net: Convolutional Networks for Biomedical Image Segmentation, Medical Image Computing and Computer-Assisted Intervention, 2015, pp. 234-241.
[0337]
[35] E. Rublee, V. Rabaud, K. Konolige, and G. Bradski, ORB: An Efficient Alternative to SIFT or SURF, Proceedings of the IEEE International Conference on Computer Vision, 2011, pp. 2564-2571.
[0338]
[36] C. Sanchez-galeana, C. Bowd, E.Z. Blumenthal, P.A. Gokhale, L.M. Zangwill, and R.N. Weinreb, Using Optical Imaging Summary Data to Detect Glaucoma, Opthamology, 2001, pp. 1812-1818.
[0339]
[37] E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, P. Fua, and F. Moreno-Noguer, Discriminative Learning of Deep Convolutional Feature Point Descriptors, IEEE International Conference on Computer Vision, April 2015, pp. 118-126.
[0340]
[38] B. Triggs, P. F. Mclauchlan, R. I. Hartley, and A. W. Fitzgibbon. Bundle Adjustment - A Modern Synthesis, Technical report.
[0341]
[39] P. Truong, S. De Zanet, and S. Apostolopoulos, Comparison of Feature Detectors for Retinal Image Alignment, ARVO, 2019.
[0342]
[40] Y. Verdie, K. M. Yi, P. Fua, and V. Lepetit, TILDE: A Temporally Invariant Learned DEtector, Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2015, pp. 5279-5288.
[0343]
[41] G. Wang, Z. Wang, Y. Chen, and W. Zhao, Biomedical Signal Processing and Control Robust Point Matching Method for Multimodal Retinal Image Registration, Biomedical Signal Processing and Control, 2015, 19:68–76.
[0344]
[42] S.A.J. Winder and M.A.Brown, Learning Local Image Descriptors, IEEE Conference on Computer Vision and Pattern Recognition, 2007.
[0345]
[43] S.A.J. Winder, G.Hua and M.A.Brown, Picking the Best DAISY, IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 178 - 185.
[0346]
[44] K.M.Yi, E.Trulls, V.Lepetit and P.Fua, LIFT: Learned Invariant Feature Transform, European Conference on Computer Vision - ECCV, 2016, pp. 467 - 483.
[0347]
[45] K.M.Yi, Y.Verdie, P.Fua and V.Lepetit, Learning to Assign Orientations to Feature Points, IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 107 - 116.
[0348]
[46] S.D.Zanet, T.Rudolph, R.Richa, C.Tappeiner and R.Sznitman, Retinal slitlamp video mosaicking, International Journal of Computer Assisted Radiology and Surgery, 2016, 11(6):1035–1041.
[0349]
[47] L. Zhou, M. S. Rzeszotarski, L. J. Singerman, and J. M. Chokreff, The Detection and Quantification of Retinopathy Using Digital Angiograms, IEEE Transactions on Medical Imaging, 1994, 13(4): 619–626.
[0350]
[48] C. L. Zitnick and K. Ramnath, Edge Foci Interest Points, International Conference on Computer Vision, 2011, pp. 359-366.
Claims
1. A method for training a model for feature point detection, the method comprising: Acquire a first image and a second image; generating a first score map for the first image and a second score map for the second image using the model; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; pair-matching a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest, the pair-matching being performed based on similarities between features detected at the first point of interest in the first image and features detected at the second point of interest in the second image; checking correctness of the pairwise matches based on a ground truth transformation between the first image and the second image to generate a reward map; combining or comparing the first score map or the second score map with the reward map; and The model is updated based on a result of the combining or comparing.
2. The method according to claim 1, wherein: Selecting the plurality of interest points includes imposing a maximum limit on the distance from any point in the image to the nearest one of the interest points.
3. The method according to claim 1, wherein: Matching is performed in a first direction by matching the first interest point with a second interest point having a feature most similar to that of the first interest point among a plurality of second interest points.
4. The method according to claim 3, wherein: Matching is further performed in the second direction by matching the second interest point with a first interest point among the plurality of first interest points that has a feature most similar to that of the second interest point.
5. The method according to claim 1, wherein: The reward map indicates a reward for successfully matched points of interest based on the ground truth data, and indicates no reward for unsuccessfully matched points of interest based on the ground truth data.
6. The method according to claim 1, wherein: The combining or comparing includes combining or comparing the score map and the reward map only for the point of interest.
7. The method according to claim 1, wherein: The combining or comparing includes balancing some of the true positive matches and some of the false positive matches by possibly randomly selecting false positive matches and combining or comparing the score map and the reward map only for the selection of the false positive matches and true positive matches, The true positive matches are points of interest that pass the correctness check, and the false positive matches are points of interest that fail the correctness check.
8. The method according to claim 1, wherein: The combining or comparing comprises calculating a sum of squared differences between the score map and the reward map.
9. A device for training a model for feature point detection, the device comprising: Control unit; as well as A memory comprising instructions for causing the control unit to perform the following steps: Acquire a first image and a second image; generating a first score map for the first image and a second score map for the second image using the model; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; pair-matching a first point of interest in the first plurality of points of interest with a second point of interest in the second plurality of points of interest, the pair-matching being performed based on similarities between features detected at the first point of interest in the first image and features detected at the second point of interest in the second image; checking correctness of the pairwise matches based on a ground truth transformation between the first image and the second image to generate a reward map; combining or comparing the first score map or the second score map with the reward map; and The model is updated based on a result of the combining or comparing.
10. A method for registering a first image to a second image, the method comprising: Acquire a first image and a second image; Using a model trained by the method of any one of claims 1 to 8 or the apparatus of claim 9, generating a first score map for the first image and a second score map for the second image; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; as well as A first point of interest in the first plurality of points of interest is matched in pairs with a second point of interest in the second plurality of points of interest.
11. The method according to claim 10, wherein: Selecting the plurality of interest points includes imposing a maximum limit on the distance from any point in the image to the nearest one of the interest points.
12. The method according to claim 10 or 11, wherein: The pairwise matching is performed based on similarities between features detected at the first interest point in the first image and features detected at the second interest point in the second image.
13. The method according to claim 10, wherein: Matching is performed in a first direction by matching the first interest point with a second interest point having a feature most similar to that of the first interest point among a plurality of second interest points.
14. The method according to claim 13, wherein: Matching is further performed in the second direction by matching the second interest point with a first interest point among the plurality of first interest points that has a feature most similar to that of the second interest point.
15. An apparatus for registering a first image to a second image, the apparatus comprising: control unit, and A memory comprising instructions for causing the control unit to perform the following steps: Acquire a first image and a second image; Using a model trained by the method of any one of claims 1 to 8 or the apparatus of claim 9, generating a first score map for the first image and a second score map for the second image; selecting a first plurality of interest points in the first image based on the first score map; selecting a second plurality of interest points in the second image based on the second score map; as well as A first point of interest in the first plurality of points of interest is matched in pairs with a second point of interest in the second plurality of points of interest. 16 . A memory having stored therein a model for feature point detection obtained by the method according to claim 1 or the device according to claim 9 .