A multi-modal fundus image registration method based on intelligent matching of key point pairs

Through the multimodal fundus image registration method based on intelligent matching of key points, using YOLOv8 neural network to train and calculate the affine transformation matrix, the problem that traditional methods are difficult to deal with background changes in multimodal fundus image registration is solved, and high-precision and high-efficiency image registration is achieved.

CN119205862BActive Publication Date: 2025-06-10NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411697216.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-10
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Traditional fundus image registration methods are difficult to effectively solve the morphological differences caused by background changes caused by different modes, cameras or retinopathies in multimodal fundus image registration, and the time cost is high.

Method used

A multimodal fundus image registration method based on intelligent matching of key point pairs is adopted. Through YOLOv8 neural network training, a key point pair is selected to calculate the affine transformation matrix to achieve efficient image registration.

Benefits of technology

High-precision and high-efficiency registration of multimodal fundus images under large differences in image brightness, contrast and viewing angles simplifies the method steps without the need for specific or complex network structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205862B_ABST
    Figure CN119205862B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal fundus image registration method based on intelligent matching of key point pairs, specifically relating to the technical field of computer image processing. The multi-modal fundus image dataset is divided and enhanced and then sent into the YOLOv8 neural network for training. The optimal weights obtained from the training are selected for prediction to obtain the prediction results. The prediction results of the YOLOv8 neural network include two parts. One part is the test set images marked with key point pairs, bounding box categories, and confidence levels. The other part is a text format file containing key point pair categories, bounding box coordinates, and key point pair coordinates. According to the prediction results, three key point pairs are selected, and an affine transformation matrix is calculated from the three key point pairs to perform an affine transformation on the image to obtain the registration results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image processing, and particularly to a multi-modal fundus image registration method based on intelligent matching of key points. Background Art

[0002] Medical image registration is a basic step in computer-aided diagnosis and image-guided surgical treatment. In fundus image registration, intensity-based registration methods were first applied. The intensity-based registration method regards the problem as an iterative optimization problem and mainly focuses on developing various similarity functions, including (normalized) cross-correlation (CC), (normalized) mutual information (MI), etc. However, the intensity-based registration method is very sensitive to changes in image illumination and contrast. Feature-based registration methods are more effective than intensity-based registration methods. The idea is to extract key features from the input images, calculate a descriptor for each feature, match the closest features in the two images, and then establish a transformation relationship. Scale-invariant feature transform (SIFT) is of pioneering significance in feature-based matching methods. Compared with SIFT, speeded-up robust features (SURF) and ORB (Oriented FAST and Rotated BRIEF) are more efficient in terms of computational efficiency. However, when the image pairs to be registered have different morphologies due to background changes caused by different cameras, modalities, or retinal diseases, the registration accuracy of these traditional methods that emerged before deep learning technology is rather unsatisfactory; at the same time, the relatively high time cost caused by iterative optimization is another problem faced by traditional methods.

[0003] The fundus image registration method based on deep learning uses a three-stage method (segmentation, detection and description, and outlier rejection). The advantage of this method is that it bypasses the intensity gap between different modalities and achieves remarkable registration accuracy, but the disadvantage is its high complexity. In other medical image registrations, most use non-linear transformations. When it comes to fundus image registration, linear transformation is the most commonly used method. Fundus image registration usually uses CF images and FFA images, which rely entirely on visible light illumination for imaging. This imaging principle, combined with the natural movement of the eyeballs of the subjects being photographed, may cause large differences in image brightness, contrast, and viewing angles when taking multiple photos. It is very difficult to obtain good registration results with non-linear transformations under such large difference conditions.

[0004] The imaging principle differences of different-modal fundus image acquisition devices result in significant differences in brightness and contrast among multi-modal fundus images. Intensity-based registration methods are very sensitive to the brightness differences of images and are difficult to be directly applied to multi-modal fundus image registration. At the same time, the main directions of feature descriptors of the same feature on different-modal fundus images vary greatly, resulting in a low matching success rate of features, making it difficult for feature-based registration methods to obtain stable registration results. In short, traditional fundus image registration methods cannot be directly applied to multi-modal fundus image registration. In addition, the time cost of traditional registration methods is relatively high, and they cannot match the efficiency of deep learning methods in terms of registration efficiency. Although deep learning technology has shown great application potential in the task of fundus image registration, the existing research on deep learning-based registration methods is still in its initial stage, and most of them rely on specific network structures or the registration models are too complex, leaving broad room for exploration. Summary of the Invention

[0005] Therefore, the present invention provides a multi-modal fundus image registration method based on intelligent matching of key point pairs to solve the problems proposed in the background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: A multi-modal fundus image registration method based on intelligent matching of key point pairs. After dividing and enhancing the multi-modal fundus image data set, it is sent to the YOLOv8 neural network for training, and the optimal weights obtained from the training are selected for prediction to obtain the prediction results. The prediction results of the YOLOv8 neural network include two parts. One part is the test set images marked with key point pairs, bounding box categories, and confidence levels. The other part is a text format file containing key point pair categories, bounding box coordinates, and key point pair coordinates. Three key point pairs are selected according to the prediction results, and an affine transformation matrix is calculated from the three key point pairs, and the image is subjected to affine transformation to obtain the registration result.

[0007] Selection of key point pairs; When calculating the linear transformation matrix, two key point pairs can be selected respectively to calculate the rigid matrix, at least three key point pairs can be selected to calculate the affine matrix, or at least four key point pairs can be selected to calculate the homography matrix. To simplify the operation, usually at least three key point pairs are selected to calculate the affine matrix. To obtain a better registration effect, the distribution of key point pairs should be as scattered as possible.

[0008] The steps to select three key point pairs from the text format file are as follows:

[0009] S1-1. Set a threshold Thre, and classify the key points P i ∈P in the text format file, where P represents the set of key points, and each key point P i contains the attributes {Class i , box_x i, box_y i , box_width i , box_height i , x i , y i} where Class i is the category of the key point pair, box_x i is the abscissa of the upper left corner of the bounding box, box_y i is the ordinate of the upper left corner of the bounding box, box_width i is the width of the bounding box, box_height i is the height of the bounding box, x i is the abscissa of the key point, y i is the ordinate of the key point, i = 0, 1, 2, …, N, representing N key points;

[0010] If the abscissa x i coordinate of the key point is less than the threshold Thre, then the corresponding key point P i will be classified into the set CFPoints; otherwise, it will be classified into the set FFAPoints, and the sets CFPoints and FFAPoints form the key point pair;

[0011] The pseudocode is as follows:

[0012] / / Traverse all key points in the text format file

[0013] For i = 0...N do

[0014] / / If the key point P i has an x i coordinate value less than the threshold Thre

[0015] If x i < Thre then

[0016] / / Put the key point information into CFPoints, and "←" means adding P i to the set CFPoints or the set FFAPoints

[0017] CFPoints ← P i

[0018] Else

[0019] FFAPoints ← P i

[0020] End

[0021] End

[0022] S1-2. Select a key point CFkp from the set CFPoints f As the reference point, calculate this key point CFkp f and other key points CFkp i 's Euclidean distance dis_CFkp fi , and find the key point CFkp f with the farthest Euclidean distance from this key point CFkp m , connect CFkp f and CFkp m , to obtain the straight line CFkpl;

[0023] Then calculate the perpendicular distance from other key points CFkp i to the straight line CFkpl, and select the key point CFkp e with the largest perpendicular distance;

[0024] Select in the set FFAPoints the three key points with the same key point pair as the key point CFkp f , CFkp m and CFkp e as FFAkp i , FFAkp f , and FFAkp m ; The three key point pairs (CFkp e , FFAkp f ), (CFkp f ), (CFkp m ), (FFAkp m ), (CFkp e ), (FFAkp e ) will be used to calculate the affine matrix;

[0025] The pseudocode is as follows:

[0026] / / Select a reference point CFkpf from CFPoints, assume the first point P is selected 0 As the reference point, P 0 contains the attributes {Class 0 , box_x 0 , box_y 0 , box_width 0 , box_height 0 , x 0 , y 0}

[0027] CFkpf = P 0

[0028] / / Initialize the maximum distance

[0029] max_distance = 0

[0030] / / Calculate the distance from other key points P in CFPoints to CFkp i to CFkp f and find the farthest point CFkp m

[0031] For i = 1...N do

[0032] / / Calculate the Euclidean distance dis_CFkpfi between CFkpf and other key points

[0033] dis_CFkpfi =

[0034] / / If the current dis_CFkp fi is greater than max_distance, then update max_distance and the farthest distance point CFkp m

[0035] If dis_CFkp fi >max_distance then

[0036] max_distance = dis_CFkp fi

[0037] CFkpm = P i

[0038] End

[0039] End

[0040] / / Connect CFkp f and CFkp m to obtain the straight line CFkpl

[0041] CFkpl : where ,

[0042] / / Initialize the maximum distance

[0043] max_verticaldis = 0

[0044] / / Calculate the distance from other key points P i to the straight line CFkpl, where (i = 1, 2,..., N, i ≠ m), and find the farthest point CFkpe

[0045] For i = 1...N do

[0046] / / Calculate the vertical distance dis_CFkp_line from other key points in CFPoints to CFkpl i

[0047] dis_CFkp_line i =

[0048] / / If the current dis_CFkp_line i is greater than max_verticaldis, then update max_verticaldis and the key point CFkp with the maximum vertical distance e

[0049] If dis_CFkp_line i >max_verticaldis then

[0050] max_verticaldis = dis_CFkp_line i

[0051] CFkp e = P i

[0052] End

[0053] End

[0054] / / Select three key points from FFAPoints that have the same Class f ,CFkp m ,CFkp e (i = f, m,e), where the classes corresponding to CFkp i (i = f, m,e) are Class f ,CFkp m ,CFkp e are Class f , Class m ,Class e

[0055] / / Traverse all key points P in FFAPoints j

[0056] For j = 0…N do

[0057] / / If the class of the key point P j matches the class of the key point CFkp f then

[0058] If Class j = Class f then

[0059] FFAkpf = P j

[0060] Else if Class j = Class m then

[0061] FFAkp m = P j

[0062] Else Classj = Class e then

[0063] FFAkp e = P j

[0064] End

[0065] End

[0066] / / (CFkp f ,FFAkp f ),(CFkp m ,FFAkp m ),(CFkp e ,FFAkp e ) are three key point pairs for calculating the affine matrix.

[0067] Image registration; According to the selected key point pairs (CFkp f ,FFAkp f ),(CFkp m ,FFAkp m ),(CFkp e ,FFAkp e ), transform the affine matrix to achieve image registration.

[0068] Preferably, the steps for making the data set are as follows:

[0069] The directly obtained multi-modal fundus image dataset cannot be directly fed into the neural network for training. It is necessary to make it into a dataset that conforms to the input format of the neural network in this task. Scale the multi-modal fundus images (such as CF images and FFA images) to the same size and splice them to obtain the original dataset. Select the vascular bifurcation points as key points, select several pairs of key points on the original dataset, enclose them with bounding boxes of appropriate size, and label each pair of key points and the corresponding bounding box as one category to obtain pictures with key point pairs, bounding boxes, and category information, that is, labels. The original dataset and the labels together constitute the dataset.

[0070] Preferably, the steps of dataset division and augmentation are as follows:

[0071] Divide the made dataset into a training set, a validation set, and a test set in a ratio of 8:1:1, and perform data augmentation on the training set and the validation set, including but not limited to rotation, mirroring, and filtering operations, to improve the accuracy of the model prediction results and the generalization ability of the model.

[0072] Preferably, send the data-augmented training set and validation set into the YOLOv8 neural network for training to obtain the optimal weights. The steps are as follows:

[0073] S4-1. Use CSPDarkNet as the Backbone, including the SPPF module and the C2f module, and send the training set into the Backbone to extract multi-level features;

[0074] S4-2. For the multi-level features from the Backbone, use NAS-FPN (a feature pyramid network generated by a neural network structure search method) to further fuse, and detect and locate key point pairs at different scales (mainly focusing on the coarse-grained, medium-grained, and fine-grained information on the image at three scales of large, medium, and small);

[0075] S4-3. Output the classification results and position coordinates of the key point pairs on the validation set, and save the optimal weights of the YOLOv8 neural network during the validation process to obtain the optimal prediction results.

[0076] Preferably, the formula for affine matrix transformation is as follows:

[0077] ;

[0078] where, is the new coordinate after the affine transformation of the moving image, is the coordinate before the affine transformation of the moving image, is the affine transformation matrix calculated according to the selected key point pairs. a, b, c, and d are all parameters of the affine transformation matrix, representing the parameters after the compound of rotation, scaling, and shear transformation respectively, , is the translation amount.

[0079] The present invention has the following advantages:

[0080] 1. The present invention provides a brand-new solution idea for multi-modal fundus image registration under large differences in image brightness, contrast, and viewing angle, etc., and is completely applicable to single-modal fundus image registration.

[0081] 2. The present invention can not only achieve high-precision and high-efficiency registration of multi-modal fundus images, but also has a simple method without the need for a specific or complex network structure. Among them, the neural network only needs to have the ability to detect key points.

[0082] The method proposed by the present invention provides a novel and simple method for realizing multi-modal fundus image registration under large differences in image brightness, contrast, and viewing angle, etc. Through the combination of key point pair detection technology and linear transformation technology (such as rigid transformation, affine transformation, homography transformation, etc.), fast and accurate registration of fundus images is achieved. This method is simple and convenient, without the need for a specific or complex network structure, and also has certain reference significance for other medical / natural image registration fields. Description of the Drawings

[0083] Figure 1 is the flow chart for multi-modal fundus image registration;

[0084] Figure 2 is the schematic diagram of the original data set (CF-FFA) provided by the present invention;

[0085] Figure 3 is the schematic diagram of the data set labels provided by the present invention;

[0086] Figure 4 is the multi-modal fundus image key point pair map predicted by the neural network provided by the present invention;

[0087] Figure 5 is the result map of the key point pair selection algorithm provided by the present invention;

[0088] Figure 6 is the color photo of the fixed image and the moving image provided by the present invention;

[0089] Figure 7 is the result map of the angiography image before and after registration provided by the present invention;

[0090] Figure 8 is the result map of the key point pair prediction and selection provided by the present invention. Detailed Embodiment

[0091] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0092] The flowchart of multi-modal fundus image registration in this embodiment is as Figure 1 shown;

[0093] Dataset production: Horizontally splice the CF image and the FFA image to obtain the original dataset, as Figure 2 shown. Use the labelme software to make labels. Select the vascular bifurcation points as key points. Select several pairs of key points on the original dataset, enclose them with bounding boxes of appropriate sizes, and mark each pair of key points and the corresponding bounding box as one class. The dataset labels are as Figure 3 shown. In the figure, the key point pairs and bounding boxes of different colors represent different classes.

[0094] Dataset division and augmentation: The multi-modal fundus image data is a private dataset. There are 226 pairs of CF and FFA images, and there are 226 images in total after being made into a dataset. Among them, 138 images are divided into the training set, 9 images are divided into the validation set, and 79 images are divided into the test set. Adopt methods such as horizontal mirroring and rotation to augment the data of the training set and the validation set. After augmentation, there are 828 images in the training set and 54 images in the validation set.

[0095] Network training: In this embodiment, the YOLOv8 neural network (but not limited to such networks) is selected as the key point detection method. Send 828 training set images and 54 validation set images into the neural network for training, and save the optimal weights.

[0096] Result prediction: Use the optimal weights to predict the key point pairs of 79 test set images. The prediction results of the key point pairs are as Figure 4 shown. In the figure, kpi (i = 0, 1, 2...) is the aforementioned Class i , indicating the predicted key point pair class. The numbers after the class represent the confidence of the prediction results. The bounding boxes of the same color represent a pair of key point pairs. It can be seen from the figure that the selected method can accurately detect multiple pairs of key point pairs, and can provide data support for linear transformations such as rigid transformation, affine transformation, and homography transformation.

[0097] Selection of key point pairs: For the convenience of calculation, select kp0 from CFPoints as the reference point CFkp f , and according to the fifth step in the technical solution, select (CFkp f, FFAkp f ), (CFkp m , FFAkp m ), (CFkp e , FFAkp e ) Three key points are used to calculate the affine matrix. The selection results of the key point pairs are as Figure 5 shown. The key points in the upper right of the red font indicate the selected key points, and the red line represents the reference point (CFkp f / FFAkp f ) and the connection line (CFkpl / FFAkpl) of the key point (CFkp m / FFAkp m ) with the farthest Euclidean distance from the reference point. The green solid line represents the perpendicular distance between the key point CFkp e / FFAkp e and CFkpl / FFAkpl, and the green dashed line represents the perpendicular distance between other key points and CFkpl / FFAkpl.

[0098] Image registration: The affine matrix is calculated based on the selected key point pairs to complete image registration. The color photos, angiography images of the fundus fixed image and the moving image before and after registration, as well as the results of key point pair prediction and selection are as Figures 6 - 8 shown. Among them, each row is a registration case; the first column of each row is the fixed image, the second column is the moving image, the third column is the result of key point pair prediction and selection, the fourth column is the checkerboard display before image registration, and the fifth column is the checkerboard display after registration using the method of this embodiment. Through the comparison of the checkerboards before and after registration, it can be found that the method of this embodiment can achieve high-precision registration of multi-modal fundus images under large difference conditions. At the same time, the fast prediction based on deep learning and the fast calculation of linear transformation make the overall time consumption of image registration relatively low.

[0099] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection of the present invention.

Claims

1. A multimodal fundus image registration method based on intelligent matching of key point pairs, characterized by: After division and enhancement, the multimodal fundus image dataset is sent to the YOLOv8 neural network for training; The steps of data set division and enhancement are as follows: Divide the prepared data set into a training set, a validation set, and a test set; and perform data enhancement on the training set and the validation set, including but not limited to rotation, mirroring, and filtering operations, to obtain the data-enhanced training set and the validation set; The data-enhanced training set and validation set are sent to the YOLOv8 neural network for training, and the optimal weights obtained from the training are used for prediction to obtain the prediction results. The steps are as follows: Step a. Use CSPDarkNet as the Backbone, including the SPPF module and the C2f module, and send the training set into the Backbone to extract multi-level features; Step b. The multi-level features from Backbone are further integrated using the NAS-FPN network to detect and locate key point pairs at different scales; Step c. Output the classification results and position coordinates of the key point pairs on the verification set, and save the optimal weights of the YOLOv8 neural network during the verification process to obtain the optimal prediction results; The prediction result of the YOLOv8 neural network includes two parts, one part is a test set image with marked key point pairs, bounding box categories and confidence levels; the other part is a text format file containing key point pair categories, bounding box coordinates and key point pair coordinates; three key point pairs are selected according to the prediction results, the affine transformation matrix is ​​calculated from the three key point pairs, and the image is affine transformed to obtain the registration result; Select three key point pairs from the text format file. The steps are as follows: Step d. Set the threshold Thre and convert the key points P in the text format file into i ∈P for classification, P represents the key point set, each key point P i Contains attributes {Class i , box_x i , box_y i ,box_width i , box_height i , x i ,y i }, where Class i is the category of the key point pair, box_x i is the horizontal coordinate of the upper left corner of the bounding box, box_y i is the ordinate of the upper left corner of the bounding box, box_width i is the bounding box width, box_height i is the bounding box height, x i is the horizontal coordinate of the key point, y i is the ordinate of the key point; i=0,1,2,…,N, representing N key points; If the horizontal coordinate x of the key point i If it is less than the threshold Thre, the corresponding key point P i It is classified into the set CFPoints; otherwise, it is classified into the set FFAPoints, and the set CFPoints and FFAPoints form a key point pair; Step e. Select a key point CFkp from the set CFPoints f As the reference point, calculate the key point CFkp f With other key points CFkp i Euclidean distance dis_CFkp fi , and find the distance to the key point CFkp f The key point with the longest Euclidean distance CFkp m , connect CFkp f With CFkp m , and we get the straight line CFkpl; Then calculate other key points CFkp i The vertical distance to the straight line CFkpl, select the key point CFkp with the largest vertical distance e ; In the set FFAPoints select the key point CFkp f , CFkp m and CFkp e Class with the same keypoint pair i Three key points as FFAkp f ,FFAkp m and FFAkp e ; The three key points (CFkp f ,FFAkp f ), (CFkp m ,FFAkp m ), (CFkp e ,FFAkp e ) is used to calculate the affine matrix; Image registration; based on the selected key point pair (CFkp f ,FFAkp f ), (CFkp m ,FFAkp m ), (CFkp e , FFAkp e ), transform the affine matrix to achieve image registration.

2. The multimodal fundus image registration method based on intelligent matching of key point pairs according to claim 1, characterized in that: The steps for making a multimodal fundus image dataset are as follows: The multimodal fundus images are scaled to the same size and spliced ​​to obtain the original dataset; the vascular bifurcation points are selected as key points, and several key point pairs are selected on the original dataset. They are framed with bounding boxes, and each pair of key point pairs and the corresponding bounding boxes are marked as a category to obtain images with key point pairs, bounding boxes and category information, namely labels; the original dataset and the labels together constitute the dataset.

3. The multimodal fundus image registration method based on intelligent matching of key point pairs according to claim 1, characterized in that: The formula for affine matrix transformation is as follows: ; in, is the new coordinate of the moving image after affine transformation, is the coordinate of the moving image before affine transformation, is the affine transformation matrix calculated based on the selected key point pairs, a, b, c, d are the parameters of the affine transformation matrix, which represent the parameters after the rotation, scaling and shearing transformation are combined, respectively. is the translation amount.

Citation Information

Patent Citations

  • Eye fundus image registration method based on key point matching network

    CN112819867A

  • Multi-temporal remote sensing image automatic registration method based on improved SIFT algorithm

    CN114494378A