Cross-modal eye fundus image registration method, device, equipment and medium
Through the combination of wide-area vascular segmentation network and key point detection model, combined with double fitting and crop alignment strategies, the problem of fundus image registration is solved in large field of view, and high-precision alignment of images such as OCTA and CFP is achieved.
Patent Information
- Application Number
- CN202510164398.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-03
AI Technical Summary
Existing cross-modal fundus image registration technology is difficult to effectively process image pairs with large field of view gaps, especially registration between optical coherence tomography (OCTA) and fundus color imaging (CFP). Traditional methods cannot achieve accurate image alignment.
A wide-area vascular segmentation network is used to convert multimodal images into vascular maps, and feature matching is performed through key point detection models, and coordinate transformation of images is achieved using a dual fitting strategy based on the matching set. In addition, the crop alignment strategy is selected according to the difference in the field of view, and the accuracy of registration is improved by detecting the position of the visual disk and macular area.
High-precision registration of large-field gap images such as OCTA and CFP is achieved, which improves the performance of cross-modal fundus image registration and can be widely used in fundus image registration under large-field gap conditions.
Smart Images

Figure CN120088302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-modal fundus image registration method, apparatus, device and medium, and relates to the field of fundus image registration. Background Art
[0002] The fundus image registration (retina image registration) technology aims to automatically calculate the mapping between two images according to the content of the two images, so that any pixel point on one image can be mapped to the other image through this transformation. The cross-modal fundus image registration technology can achieve precise alignment between fundus images in different imaging modes, and further help doctors observe features such as fundus blood vessels and lesions on different modal images, and has broad application prospects in clinical practice.
[0003] The current mainstream cross-modal fundus image registration technologies usually target modal combinations with similar field of view differences, such as fluorescein angiography images (FA) and color fundus photography (CFP), and are mainly based on traditional local descriptors or deep learning networks of image features. Usually, feature or key point detection is first performed, and then feature matching is performed to find the corresponding relationship between a set of key points. Based on this corresponding relationship, a specific spatial transformation function is then fitted, and the most commonly used choice is the homography transformation. However, when there is a large field of view difference between images, such as optical coherence tomography angiography (OCTA) and CFP, the simple feature matching and coordinate fitting processes cannot achieve effective registration. In addition, the homography transformation is applicable to linear transformations such as translation and scaling of two-dimensional plane object images, and the fundus image is a two-dimensional projection of a three-dimensional image, which does not satisfy this principle, and the homography transformation will cause accuracy loss on high-precision pictures.
[0004] The prior art discloses that a traditional scale-invariant feature point extraction algorithm is used to obtain feature points and descriptors on two images by continuously adjusting the threshold, and the expectation-maximization algorithm is used to calculate the corresponding relationship between two feature point sets. This method solves the problem that the traditional Harris feature detection algorithm cannot be effectively registered when the local regions of the images are similar, but does not discuss the case of large modal differences in cross-modal images. The existing methods can only process image pairs with similar fields of view (such as FA and CFP), and related experiments also show that the model based on feature point matching alone cannot register image pairs under large field of view difference conditions (such as OCTA and CFP). Summary of the Invention
[0005] The present invention aims to at least solve one of the technical problems existing in the prior art. To this end, in view of the above problems, the object of the present invention is to provide a cross-modal fundus image registration method, apparatus, device and medium, which can achieve image registration and transformation.
[0006] To achieve the above-mentioned invention object, the technical solution adopted by the present invention is as follows: In the first aspect, the cross-modal fundus image registration method provided by the present invention includes: Using a wide-field blood vessel segmentation network to convert the multi-modal images to be matched into blood vessel maps, where the multi-modal images refer to two different-modal fundus images of the same eye of the same patient obtained at different times and by different imaging methods; Performing feature matching on the blood vessel maps of the multi-modal images to be matched through a key point detection model to obtain a matching set; Based on the matching set, using a double fitting strategy to achieve coordinate transformation of the multi-modal images to be matched and complete image registration.
[0007] Furthermore, it also includes selecting whether to use a cropping and alignment strategy according to the field of view difference of the multi-modal images to be matched. When the field of view of the multi-modal images to be matched has a large gap, the cropping and alignment strategy is adopted. When the fields of view of the multi-modal images to be matched are close, there is no need to use the cropping and alignment strategy.
[0008] Furthermore, the specific implementation process of the cropping and alignment strategy: Using the physiological structure of the retina to detect the positions of the optic disc and macula on the image with a larger field of view through the RetinaNet network; Centering on the macula, using the distance from the optic disc to the macula as the side length to crop a region from the blood vessel image with a larger field of view that is roughly aligned with the blood vessel image with a smaller field of view.
[0009] Furthermore, performing feature matching on the blood vessel maps of the multi-modal images to be matched through a key point detection model to obtain a matching set includes: Inputting the blood vessel maps of the multi-modal images to be matched into the key point detection model to obtain a key point set, where the key points are points with obvious position features in the image; Sampling at the corresponding pixel positions of the description feature map according to the key point set to obtain the corresponding description sub-feature set; Using the KNN matching strategy to match between the description sub-feature sets corresponding to the two images, and retaining the matching pairs whose ratio of the best matching distance to the second-best matching distance meets the set conditions to obtain a matching set.
[0010] Furthermore, the key point detection model uses SuperRetina as the feature point extraction and description network.
[0011] Furthermore, the specific process of using a double fitting strategy to achieve coordinate transformation of the multi-modal images to be matched based on the matching set is as follows: Using the RANSAC algorithm to remove outliers from the matching set using homography transformation; The matching sets on the two images with outliers removed are subjected to coordinate transformation using polynomial fitting: ; wherein, p is the highest order of the polynomial fitting function, and the default value is 2, a ij and b ij are the polynomial coefficients optimized by the least squares method, ( u, v ) are the coordinates of the key points in the source image, and ([[]] x, y ) are the coordinates of the corresponding key points in the target image. ) are the coordinates of the corresponding key points in the target image.
[0012] In a second aspect, the present invention provides a cross-modal fundus image registration device, including: A blood vessel segmentation module configured to convert the multi-modal images to be matched into blood vessel maps by using a wide-field blood vessel segmentation network, wherein the multi-modal images refer to two different-modal fundus images of the same eye of the same patient obtained at different times and by different imaging methods; A key point detection module configured to perform feature matching on the blood vessel maps of the multi-modal images to be matched through a key point detection model to obtain a matching set; An image registration module configured to perform coordinate transformation on the multi-modal images to be matched based on the matching set by using a double fitting strategy to complete image registration.
[0013] Further, it further includes a cropping and alignment module configured to select whether to use a cropping and alignment strategy according to the field of view difference of the multi-modal images to be matched. When the field of view of the fundus of the multi-modal images to be matched has a large gap, the cropping and alignment strategy is adopted, and when the fields of view of the multi-modal images to be matched are close, there is no need to use the cropping and alignment strategy.
[0014] In a third aspect, the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the processor; wherein, the memory stores instructions executable by the processor, and the instructions are executed by the processor so that the processor can execute any one of the methods.
[0015] In a fourth aspect, the present invention further provides a computer program product, including a computer program, and the computer program realizes any one of the methods when executed by a processor.
[0016] Due to the above technical solutions adopted by the present invention, it has the following characteristics: 1. The present invention adopts a modular deep learning technology to construct a cross-modal fundus image registration model through multiple processing stages. In order to perform key point detection and description under cross-modal conditions, the present invention uses a wide-field blood vessel segmentation network to convert multi-modal images into blood vessel maps, and uses wide-field blood vessel segmentation technology to solve the problems that local features cannot be directly extracted due to modal differences between different modal fundus images and that the model needs to be retrained for different modal combinations.
[0017] 2. To address the problem of large field-of-view differences, the present invention proposes a simple but effective Crop&Align operation. First, using the physiological structure of the retina, the positions of the optic disc and macula on the large-field image (CFP) are detected. Then, with the macula as the center and the distance from the optic disc to the macula as the side length, a region approximately aligned with the OCTA image is cropped from the CFP blood vessel image. Finally, key point detection and description are performed on the blood vessel maps of the OCTA and the cropped CFP images using a key point detection model. Therefore, the proposed Crop&Align strategy solves the problem that existing cross-modal fundus image registration models cannot register images with large field-of-view differences.
[0018] 3. To achieve precise coordinate transformation between images, the present invention proposes a double-fitting strategy to solve the problem of low accuracy of homography transformation for high-resolution image transformation. First, the RANSAC algorithm is used to calculate the homography matrix between two images based on the coordinates of the matched feature points, and the matching point pairs that do not conform to the matrix are filtered out. Then, polynomial fitting is used for the updated matching point pairs on the two images, and the OCTA image is transformed based on the fitting function.
[0019] In summary, the registration strategy of the present invention has achieved the best current registration performance on OCTA and CFP, and its Crop&Align and double-fitting strategies have also greatly improved the registration performance of other existing models, and can be widely applied to cross-modal fundus image registration under conditions of large field-of-view differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a schematic diagram of the cross-modal fundus image registration method according to an embodiment of the present invention.
[0021] Figure 2 It is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.
[0023] Although the terms first, second, third, etc. may be used herein to describe multiple elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or section from another. Unless the context clearly indicates otherwise, terms such as "first", "second", and other numerical terms when used herein do not imply an order or sequence. Thus, a first element, component, region, layer, or section discussed below may be referred to as a second element, component, region, layer, or section without departing from the teachings of the example embodiments.
[0024] For ease of description, spatial relative relationship terms may be used herein to describe the relationship of one element or feature shown in the figures to another element or feature, such as "inside", "outside", "inner side", "outer side", "below", "above", etc. Such spatial relative relationship terms are intended to include different orientations of the device in use or operation in addition to the orientation depicted in the figures.
[0025] The cross-modal fundus image registration method, device, equipment and medium provided by the present invention include: inputting an image pair into a wide-field blood vessel segmentation network to obtain corresponding blood vessel maps; freely selecting whether to use a cropping and alignment strategy according to the difference in the field of view, and inputting the cropped blood vessel maps into a key point detection model to obtain a set of key points; sampling at the corresponding pixel positions of the description feature maps according to the set of key points to obtain corresponding description sub-feature sets, and using the KNN matching strategy to match between the description sub-feature sets corresponding to the two images, retaining the matching pairs with the ratio of the best matching distance to the second-best matching distance lower than a specific threshold (default is 0.95), thereby obtaining a matching set; using a double fitting strategy according to the matching set, first using the RANSAC algorithm to filter out the matching point pairs with mismatched homographic transformations, and then using polynomial fitting for the remaining point pairs for transformation. Therefore, the registration strategy of the present invention has achieved the best current registration performance on OCTA and CFP, and its cropping alignment and double fitting strategies also have a great improvement in the registration performance for other existing models.
[0026] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0027] Example 1: As Figure 1 shown, the cross-modal fundus image registration method provided in this embodiment is implemented using a modular deep learning network, and the specific process includes: S1. Use a wide-field blood vessel segmentation network to convert multi-modal images into blood vessel maps.
[0028] In this embodiment, the multi-modal image pair refers to two different-modal fundus images of the same eye of the same patient obtained at different times and by different imaging methods, such as color fundus photography (CFP), fluorescein angiography (FA), optical coherence tomography angiography (OCTA), etc.
[0029] In this embodiment, for blood vessel segmentation: the wide-field blood vessel segmentation network is an improved U-Net, which is trained with multi-modal fundus images and blood vessel data. Among them, the training data set is a combined set composed of multiple public data sets, and these data sets cover various modalities. The advantage of the wide-field blood vessel segmentation network is that it can be directly applied to different modality combinations without retraining for a specific combination, which not only simplifies the process but also improves the flexibility and applicability of the method, enabling it to adapt to various cross-modal fundus image registration tasks.
[0030] S2. Freely select whether to use the cropping and alignment strategy according to the difference in the visual field. When the difference in the fundus visual field of the image pairs to be matched is large (such as CFP and OCTA), the cropping and alignment strategy needs to be adopted. When the visual fields of the image pairs to be matched are close (such as CFP and FA), the cropping and alignment strategy does not need to be used.
[0031] In this embodiment, Crop&Align: The cropping and alignment strategy uses a deep learning model to detect the macula and optic disc on the original image with a large visual field (such as CFP), and performs cropping and alignment in the vascular map based on the detection results. When the cropping and alignment strategy needs to be adopted, the cropping and alignment module network uses the physiological structure of the retina to crop out a region from the target image that is roughly aligned with the source image. Here, the source image refers to the image with a smaller visual field in the image pair to be matched, such as OCTA, and the target image refers to the image with a larger visual field in the image pair to be matched, such as CFP.
[0032] In this embodiment, the cropping and alignment strategy uses the macula and optic disc regions in the target image to achieve a rough visual field alignment between the source image and the target image. The macula and optic disc regions are important regions in fundus images and have obvious features in different modalities. The OCTA image usually centers on the macula region. Since the macula and optic disc are highly correlated in the fundus space, detecting these two related regions simultaneously is more accurate and reliable than detecting them separately.
[0033] Furthermore, the cropping and alignment module network is trained by using images with manually marked optic disc and macula positions in fundus images. The cropping and alignment network uses the object detection model RetinaNet and is supervised and trained by manually marking the positions of the optic disc and macula regions on, for example, 40 CFP images. By using RetinalNet to detect the macula and optic disc regions on the target image and cropping out a roughly corresponding region based on the relative positions between the two, it lays a foundation for subsequent precise registration and ensures the accuracy of feature matching and coordinate transformation.
[0034] Furthermore, in order to handle the large difference in the visual fields between images, the simple and effective cropping and alignment operation proposed in this embodiment uses the physiological structure of the retina to crop out a region from the target image that is roughly aligned with the source image. The specific process is as follows: First, use the physiological structure of the retina to detect the positions of the optic disc and macula regions on the large visual field image (CFP) through the trained RetinaNet network; Then, with the macula region as the center and the distance from the optic disc to the macula region as the side length, crop out a region from the CFP vascular image that is roughly aligned with the OCTA image.
[0035] S3. Perform feature matching on the vascular maps of the source image and the cropped target image through the keypoint detection model.
[0036] In this embodiment, for keypoint detection: Use a wide-field vascular segmentation network to segment the blood vessels in the training image, input the segmentation result into the keypoint detection model. After setting the number of iterations and optimizer parameters, this model is automatically optimized through iterative backpropagation without manual intervention. On the vascular maps of OCTA and the cropped CFP image, use the keypoint detection model (the single-modal fundus image registration model SuperRetina) to perform keypoint detection and description. Through the cropping and alignment strategy, the problem that the existing cross-modal fundus image registration model cannot register images with a large field of view gap is solved. The training data of the keypoint detection model is the segmented vascular map, and no less than 40 keypoints are manually marked on each image. Then, input the vascular map and keypoints into the SuperRetina model for retraining. The keypoint detection model uses SuperRetina as the feature point extraction and description network and is retrained with the vascular map to match the cross-modal registration task.
[0037] In this embodiment, for keypoint matching: Given the keypoints detected from the vascular maps of the source image and the cropped target image respectively, the present invention uses the OpenCV brute-force matcher (Brute-ForceMatcher, abbreviated as BFMatcher) to obtain the set of matching keypoints P . This method can effectively find the corresponding relationship between the two images and provide a basis for further image registration. In this way, SuperRetina can not only process single-modal images but also adapt to the requirements of multi-modal image registration, improving the generalization ability and practicality of the model.
[0038] Further, when performing feature matching on the vascular maps of the source image and the cropped target image through the keypoint detection model, the keypoint detection model can detect a certain number of keypoints on the two images, including their coordinates and feature descriptions. Then, based on these feature descriptions, the corresponding keypoint pairs on the two images can be found (for example, the keypoint p on the CFP image and the keypoint q on the OCTA image are the same point), and thus a mathematical method is used to calculate the registration mapping equation of the overall image. The specific process is as follows: Input the cropped vascular map into the keypoint detection model to obtain the keypoint set. Among them, the keypoints are the points with obvious position features such as the intersection points and edge points of the blood vessels on the image. The keypoint set includes the coordinate sequence of a group of keypoints on the image { (x i , y i )}; Sampling at the corresponding pixel positions of the description feature map according to the key point set gives the corresponding descriptor feature set. The model obtains the feature map of the entire image through feature sampling at different scales ( h , w , d ), where h is the image height, w is the image width, d is the feature dimension (default value is 256), so that there is a d -dimensional descriptor at each pixel point on the entire image ( f 1 , f 2 ,…, f 256 ). This descriptor is a dense feature representation of the image area centered on this point. In theory, the distance between the descriptors of the same point on two images should be as small as possible.
[0039] The KNN matching strategy is used to match between the descriptor feature sets corresponding to the two images. The specific operation is to calculate the distance (usually using the Euclidean distance) between the descriptors of each key point in one image and all key points in the other image, and find the two key points with the closest distance, which are called the best match and the second-best match respectively. Retain the matching pairs whose ratio of the best match distance to the second-best match distance is lower than a specific threshold (default is 0.95), and thus obtain the matching set.
[0040] S4. Use the double-fitting strategy to achieve accurate coordinate transformation from the source image to the target image.
[0041] In this embodiment, the specific process of using the double-fitting strategy to achieve accurate coordinate transformation from the source image to the target image is as follows: Use the RANSAC algorithm to calculate the homography matrix between the two images based on the coordinates of the matched feature points. The calculation can be directly performed using the existing built-in algorithm through the coordinates of the matched key points. The homography matrix is a linear transformation in the homogeneous coordinates of three-element vectors, represented by a 3×3 non-singular transformation matrix, as follows:
[0042] Among them, the left side of the equal sign is the coordinate on the target image ( x , y ), the 3×3 matrix on the right side of the equal sign is the homography matrix, and the 1×3 vector is the corresponding coordinate on the source image ( u , v ). Using the above equation, any coordinate on the source image can be mapped to the target image.
[0043] Filter out the matching points that do not match the homography transformation, where the RANSAC method can be directly used for filtering.
[0044] Use polynomial fitting for the set of matching point pairs retained after RANSAC filtering on the two images, and transform the OCTA image based on the fitting function. In the experiment, the registration strategy of the present invention achieved the best current registration performance on OCTA and CFP, and its cropping alignment and double fitting strategies also greatly improved the registration performance for other existing models.
[0045] Furthermore, in order to achieve coordinate transformation between two images, the double fitting strategy proposed by the present invention. This strategy first uses the RANSAC algorithm to remove outliers from the set of matching points P using homography transformation, and then fits the following polynomial function to obtain the final coordinate transformation model:
[0046] where p is the highest order of the polynomial fitting function, with a default value of 2, i and j are specific orders ( ), a ij and b ij are the polynomial coefficients optimized by the least squares method. ( u, v ) are the coordinates of the key points in the source image, ( x, y ) are the coordinates of the corresponding key points in the target image. This method not only inherits the robustness of the RANSAC algorithm to external outliers, but also utilizes the high-precision characteristics of polynomial fitting, thereby improving the overall performance of cross-modal fundus image registration.
[0047] The method for cross-modal fundus image registration under the condition of large field-of-view difference proposed by the present invention has the following advantages: (1) The wide-field vascular segmentation model uniformly converts fundus images of different modalities into vascular maps, eliminating modality differences and also reducing the retraining requirements when different modality combinations are involved; (2) The cropping alignment strategy enables rough alignment of images in the field of view, solving the problem that traditional deep learning methods cannot handle large field-of-view differences; (3) The double fitting strategy improves the registration transformation accuracy of the model by combining the robustness of the RANSAC algorithm and the precision of polynomial fitting. The application of the cross-modal fundus image registration method of the present invention will be described in detail below through specific embodiments.
[0048] The test set in this embodiment uses real ophthalmic data, including 30 pairs of OCTA-wfCFP (wide-field fundus color photographs) images of different fundus conditions of 30 patients from the ophthalmology department of Peking Union Medical College Hospital. Among them, the OCTA images are superficial vascular plexus (SVP) images obtained by a 12mm×12mm scan at the fovea centralis using an SVision VG200 OCT device, while the wide-field color fundus photograph (wfCFP) is taken using a ZEISS CLARUS 500 fundus camera. Each pair of images has at least 10 manually marked corresponding key points.
[0049] This embodiment detects the failure rate (failed), inaccuracy rate (inaccurate), and acceptable rate (acceptable) on the test set. If the median Euclidean error (MEE) and maximum Euclidean error (MAE) between the mapped key points and manual annotations are less than 50 pixels and 20 pixels respectively, it is considered acceptable. During the test, a graph of the acceptable rate versus the threshold based on MEE with a threshold range from 1 to 25 is also used, and the normalized area (AUC) of this graph with the X-axis is calculated, and the larger this value is, the better. In addition, this embodiment uses the soft Dice coefficient (Dice s ) to reflect the degree of overlap of the vascular maps between the target image and the source image after registration transformation. The comparison with other matching methods is shown in Table 1. The experimental results show that all the indicators of the present invention are better than other matching schemes, and traditional methods cannot effectively perform registration under large field-of-view differences. In the second row of the lower part of the table, the cropping alignment strategy is cancelled, and compared with the first row, effective registration cannot be achieved at all. In addition, the performance of other existing methods in Table 2 is significantly improved after adding this strategy, indicating the effectiveness and performance improvement of this strategy for this task; in the third and fourth rows of the lower part of the table, the fitting strategy is changed to a single polynomial fitting or RANSAC homography transformation fitting, and the performance is not as good as that of the first row, indicating that the double fitting strategy can effectively improve the model performance.
[0050] Table 1 Comparative experiments and ablation experiments of the present invention and other registration methods
[0051] Table 2 Experimental results of the performance improvement effect of the cropping alignment strategy on other registration methods
[0052] Embodiment 2: The above Embodiment 1 provides a cross-modal fundus image registration method. Correspondingly, this embodiment provides a cross-modal fundus image registration device. The device provided in this embodiment can implement the cross-modal fundus image registration method of Embodiment 1, and the device can be implemented by software, hardware, or a combination of software and hardware. For the convenience of description, when describing this embodiment, various units are described separately according to their functions. Of course, in implementation, the functions of each unit can be implemented in the same or multiple software and / or hardware. For example, the device can include integrated or separate functional modules or functional units to execute the corresponding steps in each method of Embodiment 1. Since the device in this embodiment is basically similar to the method embodiment, the description process of this embodiment is relatively simple, and the relevant parts can refer to the partial description of Embodiment 1. The embodiment of the cross-modal fundus image registration device provided by the present invention is only illustrative.
[0053] Specifically, the cross-modal fundus image registration device provided in this embodiment includes: A blood vessel segmentation module, configured to convert the multi-modal image to be matched into a blood vessel map by using a wide-field blood vessel segmentation network, where the multi-modal image refers to two different-modal fundus images of the same eye of the same patient obtained at different times and by different imaging methods; A key point detection module, configured to perform feature matching on the blood vessel map of the multi-modal image to be matched through a key point detection model to obtain a matching set; An image registration module, configured to implement coordinate transformation of the multi-modal image to be matched based on the matching set by using a double fitting strategy to complete image registration.
[0054] Further, it further includes a cropping and alignment module, configured to select whether to use a cropping and alignment strategy according to the field of view difference of the multi-modal image to be matched. When the fundus field of view of the multi-modal image to be matched has a large gap, the cropping and alignment strategy is adopted, and when the fields of view of the multi-modal image to be matched are close, there is no need to use the cropping and alignment strategy.
[0055] In summary, the present invention uses modular composition. After training, modules can be freely selected for any modality combination, including a blood vessel segmentation module, a cropping and alignment module, and a feature extraction and matching module. In practical applications, it can be determined whether to use the cropping and alignment module according to the field of view difference of the image.
[0056] Embodiment 3: This embodiment provides an electronic device corresponding to the cross-modal fundus image registration method provided in Embodiment 1. The electronic device can be an electronic device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of Embodiment 1.
[0057] Such as Figure 2As shown, the electronic device includes a processor, a memory, a communication interface, and a bus. The processor, the memory, and the communication interface are connected through the bus to complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The memory stores a computer program that can run on the processor. When the processor runs the computer program, it executes the method of Embodiment 1. The implementation principle and technical effects are similar to those of Embodiment 1 and will not be elaborated here. Those skilled in the art can understand that Figure 2 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computing device to which the solution of this application is applied. The specific computing device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0058] In a preferred embodiment, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), optical discs, etc.
[0059] In a preferred embodiment, the processor can be various types of general-purpose processors such as a central processing unit (CPU) and a digital signal processor (DSP), which are not limited here.
[0060] Embodiment 4: This embodiment provides a computer-readable storage medium storing one or more programs. The one or more programs include computer instructions that, when executed by a computer, cause the computer to execute the method provided in Embodiment 1 above.
[0061] Embodiment 5: This embodiment provides a computer program product. The computer program product may be a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method provided in the above Embodiment 1. The implementation principle and technical effects are similar to those of Embodiment 1 and will not be elaborated here.
[0062] In a preferred embodiment, the computer-readable storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. For example, it may be, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above. The computer-readable storage medium stores computer program instructions that cause the computer to execute the method provided in the above Embodiment 1.
[0063] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0066] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In the description of this specification, the descriptions with reference to terms such as "a preferred embodiment", "furthermore", "specifically", "in this embodiment", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-modal fundus image registration method, characterized in that: include: A wide-area vascular segmentation network is used to convert the multimodal image to be matched into a vascular map, wherein the multimodal image refers to two fundus images of different modalities of the same fundus of the same patient obtained at different times and in different imaging modes; Performing feature matching on the vascular map of the multimodal image to be matched through a key point detection model to obtain a matching set; Based on the matching set, a double fitting strategy is used to realize the coordinate transformation of the multimodal images to be matched and complete the image registration.
2. The cross-modality fundus image registration method according to claim 1, characterized in that: It also includes selecting whether to use a cropping alignment strategy based on differences in the field of view of the multimodal images to be matched. When the fundus field of view of the multimodal images to be matched is significantly different, a cropping alignment strategy is used. When the field of view of the multimodal images to be matched is close, there is no need to use the cropping alignment strategy.
3. The cross-modality fundus image registration method according to claim 2, characterized in that: The specific implementation process of the trimming and alignment strategy: The physiological structure of the retina is used to detect the location of the optic disc and macula in images with a larger visual field through the RetinaNet network. With the macula as the center and the distance from the optic disc to the macula as the side length, a region roughly aligned with the vascular image with a smaller field of view is cut out from the vascular image with a larger field of view.
4. The cross-modality fundus image registration method according to claim 1, characterized in that: The vascular map of the multimodal image to be matched is subjected to feature matching through the key point detection model to obtain a matching set, including: Inputting the blood vessel map of the multimodal image to be matched into the key point detection model to obtain a key point set, wherein the key point is a point with obvious position characteristics in the image; According to the key point set, the corresponding pixel position of the description feature map is sampled to obtain the corresponding descriptor feature set; The KNN matching strategy is used to match the descriptor feature sets corresponding to the two images, and the matching pairs whose ratio of the best matching distance to the second best matching distance meets the set conditions are retained to obtain the matching set.
5. The cross-modality fundus image registration method according to claim 4, characterized in that: The key point detection model uses SuperRetina as the feature point extraction and description network.
6. The cross-modality fundus image registration method according to claim 1, characterized in that: The specific process of using the dual fitting strategy based on the matching set to achieve the coordinate transformation of the multimodal image to be matched is: Use the RANSAC algorithm to remove outliers from the matching set using homography transformation; Use polynomial fitting to transform the coordinates of the matched sets on the two images to remove outliers: ; in, p is the highest order of the polynomial fitting function, the default value is 2. a ij and b ij are the polynomial coefficients optimized by the least squares method, ( u,v ) are the coordinates of the key points in the source image, ( x,y ) are the coordinates of the corresponding keypoints in the target image.
7. A cross-modal fundus image registration device, characterized in that: include: A blood vessel segmentation module is configured to convert a to-be-matched multimodal image into a blood vessel map using a wide-area blood vessel segmentation network, wherein the multimodal image refers to two fundus images of different modalities of the same fundus of the same patient obtained at different times and in different imaging modes; A key point detection module is configured to perform feature matching on the vascular map of the multimodal image to be matched through a key point detection model to obtain a matching set; The image registration module is configured to use a double fitting strategy based on the matching set to achieve coordinate transformation of the multimodal image to be matched, thereby completing the image registration.
8. The cross-modality fundus image registration device according to claim 7, characterized in that: It also includes a cropping and alignment module, which is configured to select whether to use a cropping and alignment strategy based on the difference in the field of view of the multimodal images to be matched. When the difference in the fundus field of view of the multimodal images to be matched is large, the cropping and alignment strategy is adopted. When the field of view of the multimodal images to be matched is close, there is no need to use the cropping and alignment strategy.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor so that the processor can execute the method according to any one of claims 1-7.
10. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Cited By
Fundus color photo synthesis method, electronic equipment and computer readable medium
CN120318095A
Medical image matching method and device
CN121661368A
Method and apparatus for matching medical images
CN121661368B