A Multimodal Retinal Fundus Image Registration Method and System
By constructing a key point detection model and using group-like dynamic serpentine convolution modules for feature extraction and fusion, the problem of low accuracy in cross-modal registration of U-Net network is solved, and more efficient multimodal retinal fundus image registration is achieved.
Patent Information
- Application Number
- CN202411756314.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The existing U-Net network cannot effectively register images of different modes and is difficult to effectively learn key point features, resulting in reduced image registration accuracy.
A key point detection model is constructed, including the first feature encoder, the second feature encoder, the key point detection decoder and the key point decoder, and a group and other variable dynamic serpentine convolution modules are used to extract and fusion features to improve the accuracy and robustness of image registration.
Through improved feature extraction and fusion methods, the accuracy and robustness of image registration are improved, especially when processing multimodal retinal fundus images, key points characteristics can be captured more effectively and registration effect can be improved.
Smart Images

Figure CN119228861B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a multi-modal retinal fundus image registration method and system. Background Art
[0002] Image registration refers to the process of aligning similar structures or features in multiple images to the same coordinate system. Medical image registration plays an important role in the field of medical image processing and analysis. The registration of multi-modal fundus images (such as fundus color photographs and fluorescein fundus angiography (FFA) images, fundus color photographs and optical coherence tomography angiography (OCTA) images) is beneficial to integrating information of different imaging modalities. Blood vessels have obvious consistent structures in different-modal fundus images and are often used for feature detection of multi-modal fundus images.
[0003] Currently, various fundus image registration methods have been proposed in domestic and foreign research work. For example, the U-Net network not only performs well in image segmentation tasks but can also be used as a key-point feature extraction network in image registration. SuperPoint is a deep learning-based feature point detection and description method that can use the backbone network U-Net to quickly and accurately detect key points and generate corresponding local feature descriptors. These feature points and descriptors can be used to match similar regions between different images, thereby achieving image registration. The SuperRetina key-point detection network adds progressive key-point amplification on the basis of SuperPoint, greatly reducing the network's requirement for the gold standard number of key points. Moreover, the key points extracted by this network have a certain rotational equivariance. After the key points of the image pairs to be registered are obtained by the key-point feature extraction network, an affine matrix is usually fitted according to the key points. Therefore, it is a non-end-to-end rigid registration method.
[0004] However, the single-stream U-Net structure can only receive image features of a single modality during training. Therefore, when processing different-modal images, it cannot recognize and correspond to the differences between different modalities, resulting in the network being unable to effectively register images of another modality. In addition, the U-Net network mainly relies on convolutional layers for feature extraction. However, in multi-modal fundus images, due to the high blood vessel density and strong background noise, the contrast of key points in the image is low in the local area, making it difficult to accurately separate them from the background. The convolutional layers are difficult to effectively capture the detailed features in the image, especially in the key-point area. Therefore, when the U-Net network structure processes multi-modal fundus images, it is often difficult to effectively learn key-point features, resulting in a reduction in image registration accuracy. Summary of the Invention
[0005] To this end, the technical problem to be solved by the present invention is to overcome the defect that the existing U-Net network cannot effectively register images of different modalities and is difficult to effectively learn key point features, resulting in a reduction in the accuracy of image registration.
[0006] To solve the above technical problem, the present invention provides a method for registering multi-modal retinal fundus images, comprising the following steps:
[0007] Construct a key point detection model, which includes: a first feature encoder, a second feature encoder, a key point detection decoder, and a key point descriptor decoder;
[0008] Among them, each feature encoder includes four stages connected in sequence along the forward propagation direction. Each stage is a group-equivariant dynamic snake convolution module. A group-equivariant dynamic snake convolution module includes: a feature enhancement branch, a group-equivariant dynamic snake convolution block, a third group-equivariant convolution, group normalization, and a Relu activation function; the feature enhancement branch includes: a first group-equivariant convolution, group normalization, and a Relu activation function connected in sequence along the forward propagation direction; each group-equivariant dynamic snake convolution module outputs a stage of rotation-equivariant features; for the rotation-equivariant features of other stages except the rotation-equivariant features of the first stage, their sizes are adjusted by group transposed convolution to be the same as the rotation-equivariant features of the first stage; the rotation-equivariant features of all stages are concatenated along the channel dimension to obtain a total rotation-equivariant feature map;
[0009] Copy the reference image times to generate a reference image structure body. Pass the reference image structure body through the first feature encoder to extract the reference rotation-equivariant features of four stages and the total reference rotation-equivariant feature map; where is the rotation angle dimension of the group-equivariant convolution;
[0010] Copy the floating image times to generate a floating image structure body. Pass the floating image structure body through the second feature encoder to extract the floating rotation-equivariant features of four stages and the total floating rotation-equivariant feature map;
[0011] Input the reference rotation-equivariant features of four stages into the key point detection decoder for feature fusion and output a reference key point prediction map;
[0012] Input the floating rotation-equivariant features of four stages into the key point detection decoder for feature fusion and output a floating key point prediction map;
[0013] Input the total reference rotation-equivariant feature map and the total floating rotation-equivariant feature map into the key point descriptor decoder respectively to output the descriptors of each reference key point and the descriptors of each floating key point;
[0014] Input the reference image, the floating image, the reference key point prediction map, the floating key point prediction map, the descriptors of all reference key points, and the descriptors of all floating key points into the registration module to output the registered image.
[0015] Preferably, input the rotation-equivariant features of the four stages into the key point detection decoder for feature fusion and output the key point prediction map, including:
[0016] Pass the rotation-equivariant features of the fourth stage through the group-equivariant dynamic snake convolution module and group upsampling in sequence to obtain the key point decoding features of the fourth level;
[0017] After concatenating the key point decoding features of the fourth level and the rotation-equivariant features of the third stage along the channel dimension, pass them through the group-equivariant dynamic snake convolution module and group upsampling in sequence to obtain the key point decoding features of the third level;
[0018] After concatenating the key point decoding features of the third level and the rotation-equivariant features of the second stage along the channel dimension, pass them through the group-equivariant convolution module and group upsampling in sequence to obtain the key point decoding features of the second level;
[0019] After concatenating the key point decoding features of the second level and the rotation-equivariant features of the first stage along the channel dimension, obtain the key point decoding features of the first level;
[0020] For the key point decoding features of the first level, stack all the features in the dimension to the channel dimension and then pass them through the fully connected layer and sigmoid activation function in sequence to obtain the key point prediction map.
[0021] Preferably, process the input features of the group-equivariant dynamic snake convolution module through the feature enhancement branch to obtain the enhanced features;
[0022] Process the input features of the group-equivariant dynamic snake convolution module through the group-equivariant dynamic snake convolution block to capture the spatial information in different directions and obtain the group-equivariant output features in the horizontal direction and the group-equivariant output features in the vertical direction;
[0023] After concatenating the enhanced features, the group-equivariant output features in the horizontal direction and the group-equivariant output features in the vertical direction along the channel dimension, process them through the third group-equivariant convolution, group normalization, and Relu activation function in sequence to obtain the output features of the group-equivariant dynamic snake convolution module.
[0024] Preferably, passing the input features of the group-equivariant dynamic snake convolution module through the group-equivariant dynamic snake convolution block to capture spatial information in different directions, and obtaining the group-equivariant output features in the horizontal direction and the group-equivariant output features in the vertical direction, includes:
[0025] Passing the input features of the group-equivariant dynamic snake convolution module sequentially through the second group-equivariant convolution, group normalization, and tanh activation function of the convolutional kernel to obtain the bias features in the horizontal direction and the bias features in the vertical direction respectively; wherein, the value of each pixel in the bias features in the horizontal direction represents the offset in the horizontal direction, and the value of each pixel in the bias features in the vertical direction represents the offset in the vertical direction;
[0026] According to the value of each pixel in the bias features in the horizontal direction, after offsetting each pixel in the input features of the group-equivariant dynamic snake convolution module in the horizontal direction, the bias input features in the horizontal direction are obtained through four-neighborhood interpolation;
[0027] According to the value of each pixel in the bias features in the vertical direction, after offsetting each pixel in the input features of the group-equivariant dynamic snake convolution module in the vertical direction, the bias input features in the vertical direction are obtained through four-neighborhood interpolation;
[0028] Passing the bias input features in the horizontal direction sequentially through the convolutional kernel, group normalization, and Relu activation function in the horizontal direction to obtain the feature map in the horizontal direction; wherein, is a set parameter;
[0029] Passing the bias input features in the vertical direction sequentially through the convolutional kernel, group normalization, and Relu activation function in the vertical direction to obtain the feature map in the vertical direction;
[0030] Integrating the feature map in the horizontal direction and the feature map in the vertical direction along the dimension respectively to obtain the group-equivariant output features in the horizontal direction and the group-equivariant output features in the vertical direction.
[0031] Preferably, inputting the total rotation-equivariant feature map into the key point descriptor decoder to output the descriptor of each key point, includes:
[0032] Based on the key point prediction map, obtaining the coordinates of each key point; according to the coordinates of each key point, extracting the pixel values of each channel corresponding to the position of the key point in the total rotation-equivariant feature map, and constructing the initial descriptor of each key point;
[0033] Passing the initial descriptor of each key point sequentially through batch normalization, 1×1 convolutional layer, Sigmoid activation function, 3×3 convolutional layer, and fully connected layer to obtain the descriptor of each key point.
[0034] Preferably, inputting the reference image, the floating image, the reference key-point prediction map, the floating key-point prediction map, the descriptors of all reference key-points, and the descriptors of all floating key-points into the registration module, and outputting the registered image, includes:
[0035] Performing key-point matching on the reference key-point prediction map and the descriptors of all reference key-points, and the floating key-point prediction map and the descriptors of all floating key-points through the flann algorithm to obtain key-point pairs between the reference image and the floating image;
[0036] Using the random sample consensus algorithm to fit the key-point pairs between the reference image and the floating image, and removing the key-point pairs that do not conform to the transformation constraint to obtain an affine matrix;
[0037] Aligning the floating image to the reference image by using the affine matrix to obtain the registered image.
[0038] Preferably, the reference image is a fundus color photo, and the floating image is any one of a fundus fluorescein angiography image and an optical coherence tomography angiography image.
[0039] Preferably, the training method of the key-point detection model is:
[0040] Obtaining an image data set, where the image data set includes different images and their corresponding multiple gold standard key-points; dividing the image data set into a training set, a validation set, and a test set;
[0041] Training the key-point detection model by using the training set, and optimizing each parameter of the key-point detection model by using a total loss function, where during the training process, it includes: after obtaining the key-point prediction map of the image in this round and the descriptor of each key-point, freezing the network parameters, performing a random affine transformation on the image in this round, re-inputting the affine-transformed image into the key-point detection model to obtain the key-point prediction map of the affine-transformed image in this round and the descriptor of each key-point, calculating the Euclidean distance between the descriptors of the th key-point in the key-point prediction maps of the image in this round before and after the affine transformation and the Euclidean distance between the descriptor of the th key-point in the key-point prediction map of the image in this round before the affine transformation and the descriptor of the second nearest neighbor key-point of the th key-point in the key-point prediction map of the affine-transformed image in this round in the key-point prediction map of the affine-transformed image in this round , when When the ratio is greater than the set threshold, set the key point as the gold standard key point for the next round of gold standard key points;
[0042] After the key point detection model is trained, the key points detected by the key point detection model are registered with the descriptors through the registration module.
[0043] Preferably, the total loss function includes: Dice loss function and triplet loss function;
[0044] The calculation formula of the total loss function is:
[0045] ,
[0046] ,
[0047] ,
[0048] where is the total loss function of the key point detection model, is the Dice loss function, is the triplet loss function, is the key point prediction map, is the image rotation the key point prediction map corresponding to the angle, is the soft label obtained by applying the Gaussian Boolean function to the key point map of the gold standard, is element-wise multiplication, is the maximum function, is the th descriptor of the key point in the image and the th descriptor of the corresponding key point after the image is rotated by angle, is the threshold for dividing positive and negative samples, is the th descriptor of the key point in the image and the th descriptor of any other key point corresponding to the image rotated by angle, is the th descriptor of the key point in the image and the maximum Euclidean distance between the descriptors of all other corresponding key points after the image is rotated by
[0049] The present invention also provides a multi-modal retinal fundus image registration system, including:
[0050] A model construction module for constructing a key point detection model, where the key point detection model includes: a first feature encoder, a second feature encoder, a key point detection decoder, and a key point descriptor decoder;
[0051] Among them, each feature encoder includes four stages connected in sequence along the forward propagation direction. Each stage is a group-equivariant dynamic snake convolution module. A group-equivariant dynamic snake convolution module includes: a feature enhancement branch, a group-equivariant dynamic snake convolution block, a third group-equivariant convolution, group normalization, and a Relu activation function; the feature enhancement branch includes: a first group-equivariant convolution, group normalization, and a Relu activation function connected in sequence along the forward propagation direction; each group-equivariant dynamic snake convolution module outputs a stage of rotation-equivariant features; for the rotation-equivariant features of other stages except the rotation-equivariant features of the first stage, their sizes are adjusted by group transposed convolution to be the same as the rotation-equivariant features of the first stage; the rotation-equivariant features of all stages are concatenated along the channel dimension to obtain a total rotation-equivariant feature map;
[0052] A reference rotation-equivariant feature acquisition module for copying the reference image multiple times to generate a reference image structure, and passing the reference image structure through the first feature encoder to extract four stages of reference rotation-equivariant features and a total reference rotation-equivariant feature map; where is the rotation angle dimension of the group-equivariant convolution;
[0053] A floating rotation-equivariant feature acquisition module for copying the floating image multiple times to generate a floating image structure, and passing the floating image structure through the second feature encoder to extract four stages of floating rotation-equivariant features and a total floating rotation-equivariant feature map;
[0054] A reference key point prediction map acquisition module for inputting the four stages of reference rotation-equivariant features into the key point detection decoder for feature fusion and outputting a reference key point prediction map;
[0055] A floating key point prediction map acquisition module for inputting the four stages of floating rotation-equivariant features into the key point detection decoder for feature fusion and outputting a floating key point prediction map;
[0056] A key point descriptor acquisition module for inputting the total reference rotation-equivariant feature map and the total floating rotation-equivariant feature map into the key point descriptor decoder respectively, and outputting the descriptors of each reference key point and each floating key point;
[0057] An image registration module for inputting the reference image, the floating image, the reference key point prediction map, the floating key point prediction map, the descriptors of all reference key points, and the descriptors of all floating key points into the registration module and outputting the registered image.
[0058] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0059] A multimodal retinal fundus image registration method and system according to the present invention respectively extracts features of a reference image and a floating image through a first feature encoder and a second feature encoder, solving the defect that the traditional single-stream structure can only receive image features of a single modality in cross-modal registration; and replacing the traditional convolutional layer with a group-equivariant dynamic snake convolutional module, each group-equivariant dynamic snake convolutional module includes: a feature enhancement branch, a group-equivariant dynamic snake convolutional block, a third group-equivariant convolution, group normalization, and a Relu activation function; the feature enhancement branch includes: a first group-equivariant convolution, group normalization, and a Relu activation function connected in sequence along the forward propagation direction; the feature enhancement branch processes and enhances the input features through the first group-equivariant convolution, group normalization, and Relu activation function, thereby enhancing the ability to express key information in the image and obtaining enhanced features; the group-equivariant dynamic snake convolutional block significantly improves the robustness and accuracy of the model to spatial transformations between images (such as translation, rotation, etc.) through offset modeling and directional processing, and helps the feature encoding module capture directional information in the image by extracting features in the horizontal and vertical directions respectively, improving the feature extraction ability of local key point regions, thereby improving the image registration accuracy and robustness, and obtaining group-equivariant output features in the horizontal direction and group-equivariant output features in the vertical direction; after splicing the enhanced features, the group-equivariant output features in the horizontal direction, and the group-equivariant output features in the vertical direction along the channel dimension, and then processing them through a third group-equivariant convolution and group normalization in sequence, the output of the group-equivariant dynamic snake convolutional module is obtained; effectively improving the separation degree of the feature encoder for low-contrast key point regions, thereby improving the registration accuracy of multimodal retinal fundus images.
[0060] In addition, the present invention respectively inputs the reference rotation-equivariant features of four stages and the floating rotation-equivariant features of four stages into a key point detection decoder for feature fusion, and outputs a reference key point prediction map and a floating key point prediction map. The rotation-equivariant features of the fourth stage are sequentially passed through a group-equivariant dynamic snake convolutional module and group upsampling to obtain key point decoding features of the third level. In order to reduce the overfitting phenomenon caused by the network being too deep, after channel splicing the key point decoding features of the third level and the rotation-equivariant features of the second stage, and then passing through a group-equivariant convolutional module and group upsampling in sequence, key point decoding features of the second level are obtained; after channel splicing the key point decoding features of the second level after group upsampling and the rotation-equivariant features of the first stage, the key point decoding features of the first stage are obtained; for the key point decoding features of the first stage, the features are in After stacking all the features on the dimension to the channel dimension, they are successively passed through a fully connected layer and a sigmoid activation function to obtain a key point prediction map. Stacking the features on the dimension to the channel dimension enables the model to capture richer spatial and semantic information in the rotation angle dimension space and be more rotationally robust to rotational transformations. The key point detection decoder proposed in the present invention uses group-equivariant dynamic snake convolution blocks and group-equivariant convolution modules for channel splicing and high-dimensional feature stacking, enabling the model to better capture multi-scale and multi-directional feature information and making the model more rotationally robust to rotational transformations, thereby improving the robustness and accuracy of the key point detection model. Brief Description of the Drawings
[0061] To make the content of the present invention easier to understand clearly, the following further elaborates on the present invention in detail according to specific embodiments of the present invention in conjunction with the drawings, where:
[0062] Figure 1 is a structural diagram of the key point detection model and the registration module.
[0063] Figure 2 is a structural diagram of the feature encoder.
[0064] Figure 3 is a structural diagram of the key point detection decoder.
[0065] Figure 4 is a structural diagram of the key point descriptor decoder.
[0066] Figure 5 is a structural diagram of the group-equivariant dynamic snake convolution module.
[0067] Figure 6 is a structural diagram of the group-equivariant dynamic snake convolution block.
[0068] Figure 7 is a registration checkerboard diagram of a fundus color photo and a fluorescein fundus angiography image, Figure 7 where (a) in it is the fundus color photo, Figure 7 where (b) in it is the fluorescein fundus angiography image, Figure 7 where (c) in it is the registration result of the fundus color photo and the fluorescein fundus angiography image through scale-invariant feature transform, Figure 7 where (d) in it is the registration result of the fundus color photo and the fluorescein fundus angiography image through Orb, Figure 7 where (e) in it is the registration result of the fundus color photo and the fluorescein fundus angiography image through the Harris corner detector, Figure 7 where (f) in it is the registration result of the fundus color photo and the fluorescein fundus angiography image through the SuperRetina key point detection network, Figure 7In (g) is the registration result of the fundus color photo and the fluorescein fundus angiography image through the dynamic snake network, Figure 7 In (h) is the registration result of the fundus color photo and the fluorescein fundus angiography image through the key point detection model.
[0069] Figure 8 is the registration checkerboard pattern of the fundus color photo and the optical coherence tomography angiography image, Figure 8 In (a) is the fundus color photo, Figure 8 In (b) is the optical coherence tomography angiography image, Figure 8 In (c) is the registration result of the fundus color photo and the fluorescein fundus angiography image through scale-invariant feature transform, Figure 8 In (d) is the registration result of the fundus color photo and the fluorescein fundus angiography image through Orb, Figure 8 In (e) is the registration result of the fundus color photo and the fluorescein fundus angiography image through the Harris corner detector, Figure 8 In (f) is the registration result of the fundus color photo and the fluorescein fundus angiography image through the SuperRetina key point detection network, Figure 8 In (g) is the registration result of the fundus color photo and the fluorescein fundus angiography image through the dynamic snake network, Figure 8 In (h) is the registration result of the fundus color photo and the fluorescein fundus angiography image through the key point detection model. Specific implementation manner
[0070] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments given are not intended to limit the present invention.
[0071] Embodiment 1 of the present invention provides a multi-modal retinal fundus image registration method, including:
[0072] Step S11: Construct a key point detection model (G_DSCNet), and the key point detection model includes: a first feature encoder, a second feature encoder, a key point detection decoder, and a key point descriptor decoder;
[0073] As Figure 1 shown, Figure 1 is the structural diagram of the key point detection model and the registration module.
[0074] As Figure 2 shown, Figure 2It is the structural diagram of the feature encoder. Among them, each feature encoder includes four stages connected in sequence along the forward propagation direction. Each stage is a group-equivariant dynamic snake-shaped convolution module. A group-equivariant dynamic snake-shaped convolution module includes: a feature enhancement branch, a group-equivariant dynamic snake-shaped convolution block, a third group-equivariant convolution, group normalization, and a Relu activation function; the feature enhancement branch includes: a first group-equivariant convolution, group normalization, and a Relu activation function connected in sequence along the forward propagation direction; each group-equivariant dynamic snake-shaped convolution module outputs the rotation-equivariant features of one stage; for the rotation-equivariant features of other stages except the rotation-equivariant features of the first stage, their sizes are adjusted by group transposed convolution to be the same as the rotation-equivariant features of the first stage; the rotation-equivariant features of all stages are concatenated along the channel dimension to obtain the total rotation-equivariant feature map.
[0075] Step S12: Copy the reference image for a certain number of times to generate a reference image structure, and pass the reference image structure through the first feature encoder to extract the reference rotation-equivariant features of four stages and the total reference rotation-equivariant feature map; among them, is the rotation angle dimension of the group-equivariant convolution;
[0076] Step S13: Copy the floating image for a certain number of times to generate a floating image structure, and pass the floating image structure through the second feature encoder to extract the floating rotation-equivariant features of four stages and the total floating rotation-equivariant feature map;
[0077] In this embodiment, specifically, the reference image is a fundus color photo, and the floating image is any one of a fundus fluorescein angiography image and an optical coherence tomography angiography image.
[0078] Step S14: Input the reference rotation-equivariant features of four stages into the key point detection decoder for feature fusion, and output the reference key point prediction map;
[0079] Step S15: Input the floating rotation-equivariant features of four stages into the key point detection decoder for feature fusion, and output the floating key point prediction map;
[0080] As Figure 3 shown, Figure 3 is the structural diagram of the key point detection decoder. In this embodiment, preferably, input the rotation-equivariant features of four stages into the key point detection decoder for feature fusion, and output the key point prediction map, including:
[0081] Step 151: Pass the rotation-equivariant features of the fourth stage sequentially through a group-equivariant dynamic snake-shaped convolution module and group upsampling to obtain the key point decoding features of the fourth level ;
[0082] Step 152: Decode the key-point features of the fourth level and the rotation-equivariant features of the third stage After performing channel concatenation, successively pass through the group-equivariant dynamic snake convolution module and group upsampling to obtain the key-point decoding features of the third level ;
[0083] Step 153: Decode the key-point features of the third level and the rotation-equivariant features of the second stage After performing channel concatenation, successively pass through the group-equivariant convolution module and group upsampling to obtain the key-point decoding features of the second level ;
[0084] Step 154: Decode the key-point features of the second level and the rotation-equivariant features of the first stage After performing channel concatenation, obtain the key-point decoding features of the first level;
[0085] Step 155: For the key-point decoding features of the first level, stack all the features in the dimension to the channel dimension, and then successively pass through a fully connected layer and a sigmoid activation function to obtain the key-point prediction map.
[0086] The key-point detection decoder proposed by the present invention can better capture multi-scale and multi-directional feature information and is robust to rotation transformation by using group-equivariant dynamic snake convolution blocks and group-equivariant convolution modules for channel concatenation and high-dimensional feature stacking, improving the robustness and accuracy of the key-point detection model.
[0087] Step S16: Input the total reference rotation-equivariant feature map and the total floating rotation-equivariant feature map into the key-point descriptor decoder respectively, and output the descriptor of each reference key point and the descriptor of each floating key point;
[0088] As Figure 4 shown, Figure 4 is the structural diagram of the key-point descriptor decoder. In this embodiment, specifically, input the total rotation-equivariant feature map into the key-point descriptor decoder, and output the descriptor of each key point, including:
[0089] Step 161: Based on the key-point prediction map, obtain the coordinates of each key point; according to the coordinates of each key point , extract the pixel values of each channel corresponding to the position of the key point in the total rotation-equivariant feature map to construct the initial descriptor of each key point , is the number of channels of the total rotation equivariant feature map, is the initial descriptor of the th key point,
[0090] Step 162: Pass the initial descriptor of each key point through batch normalization and a 1×1 convolutional layer in sequence, flatten the initial descriptor of each key point into , and obtain the descriptor of each key point through the Sigmoid activation function, a 3×3 convolutional layer, and a fully connected layer .
[0091] Step S17: Input the reference image, floating image, reference key point prediction map, floating key point prediction map, descriptors of all reference key points, and descriptors of all floating key points into the registration module, and output the registered image.
[0092] In this embodiment, preferably, the step of inputting the reference image, floating image, reference key point prediction map, floating key point prediction map, descriptors of all reference key points, and descriptors of all floating key points into the registration module and outputting the registered image includes:
[0093] Step S171: Use the flann algorithm to perform key point matching on the reference key point prediction map and descriptors of all reference key points, and the floating key point prediction map and descriptors of all floating key points, to obtain key point pairs between the reference image and the floating image;
[0094] Step S172: Use the random sample consensus algorithm to fit the key point pairs between the reference image and the floating image, and remove the key point pairs that do not meet the transformation constraints to obtain the affine matrix;
[0095] Step S173: Align the floating image to the reference image using the affine matrix to obtain the registered image.
[0096] In multi-modal retinal fundus image registration, the registration method based on the FLANN algorithm, the Random Sample Consensus algorithm (RANSAC), and affine transformation has significant advantages. Retinal fundus images usually have different imaging modes, and images under these modes may have differences in illumination, contrast changes, and local deformations. Through efficient FLANN feature matching, corresponding key points can be accurately found between different modalities, avoiding the problem that traditional pixel intensity or gray value-based matching methods are easily interfered by illumination and contrast changes. The Random Sample Consensus algorithm (RANSAC) further eliminates incorrect matches, improves the robustness of registration, and ensures that the final registration result is more accurate. In addition, affine transformation can effectively handle situations such as rotation, translation, and scale transformation in images, ensuring the precise alignment of different-modal images. Compared with other registration methods, this method can not only stably perform image registration under different modalities but also has high computational efficiency, especially suitable for scenarios in clinical practice that require rapid and accurate registration of different types of fundus images.
[0097] In this embodiment, in addition to the above registration method based on the FLANN algorithm, the Random Sample Consensus algorithm (RANSAC), and affine transformation, a spectrum-based registration method can also be used. The spectrum-based registration method is specifically as follows: perform Fourier transforms on the reference image and the floating image respectively to obtain the spectral representation of the reference image in the frequency domain and the spectral representation of the floating image in the frequency domain, calculate the phase difference between the spectra of the reference image and the floating image, and obtain the image translation amount based on the phase difference between the spectra of the reference image and the floating image; use the image translation amount to align the floating image to the reference image to obtain the registered image. It is also possible to use a convolutional neural network (CNN) to predict the registration parameters of the reference image and the floating image, and predict the deformation field or affine matrix between the two images, and use the predicted deformation field or affine matrix to align the floating image to the reference image to obtain the registered image.
[0098] As Figure 5 shown, Figure 5 is the structural diagram of the group-equivariant dynamic snake convolution module. In this embodiment, preferably, the input feature of the group-equivariant dynamic snake convolution module is processed through a feature enhancement branch to obtain an enhanced feature , including:
[0099] The input feature of the group-equivariant dynamic snake convolution module sequentially passes through a convolutional kernel, and the first group-equivariant convolution with a padding stride of 1 to expand the number of channels to , and after group normalization and the Relu activation function, the enhanced feature is obtained; where is the input feature of the group-equivariant dynamic snake convolution module The number of channels after being expanded by the first group of equivariant convolutions.
[0100] The input features of the group-equivariant dynamic snake convolution module are processed by the group-equivariant dynamic snake convolution block to capture spatial information in different directions, and the group-equivariant output features in the horizontal direction are obtained and the group-equivariant output features in the vertical direction ;
[0101] The enhanced features , the group-equivariant output features in the horizontal direction and the group-equivariant output features in the vertical direction are concatenated along the channel dimension and then processed successively by the third group-equivariant convolution, group normalization, and Relu activation function to obtain the output features of the group-equivariant dynamic snake convolution module .
[0102] In the present invention, the traditional convolution layer in the traditional encoder is replaced by a group-equivariant dynamic snake convolution module (G_DSC convolution module), a group-equivariant convolution is used as the basic convolution kernel, and a group-equivariant pooling layer and an upsampling layer are adopted, so that the network has rotational equivariance, ensuring that the attention of the network remains consistent before and after rotation, being able to adapt to the rotation change of the input image, and ensuring the attention consistency of the network before and after rotation. The feature enhancement branch improves the sensitivity of the network to key features and reduces noise interference. By concatenating the group-equivariant output features in the horizontal and vertical directions with the enhanced features along the channel dimension and through further processing, the group-equivariant dynamic snake convolution module (G_DSC convolution module) performs excellently in terms of rotational equivariance, spatial feature capture, and multi-modal information fusion.
[0103] As Figure 6 shown, Figure 6 is the structural diagram of the group-equivariant dynamic snake convolution block. In this embodiment, preferably, the input features of the group-equivariant dynamic snake convolution module are passed through the group-equivariant dynamic snake convolution block to capture spatial information in different directions, and the group-equivariant output features in the horizontal direction and the group-equivariant output features in the vertical direction are obtained, including:
[0104] The input features of the group-equivariant dynamic snake convolution module are successively passed through the second group-equivariant convolution, group normalization, and tanh activation function of the convolution kernel to respectively obtain the bias features in the horizontal direction and the bias features in the vertical direction ; wherein, the value of each pixel in the bias feature in the horizontal direction represents the offset in the horizontal direction, and the value of each pixel in the bias feature in the vertical direction represents the offset in the vertical direction;
[0105] According to the value of each pixel in the horizontal offset feature, each pixel in the input feature of the group-equivariant dynamic snake convolution module is horizontally offset, and then the horizontally offset input feature is obtained through four-neighborhood interpolation. ;
[0106] According to the value of each pixel in the vertical offset feature, each pixel in the input feature of the group-equivariant dynamic snake convolution module is vertically offset, and then the vertically offset input feature is obtained through four-neighborhood interpolation. ;
[0107] The horizontally offset input feature is successively passed through the convolution kernel, group normalization, and Relu activation function in the horizontal direction to obtain the feature map in the horizontal direction; where is a set parameter;
[0108] The vertically offset input feature is successively passed through the convolution kernel, group normalization, and Relu activation function in the vertical direction to obtain the feature map in the vertical direction ;
[0109] The feature map in the horizontal direction and the feature map in the vertical direction are respectively integrated along the dimension to obtain the group-equivariant output feature in the horizontal direction and the group-equivariant output feature in the vertical direction .
[0110] Although the dynamic snake convolution (DSCNet) has good flexibility in capturing the spatial structure in images, traditional convolution operations lack intrinsic adaptability to geometric transformations such as rotation and scale transformation of images. When processing rotated, translated, or other deformed images, ordinary convolution may lead to feature loss or inconsistency, thus affecting the performance of the model. Especially when facing large-scale rotation, tilt, and other deformations, ordinary convolution cannot automatically adapt to the geometric transformation of the image, which limits the generalization ability of the model.
[0111] In the present invention, the ordinary convolution in the dynamic snake network is replaced by a second group-equivariant convolution. Different from ordinary convolution that can only process local features in fixed directions and scales, the second group-equivariant convolution can capture features at different rotation angles and directions by introducing group symmetries (such as geometric transformations like rotation, translation, and reflection) in the convolution operation. This enables the network to simultaneously learn feature representations in multiple directions and scales, thereby improving the diversity of features and enhancing the adaptability to complex visual patterns (such as rotation, tilt, scaling, etc.). Combining the flexibility of the dynamic snake convolution, the key point detection model can more efficiently capture diverse spatial information and improve the overall expression ability.
[0112] In the first embodiment, the blood vessels in different modal fundus images are respectively subjected to feature extraction, and the coordinates of each key point are obtained through the same key point detection decoder. The rotation equivariant features integrated by the feature encoder are used to allocate descriptors on the key point coordinates. Finally, the detected key points and descriptors are used to fit the affine matrix by the flann algorithm and the random sample consensus algorithm (RANSAC), realizing the registration of multi-modal fundus images, which can effectively improve the registration accuracy. Especially when dealing with fundus images with deformations such as rotation and scaling, stable blood vessel features can be accurately extracted and matched, enhancing the robustness of image registration.
[0113] Based on the above first embodiment, in the second embodiment of the present invention, the rotation angle dimension of the group equivariant convolution is set to 8, the feature encoder is set to four stages, and the number of channels of the rotation equivariant features in each stage is respectively set to 16, 32, 64, and 128. is set to 9; taking the reference image or the floating image as the input image, the processing process through the feature encoder and the key point decoder specifically includes the following steps:
[0114] Step S21: Duplicate the input image times to obtain the input image structure , where represents the batch size, is the initial number of channels of the input image, set to 1, and and respectively represent the height and width of the input image.
[0115] Step S22: Pass the input image structure through the feature encoder to extract the rotation equivariant features of four stages and the total rotation equivariant feature map , , is the rotation equivariant feature of the th stage, is the number of channels of the rotation equivariant feature of the th stage, is the stage index, including:
[0116] Step S221: The input image structure passes through the convolution kernel with a padding step of 1 to expand the number of channels to , and after group normalization and the Relu activation function, the enhanced feature of the first stage is obtained.
[0117] Step S222: The input image structure Through the group-equivariant dynamic snake convolution block, the group-equivariant output features in the horizontal direction of the first stage are obtained and the group-equivariant output features in the vertical direction , including:
[0118] Step S2221: Input feature structure Through The second group-equivariant convolution, group normalization, and tanh activation function respectively obtain the bias features in the horizontal and vertical directions , ;
[0119] Step S2222: According to the value of each pixel in the horizontal bias feature, each pixel in the input feature structure is offset horizontally, and the horizontally biased input feature is obtained through four-neighborhood interpolation ;
[0120] Step S2223: According to the value of each pixel in the vertical bias feature, each pixel in the input feature structure is offset vertically, and the vertically biased input feature is obtained through four-neighborhood interpolation ;
[0121] Step S2224: The horizontally biased input feature is successively passed through the convolution kernel, group normalization, and Relu activation function in the horizontal direction to obtain the feature map in the horizontal direction;
[0122] Step S2225: The vertically biased input feature is successively passed through the convolution kernel, group normalization, and Relu activation function in the vertical direction to obtain the feature map in the vertical direction;
[0123] Step S2226: The feature maps in the horizontal and vertical directions are respectively integrated along the dimension to obtain the group-equivariant output features in the horizontal direction of the first stage and the group-equivariant output features in the vertical direction .
[0124] Step S223: The enhanced features of the first stage , the group-equivariant output features in the horizontal direction of the first stage and the group-equivariant output features in the vertical direction of the first stage are concatenated along the channel dimension to obtain the feature with the number of channels , Through the third group-equivariant convolution with a 1×1 convolutional kernel and group normalization, the number of channels is changed from to , obtaining the rotation-equivariant features of the first stage . After group-equivariant downsampling, it serves as the input feature for the next group-equivariant dynamic snake convolution module.
[0125] Step S224: For the rotation-equivariant features of other stages except those of the first stage, their sizes are adjusted respectively through group transposed convolution to be the same as those of the rotation-equivariant features of the first stage; the rotation-equivariant features of all stages are concatenated along the channel dimension to obtain the total rotation-equivariant feature map.
[0126] Step S23: The rotation-equivariant features of the four stages are input into the key point detection decoder for feature fusion, and a key point prediction map is output, including:
[0127] Step S231: The rotation-equivariant features of the fourth stage After passing through the group-equivariant dynamic snake convolution module and group upsampling, the key point decoding features of the fourth level are obtained;
[0128] Step S232: The key point decoding features of the fourth level are concatenated with the rotation-equivariant features of the third stage along the channel dimension to obtain the third fusion feature , where , is the number of channels of the third fusion feature ;
[0129] Step S233: The third fusion feature is successively passed through the group-equivariant dynamic snake convolution module and group upsampling to obtain the key point decoding features of the third level ;
[0130] Step S234: The key point decoding features of the third level are concatenated with the rotation-equivariant features of the second stage along the channel dimension to obtain the second fusion feature , where , is the number of channels of the second fusion feature ;
[0131] Step S235: The second fusion feature is successively passed through the group-equivariant convolution module and group upsampling to obtain the key point decoding features of the second level ;
[0132] Step S236: Decode the key-point features of the second level and the rotation-equivariant features of the first stage to perform channel concatenation to obtain the key-point decoding features of the first level , where , is the number of channels of the key-point decoding features of the first level ;
[0133] Step S237: For the key-point decoding features of the first level, after stacking all the features of this feature in the dimension to the channel dimension, sequentially pass through a fully connected layer and a sigmoid activation function to obtain the key-point prediction map .
[0134] In the second embodiment, the rotation angle dimension of the group-equivariant convolution is set to 8. Dividing the feature encoder into four stages can effectively balance the computational complexity and the performance of the key-point detection model. The number of channels of the rotation-equivariant features of the four stages is set to 16, 32, 64, and 128, showing an exponential growth pattern, which can extract gradually rich feature information at different levels, and at the same time ensure that the expression ability of the model gradually increases. As the number of channels increases, the feature encoder can capture more detailed and high-order image information; setting to 9 can capture more extensive context information, reduce the depth of the model, and improve the efficiency and accuracy of the key-point detection model.
[0135] In this embodiment, preferably, the training method of the key-point detection model is as follows:
[0136] Obtain an image data set, where the image data set includes different images and their corresponding multiple ground-truth key points; divide the image data set into a training set, a validation set, and a test set;
[0137] Use the training set to train the key-point detection model, and use the total loss function to optimize the parameters of the key-point detection model. Among them, during the training process, it includes: after obtaining the key-point prediction map of the image of this round and the descriptor of each key point, freeze the network parameters, perform a random affine transformation on the image of this round, re-enter the affine-transformed image into the key-point detection model, and obtain the key-point prediction map corresponding to the image of this round after the affine transformation and the descriptor of each key point, and calculate the Euclidean distance between the descriptors of the th key point in the key-point prediction maps corresponding to the image of this round before and after the affine transformation and the key-point prediction map The descriptor of a key point and the predicted map of the corresponding key point of the current round of image after affine transformation in the Euclidean distance between the descriptors of the second nearest neighbor key points of the th key point. When the ratio of is greater than the set threshold, set this key point as the gold standard key point for the next round;
[0138] After the key point detection model is trained, the key points detected by the key point detection model are registered with the descriptors through the registration module.
[0139] In the multi-modal fundus image registration task, the gold standard of key points is scarce and difficult to produce, mainly because of the complex structure, large noise and differences between different modalities in fundus images. These factors make it very challenging to manually annotate key points. The present invention expands the gold standard key points by calculating the Euclidean distance based on affine transformation and key point descriptors in each round of training, and screens the key points through ratio determination, which can effectively avoid misregistration and improve the image registration accuracy.
[0140] In this embodiment, preferably, the total loss function includes: Dice loss function and triplet loss function;
[0141] The calculation formula of the total loss function is:
[0142] ,
[0143] ,
[0144] ,
[0145] wherein, is the total loss function of the key point detection model, is the Dice loss function, is the triplet loss function, is the predicted map of key points, is the key point prediction map corresponding to the image rotation angle, is the soft label obtained by applying the Gaussian Boolean function to the key point map of the gold standard, is element-wise multiplication, is the maximum value function, is the descriptor of the th key point in the image and the th key point corresponding to the image after rotation by angle, is the threshold for dividing positive and negative samples, is the Descriptor of a Key Point and Image Rotation and the descriptor of any other key point corresponding to the angle The Euclidean distance between them is the maximum Euclidean distance between the descriptor of the key point in the image and the descriptors of all other key points corresponding to the image rotation angle.
[0146] In the third embodiment of the present invention, in order to verify the effectiveness and generality of the method of the present invention, the method was verified using tasks such as key point detection and registration of multimodal fundus images, including registration of fundus color photographs and fluorescein fundus angiography (FFA) images, and registration of fundus color photographs and optical coherence tomography angiography (OCTA) images.
[0147] In this embodiment, a diabetic retinopathy data set containing 400 pairs of fundus color photographs and fluorescein fundus angiography (FFA) images was used for method verification. The data set was divided into 240 pairs of training sets, 80 pairs of validation sets, and 80 pairs of test sets. Online random contrast transformation, brightness transformation, left-right flipping, and up-down flipping were used for data augmentation. The success rate of registration was used as the evaluation index. All experiments used the flann method to pair the key points and used the random sample consensus algorithm (RANSAC) to fit the affine matrix to achieve image registration.
[0148] In the comparative experiment, the method of the present invention was compared with traditional key point detection methods, including scale-invariant feature transform (SIFT), ORB (Oriented FAST and Rotated BRIEF), Harris corner detector, and key point detection and description networks based on convolutional neural networks, such as SuperRetina key point detection network and dynamic snake convolutional network (DSCNet). Among them, the dynamic snake convolutional network (DSCNet) is a vascular segmentation network based on dynamic snake convolution rather than a key point detection and description network. In order to verify the effectiveness of the present invention and ensure fairness, this network structure was applied to the training strategy of SuperRetina for autonomous training.
[0149] Since the modality differences between fundus color photographs and fluorescein fundus angiography (FFA) images are relatively small, and they can be regarded as images with the same distribution after data augmentation by limited contrast adaptive histogram equalization (CLAHE), all deep learning methods adopt a single-stream encoder structure and are trained only using fundus color photograph images. Traditional methods and deep learning methods use the same data augmentation method. As shown in Table 1, Table 1 presents the comparative experimental results of key point detection and registration for fundus color photographs and fluorescein fundus angiography (FFA) images.
[0150] Table 1
[0151]
[0152] As can be seen from Table 1, traditional key point detection methods such as scale-invariant feature transform (SIFT), ORB (Oriented FAST and Rotated BRIEF), and Harris corner detector can detect more key points (here, the number of SIFT and Orb key points is manually set to 1000, and the number of Harris corner points is set to 200. Due to different implementation methods, the number of Harris corner points is usually less than that of SIFT and Orb key points). The registration success rates are 36.02%, 30.14%, and 37.02% respectively. Compared with traditional methods, SuperRetina learns key points on blood vessels in a self-supervised manner, uses the features of the fourth stage of the encoder to extract key points and descriptors, and detects fewer key points. However, its robustness and effectiveness are improved, and the registration success rate reaches 45.02%. The dynamic snake convolutional network (DSCNet) uses dynamic snake convolution to guide the network to effectively focus on the blood vessel area, reducing the number of key points, but further improving the robustness and effectiveness, and the registration success rate is increased to 55.00%. The present invention utilizes the rotational equivariance of group equivariant convolution to ensure the rotational robustness of the extracted descriptor features when self-supervised learning is performed with the input image rotated at random angles, and integrates the rotational equivariant features of multiple stages of the feature encoder to extract key points and descriptors. Therefore, the robustness of the key points and their descriptors is improved, and the registration success rate reaches 63.00%, demonstrating the effectiveness of the method of the present invention.
[0153] As Figure 7 shown, Figure 7 is the registration checkerboard for fundus color photographs and fluorescein fundus angiography images. Figure 7 (a) in it is the fundus color photograph, Figure 7 (b) in it is the fluorescein fundus angiography image, Figure 7Among them, (c) is the registration result of the fundus color photo and the fundus fluorescein angiography image through Scale-Invariant Feature Transform (SIFT). Figure 7 Among them, (d) is the registration result of the fundus color photo and the fundus fluorescein angiography image through Orb. Figure 7 Among them, (e) is the registration result of the fundus color photo and the fundus fluorescein angiography image through Harris corner detector. Figure 7 Among them, (f) is the registration result of the fundus color photo and the fundus fluorescein angiography image through SuperRetina key point detection network. Figure 7 Among them, (g) is the registration result of the fundus color photo and the fundus fluorescein angiography image through Dynamic Snake Network. Figure 7 Among them, (h) is the registration result of the fundus color photo and the fundus fluorescein angiography image through the key point detection model. Figure 7 The registration results of different methods are shown. Many key points are recognized by traditional methods, but they are not key points of vascular features. Therefore, the key point matching effect is poor. SuperRetina and Dynamic Snake Convolutional Network (DSCNet) successfully match the key points and fit the affine matrix. However, from the checkerboard diagram, it can be seen that there are overlapping parts of blood vessels showing connectivity, while other parts do not overlap showing interruption. Finally, the deviation of the obtained affine matrix is relatively large. The key point detection model (G_DSCNet) proposed by the present invention has the highest degree of blood vessel overlap and the most connected regions.
[0154] In this embodiment, a dataset containing 100 pairs of fundus color photo images of pathological myopia and Optical Coherence Tomography Angiography (OCTA) images is also used to evaluate the performance of the present invention. Due to the small number of samples, a five-fold cross-validation method is used. Online random contrast transformation, brightness transformation, left-right flipping, and up-down flipping are used for data augmentation. The number of detected key points and the registration success rate are used as evaluation indicators. All experiments use the flann method to pair the key points and the Random Sample Consensus (RANSAC) algorithm to fit the affine matrix to achieve image registration.
[0155] In the comparative experiment, all deep learning networks used a two-stream encoder. However, due to the large modal differences, the vascular density of Optical Coherence Tomography Angiography (OCTA) images is much higher than that of fundus color photographs and there is noise, and the image distributions are completely different. Therefore, traditional key-point detection algorithms cannot correctly fit the affine matrix and apply it to registration. The results are shown in Table 2, which is the comparative experiment results of key-point detection and registration for fundus color photographs and Optical Coherence Tomography Angiography (OCTA) images.
[0156] Table 2
[0157]
[0158] As can be seen from Table 2, for SuperRetina, because it uses the original U-shaped encoding and decoding structure, in the case of fewer training samples, it cannot effectively learn the vascular characteristics of OCTA data. The key-point detection model (G_DSCNet) proposed in the present invention, while retaining the effectiveness of key-point detection for the vascular modality of Optical Coherence Tomography Angiography (OCTA), can make the descriptor more accurate and rotationally robust compared to the Dynamic Snake-like Network (DSCNet), so the registration success rate is higher.
[0159] As Figure 8 shown, Figure 8 is the registration checkerboard diagram of fundus color photographs and Optical Coherence Tomography images. Figure 8 In (a) is the fundus color photograph, Figure 8 in (b) is the Optical Coherence Tomography image, Figure 8 in (c) is the registration result of the fundus color photograph and the fluorescein fundus angiography image by Scale-Invariant Feature Transform, Figure 8 in (d) is the registration result of the fundus color photograph and the fluorescein fundus angiography image by Orb, Figure 8 in (e) is the registration result of the fundus color photograph and the fluorescein fundus angiography image by Harris corner detector, Figure 8 in (f) is the registration result of the fundus color photograph and the fluorescein fundus angiography image by the SuperRetina key-point detection network, Figure 8 in (g) is the registration result of the fundus color photograph and the fluorescein fundus angiography image by the Dynamic Snake-like Network, Figure 8 in (h) is the registration result of the fundus color photograph and the fluorescein fundus angiography image by the key-point detection model.Figure 8 The figure shows the registration results of fundus color photos and Optical Coherence Tomography Angiography (OCTA) images using different methods. The results show that the traditional method cannot correctly match key points. SuperRetina fails to learn the vascular key point information of the OCTA modality, resulting in incorrect registration. Moreover, the key point detection model (G_DSCNet) proposed in the present invention has a smaller difference in the misaligned part compared to the Dynamic Snake Network (DSCNet), that is, the Root Mean Square Error (RMSE) of the key points is smaller.
[0160] The present invention performs well in the registration tasks of fundus color photos and Fluorescein Fundus Angiography (FFA) images, as well as fundus color photos and Optical Coherence Tomography Angiography (OCTA) images, indicating the applicability of the present method in multi-modal image registration tasks.
[0161] Embodiment 4 of the present invention provides a multi-modal retinal fundus image registration system, including:
[0162] A model construction module, configured to construct a key point detection model, where the key point detection model includes: a first feature encoder, a second feature encoder, a key point detection decoder, and a key point descriptor decoder;
[0163] Among them, each feature encoder includes four stages connected in sequence along the forward propagation direction. Each stage is a group-equivariant dynamic snake convolution module. A group-equivariant dynamic snake convolution module includes: a feature enhancement branch, a group-equivariant dynamic snake convolution block, a third group-equivariant convolution, group normalization, and a Relu activation function; the feature enhancement branch includes: a first group-equivariant convolution, group normalization, and a Relu activation function connected in sequence along the forward propagation direction; each group-equivariant dynamic snake convolution module outputs a stage of rotation-equivariant features; for the rotation-equivariant features of other stages except the rotation-equivariant features of the first stage, their sizes are adjusted through group transposed convolution to make their sizes consistent with the rotation-equivariant features of the first stage; the rotation-equivariant features of all stages are concatenated along the channel dimension to obtain a total rotation-equivariant feature map;
[0164] A reference rotation-equivariant feature acquisition module, configured to copy the reference image for a certain number of times to generate a reference image structure, and pass the reference image structure through the first feature encoder to extract the reference rotation-equivariant features of four stages and the total reference rotation-equivariant feature map; where is the rotation angle dimension of the group-equivariant convolution;
[0165] The floating rotation equivariant feature acquisition module is used to copy the floating image in multiple passes to generate a floating image structure, and extract the floating rotation equivariant features and the total floating rotation equivariant feature map in four stages through the second feature encoder for the floating image structure;
[0166] The reference key point prediction map acquisition module is used to input the reference rotation equivariant features in four stages into the key point detection decoder for feature fusion and output the reference key point prediction map;
[0167] The floating key point prediction map acquisition module is used to input the floating rotation equivariant features in four stages into the key point detection decoder for feature fusion and output the floating key point prediction map;
[0168] The key point descriptor acquisition module is used to input the total reference rotation equivariant feature map and the total floating rotation equivariant feature map into the key point descriptor decoder respectively, and output the descriptor of each reference key point and the descriptor of each floating key point;
[0169] The image registration module is used to input the reference image, the floating image, the reference key point prediction map, the floating key point prediction map, the descriptors of all reference key points, and the descriptors of all floating key points into the registration module and output the registered image.
[0170] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0171] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0172] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in one or more of the blocks.
[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in one or more of the blocks.
[0174] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A multimodal retinal fundus image registration method, characterized in that: The following steps are involved: Constructing a key point detection model, the key point detection model comprising: a first feature encoder, a second feature encoder, a key point detection decoder and a key point descriptor decoder; Among them, each feature encoder includes four stages connected in sequence along the forward propagation direction, each stage is a group equivariant dynamic snake convolution module, and each group equivariant dynamic snake convolution module outputs the rotation equivariant features of one stage respectively; for the rotation equivariant features of other stages except the rotation equivariant features of the first stage, the sizes are adjusted by group transpose convolution respectively so that their sizes are consistent with the rotation equivariant features of the first stage; the rotation equivariant features of all stages are spliced along the channel dimension to obtain the total rotation equivariant feature map; A group equivariant dynamic snake convolution module includes: The input features of the group equivariant dynamic snake convolution module are processed through a feature enhancement branch to obtain enhanced features; the feature enhancement branch includes: a first group equivariant convolution, group normalization, and Relu activation function sequentially connected along the forward propagation direction; The input features of the group equivariant dynamic snake convolution module are processed by the group equivariant dynamic snake convolution block to capture spatial information in different directions and obtain the group equivariant output features in the horizontal direction and the group equivariant output features in the vertical direction; After concatenating the enhanced features, the horizontal group equivariant output features, and the vertical group equivariant output features along the channel dimension, they are processed in turn through the third group equivariant convolution, group normalization, and Relu activation function to obtain the output features of the group equivariant dynamic snake convolution module; Copy the reference image A reference image structure is generated, and the reference image structure is passed through the first feature encoder to extract the reference rotation equivariant features of the four stages and the total reference rotation equivariant feature map; wherein, is the rotation angle dimension of the group equivariant convolution; Copy the floating image A floating image structure is generated, and the floating image structure is passed through a second feature encoder to extract floating rotation equivariant features of four stages and a total floating rotation equivariant feature map; The reference rotation equivariant features of the four stages are input into the key point detection decoder for feature fusion, and the reference key point prediction map is output; The floating rotation equivariant features of the four stages are input into the key point detection decoder for feature fusion, and the floating key point prediction map is output; The total reference rotation equivariant feature map and the total floating rotation equivariant feature map are input into the key point descriptor decoder respectively, and the descriptor of each reference key point and the descriptor of each floating key point are output; The reference image, the floating image, the reference key point prediction map, the floating key point prediction map, the descriptors of all reference key points, and the descriptors of all floating key points are input into the registration module, and the registered image is output.
2. A multimodal retinal fundus image registration method according to claim 1, characterized in that: The rotation equivariant features of the four stages are input into the key point detection decoder for feature fusion, and the key point prediction map is output, including: The rotation equivariant features of the fourth stage are sequentially passed through the group equivariant dynamic snake convolution module and group upsampling to obtain the key point decoding features of the fourth level; After channel concatenation of the key point decoding features of the fourth level and the rotation equivariant features of the third stage, the key point decoding features of the third level are obtained by sequentially passing through the group equivariant dynamic snake convolution module and group upsampling; After channel concatenation of the key point decoding features of the third level and the rotation equivariant features of the second stage, the key point decoding features of the second level are obtained by sequentially passing through the group equivariant convolution module and group upsampling; After channel concatenation of the key point decoding features of the second level and the rotation equivariant features of the first stage, the key point decoding features of the first level are obtained; For the first level of key point decoding features, the feature is After all the features in the dimension are stacked to the channel dimension, they pass through the fully connected layer and the sigmoid activation function in sequence to obtain the key point prediction map.
3. A multimodal retinal fundus image registration method according to claim 1, characterized in that: The input features of the group equivariant dynamic snake convolution module are passed through the group equivariant dynamic snake convolution block to capture spatial information in different directions, and obtain group equivariant output features in the horizontal direction and group equivariant output features in the vertical direction, including: The input features of the group equivariant dynamic snake convolution module are passed through The second group of equivariant convolution, group normalization and tanh activation function of the convolution kernel respectively obtains the horizontal bias feature and the vertical bias feature; wherein the value of each pixel in the horizontal bias feature represents the offset in the horizontal direction, and the value of each pixel in the vertical bias feature represents the offset in the vertical direction; According to the value of each pixel in the horizontal bias feature, each pixel in the input feature of the group equivariant dynamic snake convolution module is offset in the horizontal direction, and the horizontal bias input feature is obtained through four-neighborhood interpolation; According to the value of each pixel in the bias feature in the vertical direction, each pixel in the input feature of the group equivariant dynamic snake convolution module is offset in the vertical direction, and the bias input feature in the vertical direction is obtained through four-neighborhood interpolation; The horizontal bias input features are sequentially passed through the horizontal direction The convolution kernel, group normalization and Relu activation function are used to obtain the horizontal feature map; among them, To set parameters; The vertical offset input features are passed through the vertical direction in sequence The convolution kernel, group normalization and Relu activation function are used to obtain the feature map in the vertical direction; The horizontal feature map and the vertical feature map are respectively The dimensions are integrated to obtain the group equivariant output features in the horizontal direction and the group equivariant output features in the vertical direction.
4. A multimodal retinal fundus image registration method according to claim 1, characterized in that: The total rotation equivariant feature map is input into the key point descriptor decoder, and the descriptor of each key point is output, including: Based on the key point prediction map, the coordinates of each key point are obtained; according to the coordinates of each key point, the pixel value of each channel corresponding to the position of the key point in the total rotation equivariant feature map is extracted to construct the initial descriptor of each key point; The initial descriptor of each key point is passed through batch normalization, 1×1 convolution layer, Sigmoid activation function, 3×3 convolution layer and fully connected layer in sequence to obtain the descriptor of each key point.
5. A multimodal retinal fundus image registration method according to claim 1, characterized in that: The step of inputting the reference image, the floating image, the reference key point prediction map, the floating key point prediction map, the descriptors of all reference key points, and the descriptors of all floating key points into the registration module and outputting the registered image comprises: The reference key point prediction map and the descriptors of all reference key points, the floating key point prediction map and the descriptors of all floating key points are matched by the Flann algorithm to obtain the key point pairs between the reference image and the floating image; The key point pairs between the reference image and the floating image are fitted using a random sampling consensus algorithm, and the key point pairs that do not meet the transformation constraints are eliminated to obtain an affine matrix; The floating image is aligned to the reference image using an affine matrix to obtain the registered image.
6. A multimodal retinal fundus image registration method according to claim 1, characterized in that: The reference image is a fundus color photograph, and the floating image is any one of a fundus fluorescein angiography image and an optical coherence tomography angiography image.
7. A multimodal retinal fundus image registration method according to claim 1, characterized in that: The training method of the key point detection model is: Acquire an image dataset, wherein the image dataset includes different images and a plurality of corresponding gold standard key points; Divide the image dataset into training set, validation set and test set; The key point detection model is trained using the training set, and the parameters of the key point detection model are optimized using the total loss function. The training process includes: After obtaining the descriptors of each key point, freeze the network parameters, perform a random affine transformation on the image of this round, re-input the image after affine transformation into the key point detection model, and obtain the key point prediction map corresponding to the image of this round after affine transformation. and the descriptor of each key point, calculate the key point prediction map corresponding to the image in this round before and after the affine transformation The Euclidean distance of the descriptor of the key point Key point prediction map corresponding to the current round image before affine transformation Middle The descriptor of the key points and the prediction map of the key points corresponding to the image in this round after affine transformation Middle The Euclidean distance between the descriptors of the next nearest neighbor key points of a key point ,when When the ratio is greater than the set threshold, the key point is set as the gold standard key point as the gold standard key point for the next round; After the key point detection model is trained, the key points detected by the key point detection model are registered with the descriptors through the registration module.
8. A multimodal retinal fundus image registration method according to claim 7, characterized in that: The total loss function includes: Dice loss function and ternary loss function; The calculation formula of the total loss function is: , , , in, is the total loss function of the key point detection model, is the Dice loss function, is the ternary loss function, is the key point prediction map, Rotate the image Key point prediction map corresponding to the angle, is the soft label obtained by the Gaussian Boolean function of the key point graph of the gold standard, is element-wise multiplication, is the maximum value function, For the image Descriptors of key points and image rotation The corresponding angle The Euclidean distance between the descriptors of the key points is is the threshold for dividing positive samples and negative samples, For the image Descriptors of key points and image rotation Any other key point corresponding to the angle The Euclidean distance between the descriptors of For the image Descriptors of key points and image rotation The maximum Euclidean distance between the descriptors of all other key points corresponding to the angle.
9. A multimodal retinal fundus image registration system, characterized in that: include: A model construction module is used to construct a key point detection model, wherein the key point detection model includes: a first feature encoder, a second feature encoder, a key point detection decoder and a key point descriptor decoder; Among them, each feature encoder includes four stages connected in sequence along the forward propagation direction, each stage is a group equivariant dynamic snake convolution module, and each group equivariant dynamic snake convolution module outputs the rotation equivariant features of one stage respectively; for the rotation equivariant features of other stages except the rotation equivariant features of the first stage, the sizes are adjusted by group transpose convolution respectively so that their sizes are consistent with the rotation equivariant features of the first stage; the rotation equivariant features of all stages are spliced along the channel dimension to obtain the total rotation equivariant feature map; A group equivariant dynamic snake convolution module includes: The input features of the group equivariant dynamic snake convolution module are processed through a feature enhancement branch to obtain enhanced features; the feature enhancement branch includes: a first group equivariant convolution, group normalization, and Relu activation function sequentially connected along the forward propagation direction; The input features of the group equivariant dynamic snake convolution module are processed by the group equivariant dynamic snake convolution block to capture spatial information in different directions and obtain the group equivariant output features in the horizontal direction and the group equivariant output features in the vertical direction; After concatenating the enhanced features, the horizontal group equivariant output features, and the vertical group equivariant output features along the channel dimension, they are processed in turn through the third group equivariant convolution, group normalization, and Relu activation function to obtain the output features of the group equivariant dynamic snake convolution module; Reference rotation equivariant feature acquisition module, used to copy the reference image A reference image structure is generated, and the reference image structure is passed through the first feature encoder to extract the reference rotation equivariant features of the four stages and the total reference rotation equivariant feature map; wherein, is the rotation angle dimension of the group equivariant convolution; Floating rotation equivariant feature acquisition module, used to copy floating images A floating image structure is generated, and the floating image structure is passed through a second feature encoder to extract floating rotation equivariant features of four stages and a total floating rotation equivariant feature map; The reference key point prediction map acquisition module is used to input the reference rotation equivariant features of the four stages into the key point detection decoder, perform feature fusion, and output the reference key point prediction map; The floating key point prediction map acquisition module is used to input the floating rotation equivariant features of the four stages into the key point detection decoder, perform feature fusion, and output the floating key point prediction map; A key point descriptor acquisition module is used to input the total reference rotation equivariant feature map and the total floating rotation equivariant feature map into a key point descriptor decoder, and output a descriptor of each reference key point and a descriptor of each floating key point; The image registration module is used to input the reference image, the floating image, the reference key point prediction map, the floating key point prediction map, the descriptors of all reference key points, and the descriptors of all floating key points into the registration module, and output the registered image.
Citation Information
Patent Citations
Zoom multi-channel microscopic imaging system of eye retina
CN102188231A
Retina fundus image registration method and system
CN115760807A