A fundus image multi-modal registration method and system
Patent Information
- Application Number
- CN202610947677.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]为了解决现有眼底图像配准技术存在的跨模态适配性差、无医学物理约束、与临床需求脱节、低质量图像适配能力不足的核心痛点,本发明提供一种基于多尺度血管特征增强的眼底图像多模态配准方法及系统,其以全尺度血管特征增强为基础,双分支跨模态特征提取为核心,医学双约束损失函数为保障,多时序病变自动标注为临床延伸,构建了从眼底图像输入到临床辅助诊断结果输出的全流程技术闭环,实现了多模态眼底图像的高鲁棒性、高医学合理性精准配准,同时完成了从单一配准功能到临床病程辅助监测的功能升级,为糖尿病视网膜病变、黄斑病变等致盲性眼底疾病的早期诊断、病情监测、疗效评估提供了全新的技术方案,可广泛应用于各级医院眼科诊疗、大规模眼底疾病筛查、远程眼科会诊等场景
(1)本发明先通过图像标准化预处理流程,获取待配准的多模态、多时序眼底图像,完成去噪、ROI提取、灰度与尺寸归一化处理,消除不同设备、不同成像条件导致的基础图像差异,为后续处理提供标准化输入;
Smart Images

Figure CN122820780A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, specifically to a method and system for multimodal registration of fundus images based on multi-scale vascular feature enhancement. It is particularly suitable for high-precision fusion and registration of fundus images of different modalities, such as color fundus photography, optical coherence tomography angiography, and fluorescein fundus angiography, and can be used for computer-aided diagnosis, disease monitoring, and large-scale screening of ophthalmic diseases. Background Technology
[0002] Fundus diseases such as diabetic retinopathy and age-related macular degeneration are leading causes of blindness in adults in my country, and their diagnosis and treatment heavily rely on multimodal fundus imaging technology. Different modalities of imagery, including color fundus photography, OCT (optical coherence tomography), and fluorescein fundus angiography, provide complementary structural and functional information. Image registration is a core prerequisite for achieving multimodal information fusion and multi-temporal disease progression comparison. Fundus image registration technology has mainly gone through two stages: traditional manual feature matching (such as SIFT) and deep learning registration (such as VoxelMorph and TransMorph).
[0003] Patent CN109767459B discloses a fundus image registration method based on an unsupervised convolutional neural network. It enhances blood vessels using Hessian filtering, predicts the deformation field using a single-branch U-Net, and uses MSE as the loss function. However, existing technologies, including the aforementioned patent, still suffer from the following fundamental defects: First, there is the problem of poor cross-modal adaptation capability. Most of them adopt a single-branch feature extraction architecture, which makes it difficult to adapt to the huge differences in brightness and texture between color images, OCT, and imaging images at the same time, resulting in low cross-modal registration accuracy and poor robustness.
[0004] Secondly, there is the problem of lack of medical physical constraints. Existing methods mostly use pixel-level similarity measures such as mean squared error (MSE) as loss functions, focusing only on pixel-level matching and ignoring the core medical feature of fundus vascular topology. This results in registration results that appear reasonable at the pixel level, but misalignment of key structures such as vascular skeletons and bifurcation points, which does not meet the requirements of clinical diagnosis and treatment.
[0005] Another issue is the weak ability to extract microvessels. The basic Hessian filter and its single-scale / equal-weight multi-scale variants have a weak response to microvessels and peripheral vessels, and registration is prone to failure, especially on low-quality (blurred, unevenly illuminated) images.
[0006] Finally, there is the problem of the disconnect between functionality and clinical needs. Existing technology only provides a single function of image alignment, and cannot directly translate the results of accurate registration into the quantitative analysis of lesion changes required in clinical practice. Doctors still need to make manual comparisons, which is inefficient and prone to missed diagnoses. Summary of the Invention
[0007] To address the core pain points of existing fundus image registration technologies, such as poor cross-modal adaptability, lack of medical physical constraints, disconnect from clinical needs, and insufficient adaptation capability for low-quality images, this invention provides a multimodal fundus image registration method and system based on multi-scale vascular feature enhancement. It is based on full-scale vascular feature enhancement, with bi-branch cross-modal feature extraction as its core, a medical dual-constraint loss function as a guarantee, and automatic labeling of multi-temporal lesions as a clinical extension. This constructs a closed-loop technology from fundus image input to clinical auxiliary diagnostic results output, achieving highly robust and medically sound registration of multimodal fundus images. Simultaneously, it upgrades the function from single registration to clinical disease progression monitoring, providing a novel technical solution for early diagnosis, disease monitoring, and efficacy evaluation of blinding fundus diseases such as diabetic retinopathy and macular degeneration. It can be widely applied in ophthalmology diagnosis and treatment at all levels of hospitals, large-scale fundus disease screening, and remote ophthalmological consultations.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: A first aspect of the present invention provides a method for multimodal registration of fundus images based on multi-scale vascular feature enhancement, comprising the following steps: S1. Obtain the reference fundus image and the fundus image to be registered, and perform preprocessing to obtain a standardized preprocessed image; S2. Perform multi-scale vascular feature enhancement on the standardized preprocessed image. By constructing a multi-scale Hessian matrix and using inverse proportional adaptive weights to weight and fuse the filter responses at different scales, a vascular enhancement image and the corresponding vascular binarization mask are generated. S3. Construct a dual-branch feature extraction network, input the enhanced blood vessel image in parallel to the texture feature branch and the blood vessel topology feature branch, extract the local texture features of the image and the global blood vessel topology map features based on the blood vessel skeleton respectively, and fuse the features output by the two branches to obtain a multi-dimensional registration feature map. S4. Construct a dual-constraint loss function that integrates structural similarity loss and vascular Dice loss. Use multi-dimensional registration feature maps and vascular binarization masks as inputs to train and drive a registration network based on the U-Net architecture to predict a smooth and reversible deformation field. Then, based on the deformation field, complete the non-rigid and accurate registration of the image to be registered.
[0009] Furthermore, the multi-scale vascular features in step S2 are enhanced, specifically as follows: For each pixel of the input standardized preprocessed image, construct a Hessian matrix at 10 equally spaced scales within the interval [0.5, 6.0], complete the second-order partial derivative calculation and eigenvalue decomposition, and distinguish blood vessels from the background based on this characteristic; An inversely proportional adaptive weighting coefficient is set for the filter response at different scales. The weighting calculation formula is as follows: In the formula, σ represents the weighting coefficient corresponding to the scale.
[0010] Furthermore, in step S2, the single-scale vascular enhancement response is calculated based on the feature value and adaptive weight, and the pixel-level maximum value fusion is performed on the response maps of 10 scales to obtain the final vascular enhancement image. The formula for calculating single-scale vascular enhancement response is: In the formula, Set the corresponding bright background area to 0 directly to suppress background interference; The larger the ratio, the more pronounced the tubular structure features and the higher the response value of the vascular region. The multi-scale response fusion calculation formula is as follows: In the formula, To enhance the pixel values of the final blood vessel image, maximum value fusion is used to preserve the strongest response of blood vessels across the entire scale.
[0011] Furthermore, the dual-branch feature extraction network in step S3 is specifically as follows: The enhanced blood vessel image is input in parallel to the texture feature branch and the blood vessel topology feature branch by sharing an input layer; The texture feature branch adopts a cascaded structure of 5 layers of 3×3 convolutional modules to obtain local texture feature maps; The vascular topology feature branch first extracts the skeleton of the vascular enhancement image through the Zhang-Suen thinning algorithm to obtain a vascular skeleton map with a single pixel width. Then, it extracts features through a cascaded structure of 3 layers of 7×7 convolutional modules + 2 layers of 11×11 convolutional modules to obtain a global vascular topology feature map. The feature maps output from the two branches are concatenated along the channel dimension to obtain a concatenated feature map. Then, a 1×1 convolutional layer is used for feature dimensionality reduction and weighted fusion to finally output a multi-dimensional registration feature map.
[0012] Furthermore, the double-constraint loss function in step S4 is calculated using the following formula: In the formula, α This is the weighting coefficient, with a value range of 0 < α < 1. For structural similarity loss, For vascular Dice loss; Weighting coefficient αIt can be adaptively adjusted according to the registration scenario. The adjustment formula is: .
[0013] Furthermore, in step S4, the bifurcation points and intersection points of blood vessels are first extracted based on the vascular skeleton map as registration key points. The homography transformation matrix is calculated through feature matching to complete rigid coarse registration of the image to be registered. Then, non-rigid fine registration is completed based on the deformation field predicted by the registration network.
[0014] Furthermore, based on the optimal deformation field predicted by the registration network, a spatial transformation is performed on the coarsely registered image to be registered. The transformed pixel grayscale values are calculated using bilinear interpolation, and the final registered fundus image is output. The calculation formula is as follows: In the formula, For the grayscale values of the registered image pixels, The original image to be registered. for By mapping the deformation field to floating-point coordinates in the image to be registered. The floating-point coordinates are the pixel coordinates of the four integer neighboring pixels.
[0015] Furthermore, following step S4, there is also a step of automatic comparison and annotation of lesions across multiple time periods, specifically as follows: The system acquires registered images of the same object at different time points, calculates the difference matrix using a pixel-by-pixel absolute value difference algorithm, automatically calculates the difference threshold using the Otsu method to filter out candidate difference regions, removes isolated noise regions using morphological opening operations to obtain the real lesion difference regions, and performs visual annotation and quantitative statistics such as quantity and area on the real lesion difference regions to generate a fundus disease course analysis report.
[0016] Furthermore, the preprocessing in step S1 specifically includes: The algorithm sequentially performs adaptive median filtering for noise reduction, Otsu's method combined with morphological operations for ROI extraction, grayscale normalization, and size normalization to eliminate image differences caused by different devices and imaging conditions.
[0017] In a second aspect, the present invention provides a multimodal registration system for fundus images based on multi-scale vascular feature enhancement, comprising a core functional unit, an auxiliary functional unit, and a control computer; The core functional units include an image acquisition module, a preprocessing module, a vascular feature enhancement module, a bi-branch feature extraction module, an image registration module, a multi-temporal lesion comparison module, and a result output module, which are connected in sequence. The auxiliary functional unit includes a GPU hardware acceleration module and a parameter adaptive adjustment module; the GPU hardware acceleration module is connected in parallel with the dual-branch feature extraction module, the image registration module, and the multi-temporal lesion comparison module, and the parameter adaptive adjustment module is bidirectionally connected with the vascular feature enhancement module, the image registration module, and the multi-temporal lesion comparison module, respectively. The control computer is the core computing and control carrier of the system, carrying and scheduling all modules to complete instruction interaction, data transmission and computing processing; The system is used to perform the above methods.
[0018] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention first obtains multimodal and multitemporal fundus images to be registered through an image standardization preprocessing process, and completes denoising, ROI extraction, grayscale and size normalization processing to eliminate the differences in basic images caused by different devices and imaging conditions, and provides standardized input for subsequent processing; (2) An improved multi-scale Hessian filtering algorithm is used to enhance the full-scale vascular features of the preprocessed image. Microvascular and terminal vascular features are enhanced by inverse proportional adaptive weights, while background interference is suppressed simultaneously, thus solving the problems of difficult extraction of microvascular features and poor registration accuracy of low-quality images in the existing technology. (3) Construct a dual-branch feature extraction network. The texture feature branch uses small-sized convolutional kernels to extract local texture features of the image to adapt to the brightness difference of cross-modal images. The blood vessel topology feature branch uses large-sized convolutional kernels combined with blood vessel skeleton extraction to capture global blood vessel topology features such as blood vessel course, intersection points, and bifurcation points, which are used as the core anchor points for cross-modal registration. Then, the features of the two branches are normalized and fused to obtain cross-modal robust multi-dimensional registration features, which solves the problem of multi-modal image feature mismatch from the root. (4) Construct a dual-constraint loss function that fuses SSIM loss and vascular Dice loss. Use the fused multi-dimensional registration features as input to train a U-Net-based registration network to predict the optimal deformation field. First, complete the overall image alignment through rigid coarse registration, and then complete the non-rigid fine registration based on the optimal deformation field. This ensures both the overall image structure matching and the registration result meets the medical physical requirements of clinical diagnosis and treatment through the vascular skeleton overlap constraint. (5) For the registered images of the same patient at different time intervals, the adaptive threshold pixel difference algorithm is used to calculate and filter out the pixel difference area, automatically mark the newly added bleeding points, exudates and other lesion areas, generate difference annotation map and quantitative analysis report and output to the clinical terminal, extending the registration technology to the core clinical scenario of fundus disease course monitoring. (6) The four features of the improved multi-scale Hessian enhancement, bi-branch feature extraction, dual-constraint loss function, and multi-temporal lesion annotation in this invention are not simply superimposed. Ablation experiments have confirmed that the full-scale vascular feature enhancement achieved by the improved multi-scale Hessian filtering provides a clear vascular feature foundation for the bi-branch feature extraction network, solves the core problem of difficult microvascular feature extraction, and improves the global feature extraction accuracy of the vascular topology feature branch by 26%. The bi-branch feature extraction network captures both local texture differences and global vascular topology commonalities, using the vascular topology structure as a stable anchor point for cross-modal registration, fundamentally solving the problem of multi-modal image feature mismatch and improving the cross-modal registration success rate by 39%. The dual-constraint loss function simultaneously constrains the overlap between the overall image structure and the vascular skeleton, ensuring the medical rationality of the registration results and avoiding the problem of pixel matching but vascular mismatch, thus improving the vascular Dice coefficient by 9.8%. The accurate registration results provide a core prerequisite for automatic annotation of multi-temporal lesions, improving the accuracy of lesion difference detection by 37% and reducing the missed diagnosis rate by 75%. The four technical features work together in a coordinated manner to achieve a significant improvement in cross-modal registration performance. Attached Figure Description
[0019] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0020] Figure 1 This is an overall flowchart of the fundus image multimodal registration method based on multi-scale vascular feature enhancement in this invention; Figure 2 This is a schematic diagram of the dual-branch feature extraction network in this invention; Figure 3 This is a schematic diagram illustrating the calculation logic of the fusion loss function in this invention; Figure 4 This is a schematic diagram of the module structure of the fundus image multimodal registration system based on multi-scale vascular feature enhancement in this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] This invention discloses a multimodal registration method for fundus images based on multi-scale vascular feature enhancement, such as... Figure 1As shown, the overall design logic is based on full-scale vascular feature enhancement, with bi-branch cross-modal feature extraction as the core, medical dual-constraint loss function as the guarantee, and automatic labeling of multi-temporal lesions as a clinical extension. It constructs a closed-loop technology for the entire process from fundus image input to clinical auxiliary diagnostic results output, solving the core pain points of existing fundus image registration technology such as poor cross-modal adaptability, lack of medical physical constraints, disconnect from clinical needs, and insufficient adaptability to low-quality images.
[0023] It includes the following steps: 1. First, the image standardization preprocessing workflow is used to obtain multimodal and multitemporal fundus images to be registered, and denoising, ROI extraction, grayscale and size normalization are completed to eliminate the differences in basic images caused by different devices and imaging conditions, and to provide standardized input for subsequent processing; 2. An improved multi-scale Hessian filtering algorithm is adopted to enhance the full-scale vascular features of the preprocessed image. Microvascular and terminal vascular features are enhanced by inverse proportional adaptive weights, while background interference is suppressed simultaneously, which solves the problems of difficult extraction of fine vascular features and poor registration accuracy of low-quality images in the existing technology. 3. Construct a dual-branch feature extraction network: The texture feature branch uses small-sized convolutional kernels to extract local texture features of the image, adapting to the brightness differences of cross-modal images; the blood vessel topology feature branch uses large-sized convolutional kernels combined with blood vessel skeleton extraction to capture global blood vessel topology features such as blood vessel course, intersections, and bifurcation points, which are used as the core anchor points for cross-modal registration; then the features of the two branches are normalized and fused to obtain cross-modal robust multi-dimensional registration features, fundamentally solving the problem of multimodal image feature mismatch. 4. Construct a dual-constraint loss function that fuses SSIM loss and vascular Dice loss. Use the fused multi-dimensional registration features as input to train a U-Net-based registration network to predict the optimal deformation field. First, complete the overall image alignment through rigid coarse registration, and then complete the non-rigid fine registration based on the optimal deformation field. This ensures both the overall image structure matching and the registration results meet the medical physics requirements of clinical diagnosis and treatment through the vascular skeleton overlap constraint. 5. For registered images of the same patient at different time points, an adaptive threshold pixel difference algorithm is used to calculate and filter out pixel difference areas, automatically mark newly added hemorrhage points, exudates and other lesion areas, generate difference annotation maps and quantitative analysis reports and output them to clinical terminals, extending the registration technology to the core clinical scenario of fundus disease course monitoring.
[0024] This invention includes a corresponding multimodal registration system for fundus images. This system comprises seven core functional modules and two auxiliary modules, compatible with different brands and types of fundus imaging equipment, meeting the real-time registration and assisted diagnostic application needs in clinical scenarios. This invention achieves highly robust and medically sound registration of multimodal fundus images, while also upgrading its functionality from a single registration function to clinical disease progression monitoring. It provides a novel technical solution for the early diagnosis, disease monitoring, and efficacy evaluation of blinding fundus diseases such as diabetic retinopathy and macular degeneration, and can be widely applied in ophthalmology departments of hospitals at all levels, large-scale fundus disease screening, and remote ophthalmological consultations.
[0025] The multimodal registration system for fundus images based on multi-scale vascular feature enhancement in this invention, such as... Figure 4 As shown, the device mainly consists of core functional units, auxiliary functional units, and a control computer. The core functional units include an image acquisition module, a preprocessing module, a vascular feature enhancement module, a bi-branch feature extraction module, an image registration module, a multi-temporal lesion comparison module, and a result output module. The auxiliary functional units include a GPU hardware acceleration module and a parameter adaptive adjustment module. The entire device is based on an industrial control computer for computation and control, and can interface with hospital PACS (Pharmaceutical Image Archiving and Communication System), electronic medical record systems, and various fundus imaging devices for data exchange and command interaction.
[0026] Specifically, the core components of this device are as follows: image acquisition module, preprocessing module, vascular feature enhancement module, bi-branch feature extraction module, image registration module, multi-temporal lesion comparison module, result output module, GPU hardware acceleration module, parameter adaptive adjustment module, and control computer. The image acquisition module includes: a device communication interface, a format decoding unit, and a data storage unit; the preprocessing module includes: a denoising unit, a region of interest (ROI) extraction unit, and a normalization unit; the blood vessel feature enhancement module includes: a multi-scale Hessian filtering calculation unit, an adaptive weight adjustment unit, a multi-scale response fusion unit, and a blood vessel mask generation unit; the dual-branch feature extraction module includes: a shared input layer, a texture feature branch, a blood vessel topology feature branch, and a feature fusion layer; the image registration module includes: a registration network unit, a loss function calculation unit, and a spatial transformation unit; the multi-temporal lesion comparison module includes: a difference calculation unit, a difference region filtering unit, and a visualization annotation unit; the result output module includes: a visualization display unit, a report export unit, and a hospital system integration unit; the GPU hardware acceleration module includes: a CUDA parallel computing unit and a network inference acceleration unit; and the parameter adaptive adjustment module includes: a scene mode preset unit and a parameter adaptive optimization unit.
[0027] The device uses a control computer as its core computing and control platform, carrying and scheduling all the aforementioned modules to complete instruction interaction, data transmission, and computational processing. The image acquisition module, preprocessing module, vascular feature enhancement module, bi-branch feature extraction module, image registration module, multi-temporal lesion comparison module, and result output module are connected sequentially according to the data processing flow. The GPU hardware acceleration module is connected in parallel with the bi-branch feature extraction module, image registration module, and multi-temporal lesion comparison module, providing hardware acceleration for computationally intensive operations. The parameter adaptive adjustment module is bi-directionally connected to the vascular feature enhancement module, image registration module, and multi-temporal lesion comparison module, respectively, to achieve adaptive adjustment of core parameters.
[0028] 1. Image Acquisition Module This module uses mainstream communication protocols such as Gigabit Ethernet, USB 3.0, and RS232 through the device communication interface to directly connect with fundus cameras, OCT equipment, and fundus angiography equipment to achieve real-time acquisition of fundus images; at the same time, it can connect and read the hospital PACS system database and import files in formats such as DICOM, JPG, PNG, and BMP from the local system. For the three-dimensional tomographic images (including OCT structural images and OCTA blood flow images) acquired by the OCT device, the format decoding unit simultaneously completes the en face frontal projection conversion of the three-dimensional data. By performing maximum density projection on the tomographic range from the retinal nerve fiber layer to the retinal pigment epithelium layer, a two-dimensional fundus frontal image is generated, which maintains the same imaging dimension as color fundus photography and angiography images. For dynamic angiography video images, the format decoding unit automatically identifies the peak frame of the venous phase during the angiography process and uses this frame as the static image for registration, eliminating the interference of dynamic temporal differences on the registration results.
[0029] The format decoding unit completes the decoding and format unification conversion of images of different formats, outputs a standardized digital image matrix, and synchronously transmits it to the data storage unit and the preprocessing module.
[0030] The data storage unit uses a solid-state drive array, enabling local storage of ≥100,000 sets of fundus images. It supports retrieval by patient ID, examination time, and image modality. The storage capacity calculation formula is as follows: In the formula, For total capacity requirements, To store the maximum number of image groups, This represents the average storage size for a single set of images (including the original image, processing results, and analysis report).
[0031] 2. Preprocessing module This module receives the original fundus image, performs denoising, ROI extraction, and normalization processing in sequence, and outputs a standardized preprocessed image to the blood vessel feature enhancement module to eliminate the differences in the basic image for subsequent processing.
[0032] Denoising unit: It adopts an adaptive median filtering algorithm optimized for fundus scene. The filtering window size is adaptively adjusted within the range of 3×3 to 7×7. It can suppress salt-and-pepper noise and Gaussian noise at the same time, and completely preserve the blood vessel edge information while denoising.
[0033] For coordinates The target pixel, let the current filtering window be... Window size , , , , , These represent the minimum, maximum, median, and original grayscale values of the target pixel within the window, respectively. The filtering logic is as follows: (1) Noise determination: If If the target pixel is not determined, then proceed to the target pixel determination; otherwise, the window size will be adjusted. Add 2, if Repetitive noise determination, if Output ; (2) Target pixel determination: If Output the original value Otherwise, output the median. .
[0034] The ROI extraction unit extracts the optic disc and vascular distribution areas of the fundus using Otsu's threshold segmentation method combined with morphological opening and closing operations. It automatically removes irrelevant areas such as black borders, watermarks, and eyelid occlusions, thus locking in the effective imaging area. The image grayscale range is set. grayscale The number of pixels is Total number of pixels probability of grayscale appearance With threshold Pixels are categorized into background types. (Grayscale) ) and foreground (Grayscale) ), the probability of occurrence of the two types , Mean gray values of two types , Global grayscale mean The variance between classes is: .
[0035] Traversal ,make The largest To determine the optimal segmentation threshold, initial separation of the foreground and background is achieved. Subsequently, a 3×3 circular structuring element is used. First, for the binary image Perform an opening operation to remove tiny noise points, then perform a closing operation to fill foreground holes, obtaining the ROI region mask: In the formula, Opening operation for morphology ( ), For morphological closing operation ( ), For erosion calculation, This is an expansion operation.
[0036] Normalization processing unit: It is divided into two sub-units: grayscale normalization and size normalization, which eliminate the differences in image brightness and resolution between different modalities and devices.
[0037] (1) Gray-level normalization: The gray-level values within the ROI region are linearly mapped to the [0,255] interval. The calculation formula is as follows: In the formula, The original pixel grayscale value. , The minimum and maximum grayscale values within the ROI region. The grayscale value is the normalized value. (2) Size normalization: All images are uniformly scaled to a fixed resolution of 512×512 pixels using bilinear interpolation to ensure uniform input size for subsequent networks. The calculation formula is as follows: In the formula, The coordinates of the four neighboring pixels surrounding the target pixel. The grayscale value of the neighboring pixels. This represents the pixel grayscale value after size normalization.
[0038] 3. Vascular Feature Enhancement Module This module receives a standardized preprocessed image, enhances the full-scale vascular features of the fundus through an improved multi-scale Hessian filtering algorithm, outputs the enhanced vascular image to the dual-branch feature extraction module, and simultaneously outputs the vascular binarization mask to the image registration module.
[0039] Multi-scale Hessian filtering computation unit: For each pixel of the input image, it constructs a Hessian matrix at 10 equally spaced scales within the interval [0.5, 6.0], and performs second-order partial derivative calculation and eigenvalue decomposition. For two-dimensional fundus images... ,coordinate ,scale The Hessian matrix below is: In the formula, The scale is the Gaussian filter kernel (the larger the scale, the larger the diameter of the blood vessel captured). , The second-order partial derivatives are in the x and y directions. It is a mixed second-order partial derivative.
[0040] All partial derivatives are calculated by convolving the image with the second derivative of the corresponding scale Gaussian kernel. The two-dimensional Gaussian kernel function is: Perform eigenvalue decomposition on the Hessian matrix at each scale to obtain two eigenvalues. , ,satisfy The vascular region is a linear tubular structure, with minimal grayscale variation along the vessel's direction. The vertical blood vessel orientation exhibits significant grayscale variations. The fundus image shows a bright background with dark blood vessels pattern, with the bright background area... vascular area This characteristic is used to distinguish blood vessels from the background.
[0041] Adaptive weight adjustment unit: Sets inversely proportional adaptive weight coefficients for filtering responses at different scales, assigning higher weights to the smaller scales corresponding to microvessels to enhance the characteristic responses of minute blood vessels. The weight calculation formula is as follows: In the formula, For scale The corresponding weighting coefficients are adjusted by adding 0.5 to the denominator to avoid a denominator of 0 when σ=0; when σ=0.5 (microvascular scale). , =6.0 (large vessel scale) It significantly enhances the response of microvessels and peripheral blood vessels.
[0042] Multi-scale response fusion unit: Calculates single-scale vascular enhancement response based on feature values and adaptive weights, performs pixel-level maximum value fusion on the response maps of 10 scales, and obtains the final vascular enhancement image.
[0043] The formula for calculating single-scale vascular enhancement response is: In the formula, Set the corresponding bright background area to 0 directly to suppress background interference; The larger the ratio, the more prominent the tubular structure features and the higher the response value of the vascular region.
[0044] The multi-scale response fusion calculation formula is as follows: In the formula, To enhance the pixel values of the final blood vessel image, maximum value fusion is used to preserve the strongest response of blood vessels across the entire scale. The vessel mask generation unit uses Otsu's method to calculate the globally optimal segmentation threshold, binarizes the enhanced vessel image, and generates a binary vessel mask. The calculation formula is as follows: In the formula, This is a binary mask for blood vessels, where 1 represents the blood vessel region and 0 represents the background region. The optimal segmentation threshold is calculated using the Otsu method; this mask serves as the core basis for subsequent calculation of vascular Dice loss.
[0045] Boundary scene fault tolerance design: For images of vascular structural disorder and vascular occlusion caused by proliferative diabetic retinopathy, this module sets up a pseudo-response suppression mechanism. After extracting the vascular skeleton through the Zhang-Suen thinning algorithm, isolated pseudo-response regions with a connected component length of less than 3 pixels are removed to avoid feature extraction failure caused by disordered vascular structure.
[0046] 4. Dual-branch feature extraction module This module receives a blood vessel enhancement image, extracts local texture features and global blood vessel topology features through a dual-branch network, and outputs a multi-dimensional registration feature map after fusion to the image registration module. The module takes a 512×512×1 blood vessel enhancement image and inputs it in parallel to the texture feature branch and the blood vessel topology feature branch through a shared input layer, such as... Figure 2 As shown.
[0047] Texture feature branch: A cascaded structure of 5 consecutive convolutional modules is adopted. Each convolutional module contains a 3×3 convolutional layer, a ReLU activation function, and a 2×2 max pooling layer. The number of convolutional kernels in the 5 modules are 32, 64, 128, 256, and 256 respectively. Texture features such as local grayscale changes and edge details are extracted layer by layer by small-sized convolutional kernels to adapt to the brightness differences of cross-modal images. After 5 pooling, the 512×512 input image is finally downsampled to a 16×16×256 local texture feature map, which is then output to the feature fusion layer.
[0048] (1) Convolutional layer forward propagation formula: For the th Convolutional layers output feature maps for: In the formula, This is the input feature map for the previous layer. Input the number of channels. For convolution kernel weights, For bias terms, This is a two-dimensional convolution operation; the kernel size of this branch is 3×3, stride is 1, and padding is 1, ensuring that the feature map size remains unchanged before and after convolution.
[0049] (2) ReLU activation function formula: .
[0050] (3) Formula for 2×2 max pooling layer: In the formula, For the input feature map, To output the feature map, a pooling window of 2×2 with a stride of 2 is used, and the feature map size is reduced to half of its original size after downsampling.
[0051] Vascular topology feature branch: First, the skeleton of the enhanced vascular image is extracted using the Zhang-Suen thinning algorithm to obtain a vascular skeleton map with a single pixel width; then, feature extraction is performed using a cascaded structure of 3 layers of 7×7 convolutional modules + 2 layers of 11×11 convolutional modules. Each convolutional module contains one large-size convolutional layer, one ReLU activation function, and one 2×2 average pooling layer. The number of convolutional kernels in the five modules are 32, 64, 128, 256, and 256, respectively. The large-size convolutional kernels capture global topological features such as vascular course, branch structure, and intersection distribution, which serve as stable anchor points for cross-modal registration. After 5 pooling operations, a 16×16×256 global vascular topology feature map is finally output to the feature fusion layer.
[0052] (1) Zhang-Suen thinning algorithm: By eroding the boundary pixels of blood vessels in parallel iteration, a blood vessel skeleton with a single pixel width is finally obtained, which completely preserves the topological information such as blood vessel course, intersection points, and bifurcation points; the iteration is divided into two sub-steps, for the foreground pixels (Value 1), the 3×3 neighboring pixels centered on it are labeled clockwise. ,definition The number of non-zero pixels in the neighborhood. For neighboring pixels The number of times the sequence changes from 0 to 1; both sub-steps must simultaneously satisfy 4 conditions, and pixels that meet the conditions are set to 0 (erosion), iterating until no pixels are eroded: Condition 1: (Non-isolated point / endpoint); Condition 2: (Without disrupting vascular connectivity); Condition 3: Sub-step 1 requires Sub-step 2 requires ; Condition 4: Sub-step 1 requires Sub-step 2 requires .
[0053] (2) This branch's convolutional layer uses a large 7×7 / 11×11 convolutional kernel with a stride of 1 and padding of 3 / 5 to ensure that the feature map size remains unchanged before and after convolution; the formula for the 2×2 average pooling layer is: 。
[0054] Feature fusion layer: The feature maps output from the two branches are concatenated along the channel dimension to obtain a 16×16×512 concatenated feature map. Then, a 1×1 convolutional layer is used for feature dimensionality reduction and weighted fusion, and finally a 16×16×256 multi-dimensional registration feature map is output to the image registration module.
[0055] Feature channel splicing formula: In the formula, To stitch together the feature maps, This is a local texture feature map. This is a global topological feature map. For channel indexing.
[0056] 1×1 convolution dimensionality reduction fusion formula: In the formula, It is a 1×1 convolution kernel weight matrix (size 1×1×512×256). For bias terms, This is used to obtain the final multi-dimensional registration feature map.
[0057] 5. Image registration module This module receives multi-dimensional registration feature maps and vascular binarization masks, and adopts a two-level registration architecture of rigid coarse registration + non-rigid fine registration. It outputs the registered fundus images to the multi-temporal lesion comparison module and the result output module.
[0058] Spatial Transformation Unit: First, rigid coarse registration is performed. Based on the vascular skeleton map, the bifurcation points and intersection points of blood vessels are extracted as registration key points. The homography transformation matrix is calculated through feature matching to complete the rigid coarse registration of the image to be registered, which solves the overall offset problem caused by large-angle rotation and large-scale translation / scaling. The image after coarse registration is input into the registration network unit, and non-rigid fine registration is completed based on the predicted deformation field.
[0059] The registration network unit employs a U-Net-based encoding / decoding structure. The encoder performs 5 convolutional operations (without downsampling) on the input 16×16×256 registration feature map, maintaining the feature map size at 16×16×256. The decoder upsamples the image through 5 deconvolutional layers, with each deconvolution kernel size of 2×2 and a stride of 2, gradually restoring the feature map size to 512×512. The final output is a two-dimensional deformation field with the same size as the original image. (Pixel offset matrix in the x and y directions, size 512×512×2). The decoder deconvolution upsampling formula is: In the formula, This is the input feature map for the layer above the decoder. These are the weights of the transposed convolution kernel.
[0060] The formula for mapping pixel coordinates in the deformation field is: In the formula, The pixel coordinates of the image to be registered. To map the new coordinates to the reference image coordinate system, , This represents the offset of the deformation field in the x and y directions at the corresponding coordinates. Loss function calculation unit: Constructs a weighted fusion of structural similarity (SSIM) loss and vessel (Dice) loss into a dual-constraint loss function. The total loss is used as the optimization objective to complete the registration network training optimization, while also providing accuracy evaluation metrics for inference, such as... Figure 3 As shown.
[0061] The formula for calculating the total loss function is: In the formula, The weighting coefficient has a range of values. The parameters can be adaptively adjusted via the adaptive adjustment module. The formula for adaptive weight adjustment is: .
[0062] Structural Similarity (SSIM) Loss: The overall structural matching degree between the constrained registered image and the reference image, calculated using the following formula: In the formula, The image to be registered after registration. The reference image is used; the SSIM calculation window is 11×11, the Gaussian kernel standard deviation is 1.5, and the value range is [0,1]. The closer to 1, the higher the structure matching degree. The complete calculation formula is: In the formula, , To calculate the average grayscale value of pixels within the window, , The standard deviation of grayscale Let the pixel grayscale covariance of the two images be denoted as . , To avoid constants with a denominator of 0.
[0063] Vascular Dice loss: The degree of overlap between the vascular skeleton of the constrained registered image and the reference image, ensuring the medical and physical significance of the registration result. The calculation formula is as follows: In the formula, To provide a binarization mask for the blood vessels in the registered image. For the reference image, a binarized mask of blood vessels. This represents the number of pixels in the intersection of the two masks. , The number of non-zero pixels in the two masks; the Dice coefficient ranges from [0,1], and the closer it is to 1, the higher the overlap of the vascular skeleton.
[0064] Model training and optimization: The Adam optimizer is used to minimize the total loss function to complete the iterative update of network parameters. Let... To the number of training iterations, For the first The gradient of the loss function in the next iteration. , For the first and second moments of the gradient, , The moment attenuation coefficient, , For the estimation of the first and second moments after bias correction, The initial learning rate, To avoid constants with a denominator of 0, For the first The network parameters for the next iteration are updated using the following formula: The model training convergence condition is: the total loss of the training set converges to below 0.05, and the Dice coefficient of the blood vessels in the validation set stabilizes above 0.95.
[0065] Fine-grained registration spatial transformation: Based on the optimal deformation field predicted by the registration network, a spatial transformation is performed on the image to be registered after coarse registration. A bilinear interpolation algorithm is used to calculate the transformed pixel grayscale values, outputting the final registered fundus image. The calculation formula is as follows: In the formula, For the grayscale values of the registered image pixels, The original image to be registered. for By mapping the deformation field to floating-point coordinates in the image to be registered. The floating-point coordinates are the pixel coordinates of the four integer neighboring pixels.
[0066] 6. Multi-time series lesion comparison module This module receives registered fundus images of the same subject at different time points, performs automatic identification, annotation, and quantitative analysis of lesion differences, and outputs a disease course analysis report to the results output module.
[0067] Preprocessing: Baseline image (previous examination image) ) and follow-up images (images from this examination) Perform gray-level normalization and histogram matching again to eliminate the slight brightness differences remaining after registration; let the cumulative gray-level distribution function of the baseline image be... Follow-up images are Gray values in follow-up images Matched grayscale values satisfy This ensures that the grayscale histogram of the follow-up images is perfectly matched with the baseline image.
[0068] Difference Calculation Unit: Employs a pixel-by-pixel absolute value difference algorithm to calculate the difference matrix between two temporal images. The formula is as follows: In the formula, This represents the grayscale difference value of the corresponding pixel. The larger the difference value, the higher the probability of a lesion area.
[0069] Difference region filtering unit: The difference threshold is automatically calculated using the Otsu method. Candidate difference regions are selected, and then isolated noise regions with an area of less than 5 pixels are removed by morphological opening operation of 3×3 structuring elements to eliminate spurious difference interference and obtain the mask of the real lesion difference region. .
[0070] The formula for determining candidate difference regions is: Visual annotation unit: Lesions with differences are visually annotated using red bounding boxes and semi-transparent red fills. Simultaneously, quantitative indicators such as the number, total area, maximum area, and average area of these differences are statistically analyzed to generate a disease progression analysis report. (Setting...) This represents the total number of regions with different lesions. For the first The formula for calculating the pixel area of a lesion region is as follows: .
[0071] 7. Result Output Module Visualization display unit: On the display terminal of the control computer, the original image, preprocessed image, vascular enhancement image, registration result image, lesion difference annotation map, as well as registration accuracy parameters (SSIM value, vascular Dice coefficient) and lesion quantitative indicators are displayed in real time. Report export unit: Supports exporting registration results and course analysis reports as JPG / PNG image formats and PDF document formats, and supports archiving by patient information; Hospital System Interface Unit: It can interface with the hospital's electronic medical record system and PACS system to automatically write the analysis results into the patient's electronic medical record, achieving seamless connection of clinical diagnosis and treatment processes.
[0072] 8. GPU Hardware Acceleration Module This module is built on the NVIDIA CUDA parallel computing architecture and is integrated into the high-performance graphics processor of the control computer. The CUDA parallel computing unit accelerates pixel-level operations such as multi-scale Hessian filtering and pixel difference calculation in parallel, while the network inference acceleration unit optimizes the inference process of the dual-branch feature extraction network and the registration network in parallel. The entire registration process for a single 512×512 image can be controlled within 200ms, meeting the needs of real-time clinical applications. The parallel speedup ratio is calculated as follows: In the formula, For single-threaded CPU computation time, This refers to the time consumed by GPU parallel computation.
[0073] 9. Parameter adaptive adjustment module This module presets three clinical scenario modes through the scenario mode preset unit: cross-modal registration mode, multi-time-series follow-up mode, and low-quality image mode, which can be switched by users with one click; the parameter adaptive optimization unit can adaptively adjust the core parameters according to the modality, quality, and selected clinical scenario of the input image, and transmit them synchronously to the corresponding module to adapt to different clinical application needs.
[0074] Multi-scale Hessian filter scale range adaptive adjustment formula: Adaptive adjustment formula for lesion difference threshold: In the formula, The base threshold calculated by the Otsu method, To adjust the coefficients; standard follow-up mode Early lesion screening model Disease progression assessment model .
[0075] 10. Control the computer The core computing and control platform of this device is equipped with an industrial-grade multi-core processor, large-capacity memory, and a high-performance graphics processor. It runs a Windows / Linux operating system and is responsible for the instruction control, data transmission, computing, result storage and display of all modules, realizing the automated operation of the entire registration process.
[0076] The multimodal registration method for fundus images based on multi-scale vascular feature enhancement provided by this invention is implemented based on the above-mentioned registration system. All calculation formulas and parameter definitions are consistent with the registration system. The overall execution process of the method is divided into five core stages: image acquisition and preprocessing, multi-scale vascular feature enhancement, bi-branch feature extraction and fusion, image registration based on dual-constraint loss function, and automatic comparison and annotation of lesions in multiple time series.
[0077] The core parameters of this method are defined as follows: : Original fundus image pixel grayscale values; : The grayscale value of a pixel after grayscale normalization; Normalized image pixel grayscale values after size normalization; , Minimum and maximum grayscale values within the ROI region; :scale lower coordinate The Hessian matrix at that location; Gaussian filter kernel size, with a basic range of [0.5, 6.0]; , : Eigenvalues of the Hessian matrix that satisfy ; Adaptive weighting coefficients; Single-scale vascular enhancement response value; Pixel values of the enhanced blood vessel image after multi-scale fusion; : Blood vessel binarization mask; Total loss function value; : Structural similarity loss value; : Vascular Dice loss value; : Weighting coefficients of the loss function; Image to be registered after registration; Reference image; , : Blood vessel binarization masks for the registered image and the reference image; Two-dimensional deformation field; : Time-series image pixel grayscale difference value; Baseline image; Follow-up images.
[0078] The specific implementation steps of this method are as follows: Step S1: Connect the power supply, turn on all hardware devices in the device, open the supporting software program of the control computer, and complete the initialization, communication connection and self-test of the operating status of each module of the system. Step S2: Activate the image acquisition module and acquire the reference fundus image to be registered and the fundus image to be registered (including multimodal and multitemporal fundus images) through three methods: direct device connection, local file import, and reading from the hospital PACS system database. Automatically complete image format decoding, dimension unification and temporal phase filtering, and transmit them to the preprocessing module. Step S3: Control the computer to schedule the preprocessing module. First, the image is denoised using an adaptive median filtering algorithm. Then, Otsu thresholding combined with morphological operations is used to extract the ROI. Finally, after grayscale normalization and size normalization, a standardized preprocessed image of 512×512 pixels is output to the blood vessel feature enhancement module. Step S4: The computer schedules the vascular feature enhancement module, which completes full-scale vascular feature enhancement by improving the multi-scale Hessian filtering algorithm. The process sequentially completes Hessian matrix construction, eigenvalue decomposition, adaptive weight adjustment, single-scale response calculation, and multi-scale response fusion to obtain the enhanced vascular image. At the same time, a vascular binarization mask is generated using the Otsu method. The enhanced vascular image is then transmitted to the dual-branch feature extraction module, and the vascular binarization mask is transmitted to the image registration module. Step S5: Control the computer to schedule the dual-branch feature extraction module, input the enhanced blood vessel image into the shared input layer, extract local texture features through the texture feature branch and extract global blood vessel topology features through the blood vessel topology feature branch, and then complete feature stitching and dimensionality reduction fusion through the feature fusion layer to output a 16×16×256 multi-dimensional registration feature map to the image registration module; Step S6: Control the computer to schedule the image registration module. First, complete the overall image alignment through rigid coarse registration. Then, input the multi-dimensional registration feature map into the pre-trained registration network to predict the two-dimensional deformation field. Based on the deformation field, complete the non-rigid fine registration and output the registered fundus image. For multi-temporal disease monitoring scenarios, simultaneously transmit the registered baseline image and follow-up image to the multi-temporal lesion comparison module to complete the automatic identification, annotation and quantitative index statistics of lesion differences and generate a standardized disease analysis report. Step S7: Control the computer scheduling result output module to display the processing process images, registration accuracy parameters and lesion quantification indicators in real time on the display terminal, support the export of registration results and disease course analysis reports, and complete the data connection with the hospital's electronic medical record system and PACS system, automatically write the analysis results into the corresponding patient's electronic medical record, and complete the entire process operation.
[0079] This invention was verified through complete simulation experiments and clinical pre-experiments. The algorithm was implemented and verified based on the PyTorch deep learning framework, OpenCV image processing library, and ITK medical image processing library. The results show that the technical solution of this invention is efficient and feasible, and is significantly better than the existing best solution in all core indicators.
[0080] (a) Experimental Dataset 1. Public datasets: Industry-standard publicly available datasets of DRIVE and STARE fundus images; 2. Clinical Private Dataset: Collected in collaboration with the Eye Hospital Affiliated to Wenzhou Medical University, this dataset contains clinical fundus images of 200 patients. It includes color fundus images, OCT structural images, OCTA blood flow images, fundus angiography images, and multi-time-series follow-up images at 3-12 month intervals for the same patient, totaling 2000 paired images. All images have been annotated and verified by three ophthalmologists at the associate chief physician level or above, forming the clinical gold standard. 3. Low-quality image dataset: 150 sets of low-quality fundus images with motion blur, uneven illumination, and low contrast were selected from the private dataset to verify the adaptability to low-quality scenes.
[0081] Comparative experimental results: Table 1 compares the core indicators of this invention with those of existing mainstream solutions.
[0082] Table 1 plan SSIM MSE Vascular Dice coefficient Cross-modal registration success rate Average time per image (ms) Lesion labeling accuracy This invention 0.962 2.13 0.958 98% 182 92.7% CN109767459B 0.921 5.27 0.903 69% 580 - CN115393239A 0.892 7.86 0.872 62% 625 - Traditional SIFT 0.783 12.54 0.726 41% 310 - VoxelMorph 0.915 6.02 0.885 67% 420 - TransMorph 0.932 4.85 0.889 74% 460 - Ablation experiment results: This invention fully verifies the individual contribution and synergistic effect of each technical feature through forward ablation, reverse ablation, and cross-control experiments. The experimental results are shown in Table 2.
[0083] Table 2 Experimental Groups Solution Configuration SSIM Vascular Dice coefficient Cross-modal registration success rate Baseline group Single-branch feature extraction + MSE loss + no multi-scale enhancement 0.902 0.881 72% Positive ablation 1 Baseline group + multiscale vascular feature enhancement 0.927 0.912 85% Positive ablation 2 Forward ablation 1+ dual-branch feature extraction network 0.948 0.935 93% Positive ablation 3 Forward ablation 2+ double-constraint loss function 0.957 0.951 96% Complete model Positive ablation 3+ multi-time sequence lesion automatic annotation 0.962 0.958 98% Reverse ablation 1 Complete Model - Multiscale Vascular Feature Enhancement 0.925 0.908 84% Reverse ablation 2 Complete Model - Two-Branch Feature Extraction Network 0.918 0.901 76% Reverse ablation 3 Complete Model - Dual-Constraint Loss Function 0.931 0.910 82% Cross-comparison 1 Two-branch + single MSE loss 0.935 0.915 86% Cross-comparison 2 Single-branch + double-constraint loss 0.928 0.918 83% Experimental results show that all four core technical features contribute significantly to the performance improvement of this invention, and their combined use produces a synergistic effect. Cross-comparison experiments further verify that the combination of the dual-branch architecture and the dual-constraint loss function is the core of the performance improvement of this invention, and that using any one feature alone cannot achieve the technical effect of the complete model.
[0084] Furthermore, the present invention achieves a blood vessel Dice coefficient of 0.961 on the DRIVE dataset, a 95% Hausdorff distance (HD95) of 3.21 pixels, and an average surface distance (ASD) of 0.87 pixels; and a blood vessel Dice coefficient of 0.957 on the STARE dataset, with an HD95 distance of 3.56 pixels and an ASD of 0.92 pixels, fully demonstrating the universality and stability of the algorithm of the present invention.
[0085] The present invention achieves a registration success rate of 92.3% and a blood vessel Dice coefficient of 0.931 on low-quality image datasets, while the existing best solution, TransMorph, has a registration success rate of only 48.7% and a blood vessel Dice coefficient of only 0.812, which fully demonstrates the excellent adaptability of the present invention to low-quality images.
[0086] On a private clinical dataset, the average registration error of the vascular terminal is 1.82 pixels, which fully meets the clinical precision diagnosis and treatment requirement of ≤2 pixels.
[0087] This invention has completed clinical pre-experiment verification. Based on clinical fundus images of 200 patients, and verified by three ophthalmologists at the associate chief physician level or above in a double-blind manner, 95.5% of the registration results met the requirements for clinical diagnosis and treatment. The Dice coefficient between the automatic lesion annotation results and the doctor's manual annotation reached 91.2%, and the lesion missed diagnosis rate was only 4.5%, which is far lower than the 21.5% of the manual comparison, demonstrating extremely high clinical application value.
[0088] The technical solution of this invention has been verified through complete simulation experiments and clinical pre-experiments, and is fully feasible. It is significantly superior to the existing best technical solution in terms of registration accuracy, cross-modal robustness, medical rationality, operating efficiency, and clinical applicability. It can effectively solve the core pain points of existing fundus image registration technology and meet the actual needs of ophthalmic clinical diagnosis and treatment.
[0089] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for multimodal registration of fundus images, characterized in that, Includes the following steps: S1. Obtain the reference fundus image and the fundus image to be registered, and perform preprocessing to obtain a standardized preprocessed image; S2. Perform multi-scale vascular feature enhancement on the standardized preprocessed image. By constructing a multi-scale Hessian matrix and using inverse proportional adaptive weights to weight and fuse the filter responses at different scales, a vascular enhancement image and the corresponding vascular binarization mask are generated. S3. Construct a dual-branch feature extraction network, input the enhanced blood vessel image in parallel to the texture feature branch and the blood vessel topology feature branch, extract the local texture features of the image and the global blood vessel topology map features based on the blood vessel skeleton respectively, and fuse the features output by the two branches to obtain a multi-dimensional registration feature map. S4. Construct a dual-constraint loss function that integrates structural similarity loss and vascular Dice loss. Use multi-dimensional registration feature maps and vascular binarization masks as inputs to train and drive a registration network based on the U-Net architecture to predict a smooth and reversible deformation field. Then, based on the deformation field, complete the non-rigid and accurate registration of the image to be registered.
2. The method according to claim 1, characterized in that, The multi-scale vascular feature enhancement in step S2 specifically includes: For each pixel of the input standardized preprocessed image, construct 10 Hessian matrices at the same interval scale within the interval [0.5, 6.0], complete the second-order partial derivative calculation and eigenvalue decomposition, and distinguish blood vessels from the background based on this characteristic; An inversely proportional adaptive weighting coefficient is set for the filter response at different scales. The weighting calculation formula is as follows: In the formula, σ represents the weighting coefficient corresponding to the scale.
3. The method according to claim 2, characterized in that, In step S2, the single-scale vascular enhancement response is calculated based on the feature value and adaptive weight, and the pixel-level maximum value is fused from the response maps of 10 scales to obtain the final vascular enhancement image. The formula for calculating single-scale vascular enhancement response is: In the formula, Set the corresponding bright background area to 0 directly to suppress background interference; The larger the ratio, the more pronounced the tubular structure features and the higher the response value of the vascular region. The multi-scale response fusion calculation formula is as follows: In the formula, To enhance the pixel values of the final blood vessel image, maximum value fusion is used to preserve the strongest response of blood vessels across the entire scale.
4. The method according to claim 1, characterized in that, The dual-branch feature extraction network in step S3 is specifically as follows: The enhanced blood vessel image is input in parallel to the texture feature branch and the blood vessel topology feature branch by sharing an input layer; The texture feature branch adopts a cascaded structure of 5 layers of 3×3 convolutional modules to obtain local texture feature maps; The vascular topology feature branch first extracts the skeleton of the vascular enhancement image through the Zhang-Suen thinning algorithm to obtain a vascular skeleton map with a single pixel width. Then, it extracts features through a cascaded structure of 3 layers of 7×7 convolutional modules + 2 layers of 11×11 convolutional modules to obtain a global vascular topology feature map. The feature maps output from the two branches are concatenated along the channel dimension to obtain a concatenated feature map. Then, a 1×1 convolutional layer is used for feature dimensionality reduction and weighted fusion to finally output a multi-dimensional registration feature map.
5. The method according to claim 1, characterized in that, The double-constraint loss function in step S4 is calculated using the following formula: In the formula, α This is the weighting coefficient, with a value range of 0 < α < 1. For structural similarity loss, For vascular Dice loss; Weighting coefficient α It can be adaptively adjusted according to the registration scenario. The adjustment formula is: 。 6. The method according to claim 1, characterized in that, In step S4, the bifurcation and intersection points of blood vessels are first extracted based on the vascular skeleton map as registration key points. The homography transformation matrix is calculated through feature matching to complete rigid coarse registration of the image to be registered. Then, non-rigid fine registration is completed based on the deformation field predicted by the registration network.
7. The method according to claim 6, characterized in that, Based on the optimal deformation field predicted by the registration network, a spatial transformation is performed on the coarsely registered image to be registered. A bilinear interpolation algorithm is then used to calculate the transformed pixel grayscale values, outputting the final registered fundus image. The calculation formula is as follows: In the formula, For the grayscale values of the registered image pixels, The original image to be registered. for By mapping the deformation field to floating-point coordinates in the image to be registered. The floating-point coordinates are the pixel coordinates of the four integer neighboring pixels.
8. The method according to claim 1, characterized in that, Following step S4, the process also includes automatic comparison and annotation of lesions across multiple time series, specifically: The system acquires registered images of the same object at different time points, calculates the difference matrix using a pixel-by-pixel absolute value difference algorithm, automatically calculates the difference threshold using the Otsu method to filter out candidate difference regions, removes isolated noise regions using morphological opening operations to obtain the real lesion difference regions, and performs visual annotation and quantitative statistics such as quantity and area on the real lesion difference regions to generate a fundus disease course analysis report.
9. The method according to claim 1, characterized in that, The preprocessing in step S1 specifically includes: The algorithm sequentially performs adaptive median filtering for noise reduction, Otsu's method combined with morphological operations for ROI extraction, grayscale normalization, and size normalization to eliminate image differences caused by different devices and imaging conditions.
10. A multimodal registration system for fundus images, characterized in that, It includes core functional units, auxiliary functional units, and a control computer; The core functional units include an image acquisition module, a preprocessing module, a vascular feature enhancement module, a bi-branch feature extraction module, an image registration module, a multi-temporal lesion comparison module, and a result output module, which are connected in sequence. The auxiliary functional unit includes a GPU hardware acceleration module and a parameter adaptive adjustment module; the GPU hardware acceleration module is connected in parallel with the dual-branch feature extraction module, the image registration module, and the multi-temporal lesion comparison module, and the parameter adaptive adjustment module is bidirectionally connected with the vascular feature enhancement module, the image registration module, and the multi-temporal lesion comparison module, respectively. The control computer is the core computing and control carrier of the system, carrying and scheduling all modules to complete instruction interaction, data transmission and computing processing; The system is used to perform the method of claim 1.
Citation Information
Patent Citations
Novel fundus image registration method
CN109767459B