Fundus blood vessel segmentation method based on multi-modal data

Through the simultaneous acquisition, preprocessing and deep learning segmentation network model of multimodal data, combined with confidence assessment and interactive correction, the problems of insufficient accuracy and completeness in existing fundus vascular segmentation methods are solved, high-precision fundus vascular segmentation is achieved, and the accuracy and efficiency of disease diagnosis are improved.

CN120726073AInactive Publication Date: 2025-09-30MIANYANG THIRD PEOPLES HOSPITAL

Patent Information

Application Number
CN202511203055.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing fundus vessel segmentation methods have problems with accuracy and completeness when processing multimodal data, especially the low contrast and high noise interference of single-modality images. In addition, the differences in parameters of different devices make data integration more difficult, affecting the clinical application of segmentation results.

Method used

High-precision fundus vascular segmentation results are generated by using a method of multimodal data synchronous acquisition, preprocessing, feature fusion and deep learning segmentation network model combined with confidence assessment and interactive correction.

Benefits of technology

Through the comprehensive utilization of multimodal data and the training of deep learning models, the accuracy and precision of fundus blood vessel segmentation have been significantly improved, providing a more reliable diagnostic basis and improving the efficiency and accuracy of disease diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726073A_ABST
    Figure CN120726073A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of ophthalmologic image processing, and discloses a fundus blood vessel segmentation method based on multi-modal data. The method comprises the following steps: firstly, synchronously acquiring an optical coherence tomography image, a color fundus photographic image and a fluorescent angiography image to form multi-modal fundus data; preprocessing the eye fundus image data set to generate a standardized multi-mode eye fundus image data set; fusing heterogeneous features based on the data set, and constructing a multi-dimensional fundus feature space; using the space to train a deep learning segmentation network model, and generating an initial fundus blood vessel segmentation mask; carrying out confidence evaluation analysis on the initial mask to obtain a blood vessel segmentation result confidence distribution map; dividing blood vessel segmentation quality grades according to the distribution diagram, and defining a judgment rule; and performing interactive correction on the low-confidence region based on a rule to generate a final optimized fundus blood vessel segmentation result. And the final result is transmitted to a visual terminal for three-dimensional topology reconstruction. According to the method, the advantages of multi-modal data are integrated, and a more comprehensive fundus blood vessel segmentation result is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ophthalmic image processing, and in particular to a fundus blood vessel segmentation method based on multimodal data. Background Art

[0002] The morphology and functional status of retinal blood vessels play a crucial role in the diagnosis of various diseases. Subtle changes in retinal blood vessels, such as abnormal branching, changes in vessel diameter, and neovascularization, are often associated with systemic diseases such as diabetic retinopathy and vascular damage caused by hypertension. Observing the status of retinal blood vessels can help determine the stage of disease progression and assess the effectiveness of interventions. Currently, fundus imaging technologies are developing in a diverse range of fields. Color fundus photography can produce high-definition two-dimensional images, visually demonstrating the overall distribution of fundus blood vessels and is a commonly used examination method in clinical practice. Optical coherence tomography provides three-dimensional information on the structure of each retinal layer, helping to observe the spatial relationship between blood vessels and surrounding tissues. Fluorescent angiography, through the injection of contrast agents, can clearly display the hemodynamic characteristics of blood vessels, helping to detect subtle vascular lesions. Existing fundus vessel segmentation methods have certain limitations. Single-modality images contain limited information. For example, color fundus images have difficulty accurately distinguishing vascular boundaries when the contrast between blood vessels and the background is low; optical coherence tomography images are easily affected by noise, which affects the identification of small blood vessels. Traditional segmentation algorithms often find it difficult to balance accuracy and completeness when dealing with complex vascular structures. For twisted and intersecting vascular areas, segmentation breaks or misjudgments are prone to occur. In addition, parameter differences between different imaging devices lead to inconsistencies in the format and resolution of the collected data, increasing the difficulty of cross-modal data integration. These factors work together to make the existing segmentation results have certain limitations in clinical applications. Summary of the Invention

[0003] The purpose of the present invention is to provide a fundus blood vessel segmentation method based on multimodal data to solve the problems raised in the above background technology.

[0004] To achieve the above objectives, the present invention provides a fundus blood vessel segmentation method based on multimodal data, the method comprising: Synchronously acquire multimodal fundus data including optical coherence tomography images, color fundus photography images, and fluorescein angiography images; Perform preprocessing operations on the acquired multimodal fundus data to generate a standardized multimodal fundus image dataset; Based on a standardized multimodal fundus image dataset, the heterogeneous features of fundus image data of different modalities are fused to construct a multi-dimensional fundus feature space; A deep learning segmentation network model is trained based on the multi-dimensional fundus feature space to generate an initial fundus vessel segmentation mask; Perform confidence evaluation analysis on the initial fundus vessel segmentation mask to obtain a confidence distribution map of the vessel segmentation result; According to the confidence distribution map of the blood vessel segmentation results, the blood vessel segmentation quality level is divided and the quality level determination rules are defined; Perform interactive correction processing on low-confidence areas based on quality level judgment rules to generate the final optimized fundus vessel segmentation results; The final optimized fundus vessel segmentation results are transmitted to the visualization terminal for three-dimensional topology reconstruction.

[0005] Preferably, the operation of synchronously collecting multimodal fundus data specifically includes: deploying a multispectral imaging sensor array, which is configured with collection devices of different spectral bands according to the anatomical structure of the retina; setting the optical coherence tomography image collection depth range to cover the inner limiting membrane of the retina to the suprachoroidal space; adjusting the exposure parameters of the color fundus photography image to adapt to different degrees of fundus pigment deposition; configuring the time series collection frequency of the fluorescent angiography image to fully capture the dynamic perfusion process of the contrast agent in the arterial phase, venous phase and late phase; and aligning the collection time points of different modal images through a timestamp synchronization mechanism.

[0006] Preferably, the operation of fusing heterogeneous features of fundus image data of different modalities specifically includes: extracting the layered structure feature vector of the optical coherence tomography image, which contains the reflection intensity distribution characteristics of each layered tissue of the retina; parsing the color space feature vector of the color fundus photography image, which records the chromatic difference information between blood vessels and background tissues; quantizing the time series feature vector of the fluorescent angiography image, which describes the filling dynamics parameters of the contrast agent in the vascular network; inputting the layered structure feature vector, the color space feature vector and the time series feature vector into the feature fusion engine, which adopts the cross-modal attention mechanism to weightedly integrate different feature vectors; and outputting a multi-dimensional fundus feature space with spatial topological correlation.

[0007] Preferably, the operation of training a deep learning segmentation network model based on a multi-dimensional fundus feature space to generate an initial fundus vascular segmentation mask specifically includes: constructing a vascular segmentation network model with an encoder-decoder architecture, the encoder of the vascular segmentation network model including a multi-scale feature extraction module; using a multi-dimensional fundus feature space as a training sample set, the training sample set has been labeled with real vascular boundary information; optimizing the weight parameters of the vascular segmentation network model through a back propagation algorithm; using the optimized vascular segmentation network model to infer unlabeled multimodal fundus data to generate an initial fundus vascular segmentation mask containing probability values.

[0008] Preferably, the operation of performing confidence assessment analysis on the initial fundus vascular segmentation mask specifically includes: calculating the probability variance value of each pixel point in the initial fundus vascular segmentation mask to generate an original confidence heat map; analyzing the continuity index of the vascular morphological feature, which includes the branch connectivity and the uniformity of the tube diameter; combining the original confidence heat map with the continuity index of the vascular morphological feature, and generating a confidence distribution map of the vascular segmentation result through an adaptive weighting function.

[0009] Preferably, the operation of dividing the blood vessel segmentation quality level specifically includes: setting a high confidence threshold and a low confidence threshold, and the high confidence threshold and the low confidence threshold are dynamically adjusted according to clinical diagnosis requirements; marking the area above the high confidence threshold in the blood vessel segmentation result confidence distribution map as a high-quality segmentation area; marking the area between the high confidence threshold and the low confidence threshold as a medium-quality segmentation area; marking the area below the low confidence threshold as a low-quality segmentation area; and generating a fundus blood vessel quality zoning map containing three-level quality grade annotations.

[0010] Preferably, the operations for performing interactive correction processing specifically include: receiving coordinate information of low-quality segmentation areas in the fundus vascular quality zoning map; overlaying and displaying the original multimodal fundus data corresponding to the low-quality segmentation areas on a visualization terminal; receiving vascular boundary correction instructions annotated by a physician through a human-computer interaction interface; using an adaptive deformation algorithm to fuse the vascular boundary correction instructions annotated by the physician into the initial fundus vascular segmentation mask; and updating and generating the final optimized fundus vascular segmentation result.

[0011] Preferably, the operation of three-dimensional topology reconstruction specifically includes: parsing the coordinates of the vascular centerline in the final optimized fundus vascular segmentation result; reconstructing the three-dimensional spatial coordinate mapping relationship of the vascular centerline, which is associated with the retinal surface curvature parameters; generating a three-dimensional mesh model of the vascular wall based on the vascular diameter data; and rendering the three-dimensional topological structure of the vascular network including hemodynamic simulation on a visualization terminal.

[0012] Preferably, the generation process of the confidence distribution map of the vascular segmentation result includes a feature interaction mechanism: the layered structure feature vector in the multidimensional fundus feature space is input into the confidence assessment module; the time series feature vector in the multidimensional fundus feature space is transmitted to the confidence assessment module; the confidence assessment module calculates the association weight coefficient between the layered structure feature vector and the time series feature vector through a cross-validation algorithm; and the confidence distribution map of the vascular segmentation result is optimized based on the association weight coefficient.

[0013] Preferably, the preprocessing operation includes a cross-modal data verification mechanism: performing inter-layer structural alignment processing on the optical coherence tomography image to generate a retinal layered topology map; performing illumination uniformity correction on the color fundus photography image to eliminate retinal reflection artifacts; performing time domain sequence alignment on the fluorescent angiography image to compensate for the patient's eye micro-movement; performing spatial coordinate matching verification on the retinal layered topology map and the color fundus photography image after illumination uniformity correction; using the time domain sequence alignment result to correct the vascular position offset of the retinal layered topology map; verifying the spatial consistency of the multimodal fundus data based on the medical gold standard annotation, and outputting a standardized multimodal fundus image dataset with anatomical structure correspondence.

[0014] Compared with the prior art, the present invention has the following beneficial effects: During the data acquisition phase, three types of multimodal fundus data, optical coherence tomography (OCT) images, color fundus photography images, and fluorescein angiography (FACE) images, are collected simultaneously to reflect the condition of fundus blood vessels from different angles. Optical coherence tomography (OCT) images can present structural information of each layer of the retina, helping to understand the relationship between blood vessels and surrounding tissues; color fundus photography images intuitively display the overall morphology and distribution of fundus blood vessels; and FACE images can highlight details such as blood flow conditions and the integrity of vessel walls. Multiple data types complement each other, providing more comprehensive and richer information for subsequent segmentation than single-modality data, effectively avoiding segmentation errors caused by the limitations of a single data set and greatly improving segmentation accuracy. Preprocessing acquired multimodal fundus data and generating a standardized multimodal fundus image dataset is crucial. Data collected from different devices and environments may vary in format, resolution, brightness, and other factors. Preprocessing eliminates these discrepancies, making the data consistent and comparable, facilitating subsequent unified processing and analysis. Standardized datasets help deep learning models better learn data features, reduce interference caused by non-standardized data, and improve the stability and convergence speed of model training, thus laying a solid foundation for accurate segmentation. A major innovation of this method is the fusion of heterogeneous features from fundus image data of different modalities and the construction of a multi-dimensional fundus feature space. Data from different modalities contain different features. By fusing these heterogeneous features, the multifaceted characteristics of fundus blood vessels can be fully explored to form a more comprehensive and accurate feature description. The multi-dimensional fundus feature space can more finely distinguish between blood vessels and background, different types of blood vessels, and diseased blood vessels, greatly improving the ability to express fundus vascular characteristics. This enables subsequent deep learning segmentation network models to more accurately identify and segment blood vessels, significantly improving segmentation accuracy and providing more reliable blood vessel segmentation results for disease diagnosis. A deep learning segmentation network model is trained based on the multi-dimensional fundus feature space to generate an initial fundus vascular segmentation mask, fully leveraging the powerful feature learning and pattern recognition capabilities of deep learning. Deep learning models can automatically learn complex vascular features and patterns from large amounts of data. Compared to traditional methods, they do not require manual design of complex feature extraction rules and have greater adaptability and generalization capabilities. Supported by the multi-dimensional fundus feature space, the model can more accurately capture the detailed information of the blood vessels. The generated initial fundus vascular segmentation mask can better reflect the true morphology of the vessels, and can achieve relatively accurate segmentation performance even in areas with complex vascular structures and low contrast. A confidence assessment analysis is performed on the initial fundus vessel segmentation mask to generate a confidence distribution map of the vessel segmentation results. This provides a quantitative basis for evaluating segmentation quality. This confidence assessment clearly demonstrates the reliability of different regions within the segmentation results. Regions with high confidence levels indicate relatively accurate and reliable segmentation results, while regions with low confidence levels indicate potential segmentation inaccuracies and require further processing. This quantitative assessment helps doctors or researchers quickly assess the quality of segmentation results, enabling targeted subsequent analysis and processing, thereby improving work efficiency and diagnostic accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a diagram showing the working principle of the fundus blood vessel segmentation method based on multimodal data according to the present invention; Figure 2 Flowchart for simultaneous acquisition of multimodal fundus data; Figure 3 Flowchart for deep learning segmentation network training and inference; Figure 4 Flowchart for grading the quality of blood vessel segmentation; Figure 5 Flowchart for 3D topological reconstruction of blood vessels. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] See also Figure 1 The present invention provides a method for segmenting fundus blood vessels based on multimodal data, the method comprising: Synchronously acquire multimodal fundus data including optical coherence tomography images, color fundus photography images, and fluorescein angiography images; Perform preprocessing operations on the acquired multimodal fundus data to generate a standardized multimodal fundus image dataset; Based on a standardized multimodal fundus image dataset, the heterogeneous features of fundus image data of different modalities are fused to construct a multi-dimensional fundus feature space; A deep learning segmentation network model is trained based on the multi-dimensional fundus feature space to generate an initial fundus vessel segmentation mask; Perform confidence evaluation analysis on the initial fundus vessel segmentation mask to obtain a confidence distribution map of the vessel segmentation result; According to the confidence distribution map of the blood vessel segmentation results, the blood vessel segmentation quality level is divided and the quality level determination rules are defined; Perform interactive correction processing on low-confidence areas based on quality level judgment rules to generate the final optimized fundus vessel segmentation results; The final optimized fundus vessel segmentation results are transmitted to the visualization terminal for three-dimensional topology reconstruction.

[0018] Example 1: See Figure 2 When synchronously collecting multimodal fundus data, a multispectral imaging sensor array is deployed. The array is configured with acquisition devices of different spectral bands according to the different partitions of the retinal anatomical structure. For the macular area of ​​the retina, a 480-520nm blue light band acquisition module is configured. This band can better present the subtle structure of the macular area; for the optic disc area, a 530-580nm green light band acquisition module is used to highlight the contrast between the optic disc and the surrounding tissue. The acquisition depth range of the optical coherence tomography image is set to extend from the inner limiting membrane of the retina to the suprachoroidal space. By adjusting the parameters of the scanning equipment, it is ensured that all tissue structures within this depth range can be fully covered, including the cells of each layer of the retina and the boundary of the suprachoroidal space.

[0019] When adjusting the exposure parameters for color fundus photography, settings are flexibly adjusted based on the individual subject's degree of fundus pigmentation. When encountering areas with heavy fundus pigmentation, the exposure intensity is appropriately increased to clearly display vascular structures obscured by pigmentation. For areas with less pigmentation, the exposure intensity is reduced to avoid overbrightening the image and losing vascular detail. The acquisition frequency for fluorescein angiography is configured. During the arterial phase, the acquisition frequency is set to 30 frames per second to capture the rapid flow of contrast agent immediately after entering the artery. During the venous phase, the frequency is adjusted to 20 frames per second to accommodate the relatively stable flow of contrast agent in the veins. In the late phase, the acquisition frequency is reduced to 10 frames per second, sufficient to record the gradual disappearance of the contrast agent, thereby fully capturing the dynamic perfusion process of the contrast agent during the arterial, venous, and late phases. A timestamp synchronization mechanism precisely aligns the acquisition time points of images from different modalities, ensuring temporal consistency of the data from each modality and ensuring data relevance during subsequent processing.

[0020] When fusing heterogeneous features from fundus image data of different modalities, the laminar structural feature vector of the optical coherence tomography image is first extracted. Using a dedicated image analysis algorithm, the various retinal layers, such as the internal limiting membrane, nerve fiber layer, and outer plexiform layer, are identified. The reflection intensity values ​​of each layer are calculated, and the distribution characteristics of the reflection intensity within each layer are then derived. These characteristics together form the laminar structural feature vector, which contains the reflection intensity distribution information of each retinal layer.

[0021] Next, the color space feature vector of the color fundus photography image is analyzed. The color image is converted from the common RGB color space to the HSV color space. The numerical differences between the vascular region and the background tissue in the three channels of H (hue), S (saturation), and V (value) are analyzed. The distribution of these differences is recorded to form a color space feature vector. This vector fully records the chromatic difference information between the blood vessels and the background tissue.

[0022] The time series feature vectors of the fluorescence angiography images are then quantified. By analyzing images acquired at different time points, the kinetic parameters of the contrast agent filling velocity and time in the vascular network are calculated. These parameters are arranged in chronological order to form a time series feature vector, which describes the filling kinetic parameters of the contrast agent in the vascular network.

[0023] The layered structure feature vector, color space feature vector, and time series feature vector obtained above are input into the feature fusion engine. This feature fusion engine uses a cross-modal attention mechanism to weight and integrate different feature vectors based on their importance in describing fundus vascular characteristics. For example, when describing the layered distribution of blood vessels, the layered structure feature vector is given a relatively higher weight; when distinguishing blood vessels from the background, the color space feature vector is given a higher weight; and when analyzing vascular perfusion, the time series feature vector is given a higher weight.

[0024] Through this weighted integration process, the feature fusion engine outputs a multi-dimensional fundus feature space with spatial topological correlations. This space not only contains the feature information carried by each modality data, but also establishes the spatial correlation between these features, allowing the features of different modalities to complement each other and provide a comprehensive and accurate feature foundation for subsequent fundus vessel segmentation.

[0025] Example 2: See Figure 3 When training a deep learning segmentation network model based on the multi-dimensional fundus feature space, a vascular segmentation network model with an encoder-decoder architecture is constructed. The encoder part contains multiple multi-scale feature extraction modules. Each module has a 3×3 convolutional layer, a 5×5 convolutional layer, and a 7×7 convolutional layer. These convolutional layers work in parallel to extract feature information at different scales in the image. The 3×3 convolutional layer focuses on capturing local subtle features, such as the edges of small blood vessels; the 5×5 convolutional layer is used to extract medium-scale features, such as the branching structure of blood vessels; and the 7×7 convolutional layer is responsible for obtaining larger-scale features, such as the overall distribution trend of the vascular network. The high-level semantic features of the encoder are fused with the low-level detailed features through skip connections, allowing the network to learn global features without losing local detailed information.

[0026] The training sample set uses a multi-dimensional fundus feature space, which has been annotated with pixel-level vascular boundaries using professional medical image annotation tools. This annotation process, performed by experienced ophthalmologists, covers a wide range of normal vascular morphologies, such as straight segments, bifurcations, and tortuosity. It also includes a variety of abnormal vascular morphologies, such as hemangiomas, vascular stenosis, and tortuosity, ensuring comprehensive and representative training data.

[0027] The network's weight parameters are optimized using the backpropagation algorithm. The initial learning rate is set, and an appropriate optimizer, such as the Adam optimizer, is selected. During training, the learning rate is dynamically adjusted based on changes in the loss function. For example, after each training epoch, the learning rate is decayed to a certain percentage of the previous value to promote convergence of the loss function. Training continues until the loss function stabilizes within a small range and no longer decreases significantly. At this point, the network model is considered to have achieved satisfactory training results.

[0028] The optimized vessel segmentation network model is used to infer unlabeled multimodal fundus data. The network calculates the probability of each pixel in the input image belonging to a vessel region, ranging from 0 to 1. These probabilities are presented as grayscale images to generate an initial fundus vessel segmentation mask. The higher the grayscale value of a pixel, the greater the probability that the location belongs to a vessel region, and vice versa.

[0029] When performing confidence assessment analysis on the initial fundus vessel segmentation mask, the probability variance of each pixel in the mask within its 3×3 neighborhood is first calculated. The probability variance reflects the segmentation consistency of the pixel and its surrounding neighborhood. Smaller variances indicate more stable and consistent segmentation results in that area; larger variances indicate greater volatility and lower reliability. Based on the calculated probability variances, a raw confidence heatmap is generated. The heatmap uses different colors to visually display the distribution of confidence levels. Generally, red areas represent high variance (low confidence) and blue areas represent low variance (high confidence).

[0030] Continuity indices of vascular morphological characteristics were analyzed, including branch connectivity and diameter uniformity. Branch connectivity assesses the rationality and continuity of connections between branches at a vascular bifurcation. It is calculated by multiplying the connection probabilities of each branch at the bifurcation point. A product closer to 1 indicates a more reliable branch connection; a lower value indicates uncertainty in the branch connection. Diameter uniformity measures the variation in diameter within the same vascular segment. This is achieved by evenly selecting multiple measurement points along the segment, measuring the diameter at each point, and then calculating the standard deviation of these diameter values. A smaller standard deviation indicates a more uniform diameter and a more regular vascular morphology. A larger standard deviation indicates greater diameter variation and possible segmentation errors.

[0031] Combining the original confidence heatmap with the continuity index of vascular morphology, an adaptive weighting function is used to generate a confidence distribution map of the vascular segmentation results. This adaptive weighting function dynamically adjusts the weights of various indicators based on the complexity of the vascular morphology within the region. For regions with numerous bifurcations and complex morphology, the weight of the continuity index is increased, as the rationality of the morphology in these regions has a greater impact on the reliability of the segmentation results. For regions with relatively simple, straight vascular morphology, the weight of the probability variance is increased, as the consistency of the probability in these regions is more representative of the reliability of the segmentation. Through this weighted integration, a confidence distribution map of the vascular segmentation results is ultimately obtained that accurately reflects the reliability of the segmentation results in each region.

[0032] Example 3: See Figure 4 When dividing the quality level of blood vessel segmentation, a high confidence threshold and a low confidence threshold are set. The initial values ​​of the high confidence threshold and the low confidence threshold are determined according to the common clinical blood vessel segmentation accuracy requirements, and support flexible adjustment according to specific diagnostic scenarios. For example, in the diagnosis of early-stage microvascular lesions, due to the higher requirements for segmentation accuracy, the high confidence threshold and the low confidence threshold can be appropriately increased; when performing large-scale fundus blood vessel screening, the threshold can be appropriately lowered to improve processing efficiency.

[0033] Regions in the confidence distribution map of the vessel segmentation results with probability values ​​above the high confidence threshold are marked as high-quality segmentation regions. Within this region, the vessel boundaries are clearly discernible, and the morphological features of the vessels, such as their course and branching, are continuous and consistent with normal physiological structures. Whether it's changes in vessel diameter or bifurcation angles, the transitions are natural, eliminating the need for additional corrections.

[0034] Regions with probability values ​​between the high and low confidence thresholds are marked as medium-quality segmentation regions. The vessel segmentation results in this region can generally reflect the basic morphology and distribution trend of the vessels, but there may be local boundary ambiguity or slight morphological irregularities. For example, the boundaries at the ends of small branches of a vessel may be unclear, or there may be slight morphological deviations in areas with large vessel bends. However, these problems do not affect the judgment of the overall vascular structure.

[0035] Regions with probability values ​​below the low confidence threshold are marked as low-quality segmentation regions. The vessel segmentation results in these regions have obvious defects, such as missing vessels, false vessel segments, and vessel boundaries that deviate significantly from their actual positions. These problems interfere with the accurate determination of vascular structure and require further processing.

[0036] Generate a fundus vascular quality zoning map with three-level quality grade annotations, and distinguish and display the three levels of areas with different colors, so that areas of different quality grades can be intuitively identified.

[0037] During interactive correction, the system automatically analyzes the fundus vascular quality zoning map, identifies low-quality segments, and accurately extracts the coordinate range of these regions. The coordinate range is expressed in pixels and includes the pixel coordinates of the upper left and lower right corners of the region, accurately locating the area requiring correction.

[0038] The visualization terminal displays the original multimodal fundus data corresponding to the low-quality segmented areas simultaneously. These data, including optical coherence tomography (OCT) images, color fundus photography, and fluorescein angiography (FA) images, are arranged in a split-screen format. The initial segmentation results for each area are superimposed on each image, allowing physicians to compare and analyze the segmentation results with the original images.

[0039] Physicians use a variety of tools provided by the human-computer interface to correct vascular boundaries. For example, they can use the brush tool to manually draw missing vascular segments, the eraser tool to remove false vascular parts, and the boundary adjustment tool to drag and correct vascular boundaries that deviate from their actual positions. These operations input vascular boundary correction instructions.

[0040] The system uses an adaptive deformation algorithm to process the correction instructions entered by the physician. This algorithm first converts the physician's correction operations into a series of mathematical morphological operations, selecting the appropriate operation method for different correction requirements. For example, dilation is used to expand the vessel boundary, erosion is used to shrink the vessel boundary, and curve fitting is used to smooth the vessel boundary. Through these operations, the corrected vessel boundary maintains a natural connection with the surrounding vessel morphology, avoiding abrupt turns or breaks. Ultimately, the correction results are integrated into the initial fundus vessel segmentation mask.

[0041] After the correction is complete, the system generates the final optimized fundus vessel segmentation results. At the same time, the system automatically saves the segmentation results before and after the correction, as well as all the doctor's operation records during the correction process, including the time of the operation, the tools used, the modified area, and other information. This data can be used for subsequent tracing and analysis.

[0042] During the entire interactive correction process, in order to more accurately determine the scope of the correction area, the calculation of the regional impact factor is introduced, and its calculation formula is: , in, represents the regional impact factor, Indicates the first The distance between a pixel and the surrounding normal blood vessel pixels, Indicates the The weight of each pixel is determined by the probability value of the pixel. The lower the probability value, the greater the weight. Represents the total number of pixels within the low-quality segmentation area. The regional impact factor calculated by this formula can help the system determine the range that requires extended correction, making the correction operation more accurate and efficient.

[0043] Example 4: See Figure 5 During 3D topology reconstruction, the final optimized fundus vessel segmentation results are first analyzed. A skeleton extraction algorithm is used to traverse the vascular regions in the segmentation results, extracting centerline coordinates from each connected vascular branch. These coordinates are presented as a continuous sequence of pixel points, with each point containing positional information on a 2D plane. For example, the centerline coordinates of a vascular branch extending from the optic disc to the macula would form an ordered set of points from its starting point to its end point, recording the direction and curvature of the vessel.

[0044] Based on the depth data provided by the optical coherence tomography image, the two-dimensional centerline coordinates are converted into three-dimensional spatial coordinates. Specifically, by analyzing the structural information of different depth layers in the optical coherence tomography image, the depth value corresponding to each two-dimensional coordinate point is determined, thereby constructing the three-dimensional spatial coordinates (x, y, z). Among them, the calculation of the z value is related to the curvature parameter of the retinal surface. By performing surface fitting on the retinal surface, the curvature characteristics of different areas are obtained, so that the three-dimensional coordinates can reflect the actual burial depth of the blood vessels in the retinal tissue. For example, the depth of blood vessels in the peripheral area of ​​the retina is different from the depth of blood vessels near the macular area, and the three-dimensional coordinates will accurately reflect this spatial position relationship.

[0045] Based on the 3D coordinates of the vessel centerline and the diameter data obtained from the segmentation results, a 3D mesh model of the vessel wall is generated. For each vascular branch, a cylindrical surface fitting method is used to construct the vessel wall contour based on the direction of its centerline and the diameter values ​​at each point. For example, if the diameter of a certain section of the vessel is relatively uniform, the corresponding cylindrical surface radius will vary slightly. However, at the bifurcation of the vessel, the diameter will change as the branch extends, and the mesh model will adjust the cylindrical radius accordingly to match the actual shape. Adjacent cylindrical surfaces are connected using a triangulation algorithm to form a continuous 3D mesh of the vessel wall, ensuring the surface of the mesh model is smooth and free of cracks.

[0046] In the visualization terminal, volume rendering technology is used to render the three-dimensional topological structure of the vascular network. During the rendering process, the spatial distribution of the blood vessels is highlighted by adjusting the transparency parameters, and the simulation data related to hemodynamics is superimposed. For example, blood flow velocity information is displayed using a color gradient, with different colors corresponding to different blood flow velocity ranges, allowing physicians to intuitively observe the dynamic changes in blood flow within the blood vessels. The terminal supports a variety of interactive operations, such as rotating the model around any axis to observe the side morphology of the blood vessels, using the zoom function to view the details of small blood vessel branches, and using the sectioning tool to display the cross-sectional structure of the blood vessels at a specific depth plane.

[0047] The process of generating the confidence distribution map of the vessel segmentation results involves a feature interaction mechanism. The layered structural feature vector in the multidimensional fundus feature space is input into the confidence assessment module, which analyzes the reflection intensity distribution characteristics of each retinal layer contained in the vector, such as whether the reflection intensity of the nerve fiber layer is uniform and whether the reflection boundary of the pigment epithelium is clear. Simultaneously, the time series feature vector is transmitted to the module to extract dynamic parameters of the contrast agent filling process, such as the change in filling time from the arterial phase to the venous phase and the difference in contrast agent concentration at different locations within the same vessel segment.

[0048] The confidence assessment module analyzes these two eigenvectors using a cross-validation algorithm and calculates the correlation weight coefficient between them. For example, for superficial retinal vessels, the contrast agent filling rate in the time series eigenvector is more strongly correlated with the vessel segmentation results and is therefore assigned a higher weight. For deeper vessels, the reflection intensity distribution of each layer of tissue in the layered structure eigenvector has a greater impact on the segmentation results, and its weight is increased accordingly. Based on these correlation weight coefficients, the module adjusts the initially generated confidence heat map so that the final output vessel segmentation result confidence distribution map is more consistent with the actual segmentation of vessels at different depths.

[0049] The following is an example table of weight distribution of different feature vectors in confidence assessment: Blood vessel type Weight ratio of layered structure feature vector Weight ratio of time series feature vector Superficial blood vessels 30%-40% 60%-70% mid-layer blood vessels 50% 50% deep blood vessels 60%-70% 30%-40% This table reflects the approximate distribution range of the weights of the two feature vectors when evaluating the segmentation confidence of blood vessels at different depths. The module will dynamically adjust the weight value within this range according to the morphological characteristics of the specific blood vessels.

[0050] Example 5: The preprocessing operation includes a cross-modal data verification mechanism to perform inter-layer structure alignment on the optical coherence tomography image. Through the elastic registration algorithm, the spatial position of each layer of structure in the image is adjusted with the inner limiting membrane of the retina and the retinal pigment epithelium as the reference layer. During the adjustment process, the algorithm will identify the edge feature points of each layer of structure, calculate the spatial offset between the feature points, and perform deformation correction on each layer according to the offset, so that the relative position relationship between different layers conforms to the normal retinal anatomical structure, and finally generate a retinal layered topology map. The map records in detail the boundary coordinates, thickness information and relative distance between each layer of structure.

[0051] Color fundus photography images are corrected for illumination uniformity and processed using a homomorphic filtering algorithm. The algorithm first decomposes the image into an illumination component and a reflection component. The illumination component primarily reflects the overall brightness distribution of the image, while the reflection component contains detailed information about blood vessels and the background. By suppressing high-frequency variations in the illumination component, artifacts caused by reflections on the retinal surface are eliminated. Meanwhile, vascular features in the reflection component are enhanced, improving the contrast between blood vessels and background tissue. The corrected image maintains overall brightness balance while clearly displaying vascular structures of varying thicknesses, particularly the details of tiny vessels.

[0052] Time-domain sequence registration is performed on fluorescein angiography images, and images from different time frames are processed using a feature point matching algorithm. The algorithm automatically identifies stable feature points in the image, such as the edge of the optic disc and the bifurcation points of major blood vessels. These feature points have minimal positional changes throughout the angiography process. Based on the coordinate changes of these feature points in different frames, the translation, rotation, and scaling parameters between frames are calculated. These parameters are used to perform geometric transformations on subsequent frames to compensate for vascular position shifts caused by micro-movements of the patient's eye. This ensures that the spatial position of the same vascular structure remains consistent in the time-series images, ensuring accurate tracking of the dynamic flow of contrast agent within the vessels.

[0053] The spatial coordinates of the retinal layer topology map and the color fundus photography images corrected for illumination homogenization were matched and verified. A two-dimensional coordinate system was established with the center of the optic disc as the origin. The characteristic coordinates of the bifurcation points and turning points of major blood vessels in the two images were extracted, and the positional deviation of the corresponding characteristic points in the coordinate system was calculated. The spatial alignment accuracy of the two images was determined by calculating the deviation values ​​of multiple characteristic points. If the deviation value was within the preset range, the matching result was considered to meet the requirements. If the deviation was too large, the registration process was repeated until the spatial consistency requirements were met.

[0054] The time-domain sequence registration results are used to correct the offset of the vascular position in the retinal layer topology map. Based on the position correction parameters obtained from the fluorescence angiography image during the registration process, the position of the vascular annotations in the retinal layer topology map is adjusted accordingly. For example, if the registration results show that the blood vessels in a certain area are offset to the right, the vascular coordinates of that area in the map are shifted to the right by the same distance to ensure that the vascular position in the map is consistent with the actual vascular position in the fluorescence angiography image, ensuring the spatial correspondence of the vascular structures in the different modal data.

[0055] Verification of spatial consistency of multimodal fundus data based on medical gold standard annotation: The "medical gold standard" referred to in this invention refers to a standardized reference dataset formed by independently manually annotating vascular boundaries and anatomical structural features in multimodal fundus data by three ophthalmologists with at least five years of experience in fundus imaging diagnosis. The annotation results were verified for consistency using a Kappa coefficient test (Kappa ≥ 0.85). For verification, eight key anatomical features were extracted (including the optic disc center, the fovea, the origin of the superior temporal vascular arch, the origin of the inferior temporal vascular arch, the bifurcation of the nasal vascular trunk, the bifurcation of the temporal vascular trunk, the apex of the superior nasal vascular arch, and the apex of the inferior nasal vascular arch). The 3D coordinate offsets of the corresponding feature points were calculated between the multimodal data and the gold standard data. If the average offset was ≤ 2 pixels (corresponding to an actual physical distance ≤ 0.04 mm), the spatial consistency was considered satisfactory. Otherwise, the multimodal data was spatially corrected using an affine transformation algorithm. Through the above series of processing and verification steps, a standardized multimodal fundus image dataset with anatomical structure correspondence is finally output. This dataset contains corrected image data of each modality, detailed annotation information and a complete verification report, providing a reliable data foundation for subsequent feature fusion and vascular segmentation.

[0056] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0057] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A fundus blood vessel segmentation method based on multimodal data, characterized in that: This is achieved by following these steps: Synchronously acquire multimodal fundus data including optical coherence tomography images, color fundus photography images, and fluorescein angiography images; Perform preprocessing operations on the acquired multimodal fundus data to generate a standardized multimodal fundus image dataset; Based on a standardized multimodal fundus image dataset, the heterogeneous features of fundus image data of different modalities are fused to construct a multi-dimensional fundus feature space; A deep learning segmentation network model is trained based on the multi-dimensional fundus feature space to generate an initial fundus vessel segmentation mask; Perform confidence evaluation analysis on the initial fundus vessel segmentation mask to obtain a confidence distribution map of the vessel segmentation result; According to the confidence distribution map of the blood vessel segmentation results, the blood vessel segmentation quality level is divided and the quality level determination rules are defined; Perform interactive correction processing on low-confidence areas based on quality level judgment rules to generate the final optimized fundus vessel segmentation results; The final optimized fundus vessel segmentation results are transmitted to the visualization terminal for three-dimensional topology reconstruction.

2. The method for segmenting fundus blood vessels based on multimodal data according to claim 1, characterized in that: The operations for synchronously acquiring multimodal fundus data specifically include: deploying a multispectral imaging sensor array, which is configured with acquisition devices of different spectral bands according to the anatomical structure of the retina; setting the optical coherence tomography image acquisition depth range to cover the inner limiting membrane of the retina to the suprachoroidal space; adjusting the exposure parameters of the color fundus photography image to adapt to different degrees of fundus pigment deposition; configuring the time-series acquisition frequency of the fluorescent angiography image to fully capture the dynamic perfusion process of the contrast agent in the arterial phase, venous phase and late phase; and aligning the acquisition time points of different modal images through a timestamp synchronization mechanism.

3. The method for segmenting fundus blood vessels based on multimodal data according to claim 2, characterized in that: The operations of fusing heterogeneous features of fundus image data of different modalities specifically include: extracting the layered structure feature vector of the optical coherence tomography image, which contains the reflection intensity distribution characteristics of each layered tissue of the retina; analyzing the color space feature vector of the color fundus photography image, which records the chromatic difference information between blood vessels and background tissues; quantizing the time series feature vector of the fluorescence angiography image, which describes the filling dynamics parameters of the contrast agent in the vascular network; inputting the layered structure feature vector, color space feature vector and time series feature vector into the feature fusion engine, which uses a cross-modal attention mechanism to weightedly integrate different feature vectors; and outputting a multi-dimensional fundus feature space with spatial topological correlation.

4. The method for segmenting fundus blood vessels based on multimodal data according to claim 3, characterized in that: The operations of training a deep learning segmentation network model based on the multi-dimensional fundus feature space and generating an initial fundus vascular segmentation mask specifically include: constructing a vascular segmentation network model with an encoder-decoder architecture, wherein the encoder of the vascular segmentation network model includes a multi-scale feature extraction module; using the multi-dimensional fundus feature space as a training sample set, which has been annotated with real vascular boundary information; optimizing the weight parameters of the vascular segmentation network model through the backpropagation algorithm; and using the optimized vascular segmentation network model to infer unlabeled multimodal fundus data to generate an initial fundus vascular segmentation mask containing probability values.

5. The method for segmenting fundus blood vessels based on multimodal data according to claim 4, characterized in that: The confidence assessment and analysis of the initial fundus vascular segmentation mask specifically includes: calculating the probability variance value of each pixel in the initial fundus vascular segmentation mask to generate an original confidence heat map; analyzing the continuity index of the vascular morphological characteristics, which includes branch connectivity and vessel diameter uniformity; combining the original confidence heat map with the vascular morphological continuity index, and generating a confidence distribution map of the vascular segmentation results through an adaptive weighting function.

6. The method for segmenting fundus blood vessels based on multimodal data according to claim 5, characterized in that: The operations for dividing the quality levels of vascular segmentation specifically include: setting a high confidence threshold and a low confidence threshold, which are dynamically adjusted according to clinical diagnostic requirements; marking the areas above the high confidence threshold in the confidence distribution map of the vascular segmentation results as high-quality segmentation areas; marking the areas between the high confidence threshold and the low confidence threshold as medium-quality segmentation areas; marking the areas below the low confidence threshold as low-quality segmentation areas; and generating a fundus vascular quality zoning map containing three-level quality grade annotations.

7. The method for segmenting fundus blood vessels based on multimodal data according to claim 6, characterized in that: The operations of performing interactive correction processing specifically include: receiving the coordinate information of the low-quality segmentation area in the fundus vascular quality zoning map; superimposing and displaying the original multimodal fundus data corresponding to the low-quality segmentation area on the visualization terminal; receiving the vascular boundary correction instructions marked by the physician through the human-computer interaction interface; using the adaptive deformation algorithm to fuse the vascular boundary correction instructions marked by the physician into the initial fundus vascular segmentation mask; and updating and generating the final optimized fundus vascular segmentation result.

8. The method for segmenting fundus blood vessels based on multimodal data according to claim 7, characterized in that: The three-dimensional topology reconstruction operations specifically include: parsing the coordinates of the vascular centerline in the final optimized fundus vascular segmentation results; reconstructing the three-dimensional spatial coordinate mapping relationship of the vascular centerline, which is associated with the retinal surface curvature parameters; generating a three-dimensional mesh model of the vascular wall based on the vascular diameter data; and rendering the three-dimensional topological structure of the vascular network including hemodynamic simulation on the visualization terminal.

9. The method for segmenting fundus blood vessels based on multimodal data according to claim 5, characterized in that: The generation process of the confidence distribution map of the vascular segmentation results includes a feature interaction mechanism: the layered structure feature vector in the multidimensional fundus feature space is input into the confidence assessment module; the time series feature vector in the multidimensional fundus feature space is transmitted to the confidence assessment module; the confidence assessment module calculates the association weight coefficient between the layered structure feature vector and the time series feature vector through a cross-validation algorithm; and the confidence distribution map of the vascular segmentation results is optimized based on the association weight coefficient.

10. The method for segmenting fundus blood vessels based on multimodal data according to claim 1, characterized in that: The preprocessing operation includes a cross-modal data verification mechanism: performing inter-layer structural alignment processing on the optical coherence tomography image to generate a retinal layer topology map; performing illumination uniformity correction on the color fundus photography image to eliminate retinal reflection artifacts; performing time domain sequence registration on the fluorescein angiography image to compensate for the patient's eye movement; performing spatial coordinate matching verification on the retinal layer topology map and the color fundus photography image after illumination uniformity correction; and using the time domain sequence registration result to correct the vascular position offset of the retinal layer topology map; The spatial consistency of multimodal fundus data is verified based on medical gold standard annotation, and a standardized multimodal fundus image dataset with anatomical structure correspondence is output.

Citation Information

Patent Citations

  • Deep network segmentation method for coronary artery lumen contour under OCT image

    CN113362332A

  • Liver blood vessel segmentation method

    CN118379302A

  • Cerebrovascular three-dimensional model construction method and system based on cerebrovascular image segmentation

    CN119359930A

  • Three-dimensional hepatic duct, hepatic artery, portal vein and hepatic vein image segmentation method and system

    CN120107601A

  • Retinochoroidal disease course structure change monitoring system based on artificial intelligence

    CN120318586A

Cited By

  • Arthroscope image real-time anomaly recognition system based on deep learning

    CN121074376A

  • Arthroscopic image real-time abnormality recognition system based on deep learning

    CN121074376B

  • Three-dimensional reconstruction and quantitative analysis consistency control method and system for multi-center CT angiography data, electronic equipment and storage medium

    CN121392150A

  • Method and device for segmenting retinal blood vessels in fundus image

    CN121582274A

  • Cardiovascular image processing method and electronic equipment

    CN121708636A