Diabetic retinopathy grading diagnosis system and method

By constructing a deeply coupled closed-loop collaborative system that integrates multi-lesion collaborative detection, vascular topology analysis, and hierarchical decision-making, the problems of insufficient utilization of lesion spatial correlation information and lack of vascular topology features in existing technologies have been solved, thereby improving the accuracy and reliability of hierarchical diagnosis of diabetic retinopathy.

CN121982008APending Publication Date: 2026-05-05SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
Filing Date
2026-01-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for the graded diagnosis of diabetic retinopathy lack sufficient utilization of spatial correlation information of lesions, lack of vascular topology feature mining, and insufficient adaptive optimization capabilities, resulting in insufficient diagnostic accuracy and reliability.

Method used

A multi-lesion collaborative detection module, a vascular topology map construction module, a hierarchical decision-making module, and an adaptive feedback optimization module are constructed. Through a deeply coupled closed-loop collaborative system, the detection parameters are dynamically adjusted using the spatial distribution characteristics of lesions and vascular topology information to achieve lesion grading diagnosis.

Benefits of technology

It significantly improves the accuracy and reliability of diabetic retinopathy grading, reduces the rate of missed diagnoses and misdiagnoses, and meets the actual needs of primary healthcare institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982008A_ABST
    Figure CN121982008A_ABST
Patent Text Reader

Abstract

The invention discloses a diabetic retinopathy grading diagnosis system and method, and belongs to the technical field of medical image processing and computer vision, and the system comprises a multi-focus cooperative detection module, a vascular topological graph construction module, a hierarchical grading decision module and a self-adaptive feedback optimization module. The multi-focus cooperative detection module identifies five types of typical focuses and calculates focus cooperative perception indexes; the blood vessel topological graph construction module learns blood vessel topological features through a graph neural network; the hierarchical grading decision-making module fuses the focus features and the vascular topological features to generate a grading result and a confidence score; the adaptive feedback optimization module dynamically adjusts detection parameters according to confidence scores to form a closed-loop optimization mechanism, focus collaborative information and a vascular topological structure are fully utilized, accurate and reliable diabetic retinopathy grading diagnosis is achieved through a deep coupling closed-loop collaborative system, and the method is suitable for fundus screening of primary medical institutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical image processing and computer vision technology, specifically to a grading diagnostic system and method for diabetic retinopathy. Background Technology

[0002] Diabetic retinopathy is one of the most common microvascular complications of diabetes and a leading cause of blindness in working-age individuals. According to the International Clinical Diabetic Retinopathy Classification Standard, diabetic retinopathy can be classified into five stages: no obvious lesions, mild non-proliferative stage, moderate non-proliferative stage, severe non-proliferative stage, and proliferative stage. Early detection and accurate classification are crucial for timely intervention and slowing disease progression.

[0003] Currently, primary healthcare institutions generally lack specialized ophthalmologists, and relying on manual image interpretation for diabetic retinopathy screening suffers from problems such as high workload, strong subjectivity, and high rates of missed and misdiagnosed cases. In recent years, deep learning technology has made significant progress in the field of medical image analysis, providing a new solution for the automated diagnosis of diabetic retinopathy.

[0004] Existing automated diagnostic methods for diabetic retinopathy have the following main shortcomings: Chinese patent application No. 202510208120.9 discloses a lesion grading method. This method employs a multi-branch network architecture, including a feature extractor, a multi-branch classifier, and a visual converter. This method processes features at different scales through a multi-branch structure and dynamically adjusts the weights of each branch using the visual converter to address the problem of imbalanced samples. However, this method has the following shortcomings: First, it only focuses on image-level global feature learning, failing to fully utilize the local spatial distribution information of lesions and the collaborative relationships between lesions; second, it ignores the important influence of fundus vascular topology on the grading of diabetic retinopathy, as vascular morphology and topological changes are key indicators for judging the severity of lesions; third, it lacks an effective closed-loop feedback mechanism, failing to dynamically adjust detection parameters based on the confidence level of the diagnostic results, resulting in low diagnostic accuracy for low-confidence samples.

[0005] In summary, existing technologies for the graded diagnosis of diabetic retinopathy suffer from problems such as insufficient utilization of lesion spatial correlation information, lack of vascular topology feature mining, and insufficient adaptive optimization capabilities. There is an urgent need for a graded diagnostic system and method for diabetic retinopathy that can fully integrate lesion collaborative information and vascular topology features and has closed-loop feedback optimization capabilities. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the present invention aims to provide a grading and diagnostic system and method for diabetic retinopathy. By constructing a deeply coupled closed-loop collaborative system that integrates multi-lesion collaborative detection, vascular topology analysis, and hierarchical grading decision-making, the system fully utilizes the spatial distribution characteristics of lesions and vascular topology information. Furthermore, by dynamically adjusting detection parameters through an adaptive feedback optimization mechanism, the system significantly improves the accuracy and reliability of diabetic retinopathy grading, thus meeting the actual needs of fundus screening in primary healthcare institutions.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: A grading and diagnostic system for diabetic retinopathy includes a multi-lesion collaborative detection module, a vascular topology mapping module, a hierarchical grading decision module, and an adaptive feedback optimization module. The multi-lesion collaborative detection module extracts features from fundus images based on a shared feature encoder, identifies five typical lesion types, and outputs a spatial distribution heatmap and a lesion collaborative perception index. The vascular topology mapping module performs vessel segmentation and topology extraction on fundus images, learning vascular topology features through a graph neural network. The hierarchical grading decision module fuses lesion features and vascular topology features, generating grading results and confidence scores through an attention mechanism. The adaptive feedback optimization module dynamically adjusts the lesion detection threshold based on the confidence score, forming a closed-loop optimization mechanism.

[0008] A method for grading and diagnosing diabetic retinopathy includes four main steps: multi-lesion collaborative detection, vascular topology map construction, hierarchical grading decision-making, and adaptive feedback optimization. It achieves accurate grading and diagnosis of diabetic retinopathy through a deeply coupled closed-loop collaborative mechanism.

[0009] Compared with the prior art, the present invention has the following beneficial effects: First, by designing a multi-lesion collaborative detection module, this invention not only detects the location and number of five typical lesions, but also calculates the spatial proximity and co-occurrence patterns between lesions, generating a lesion collaborative perception index. This fully explores the collaborative information of lesions. Compared with existing technologies that only focus on the detection of a single lesion, this invention can more comprehensively reflect the complexity of the lesion and significantly improve the accuracy of grading.

[0010] Secondly, this invention introduces a vascular topology map construction module, which learns the topological structural features of retinal vessels through graph neural networks, thus overcoming the shortcomings of existing technologies that neglect vascular morphology and topological changes. As an important biomarker for diabetic retinopathy, the fusion of vascular topology features with lesion features can provide richer information for grading decisions, improving the comprehensiveness and accuracy of diagnosis.

[0011] Third, this invention designs an adaptive feedback optimization module that dynamically adjusts the lesion detection threshold based on the confidence score of the grading results, forming a closed-loop feedback mechanism. Compared with existing technologies that use fixed detection parameters, this invention can automatically enhance detection sensitivity for low-confidence samples, effectively reducing the false negative rate, while appropriately reducing sensitivity for high-confidence samples to reduce false positives. This achieves adaptive optimization of detection parameters and significantly improves the robustness of the system.

[0012] Fourth, this invention integrates lesion co-sensing index, spatial distribution heatmap and vascular topology features at multiple scales through a hierarchical decision-making module. It uses an attention mechanism to learn the importance weights of different features, enabling the system to automatically focus on the information most valuable for grading. Compared with the simple feature splicing or weighted fusion of existing technologies, the fusion strategy of this invention is more intelligent and efficient, and the grading decision is more accurate and reliable.

[0013] Fifth, through the deep coupling and closed-loop collaboration of four modules, this invention achieves synergistic optimization of lesion detection and hierarchical decision-making, feature complementarity between vascular topology analysis and lesion collaborative perception, and dynamic adjustment of adaptive feedback optimization. Multiple modules promote each other, forming a synergistic effect of 1+1>2, and the overall diagnostic performance is significantly better than the existing technology. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the overall structure of the diabetic retinopathy grading and diagnostic system of the present invention; Figure 2 This is a schematic diagram of the multi-lesion collaborative detection module of the present invention; Figure 3 This is a schematic diagram of the structure of the blood vessel topology map construction module of the present invention; Figure 4 This is a schematic diagram of the hierarchical decision-making module of the present invention; Figure 5 This is a flowchart illustrating the grading and diagnostic method for diabetic retinopathy of the present invention. Detailed Implementation

[0015] Please refer to the attached document. Figures 1-5 To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of protection of this invention.

[0016] Reference Figure 1This invention provides a grading and diagnostic system for diabetic retinopathy, comprising a multi-lesion collaborative detection module 1, a vascular topology map construction module 2, a hierarchical grading decision module 3, and an adaptive feedback optimization module 4. These four modules work collaboratively through deep coupling and a closed-loop feedback mechanism to achieve accurate grading and diagnosis of fundus images.

[0017] Reference Figure 2 The multi-lesion collaborative detection module 1 is the first-level module of the system. Its main function is to detect five typical lesions in fundus images and calculate the collaborative relationships between lesions. This module includes a shared feature encoder, a lesion detection sub-network, and a spatial correlation calculation unit.

[0018] The shared feature encoder uses an improved EfficientNet-B4 as its backbone network. EfficientNet optimizes network depth, width, and resolution simultaneously through a composite scaling method, achieving excellent feature extraction capabilities while maintaining a relatively small number of parameters. The input fundus image has a resolution of 512×512 pixels. After multi-scale feature extraction through seven convolutional blocks, feature maps at different levels are generated. The third-layer feature map has a resolution of 64×64 pixels and a receptive field of 32 pixels, suitable for capturing small lesions such as microaneurysms and small hemorrhages; the fifth-layer feature map has a resolution of 16×16 pixels and a receptive field of 128 pixels, suitable for capturing cotton wool spots and large areas of hard exudates; the seventh-layer feature map has a resolution of 8×8 pixels and a receptive field of 256 pixels, suitable for capturing large-scale lesions such as neovascularization. This multi-scale feature extraction strategy ensures comprehensive perception of lesions of different sizes.

[0019] The lesion detection subnetwork detects five types of lesions based on multi-scale feature maps. For each lesion type, an independent detection head is designed, consisting of two convolutional layers and one classification layer. The detection head outputs the lesion location, class confidence score, and bounding box parameters. Focal Loss is used as the classification loss function, which reduces the weight of easily classified samples, making the network focus more on difficult-to-classify samples, effectively addressing the lesion class imbalance problem. For the five lesion types—microaneurysms, hard exudates, cotton wool spots, hemorrhages, and neovascularization—corresponding detection results are output. The detection results are labeled with bounding boxes indicating the lesion location, with each bounding box containing a class label and a confidence score.

[0020] A spatial distribution heatmap is generated by overlaying all detected lesion locations. The pixel values ​​in the heatmap represent the probability of a lesion being present at that location, with highlighted areas indicating densely distributed lesions. Simultaneously, the number of lesions of each type is counted to form a statistical measure. The spatial distribution heatmap is generated using a Gaussian kernel weighted method, and for each detected lesion's center location... Apply Gaussian weights to the heatmap using the following formula: , in, Coordinates on the heatmap Pixel value at that location, This represents the total number of lesions detected. For the first Confidence weights for individual lesions The standard deviation of the Gaussian kernel is set to 20 pixels in this embodiment, corresponding to an actual distance of approximately 0.1 mm in the fundus image. After the heatmap is generated, it is normalized to the range of 0-1 to facilitate subsequent feature fusion.

[0021] The spatial correlation calculation unit is an innovative design of this module, used to uncover the synergistic relationships between lesions. The clinical manifestations of diabetic retinopathy are often not the independent appearance of a single lesion, but rather the co-occurrence of multiple lesions in a specific area. The co-occurrence pattern of different lesions is of significant indicative value for graded diagnosis. This unit first extracts the center coordinates of all detected lesions and calculates the Euclidean distance between any two lesions. For each lesion... and lesions The formula for calculating spatial distance is: , in, lesion and lesions The Euclidean distance between them lesion The center coordinates, lesion The center coordinates.

[0022] Distance threshold The settings are based on the following clinical and statistical criteria: First, according to the standard field of view for fundus images (usually 45 degrees), a 512×512 pixel resolution corresponds to an actual retinal area of ​​approximately 10mm×10mm, therefore each pixel corresponds to approximately 0.02mm. Second, clinical medical research shows that in diabetic retinopathy, lesions with a spatial distance of less than 0.5mm have a strong pathological correlation and often share the same microvascular damage area. Third, based on statistical analysis of 88,702 images in the EyePACS training set, this invention quantifies the correlation between lesion co-occurrence patterns and grading labels at different distance thresholds. Statistical results show that when... At a pixel size (corresponding to 0.25 mm), the correlation between lesion co-occurrence frequency and mild lesions was 0.42, while the correlation with severe lesions was 0.38, indicating insufficient discrimination. At a pixel size (corresponding to 0.5 mm), the correlation between lesion co-occurrence frequency and mild lesions decreased to 0.18, while the correlation with severe lesions increased to 0.67. The Pearson correlation coefficient difference reached its maximum of 0.49, indicating that the lesion co-sensing index had the strongest ability to distinguish grading at this threshold. At a pixel size (corresponding to 0.75mm), the correlation coefficient difference drops to 0.41, indicating a weakening of discriminative ability. Fourth, a threshold sensitivity experiment was conducted on 1,748 images in the Messidor-2 validation set. At the pixel level, the grading accuracy reached a maximum of 87.6%, with a Kappa coefficient of 0.852. Based on comprehensive clinical evidence and statistical experimental results, this invention sets the distance threshold to... Pixel.

[0023] Statistical distance less than The number of lesion pairs is used to calculate the lesion co-sensing index. Specifically, the lesion co-sensing index is defined. for: , in, This refers to the lesion co-sensing index. This represents the total number of lesions detected. lesion and lesions The Euclidean distance between them This is a distance threshold, with a value of 100 pixels. This is an indicator function; it takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. This is a weighting factor for lesion type. When the lesion... and lesions When they belong to different types, Because the co-occurrence of different types of lesions usually indicates a more severe condition; when lesions and lesions When they belong to the same type, .

[0024] The lesion co-sensing index reflects the spatial clustering and diversity of lesions. A higher index indicates a denser and more diverse distribution of lesions, typically corresponding to a more severe lesion grade. In a preferred embodiment, when... A value less than 0.1 indicates sparse distribution of lesions, usually corresponding to no lesions or mild lesions; when A value between 0.1 and 0.3 indicates moderate aggregation of lesions, typically corresponding to moderate lesions; when... A value greater than 0.3 indicates highly dense lesions, usually corresponding to severe or proliferative lesions.

[0025] The output of the multi-lesion collaborative detection module 1 includes spatial distribution heatmaps of five types of lesions, count statistics, and lesion collaborative perception index. This information will be passed to the hierarchical decision-making module 3 for further analysis. Simultaneously, the detection threshold parameters of this module are dynamically adjusted by the adaptive feedback optimization module 4, forming a closed-loop optimization mechanism.

[0026] Reference Figure 3 The vascular topology mapping module 2 and the multi-lesion collaborative detection module 1 process fundus images in parallel. Their main function is to extract the topological features of fundus vessels. Diabetic retinopathy leads to significant changes in the morphology and topology of retinal vessels, including vascular stenosis, increased tortuosity, and abnormal branching. These changes are important indicators for assessing the severity of the disease. This module includes a vessel segmentation unit, a topology extraction unit, and a graph embedding unit.

[0027] The vessel segmentation unit employs an improved U-Net network for vessel segmentation. U-Net is one of the most classic network architectures in the field of medical image segmentation, and its encoder-decoder structure and skip connection design can effectively fuse multi-scale features. This invention introduces an attention gating mechanism based on U-Net, adding an attention gating module at the skip connections of the decoder. This module dynamically adjusts the weights of the encoder features based on the decoder features, suppressing interference from background regions and enhancing the feature representation of the vessel region. After inputting a fundus image, the vessel segmentation unit outputs a binary mask of vessels, where pixels with a value of 1 represent vessel regions and pixels with a value of 0 represent background regions.

[0028] The topology extraction unit performs skeletonization on the binary mask of blood vessels, extracting the vessel centerline and branch nodes. The skeletonization employs the Zhang-Suen thinning algorithm, which iteratively removes boundary pixels while preserving the single-pixel centerline of the blood vessel, maintaining its connectivity and topological structure. The thinned vessel skeleton retains the main direction and branching relationships of the blood vessels. Branch nodes and endpoints are identified on the vessel skeleton. Branch nodes are defined as pixels with three or more neighboring connections, and endpoints are defined as pixels with only one neighboring connection. The coordinates of all branch nodes and endpoints are extracted, forming the node set of the vessel topology graph. The vessel segments between adjacent nodes form the edge set of the topology graph. Each edge's attributes include the segment's length, average width, and tortuosity. The segment length is calculated by accumulating the number of pixels. The average width is obtained by sampling every 10 pixels along the segment, calculating the width of the blood vessel at each sampling point, and averaging the results. Tortuosity is defined as the ratio of the actual length of the vessel segment to the straight-line distance between its two endpoints; a higher tortuosity indicates a more curved blood vessel. The vascular topology graph represents the vascular structure using nodes and edges. Nodes are labeled with circles, and edges are connected by line segments.

[0029] The graph embedding unit employs a graph neural network to learn features from the vascular topology graph. Graph neural networks are a class of deep learning models specifically designed for processing graph-structured data, capable of learning low-dimensional representations of graph nodes while preserving the graph's topological information. This invention uses a Graph Convolutional Network (GCN) as the core component of the graph embedding unit. For the vascular topology graph, each node is initialized as a feature vector, with node features including three dimensions: node degree, local vessel width, and vessel branch angle at the node. Node degree represents the number of edges connected to the node, reflecting the complexity of the vessel branches; local vessel width is the average width of all vessel segments connected to the node; and vessel branch angle is the angle between adjacent vessel segments at the node. For nodes with three or more branches, the average angle between all branch pairs is taken.

[0030] Graph convolutional networks aggregate feature information from neighboring nodes through multiple layers of convolutional operations. In the first... In layer graph convolution, nodes The feature update formula is: , in, For the first Layer nodes eigenvectors, For nodes The set of neighboring nodes, For nodes The degree, For nodes The degree, For the first The weight matrix of the layer, For the first Layer nodes eigenvectors, For the first The layer's bias vector, The ReLU function is used as the activation function. This formula uses normalized coefficients. Balance the feature contributions of nodes with different degrees to avoid nodes with higher degrees dominating the feature aggregation process.

[0031] This invention employs a 3-layer graph convolutional network. The first layer maps node features from 3D to 64D, the second layer maintains 64D, and the third layer maps features to 128D. After these 3 layers of graph convolution, each node's feature vector contains information about that node and its 3-hop neighborhood nodes, fully integrating local and global features of the vascular topology. Finally, global pooling is used to aggregate all node features into a fixed-length vascular topology feature vector. Global pooling uses a concatenation of average pooling and max pooling, with the following formula: , in, This represents the topological feature vector of blood vessels. This is the average vector of features from all nodes. The total number of nodes. This is a vector of the maximum values ​​of all node features (maximum values ​​are taken element by element). This represents a vector concatenation operation. Average pooling captures the overall statistical properties of vascular topology, while max pooling captures the most salient local features. Combining the two can provide a more comprehensive characterization of vascular topology.

[0032] The output of the vascular topology map construction module 2 is a vascular topology feature vector. This vector has a dimension of 256 (128 dimensions from average pooling + 128 dimensions from max pooling), containing rich information such as the topological structure, branching patterns, vascular width distribution, and tortuosity of the fundus vessels. This information is of great value for the graded diagnosis of diabetic retinopathy. The vascular topology feature vector, together with the lesion features output from the multi-lesion collaborative detection module 1, is passed to the hierarchical grading decision module 3.

[0033] Reference Figure 4 The hierarchical decision-making module 3 is the core of the system's decision-making process. Its main function is to fuse lesion features and vascular topology features to generate the final grading result and confidence score. This module includes a feature fusion unit, an attention weighting unit, and a hierarchical classification unit.

[0034] The feature fusion unit receives the lesion collaborative perception index, spatial distribution heatmap, and vascular topology feature vector output by the multi-lesion collaborative detection module 1 and the vascular topology map construction module 2, respectively. The initial forms and dimensions of the three types of features are different, and they need to be unified into the same feature space for effective fusion.

[0035] For the spatial distribution heatmap, the original resolution is 512×512, containing 262,144 pixel values. To reduce computational complexity and retain key spatial information, adaptive average pooling is first used to downsample the heatmap to 16×16 resolution, retaining the average lesion distribution intensity of 256 regions, and then flattening it into a 256-dimensional vector. The rationale for downsampling to 16×16 is as follows: First, 16×16 corresponds to a receptive field of approximately 32×32 pixels in the fundus image, which matches the scale of important anatomical regions such as the macula and optic disc in clinical practice. Second, the experiment compared three downsampling resolutions: 8×8, 16×16, and 32×32. On the Messidor-2 validation set, the grading accuracy was highest at 16×16 resolution, reaching 87.6%, while the accuracy at 8×8 resolution was only 84.3%, indicating excessive loss of spatial information. The accuracy at 32×32 resolution was 86.1%, indicating that the excessively high feature dimension led to overfitting.

[0036] For the lesion co-perception index It is a scalar value. The value range is typically between 0 and 1, making scalars difficult to directly participate in high-dimensional feature fusion. To enhance their expressive power and match the dimensions of other features, a fully connected layer maps the scalar to a 64-dimensional vector space. The mapping formula is: , in, This is the extended feature vector of the lesion co-sensing index. The weight matrix is ​​a learnable matrix. For learnable bias vectors, is the ReLU activation function. This mapping method allows the network to learn the semantic extension of the lesion co-sensing index in different dimensions. For example, dimensions 1-16 may learn the effect of lesion density on mild lesions, dimensions 17-32 may learn the effect of lesion diversity on severe lesions, dimensions 33-48 may learn the effect of lesion distribution heterogeneity, and dimensions 49-64 may learn the correlation between lesion co-sensing patterns and vascular injury. The rationale for extending to 64 dimensions is as follows: First, 64 dimensions provide sufficient expression space, allowing scalar information to be expanded in multiple semantic directions, while avoiding overfitting due to excessive dimensionality. Second, experiments compared three extended dimensions: 32, 64, and 128. On the APTOS 2019 test set, the 64-dimensional extension achieved the highest hierarchical accuracy at 87.6%, with a Kappa coefficient of 0.852. In contrast, the 32-dimensional extension achieved an accuracy of 86.2%, indicating insufficient expressive power, while the 128-dimensional extension achieved an accuracy of 86.9%, suggesting slight overfitting due to excessive parameters. Third, the 64-dimensional extension maintains a balance in magnitude with the 256-dimensional spatial distribution heatmap and the 256-dimensional vascular topology features, preventing any one type of feature from dominating the fusion process.

[0037] For blood vessel topological feature vectors It is already a high-dimensional vector of 256 dimensions, containing rich information about the topology of blood vessels, and can be directly used for feature fusion.

[0038] The three features mentioned above are concatenated to form a fused feature vector. : , in, To fuse the feature vectors, the total dimension is 256 + 64 + 256 = 576. This indicates a vector concatenation operation.

[0039] The rationality and advantages of this feature fusion strategy are reflected in the following aspects: First, the three types of features represent different levels of the lesion. The spatial distribution heatmap reflects the spatial aggregation pattern and density distribution of lesions, the lesion co-perception index quantifies the co-perception relationship and diversity between lesions, and the vascular topology features represent the structural integrity and morphological changes of the vascular network. The three complement each other to form a comprehensive lesion representation. Second, the feature dimension design follows the principle of information capacity matching. The spatial distribution heatmap and vascular topology features are both 256-dimensional, maintaining equal information capacity. Although the lesion co-perception index is expanded to 64 dimensions, as a global statistic, it does not need to be too high. Third, the splicing and fusion method preserves the independence of various features, avoiding premature information mixing that leads to semantic ambiguity. Subsequent attention weighting units will learn the importance weights of different feature dimensions to achieve intelligent feature selection and weighting. Fourth, ablation experiments verified the effectiveness of this fusion strategy. On the APTOS2019 test set, the accuracy dropped to 84.2% after removing the lesion co-sensing index, 83.8% after removing the spatial distribution heatmap, and 83.7% after removing vascular topology features. All three types of features are indispensable and contribute to the final accuracy of 87.6%.

[0040] The attention-weighted unit employs a self-attention mechanism to calculate the importance weights of different feature dimensions in the fused feature set. The self-attention mechanism learns the dependencies within the features through three linear transformations: query, key, and value. Specifically, the fused feature vector... Through three weight matrices , and Transform into query vectors respectively Key vector Sum value vector : , in, To fuse feature vectors, The weight matrix is ​​a learnable matrix. These are the query vector, key vector, and value vector, respectively.

[0041] Attention weights are calculated by the dot product of the query vector and the key vector, and then scaled and normalized using softmax. , in, This is the attention weight matrix. This is the transpose of the key vector. Let be the dimension of the key vector. This is a scaling factor used to prevent the gradient from vanishing due to excessively large dot product values. The function is a normalization function to ensure that the sum of the attention weights is 1.

[0042] The weighted features are obtained by multiplying the attention weight matrix by the value vector: , in, This is the weighted fused feature vector.

[0043] The self-attention mechanism enables the network to automatically learn the correlation between different feature dimensions, assigning higher weights to features that are highly correlated with the hierarchy and lower weights to redundant features, thereby improving the effectiveness of feature representation.

[0044] The hierarchical classification unit generates hierarchical results and confidence scores based on the weighted fused features. This unit adopts a Multi-Layer Perceptron (MLP) structure, consisting of two fully connected layers and an output layer. The first fully connected layer maps the 576-dimensional features to 256 dimensions, using the ReLU activation function, as shown in the formula: , in, This is the first layer of hidden features. This is the weight matrix. This is the bias vector.

[0045] The second fully connected layer maps the 256-dimensional features to 128 dimensions, also using the ReLU activation function, as shown in the formula: , in, This is the second layer of hidden features. This is the weight matrix. This is the bias vector.

[0046] The output layer is a 5-dimensional fully connected layer, corresponding to the five grades of diabetic retinopathy: no obvious lesions (Category 0), mild non-proliferative stage (Category 1), moderate non-proliferative stage (Category 2), severe non-proliferative stage (Category 3), and proliferative stage (Category 4). The calculation formula for the output layer is as follows: , in, This is the original 5-dimensional output vector (logits). This is the weight matrix. This is the bias vector.

[0047] To map the original 5-dimensional output to a probability distribution, the output layer uses the softmax activation function. The softmax function maps any real-valued vector to a probability distribution, ensuring that the sum of the probabilities of all classes is 1, and that each probability value is between 0 and 1. The softmax function is defined as follows: , in, For category The predicted probability, For category The original output value (logit). It is an exponential function. Softmax amplifies the differences in output values ​​between different categories through exponential operations, and then normalizes to ensure that the sum of probabilities is 1. After the softmax transformation, a 5-dimensional probability distribution vector is obtained. ,satisfy and .

[0048] The final classification result is obtained by selecting the category with the highest probability: , in, The predicted classification category corresponds to the category with the highest probability among the five classifications.

[0049] The confidence score is defined as the probability value for predicting the classification category, that is: , The confidence score reflects the system's degree of certainty about the grading result. When the probability value is close to 1 (e.g., This indicates that the system is very confident in the classification result, and the probability distribution of the five categories is highly concentrated on the predicted category; when the probability value is low (e.g. A confidence score of 0 indicates that the system is not entirely confident in the classification result, and there may be a high probability that other categories also exist, indicating uncertainty in the classification decision. The confidence score will be passed to the adaptive feedback optimization module 4 for dynamically adjusting the detection parameters.

[0050] The output of the hierarchical decision module 3 includes the classification results of diabetic retinopathy. and confidence score The grading result is an integer from 0 to 4, corresponding to the five grading categories. The confidence score is a real number between 0 and 1, with a higher value indicating a more reliable grading result.

[0051] The adaptive feedback optimization module 4 is the core of the system's closed-loop control. Its main function is to dynamically adjust the detection threshold parameters of the multi-lesion collaborative detection module 1 based on the confidence score of the grading results, forming an adaptive optimization mechanism. This module realizes the closed-loop feedback of the system, enabling lesion detection and grading decisions to promote each other and achieve collaborative optimization.

[0052] This module first receives the confidence score output by the hierarchical decision-making module 3. Compare it with a preset high confidence threshold. and low confidence threshold Comparison. High confidence threshold. and low confidence threshold The settings are based on the following statistical and experimental data: First, the confidence threshold setting needs to balance the risks of false negatives (missed diagnoses) and false positives (misdiagnoses). Statistical analysis was performed on 88,702 images in the EyePACS training set, and the relationship curve between confidence score and grading accuracy was plotted. The results show that when the confidence score is higher than 0.85, the grading accuracy reaches 92.8%, indicating high reliability of the system's judgment, and the detection sensitivity can be appropriately reduced to decrease false positives. When the confidence score is lower than 0.60, the grading accuracy is only 76.3%, indicating a higher risk of misjudgment, and the detection sensitivity needs to be increased to reduce the missed diagnoses rate. When the confidence score is between 0.60 and 0.85, the grading accuracy smoothly transitions between 82.5% and 89.1%, indicating moderate reliability of the system's judgment, and the default detection parameters can be maintained.

[0053] Second, threshold sensitivity experiments were conducted on 1,748 images in the Messidor-2 validation set to test the system performance under different combinations of high-confidence thresholds (0.75, 0.80, 0.85, 0.90) and low-confidence thresholds (0.50, 0.55, 0.60, 0.65). Experimental results show that when... At this time, the overall classification accuracy was the highest, reaching 87.6%, with a Kappa coefficient of 0.852. Simultaneously, the sensitivity for severely non-proliferative and proliferative phases reached 85.3% and 89.7%, respectively, effectively balancing accuracy and sensitivity. When the value is set too high (e.g., 0.90), the proportion of high-confidence samples is too small (only 58.7%), the closed-loop feedback mechanism is triggered too frequently, and the optimization effect is limited; when When the sensitivity is set too low (e.g., 0.75), the proportion of high-confidence samples is too high (83.4%), but this includes a large number of samples that would actually be misclassified. Reducing the detection sensitivity will lead to an increase in the false negative rate. Similarly, when When the setting is too low (e.g., 0.50), the proportion of low-confidence samples is too high (accounting for 24.6%), and increasing the detection sensitivity will introduce too many false positives; when When the value is set too high (e.g., 0.65), the proportion of low-confidence samples is too small (only 6.8%), and the effect of closed-loop optimization on improving missed diagnoses is weakened.

[0054] Third, the clinical validation experiment was conducted on real screening data from a top-tier hospital. This dataset contained 2,134 fundus images, with grading results independently annotated by five senior ophthalmologists. When using... At that time, the consistency between the system and the doctor's diagnosis reached 91.3%, with a Cohen's Kappa coefficient of 0.876. For 312 images with a confidence score below 0.60, the system output a low-confidence marker, indicating that manual review was required. After manual review, 87 of these 312 images (27.9%) were actually misdiagnosed. The system's low-confidence marker effectively identified uncertain samples and reduced the risk of misdiagnosis.

[0055] Based on the statistical analysis, threshold sensitivity experiment, and clinical validation results from the above three aspects, this invention sets the high confidence threshold as follows: The low confidence threshold is set to .

[0056] If the confidence score is higher than the high confidence threshold of 0.85, that is... This indicates that the system is highly confident in the current classification results. At this point, the detection sensitivity of the multi-lesion collaborative detection module 1 can be appropriately reduced to decrease false positives. The specific adjustment strategy is as follows: increase the classification threshold of the lesion detection sub-network from the default 0.5 to 0.6, and increase the IoU threshold of Non-Maximum Suppression (NMS) from the default 0.5 to 0.6. Increasing the classification threshold means that only detection boxes with a confidence level greater than 0.6 are retained, making the detection more conservative and only retaining high-confidence lesion detection results. Increasing the IoU threshold of NMS means that only detection boxes with an overlap greater than 0.6 are suppressed, which reduces over-detection of dense lesion regions and avoids repeatedly detecting the same lesion.

[0057] If the confidence score is lower than the low confidence threshold of 0.60, that is... This indicates that the system is not confident enough in the current classification results, and there may be issues with missed lesion detection or inaccurate detection. In this case, it is necessary to improve the detection sensitivity of the multi-lesion collaborative detection module 1. The specific adjustment strategy is as follows: reduce the classification threshold of the lesion detection sub-network from the default 0.5 to 0.4, and reduce the IoU threshold of non-maximum suppression from the default 0.5 to 0.4. The reduction in the classification threshold makes the detection more aggressive, and detection boxes with a confidence greater than 0.4 are retained, which can capture more potential lesions; the reduction in the IoU threshold of NMS means that only detection boxes with an overlap greater than 0.4 are suppressed, which retains more possible lesion detection results and reduces the missed detection rate.

[0058] If the confidence score is between the low confidence threshold of 0.60 and the high confidence threshold of 0.85, that is... This indicates that the system's grading results have moderate reliability. At this point, the detection parameters of the multi-lesion collaborative detection module 1 should remain unchanged, and the classification threshold and the IoU threshold of NMS should both remain at their default values ​​of 0.5.

[0059] After parameter adjustment, the system re-executes the processing flow of the multi-lesion collaborative detection module 1 and the hierarchical grading decision module 3. Based on the adjusted detection parameters, it re-detects lesions, calculates new lesion collaborative perception indices and spatial distribution heatmaps, and regenerates grading results and confidence scores. This feedback adjustment process executes a maximum of 3 iterations. If the confidence score is still lower than the low confidence threshold of 0.60 after 3 adjustments, the system will output the current grading result and mark it as a low confidence result, indicating that manual review is required. On the APTOS 2019 test set, after adaptive feedback optimization, 73.2% of the samples reached high confidence after the first iteration. ), 18.6% of the samples reached a moderate confidence level after the second iteration ( 5.4% of the samples reached medium or high confidence after the third iteration, while only 2.8% of the samples remained at low confidence after three iterations. The system outputs low-confidence labels for these samples and recommends manual review.

[0060] The adaptive feedback optimization module 4 implements a closed-loop optimization mechanism for the system, enabling dynamic adjustment of lesion detection parameters based on the reliability of the grading results. This effectively addresses complex situations such as differences in image quality and blurred lesion features, significantly improving the system's robustness and diagnostic accuracy. This module forms a close coupling relationship with the multi-lesion collaborative detection module 1 and the hierarchical grading decision module 3. The three modules promote each other and optimize collaboratively, achieving a synergistic effect of 1+1>2.

[0061] In a preferred embodiment, the system further includes an image quality assessment module, which performs a quality assessment on the input fundus image as a preprocessing step. This module detects three key indicators: pupil size, refractive medium clarity, and image contrast.

[0062] Pupil size detection uses the Hough circle detection algorithm to identify the circular pupil region in the center of the fundus image and calculate the ratio of pupil diameter to image size. If the pupil is too small, with a ratio below 0.6, it indicates insufficient visual field and potential missed peripheral lesions. The system outputs a re-image prompt, suggesting that the image be re-acquired.

[0063] The sharpness assessment of the refractive media uses the Laplacian operator to calculate the image sharpness score. The Laplacian operator performs a second-order differential operation on the image; the variance of the differential value is larger for sharp images and smaller for blurry images. The Laplacian variance of the entire image is calculated. If the variance is lower than a preset threshold of 100, it indicates that the image is blurry, possibly due to turbidity of the refractive media or inaccurate focusing, and the system outputs a retake prompt.

[0064] Image contrast assessment calculates the grayscale histogram of an image and statistically analyzes the distribution range of pixel values. If pixel values ​​are concentrated within a narrow range, the contrast is low, which is detrimental to the accurate identification of lesions. The image contrast index is calculated as follows: ,in and These are the maximum and minimum pixel values ​​of the image, respectively. If this value is below 0.3, the system will output a retake prompt, suggesting adjusting shooting parameters or enhancing image contrast.

[0065] The image quality assessment module can identify substandard images before diagnosis, avoiding invalid analysis of low-quality images, reducing the risk of misdiagnosis, and improving the system's practicality and reliability.

[0066] Reference Figure 5 The present invention also provides a method for grading and diagnosing diabetic retinopathy, which is based on the above system and includes the following steps: Step S1: Based on the shared feature encoder, feature extraction is performed on the fundus image to identify five types of lesions: microaneurysms, hard exudates, cotton wool spots, hemorrhages, and neovascularization. Spatial distribution heatmaps and count statistics of each type of lesion are output, and the lesion co-sensing index is calculated based on the spatial proximity and co-occurrence patterns between lesions.

[0067] Specifically, fundus images are input into a shared feature encoder, and after multi-scale feature extraction, feature maps of different levels are generated. Based on the multi-scale feature maps, a lesion detection subnetwork detects the location and category of five types of lesions, and outputs the detection results. The detected lesion locations are superimposed to generate a spatial distribution heatmap, and the number of lesions of each type is counted. The center coordinates of all lesions are extracted, and the Euclidean distance between any two lesions is calculated. Lesions with distances less than a threshold are counted. The number of lesion pairs per pixel is used to calculate the lesion co-sensing index based on the lesion type weighting factor. .

[0068] Step S2: Perform vascular segmentation on the fundus image, extract the vascular centerline and branch nodes, construct a vascular topology map, use vascular nodes as graph nodes and vascular segments as graph edges, and use a graph neural network to embed and learn the vascular topology map to generate vascular topology feature vectors.

[0069] Specifically, a vessel segmentation unit is used to segment blood vessels in the fundus image, outputting a binary mask of blood vessels. The binary mask is then skeletonized to extract the vessel centerline. Branch nodes and endpoints are identified on the vessel skeleton, forming a node set for the vessel topology graph, with vessel segments between adjacent nodes forming an edge set. A feature vector is initialized for each node, including node degree, local vessel width, and vessel branch angle. Multi-layer feature learning is performed using a graph convolutional network to aggregate features from neighboring nodes. Finally, global pooling is used to generate the vessel topology feature vector. .

[0070] Step S3: Multi-scale feature fusion of lesion co-perception index, spatial distribution heat map and vascular topology feature vector, and learning the mapping relationship between different lesion combination patterns and international clinical grading standards through attention mechanism to generate diabetic retinopathy grading results and confidence scores.

[0071] Specifically, the spatial distribution heatmap is downsampled and flattened to obtain a 256-dimensional vector. The lesion co-sensing index Expanded to a 64-dimensional vector through a fully connected layer. . with vascular topological feature vector The features are concatenated to form a 576-dimensional fused feature vector. The importance weights of different feature dimensions in the fused features are calculated using a self-attention mechanism, and the fused features are then weighted to obtain the weighted fused features. Based on the weighted fusion features, a 5-dimensional raw output is generated through a multilayer perceptron. After being mapped to a 5-dimensional probability distribution by the softmax function, The category with the highest probability is selected as the classification result. The probability value of this category is used as the confidence score. .

[0072] Step S4: Dynamically adjust the lesion detection threshold parameter based on the confidence score. When the confidence score is lower than the preset threshold, enhance the sensitivity of lesion detection and form a closed-loop feedback adjustment mechanism.

[0073] Specifically, confidence scores With high confidence threshold and low confidence threshold Compare. If To reduce the sensitivity of lesion detection, the classification threshold was increased to 0.6, and the IoU threshold of NMS was increased to 0.6. If To improve the sensitivity of lesion detection, the classification threshold was lowered to 0.4, and the IoU threshold of NMS was also lowered to 0.4. If The default detection parameters remain unchanged. After adjusting the parameters, steps S1 and S3 are re-executed to re-detect lesions based on the adjusted parameters and generate grading results. This feedback adjustment process is performed a maximum of three iterations to ensure that the system outputs highly reliable grading results.

[0074] In a preferred embodiment, the method further includes an image quality assessment step, in which the fundus image is assessed before step S1, and pupil size, refractive medium clarity and image contrast are detected. When the image quality does not meet the preset standard, a retake prompt is output and no further analysis is performed.

[0075] Example: The complete calculation process for actual test samples.

[0076] To clearly demonstrate the effectiveness of the technical means and specific calculation process of the system of the present invention, the following uses an actual test sample as an example to explain in detail the complete processing flow from input fundus image to output grading result, including the specific values ​​of all key technical parameters and intermediate calculation results.

[0077] This embodiment uses a fundus image from the APTOS 2019 test set as the input sample. The image is numbered Sample_3257.png, has a resolution of 512×512 pixels, and is labeled by a professional ophthalmologist as a moderate non-proliferative stage (category 2).

[0078] Step 1: Collaborative detection of multiple lesions.

[0079] Input Sample.png into the multi-lesion collaborative detection module 1, and the shared feature encoder EfficientNet-B4 extracts multi-scale features. The lesion detection subnetwork detects the following lesions: Microaneurysms: 17 were detected, with center coordinates as follows: , , , , , , , , , , , , , , , , The confidence scores were all between 0.52 and 0.79.

[0080] Hard exudates: 9 were detected, with center coordinates as follows: , , , , , , , , The confidence scores were all between 0.58 and 0.84.

[0081] Cotton lint spots: 4 were detected, with center coordinates as follows: , , , The confidence scores were all between 0.61 and 0.73.

[0082] Hemorrhage foci: 6 were detected, with center coordinates as follows: , , , , , The confidence scores were all between 0.55 and 0.76.

[0083] New blood vessels: 0 detected.

[0084] Total detected Individual lesions. Different types of lesions are marked with different colored borders.

[0085] Based on the detected lesion locations, a spatial distribution heatmap is generated using a Gaussian kernel weighting method. For the center location of each lesion... Apply Gaussian distribution weights and Gaussian kernel standard deviations to the heatmap. Pixels. The generated heatmap has a resolution of 512×512, and the densely populated areas of lesions in the heatmap (such as...) The pixel value in the vicinity is as high as 0.87, and the sparse lesion areas (such as...) The pixel value in the vicinity is only 0.03. After the heatmap is normalized to the 0-1 range, the highlighted areas represent densely distributed lesions.

[0086] Calculate the lesion co-sensing index First, calculate the Euclidean distance between any two lesions. The distance threshold is set to Pixels. Count the number of lesion pairs with a distance of less than 100 pixels. For example, microaneurysms. With hard exudate The distance between them is: , The lesion is less than the threshold distance and belongs to a different type, therefore the weighting factor... Traverse all Among the lesion pairs, statistics showed that 214 lesion pairs were less than 100 pixels apart, of which 168 pairs belonged to different types (weight 1.5), and 46 pairs belonged to the same type (weight 1.0). The lesion co-sensing index was calculated as follows: , The lesion co-sensing index of this sample A high level (greater than 0.3) indicates highly dense and diverse lesions, usually corresponding to severe or proliferative lesions.

[0087] Step 2: Construction of blood vessel topology map.

[0088] Input Sample_3257.png into the blood vessel topology construction module 2. The blood vessel segmentation unit outputs a binary mask of blood vessels based on the improved U-Net network. In the binary mask, white areas represent blood vessels, and black areas represent the background. The binary mask is then processed using Zhang-Suen skeletonization to extract the center lines of the blood vessels.

[0089] The topology extraction unit identifies on the vascular skeleton There are 43 branch nodes (degree ≥ 3) and 84 endpoint nodes (degree = 1). 156 vascular segments form the edges of the topological graph between adjacent nodes. The vascular topological graph represents the vascular structure using nodes and edges.

[0090] Initialize a 3D feature vector for each node, including node degree, local vessel width, and vessel branch angle. (Based on node...) (coordinate Taking this node as an example, it is a branch node with a degree of 3, connecting 3 blood vessel segments. The average widths of the 3 blood vessel segments are 5.2 pixels, 4.8 pixels, and 6.1 pixels, respectively. The local blood vessel width is calculated as follows: Pixels. The angles between each pair of the three vessel segments are 118 degrees, 127 degrees, and 115 degrees, respectively. The average value of the vessel branch angles is taken. Degree. Therefore, the node The initial feature vector is .

[0091] Feature learning is performed using a 3-layer graph convolutional network. The first layer maps node features from 3D to 64D, the second layer maintains 64D, and the third layer maps features to 128D. (Based on nodes...) For example, its neighborhood node set is In the first layer of graph convolution, nodes The features are updated as follows: , Among them, nodes The degree is 2, node The degree is 3, node The degree is 4. After normalization and activation function, the node... Layer 1 features It is a 64-dimensional vector. After three layers of graph convolution, the nodes... Features It is a 128-dimensional vector that integrates information about the node and its 3-hop neighborhood nodes.

[0092] Global pooling aggregates the features of all 127 nodes into a vascular topology feature vector. Average pooling calculates the average of the features of all nodes. , Max pooling retrieves the maximum value from each element: , By concatenating the two, a 256-dimensional vascular topological feature vector is obtained. .

[0093] Step 3: Hierarchical decision-making.

[0094] The feature fusion unit receives three types of features: spatial distribution heatmap (512×512), lesion co-sensing index (scalar 0.473), and vascular topology feature vector (256-dimensional).

[0095] Adaptive average pooling is applied to the spatially distributed heatmap, downsampling to 16×16 resolution, and then flattening it into a 256-dimensional vector. After downsampling, the typical region values ​​of the 16×16 heatmap are as follows: upper left region The value is 0.12, in the central area. The value is 0.68, in the lower right corner area. The value is 0.34.

[0096] Lesion Co-perception Index Expanded to a 64-dimensional vector through a fully connected layer. : , After weight matrix and bias After transformation and ReLU activation, a 64-dimensional vector is obtained. Typical element value .

[0097] The three types of features are concatenated to form a 576-dimensional fused feature vector: , The attention-weighted unit calculates the importance weights of the fused features through a self-attention mechanism. The fused feature vector undergoes three linear transformations: query, key, and value, to obtain... Calculate attention weights: , Attention weight matrix This reflects the correlation between different feature dimensions. The weighted fusion features are: , The hierarchical classification unit generates hierarchical results based on the weighted fused features. The first fully connected layer outputs 256-dimensional hidden features. The second fully connected layer outputs 128-dimensional hidden features. The output layer generates a 5-dimensional raw output. (These values ​​are the actual output of this sample).

[0098] The 5-dimensional raw output is mapped to a probability distribution using the softmax function: Calculate the index value: , , , , .

[0099] The normalization factor is .

[0100] We obtain a 5-dimensional probability distribution: , The predicted classification result is the category with the highest probability: , The probability value of the confidence score being in this category: , Step 4: Adaptive feedback optimization.

[0101] The adaptive feedback optimization module receives confidence scores. With high confidence threshold and low confidence threshold Compare them.

[0102] because The confidence score falls between the low and high confidence thresholds, indicating that the system's grading results have moderate reliability. At this point, the detection parameters of the multi-lesion collaborative detection module remain unchanged, with the classification threshold and NMS IoU threshold both maintaining their default values ​​of 0.5. The system does not adjust parameters or re-detect, and directly outputs the grading results.

[0103] The system output a grade of diabetic retinopathy as moderate non-proliferative stage (Category 2), with a confidence score of 0.791 (79.1%). This grade is consistent with the annotation results of professional ophthalmologists, verifying the diagnostic accuracy of the system of this invention.

[0104] Intermediate results of the complete processing flow include lesion detection results, spatial distribution heatmap, vascular segmentation nodes, vascular topology map, and final grading results and confidence scores.

[0105] The complete calculation process of this actual test sample clearly demonstrates the specific values ​​of each key technical parameter of the system of the present invention, the intermediate calculation steps and the final output results, proving the effectiveness and feasibility of the technical means of the present invention.

[0106] The system of this invention was trained and validated on publicly available datasets. The EyePACS dataset, containing 88,702 fundus images labeled with five hierarchical categories, was used as the training set. The Messidor-2 dataset, containing 1,748 fundus images, was used as the validation set. The APTOS 2019 dataset, containing 3,662 fundus images, was used as the test set.

[0107] Data preprocessing includes image size normalization, contrast enhancement, and data augmentation. All images are resized to 512×512 pixels, and the CLAHE algorithm is used for contrast enhancement. Data augmentation includes random rotation, flipping, brightness adjustment, and Gaussian noise addition to expand the training data and improve the model's generalization ability.

[0108] The system training employs an end-to-end supervised learning approach. The loss function comprises three parts: lesion detection loss, hierarchical classification loss, and adaptive feedback loss. The lesion detection loss uses a combination of Focal Loss and Smooth L1 Loss, the hierarchical classification loss uses cross-entropy loss, and the adaptive feedback loss uses confidence loss. The total loss function is a weighted sum of the three loss parts. The optimizer used is Adam, with an initial learning rate of 0.001, which decays to 0.1 every 30 epochs, for a total of 100 epochs of training.

[0109] On the APTOS 2019 test set, the system of this invention achieved significant performance improvements. The five-category classification accuracy reached 87.6%, an improvement of 6.4 percentage points compared to the 81.2% of the comparative method. The Kappa coefficient reached 0.852, an improvement of 0.060 compared to 0.792 of the comparative method. For the two clinically most important categories, severe non-proliferative and proliferative phases, the sensitivities reached 85.3% and 89.7%, respectively, significantly higher than the 78.6% and 82.1% of the comparative method. The proportion of samples with a confidence score higher than 0.85 reached 73.2%, and the classification accuracy of these samples reached 92.8%, indicating that the system's confidence score can effectively reflect the reliability of the classification results.

[0110] Ablation experiments validated the contributions of each module. Removing the lesion co-sensing index reduced the accuracy to 84.2%, indicating that lesion co-sensing information is of significant value for grading decisions. Removing vascular topological features reduced the accuracy to 83.7%, demonstrating that vascular topology is indispensable for assessing lesion severity. Removing the adaptive feedback optimization mechanism reduced the accuracy to 85.1%, indicating that closed-loop optimization can effectively improve system robustness.

[0111] The system of this invention also performs excellently in actual clinical applications. On real screening data from a top-tier hospital, the system's processing time for a single image is 0.8 seconds, meeting the needs of high-throughput clinical screening. Compared with the diagnostic results of five senior ophthalmologists, the system's consistency reached 91.3%. For cases with discrepancies, the system can output a low-confidence marker, indicating the need for manual review, effectively reducing the risk of misdiagnosis.

[0112] In summary, this invention, by constructing a deeply coupled closed-loop collaborative system for multi-lesion collaborative detection, vascular topology analysis, and hierarchical decision-making, fully utilizes the spatial distribution characteristics of lesions and vascular topology information, and dynamically adjusts detection parameters through an adaptive feedback optimization mechanism, significantly improving the accuracy and reliability of diabetic retinopathy grading, and providing effective technical support for fundus screening in primary healthcare institutions.

[0113] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A grading and diagnostic system for diabetic retinopathy, characterized in that, include: The multi-lesion collaborative detection module is used to extract features from fundus images based on a shared feature encoder, identify five types of lesions: microaneurysms, hard exudates, cotton wool spots, hemorrhages, and neovascularizations, output spatial distribution heatmaps and count statistics of various lesions, and calculate the lesion collaborative perception index based on the spatial proximity and co-occurrence patterns between lesions. The vascular topology map construction module is connected to the multi-lesion collaborative detection module. It is used to segment blood vessels in fundus images, extract the vascular centerline and branch nodes, construct a vascular topology map, use vascular nodes as graph nodes and vascular segments as graph edges, and perform embedding learning on the vascular topology map through a graph neural network to generate vascular topology feature vectors. The hierarchical decision-making module is connected to the multi-lesion collaborative detection module and the vascular topology map construction module. It is used to perform multi-scale feature fusion of the lesion collaborative perception index, the spatial distribution heat map and the vascular topology feature vector. Through the attention mechanism, it learns the mapping relationship between different lesion combination patterns and international clinical grading standards to generate diabetic retinopathy grading results and confidence scores. An adaptive feedback optimization module, connected to the hierarchical decision-making module and the multi-lesion collaborative detection module, is used to dynamically adjust the detection threshold parameters of the multi-lesion collaborative detection module according to the confidence score. When the confidence score is lower than the preset threshold, the sensitivity of lesion detection is enhanced, forming a closed-loop feedback adjustment mechanism.

2. The diabetic retinopathy grading and diagnostic system according to claim 1, characterized in that, The multi-lesion collaborative detection module includes: A shared feature encoder is used for multi-scale feature extraction from fundus images; The lesion detection subnetwork, connected to the shared feature encoder, is used to detect the location and category of five types of lesions based on the multi-scale features. The spatial correlation calculation unit is connected to the lesion detection sub-network and is used to calculate the spatial proximity between different lesion types and generate a lesion collaborative perception index.

3. The diabetic retinopathy grading and diagnostic system according to claim 2, characterized in that, The spatial correlation calculation unit calculates the Euclidean distance between any two lesions based on the coordinates of the lesion center, counts the co-occurrence frequency of lesions within a preset distance threshold, and determines the lesion co-perception index based on the ratio of the lesion co-occurrence frequency to the total number of lesions.

4. The diabetic retinopathy grading and diagnostic system according to claim 1, characterized in that, The blood vessel topology construction module includes: The blood vessel segmentation unit is used to segment blood vessels in fundus images and generate a binary mask of blood vessels. The topology extraction unit, connected to the blood vessel segmentation unit, is used to perform skeletonization processing on the binary mask of the blood vessel, extract the blood vessel centerline, and identify the blood vessel nodes on the blood vessel centerline. The blood vessel nodes include branch nodes and endpoints, and the blood vessel centerline segments between adjacent blood vessel nodes are regarded as blood vessel segments. The graph embedding unit, connected to the topology extraction unit, is used to map the blood vessel nodes and the blood vessel segments into a graph structure, and to generate the blood vessel topology feature vector through feature learning via a graph convolutional network.

5. The diabetic retinopathy grading and diagnostic system according to claim 4, characterized in that, The graph embedding unit assigns a node feature vector to each blood vessel node. The node feature vector includes node degree, local blood vessel width, and blood vessel branch angle at the node. The features of neighboring nodes are aggregated through multi-layer graph convolution operations, and finally the blood vessel topology feature vector is generated through global pooling.

6. The diabetic retinopathy grading and diagnostic system according to claim 1, characterized in that, The hierarchical decision-making module includes: The feature fusion unit is used to concatenate the lesion collaborative perception index, the spatial distribution heat map, and the vascular topology feature vector to form a fused feature; An attention weighting unit, connected to the feature fusion unit, is used to calculate the importance weights of different feature dimensions in the fused features through a self-attention mechanism. A hierarchical classification unit, connected to the attention weighting unit, is used to generate the diabetic retinopathy grading result and the confidence score based on the weighted fusion features.

7. The diabetic retinopathy grading and diagnostic system according to claim 6, characterized in that, The attention weighting unit calculates attention weights by performing a dot product operation on the query vector, key vector, and value vector, and weights different dimensions of the fused features, highlighting feature dimensions that are highly relevant to the hierarchical classification.

8. The diabetic retinopathy grading and diagnostic system according to claim 1, characterized in that, The adaptive feedback optimization module compares the confidence score with preset high confidence thresholds and low confidence thresholds. If the confidence score is higher than the high confidence threshold, the detection sensitivity of the multi-lesion collaborative detection module is reduced. If the confidence score is lower than the low confidence threshold, the detection sensitivity of the multi-lesion collaborative detection module is increased.

9. The diabetic retinopathy grading and diagnostic system according to claim 1, characterized in that, Also includes: The image quality assessment module is used to assess the quality of fundus images, detect pupil size, refractive medium clarity, and image contrast, and output a retake prompt when the image quality does not meet the preset standards.

10. A method for grading and diagnosing diabetic retinopathy, using the grading and diagnosing system for diabetic retinopathy as described in any one of claims 1-9, characterized in that, Includes the following steps: Based on the shared feature encoder, feature extraction is performed on fundus images to identify five types of lesions: microaneurysms, hard exudates, cotton wool spots, hemorrhages, and neovascularization. Spatial distribution heatmaps and count statistics of each type of lesion are output, and the lesion co-sensing index is calculated based on the spatial proximity and co-occurrence patterns between lesions. Blood vessel segmentation is performed on fundus images, and the center line and branch nodes of blood vessels are extracted to construct a blood vessel topology map. Blood vessel nodes are used as graph nodes and blood vessel segments are used as graph edges. The blood vessel topology map is embedded and learned through graph neural network to generate blood vessel topology feature vectors. The lesion co-sensing index, the spatial distribution heatmap, and the vascular topology feature vector are fused at multiple scales. The mapping relationship between different lesion combination patterns and international clinical grading standards is learned through an attention mechanism to generate diabetic retinopathy grading results and confidence scores. The lesion detection threshold parameter is dynamically adjusted based on the confidence score. When the confidence score is lower than the preset threshold, the sensitivity of lesion detection is enhanced, forming a closed-loop feedback adjustment mechanism.

Citation Information

Patent Citations

  • Lesion grading method, device, equipment and storage medium

    CN119693375B