High-precision classification method for diabetic retinopathy based on hybrid deep learning model

CN122821189APending Publication Date: 2026-09-25GUIZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345164.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明的技术任务是提供一种基于混合深度学习模型的糖尿病视网膜病变高精度分类方法,来解决对眼底图像进行分类,提高对糖尿病视网膜病变分类的准确率的问题

Benefits of technology

[0016]与现有技术相比,本发明具有以下有益效果:提出了一种基于混合深度学习模型的糖尿病视网膜病变高精度分类方法。首先结合多任务损失函数(Multi-Task LossFunction, MTLF),有效利用了图像中的多层次信息,增强了模型对病变特征的识别能力;其次,集成了多种数据增强(Data Augmentation, DA)策略,通过增加训练数据的多样性和复杂性,提高了模型的泛化能力;最后,结合NAM(Normalization-based AttentionModule, NAM)注意力机制,使模型能够聚焦于图像中的关键病变区域,进一步提升了分类的准确性。通过在Kaggle上的糖尿病视网膜病变数据集IDRiD上进行实验验证,本发明的算法相比基准算法MobileNetV3,在Accuracy上提高了6%,达到了84.6%,显著提升了糖尿病视网膜病变图像分类的精度,展示了所提方法的有效性和优越性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821189A_ABST
    Figure CN122821189A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of medical image processing, and a high-precision classification method for diabetic retinopathy based on a hybrid deep learning model, which solves the technical problem that existing diabetic retinopathy grading methods are too dependent on personal experience and subjective judgment and have low grading accuracy, and is based on MobileNetV3 architecture for model building, wherein a multi-task loss function (MTLF) is designed to combine the losses of classification tasks and segmentation tasks to jointly optimize MobileNetV3, a data augmentation strategy (DA) is combined with MobileNetV3 to improve the generalization ability of the algorithm, and a NAM attention mechanism is combined to improve the feature expression ability of MobileNetV3. The deep learning model is used to automatically classify diabetic retinopathy in the test set, which can provide reliable clinical auxiliary decision-making imaging support for doctors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to a deep learning-based method for classifying fundus images of diabetic retinopathy. Background Technology

[0002] Diabetes, a global chronic disease, poses a serious threat to patients' lives and health due to its complications. Diabetic retinopathy (DR), one of the most common microvascular complications, is a leading cause of blindness in adults. With the continued increase in the number of diabetes patients, early detection and accurate diagnosis of diabetic retinopathy have become particularly important. However, traditional diagnostic methods rely on the experience of professional ophthalmologists, which is not only time-consuming and labor-intensive but also highly subjective. Therefore, developing efficient and accurate automatic classification algorithms for diabetic retinopathy images using advanced computer vision and machine learning technologies has become a research hotspot in the intersection of medical imaging and artificial intelligence.

[0003] However, traditional DR lesion identification still has the following problems that need improvement: imbalanced datasets, differences in retinal image quality, model interpretability, and its generalization ability across different populations. Summary of the Invention

[0004] The technical objective of this invention is to provide a high-precision classification method for diabetic retinopathy based on a hybrid deep learning model, in order to solve the problem of classifying fundus images and improving the accuracy of classifying diabetic retinopathy.

[0005] The technical objective of this invention is achieved as follows: a high-precision classification method for diabetic retinopathy based on a hybrid deep learning model, the specific method of which is as follows.

[0006] Data Acquisition: Preprocessed 224x224 pixel images were acquired and divided into five levels from 0 to 4, namely Mild, Moderate, No DR, Proliferate DR, and Severe, containing thousands of fundus images of diabetic retinopathy.

[0007] A multi-task learning framework is employed, combining diabetic retinopathy (DR) classification and lesion segmentation tasks. An MTLF (Mean Transmission-Lowering Failure) loss function is defined, which considers both classification and segmentation losses. The classification loss evaluates the model's accuracy in predicting DR grades, while the segmentation loss evaluates the model's performance in segmenting lesion regions.

[0008] Segmentation task input image ( The processed output is a predicted segmentation map. ( ) ,That , , and These are the length, width, number of channels, and segmentation type, respectively. To evaluate prediction accuracy, [the following parameters will be used]. Compared with the real segmentation mask For comparison, a multi-class cross-entropy loss function is used to quantify the differences between the two classifications; the smaller the loss, the better the segmentation performance. As shown in the formula: (1) In a multi-task learning framework, for scenarios with multiple outputs where each output performs the same multi-classification task, the same loss function is applied independently to each output. The final loss is defined by averaging all the losses, thereby comprehensively evaluating and optimizing the overall performance of the model, as shown in the formula: Based on the MobileNetV3 algorithm, MTLF is combined to effectively solve the problems of overfitting and insufficient feature representation that may be encountered during single-task training. This enables the model to more accurately identify lesion grades and more finely segment lesion areas in DR image classification tasks, providing more comprehensive and reliable auxiliary information for medical diagnosis.

[0009] To address issues such as image defects, out-of-focus images, underexposure, and overexposure in the dataset, a series of Data Analysis (DA) strategies are employed. These strategies include cropping black borders, contrast enhancement, random rotation, scaling, and flipping. DA increases the diversity of training samples, improves the model's generalization ability, and reduces the risk of overfitting. The DA strategy effectively alleviates the overfitting problem caused by the limited dataset, further improving classification accuracy and efficiency.

[0010] NAM, as an innovative attention mechanism, makes models more efficient while maintaining similar performance by suppressing less significant weights and features. It combines channel attention and spatial attention, using a scaling factor of Batch Normalization (BN) to measure the importance of channels and pixels, thereby achieving effective feature recognition and utilization.

[0011] The channel attention module of NAM uses the scaling factor of Batch Normalization (BN) to measure channel importance. By calculating the variance of the feature response of each channel (the standard deviation in BN), the magnitude of feature change and information richness of each channel are obtained. The scaling factor, as a trainable parameter, adjusts the feature weights, increasing the weight of important channels and decreasing the weight of insignificant channels. In the channel attention submodule, the module first passes the input features through a BN layer, and then uses the scaling factor of the BN layer to weight the feature map, effectively suppressing insignificant channels and improving model efficiency and performance.

[0012] and Small batches The mean and standard deviation, and For the trainable scale and bias, the BN scaling factor is as shown in the formula: The calculation process for the channel attention submodule is shown in the formula: in For output features, This is the channel scaling factor, and its weight is... As shown in the formula: The spatial attention submodule of NAM also utilizes the Batch Normalization (BN) scaling factor to measure pixel importance and performs pixel normalization. It assesses the importance of each pixel by calculating the variance of its feature response at each location. In the spatial attention submodule, the feature map is first normalized at the pixel level, and then weighted according to the normalization results, strengthening the weights of important pixel regions and weakening the weights of unimportant regions. In this way, NAM can further suppress insignificant pixel regions in the image, allowing the model to focus more on significant image features, thereby improving the model's accuracy and robustness.

[0013] The BN scaling factor is applied to the spatial dimension to measure pixel importance. The calculation process of the spatial attention submodule is shown in Equation 6: The output feature has the following weights: As shown in the formula: To suppress less significant weights, a regularization term is added to the loss function, as shown in the formula:

[0014] The NAM attention mechanism, by efficiently utilizing batch normalization scaling factors, improves model performance and aims to address the challenge of reducing computational costs while maintaining high accuracy in deep learning models. In MobileNetV3, NAM is integrated with the model's NAM attention mechanism. By evaluating the importance of channels and pixels in feature maps and applying intelligent weighting and suppression, lightweight networks like MobileNetV3 can more accurately capture fine-grained features of DR images, significantly improving classification performance and robustness while maintaining the advantages of high computational efficiency and fast training.

[0015] The training and evaluation module is used to input preprocessed data into the training model and to validate and evaluate the model.

[0016] Compared with existing technologies, this invention has the following advantages: It proposes a high-precision classification method for diabetic retinopathy based on a hybrid deep learning model. First, by combining a multi-task loss function (MTLF), it effectively utilizes multi-level information in the image, enhancing the model's ability to identify lesion features. Second, it integrates various data augmentation (DA) strategies, increasing the diversity and complexity of training data to improve the model's generalization ability. Finally, by combining a NAM (Normalization-based Attention Module) attention mechanism, the model can focus on key lesion regions in the image, further improving classification accuracy. Experiments on the IDRiD diabetic retinopathy dataset on Kaggle demonstrate that the algorithm of this invention improves accuracy by 6% compared to the benchmark algorithm MobileNetV3, reaching 84.6%, significantly improving the accuracy of diabetic retinopathy image classification and showcasing the effectiveness and superiority of the proposed method. Attached Figure Description

[0017] Figure 1 Example graph for the IDRID dataset.

[0018] Figure 2 This is an overview of hybrid deep learning models.

[0019] Figure 3 For comparison, a point-line diagram of the experiment is provided. Detailed Implementation

[0020] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings and preferred embodiments.

[0021] The purpose of this invention is to propose an innovative image classification method for diabetic retinopathy (DR) based on a hybrid deep learning model. First, it combines a multi-task loss function (MTLF) to effectively utilize multi-level information in the image, enhancing the model's ability to identify lesion features. Second, it integrates various data augmentation (DA) strategies to improve the model's generalization ability by increasing the diversity and complexity of the training data. Finally, it incorporates a normalization-based attention module (NAM) mechanism, enabling the model to focus on key lesion regions in the image, further improving classification accuracy. This invention can locate and identify five types of diabetic retinopathy, improving the efficiency and accuracy of DR lesion identification, and reducing the overall cost and time compared to manual identification.

[0022] Data Acquisition: Images were taken by technicians at Aravind Eye Hospital in India in rural areas with limited medical resources and were classified and labeled by experienced ophthalmologists. The dataset consists of training and testing sets. It contains pre-processed 224x224 pixel images, categorized into five levels from 0 to 4: Mild, Moderate, No DR, Proliferate DR, and Severe. It includes thousands of fundus images of diabetic retinopathy. An example image from the IDRiD dataset is shown below. Figure 1 As shown.

[0023] Model Implementation: A high-precision classification method for diabetic retinopathy based on a hybrid deep learning model. The overall diagram of the hybrid deep learning model is shown below. Figure 2 As shown, firstly, a multi-task loss function (MTLF) is designed to combine the losses of classification and segmentation tasks to jointly optimize MobileNetV3; secondly, a number augmentation strategy (DA) is combined with MobileNetV3 to improve the generalization ability of the algorithm; and thirdly, the NAM attention mechanism is combined to enhance the feature representation ability of MobileNetV3.

[0024] A multi-task learning framework is employed, combining diabetic retinopathy (DR) classification and lesion segmentation tasks. An MTLF (Mean Transmission-Lowering Failure) loss function is defined, which considers both classification and segmentation losses. The classification loss evaluates the model's accuracy in predicting DR grades, while the segmentation loss evaluates the model's performance in segmenting lesion regions.

[0025] Segmentation task input image ( The processed output is a predicted segmentation map. ( ) ,That , , and These are the length, width, number of channels, and segmentation type, respectively. To evaluate prediction accuracy, [the following parameters will be used]. Compared with the real segmentation mask For comparison, a multi-class cross-entropy loss function is used to quantify the differences between the two classifications; the smaller the loss, the better the segmentation performance. As shown in the formula: (1) In a multi-task learning framework, for scenarios with multiple outputs where each output performs the same multi-classification task, the same loss function is applied independently to each output. The final loss is defined by averaging all the losses, thereby comprehensively evaluating and optimizing the overall performance of the model, as shown in the formula: Based on the MobileNetV3 algorithm, MTLF is combined to effectively solve the problems of overfitting and insufficient feature representation that may be encountered during single-task training. This enables the model to more accurately identify lesion grades and more finely segment lesion areas in DR image classification tasks, providing more comprehensive and reliable auxiliary information for medical diagnosis.

[0026] To address issues such as image defects, out-of-focus images, underexposure, and overexposure in the dataset, a series of Data Analysis (DA) strategies are employed. These strategies include cropping black borders, contrast enhancement, random rotation, scaling, and flipping. DA increases the diversity of training samples, improves the model's generalization ability, and reduces the risk of overfitting. The DA strategy effectively alleviates the overfitting problem caused by the limited dataset, further improving classification accuracy and efficiency.

[0027] NAM, as an innovative attention mechanism, makes models more efficient while maintaining similar performance by suppressing less significant weights and features. It combines channel attention and spatial attention, using a scaling factor of Batch Normalization (BN) to measure the importance of channels and pixels, thereby achieving effective feature recognition and utilization.

[0028] The channel attention module of NAM uses the scaling factor of Batch Normalization (BN) to measure channel importance. By calculating the variance of the feature response of each channel (the standard deviation in BN), the magnitude of feature change and information richness of each channel are obtained. The scaling factor, as a trainable parameter, adjusts the feature weights, increasing the weight of important channels and decreasing the weight of insignificant channels. In the channel attention submodule, the module first passes the input features through a BN layer, and then uses the scaling factor of the BN layer to weight the feature map, effectively suppressing insignificant channels and improving model efficiency and performance.

[0029] and Small batches The mean and standard deviation, and For the trainable scale and bias, the BN scaling factor is as shown in the formula: The calculation process for the channel attention submodule is shown in the formula: in For output features, This is the channel scaling factor, and its weight is... As shown in the formula: The spatial attention submodule of NAM also utilizes the Batch Normalization (BN) scaling factor to measure pixel importance and performs pixel normalization. It assesses the importance of each pixel by calculating the variance of its feature response at each location. In the spatial attention submodule, the feature map is first normalized at the pixel level, and then weighted according to the normalization results, strengthening the weights of important pixel regions and weakening the weights of unimportant regions. In this way, NAM can further suppress insignificant pixel regions in the image, allowing the model to focus more on significant image features, thereby improving the model's accuracy and robustness.

[0030] The BN scaling factor is applied to the spatial dimension to measure pixel importance. The calculation process of the spatial attention submodule is shown in Equation 6: The output feature has the following weights: As shown in the formula: To suppress less significant weights, a regularization term is added to the loss function, as shown in the formula:

[0031] The NAM attention mechanism, by efficiently utilizing batch normalization scaling factors, improves model performance and aims to address the challenge of reducing computational costs while maintaining high accuracy in deep learning models. In MobileNetV3, NAM is integrated with the model's NAM attention mechanism. By evaluating the importance of channels and pixels in feature maps and applying intelligent weighting and suppression, lightweight networks like MobileNetV3 can more accurately capture fine-grained features of DR images, significantly improving classification performance and robustness while maintaining the advantages of high computational efficiency and fast training.

[0032] The high-precision classification method for diabetic retinopathy based on a hybrid depth model described in this invention was compared with several advanced models, and the experimental results are as follows: Figure 3 As shown, the effectiveness of the high-precision classification method for diabetic retinopathy based on a hybrid depth model described in this invention is verified.

[0033] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the present invention. Although detailed descriptions have been provided with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered within the protection scope of the claims.

Claims

1. A high-precision classification method for diabetic retinopathy based on a hybrid deep learning model, characterized in that... This includes the following steps: S1. Data Acquisition: Preprocessed 224x224 pixel images were acquired and divided into five levels from 0 to 4, namely Mild, Moderate, No DR, Proliferate DR, and Severe, containing thousands of diabetic retinopathy fundus images. S2. Model Implementation: A high-precision classification method for diabetic retinopathy using a hybrid deep learning model is built on the MobileNetV3 architecture. First, a multi-task loss function (MTLF) is designed to combine the losses of classification and segmentation tasks to jointly optimize MobileNetV3. Second, a number augmentation strategy (DA) is combined with MobileNetV3 to improve the generalization ability of the algorithm. Third, the NAM attention mechanism is combined to enhance the feature representation ability of MobileNetV3. S3. Model Application: The deep learning model built in step S2 is trained and validated using the dataset obtained in step S1. Finally, the deep learning model is used to automatically classify diabetic retinopathy on the test set.

2. The high-precision classification method for diabetic retinopathy based on a hybrid deep learning model according to claim 1, characterized in that... The MobileNetV3 architecture in step S2 is the third-generation MobileNet architecture introduced by Google's research team, designed to optimize deep learning computation on mobile and edge devices. This invention selects it as the basic model framework and conducts in-depth research on the DR image classification task.

3. The high-precision classification method for diabetic retinopathy based on a hybrid deep learning model according to claim 1, characterized in that... The processing procedure of the multi-task loss function module in step S2 is as follows: Define an MTLF loss function that considers both classification loss and segmentation loss. The segmentation task input image ( The processed output is a predicted segmentation map. ( ) ,That , , and These are the length, width, number of channels, and segmentation type, respectively. To evaluate prediction accuracy, [the following parameters will be used]. Compared with the real segmentation mask For comparison, a multi-class cross-entropy loss function is used to quantify the differences between the two classifications; the smaller the loss, the better the segmentation performance. As shown in the formula: In a multi-task learning framework, for scenarios with multiple outputs where each output performs the same multi-classification task, the same loss function is applied independently to each output. The final loss is defined by averaging all the losses, thereby comprehensively evaluating and optimizing the overall performance of the model, as shown in the formula:

4. The high-precision classification method for diabetic retinopathy based on a hybrid deep learning model according to claim 1, characterized in that... In step S2, the DA strategy is combined with MobileNetV3 as follows: by cropping the black border, enhancing contrast, random rotation, scaling, flipping, etc., the diversity of training samples is increased, the generalization ability of the model is improved, and the risk of overfitting is reduced.

5. The high-precision classification method for diabetic retinopathy based on a hybrid deep learning model according to claim 1, characterized in that... In step S2, the NAM attention mechanism is used to calculate the variance of the feature response of each channel (the standard deviation in BN) to obtain the magnitude of the feature change and the information richness of each channel. In the channel attention submodule, the module first passes the input features through the BN layer, and then uses the scaling factor of the BN layer to weight the feature map. and Small batches The mean and standard deviation, and For the trainable scale and bias, the BN scaling factor is as shown in the formula:

6. The calculation process for the channel attention submodule is shown in the formula: For output features, This is the channel scaling factor, and its weight is... As shown in the formula: The BN scaling factor is applied to the spatial dimension to measure pixel importance. The calculation process of the spatial attention submodule is shown in the formula:

7. The output feature has the following weights: As shown in the formula: To suppress less significant weights, a regularization term is added to the loss function, as shown in the formula:

8. In MobileNetV3, the NAM attention mechanism is combined. NAM can evaluate the importance of channels and pixels in the feature map and perform intelligent weighting and suppression, enabling lightweight networks such as MobileNetV3 to capture fine-grained features of DR images more accurately.

9. A high-precision classification method for diabetic retinopathy based on a hybrid deep learning model, characterized in that... The system first combines a multi-task loss function to effectively utilize multi-level information in the image, enhancing the model's ability to identify lesion features. Second, it integrates various data augmentation strategies to improve the model's generalization ability by increasing the diversity and complexity of the training data. Finally, it combines the NAM attention mechanism to enable the model to focus on key lesion areas in the image, further improving the accuracy of classification.