Diabetic retina image classification method and system based on deep learning

By combining the deep learning method of the EfficientNet network, the feature pyramid network and the lesion attention module, the problems of insufficient feature extraction and model generalization capabilities in DR automatic diagnosis technology are solved, and high-accuracy identification of early lesions and real-time diagnosis of large-scale screening are achieved, which improves the robustness and computational efficiency of the model.

CN120726409AInactive Publication Date: 2025-09-30ZHEJIANG NORMAL UNIV

Patent Information

Application Number
CN202511232563.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning-based automatic diagnosis technology for diabetic retinopathy (DR) has limited feature extraction capabilities, making it difficult to identify subtle lesions, and insufficient model generalization capabilities, resulting in low classification accuracy, especially in early or mild lesions, and limited applicability to different devices, populations, and imaging conditions.

Method used

The EfficientNet network is used as the backbone network, combined with the feature pyramid network and the lesion attention module. Through data enhancement and mixed precision training, the feature extraction capability is improved. The robustness of the model is enhanced by the lesion attention module. The data enhancement strategy is used to reduce the dependence on specific data sources, and the model training efficiency is improved by combining mixed precision training.

Benefits of technology

It significantly improves the ability to identify early lesions such as microaneurysms in retinal images, improves classification accuracy and model generalization ability, meets the real-time diagnosis needs of large-scale screening, and reduces computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726409A_ABST
    Figure CN120726409A_ABST
Patent Text Reader

Abstract

The invention discloses a diabetes retina image classification method and system based on deep learning, and relates to the technical field of image classification, and the specific steps are as follows: collecting and preprocessing a retina image data set, and constructing a training data set; constructing an initial image classification model based on a deep neural network, and performing iterative training on the initial image classification model by using the training data set to obtain an image classification model; the initial image classification model takes a pre-trained OfficientNet network as a backbone network, and integrates a feature pyramid network structure, a lesion attention module and a classification module; and obtaining a to-be-classified image, preprocessing the to-be-classified image, inputting the preprocessed to-be-classified image into the image classification model, and outputting a classification result. According to the method, the dependence of the model on a specific data source is reduced by using a data enhancement strategy, the diagnostic performance and stability of the model on unseen retina images are improved, and the universality of clinical application is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image classification technology, and more particularly to a deep learning-based diabetic retinal image classification method and system. Background Art

[0002] Currently, diabetic retinopathy (DR) is a major chronic complication of diabetes, and its early diagnosis is crucial for preventing irreversible vision loss. However, due to the hidden symptoms of early lesions, patients find it difficult to detect them themselves, often leading to delayed treatment. The current standard diagnostic method relies heavily on professional ophthalmologists, who perform manual assessments through analysis of fundus color photographs and invasive procedures such as fluorescein angiography. This method is not only time-consuming and labor-intensive, but also requires strict professional experience from physicians, making it difficult to support the needs of large-scale population screening. At the same time, invasive examinations such as fluorescein angiography have potential risks, limiting their use as routine screening methods. With advances in medical imaging technology and the development of artificial intelligence, especially the breakthroughs of deep learning (such as convolutional neural networks (CNN)) in the field of medical image recognition, image-based automatic DR diagnosis technology has become a research hotspot, providing a new technical path to overcome the aforementioned bottlenecks in manual diagnosis.

[0003] However, existing deep learning-based automatic DR diagnosis technologies still face significant challenges and shortcomings: (1) Limited feature extraction capabilities: Traditional methods or some existing models are unable to fully capture the subtle and critical lesion features (such as microaneurysms, exudates, neovascularization, etc.) in retinal lesions, resulting in low classification accuracy, especially insufficient recognition of early or mild lesions; (2) Insufficient model generalization capabilities: Model training relies on specific sources or a limited number of data sets. When faced with "new" fundus images collected by different devices, different populations, and different imaging conditions, the diagnostic performance often decreases significantly, and the clinical applicability is limited. Therefore, how to effectively improve the deep learning model's feature extraction capabilities for subtle DR lesions and enhance the model's generalization robustness on diverse and unseen data is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0004] In view of this, the present invention provides a diabetic retinal image classification method and system based on deep learning, which overcomes the above-mentioned defects.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A deep learning-based diabetic retinal image classification method, the specific steps are as follows:

[0007] Collect and preprocess retinal image datasets to construct training datasets;

[0008] An initial image classification model is constructed based on a deep neural network, and the initial image classification model is iteratively trained using the training data set to obtain an image classification model; the initial image classification model uses a pre-trained EfficientNet network as a backbone network and integrates a feature pyramid network, a lesion attention module, and a classification module;

[0009] Obtain an image to be classified, pre-process the image to be classified, input it into the image classification model, and output a classification result.

[0010] Optionally, the preprocessing includes:

[0011] performing contrast enhancement processing on the retinal images in the retinal image dataset using a contrast-limited adaptive histogram equalization technique to generate a contrast-enhanced image;

[0012] Performing a data enhancement operation on the contrast-enhanced image to obtain a data-enhanced image; the data enhancement operation includes: random affine transformation, random flipping, random color adjustment, random occlusion, and noise injection;

[0013] The data augmented image is subjected to pixel value normalization and standardization processing to generate a preprocessed image.

[0014] Optionally, the specific steps of the standardization process are:

[0015] Generating a three-channel image from the data-augmented image after pixel value normalization according to a channel processing rule;

[0016] Converting the three-channel image into a grayscale image;

[0017] Based on a predetermined central reference point, cropping and filling operations are performed on the grayscale image to generate a preprocessed image.

[0018] Optionally, the channel processing rule is:

[0019] If the data-augmented image is a single-channel image, copying the single-channel image to generate a three-channel image;

[0020] If the data-augmented image is a four-channel image, retain the data of the first three channels to generate a three-channel image;

[0021] If the data-enhanced image is a three-channel image, the data-enhanced image is directly output.

[0022] Optionally, the data processing steps of the image classification model are:

[0023] Extracting a multi-scale feature map of the preprocessed image based on the EfficientNet network and the feature pyramid network;

[0024] Performing cross-scale fusion on the multi-scale feature map to generate a multi-scale fusion feature, and inputting the multi-scale fusion feature into the lesion attention module to generate a lesion area attention map;

[0025] Performing feature enhancement on the multi-scale fusion feature using the lesion area attention map to obtain a lesion enhancement feature map;

[0026] The classification module performs category recognition on the lesion enhancement feature map and outputs a classification result.

[0027] Optionally, the lesion attention module includes: a first convolutional layer, a ReLU activation layer, a second convolutional layer, a Sigmoid activation layer and a weighted fusion layer.

[0028] Optionally, the calculation expression of the lesion enhancement feature map is:

[0029] ;

[0030] in, ;

[0031] Where, It is a multi-scale fusion feature; is the attention map of the lesion area; is element-wise multiplication; is the Sigmoid activation function; for convolution; for Activation function; for convolution; is the input feature.

[0032] Optionally, the classification module includes a first fully connected layer, a batch normalization layer, a ReLU activation layer, a Dropout layer and a second fully connected layer connected in sequence.

[0033] Optionally, the training method of the initial image classification model is:

[0034] A mixed precision training system is adopted. The calculation of the initial image classification model is converted to half precision through the autocast context manager, and the loss value is gradient scaled using GradScaler. The kappa coefficient of the initial image classification model is obtained in real time. The optimal parameters are saved based on the kappa coefficient comparison trigger mechanism to obtain the image classification model.

[0035] A deep learning-based diabetic retinal image classification system, comprising:

[0036] A training set construction module is used to collect and preprocess retinal image datasets to construct training datasets;

[0037] A model training module is used to construct an initial image classification model based on a deep neural network and iteratively train the initial image classification model using the training data set to obtain an image classification model; the initial image classification model uses the pre-trained EfficientNet network as the backbone network and integrates a feature pyramid network, a lesion attention module, and a classification module;

[0038] An image processing module, used for collecting the image to be classified and performing preprocessing on it;

[0039] The diagnosis module is used to input the pre-processed image to be classified into the image classification model and output the classification result.

[0040] From the above technical solutions, it can be seen that the present invention provides a diabetic retinal image classification method and system based on deep learning, which has the following beneficial effects compared with the existing technology:

[0041] Significantly improved ability to identify subtle lesions: Through an enhanced feature extraction mechanism, the ability to characterize early and subtle lesion features such as microaneurysms, small hemorrhages, and hard exudates in retinal images has been significantly improved, thereby improving the detection rate and classification accuracy of early and mild DR.

[0042] Enhanced model generalization and robustness: Data augmentation strategies are used to reduce the model's dependence on specific data sources, improve the model's diagnostic performance and stability on unseen retinal images, and enhance the universality of clinical applications.

[0043] Improved computing efficiency: The mixed-precision training system and dynamic model preservation strategy improve the efficiency of model training and reduce the demand for computing resources, which can meet the urgent needs of large-scale screening and clinical real-time or near-real-time diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0045] Figure 1 A schematic flow chart of the method provided by the present invention;

[0046] Figure 2A schematic diagram of the image preprocessing process provided by the present invention;

[0047] Figure 3 Schematic diagram of the EfficientNet feature extraction and multi-scale feature fusion process provided by the present invention;

[0048] Figure 4 Schematic diagram of the attention mechanism processing flow provided by the present invention. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0050] The embodiment of the present invention discloses a method for classifying diabetic retinal images based on deep learning, the steps are as follows: Figure 1 As shown, specifically:

[0051] Step 1: Collect and preprocess the retinal image dataset to construct a training dataset;

[0052] Step 2: Build an initial image classification model based on a deep neural network and iteratively train the initial image classification model using the training dataset to obtain an image classification model. The initial image classification model uses the pre-trained EfficientNet network as the backbone network and integrates the feature pyramid network, lesion attention module, and classification module.

[0053] Step 3: Obtain the image to be classified, pre-process the image to be classified, input it into the image classification model, and output the classification result.

[0054] Furthermore, in this embodiment, first, a retinal image dataset is collected. The retinal image dataset is sourced from public datasets or medical institutions and covers retinal images of five categories: normal, mild, moderate, severe, and proliferative. The retinal images are then preprocessed, including CLAHE preprocessing to enhance contrast, data augmentation operations to expand the retinal image dataset, and normalization processing to unify the retinal image scale. The preprocessed retinal images are input into the EfficientNet-B5 model to extract features. A multi-scale feature map is generated using a feature pyramid network (FPN). The multi-scale feature map is fused and input into the lesion attention module. The output weighted features are finally input into the classification module (in this embodiment, the classification module uses a fully connected classifier), which outputs the classification results, thereby achieving multi-classification recognition of diabetic retinopathy.

[0055] In one embodiment, the specific steps of preprocessing are:

[0056] The contrast-limited adaptive histogram equalization technique is used to perform contrast enhancement on the retinal images in the retinal image dataset to generate contrast-enhanced images.

[0057] Perform data augmentation operations on the contrast-enhanced image to obtain a data-enhanced image; the data augmentation operations include: random affine transformation, random flipping, random color adjustment, random occlusion, and noise injection;

[0058] The data augmented image is normalized and standardized to generate a preprocessed image.

[0059] Furthermore, if Figure 2 As shown in the figure, the original fundus image (i.e., the original input retinal image) is first converted into color space (RGB→LAB), and then the contrast-limited adaptive histogram equalization technique (CLAHE, clipLimit=3.0) is applied to enhance the image contrast before converting it back to the RGB color space. Next, data augmentation operations such as random affine transformation, random flipping, random color adjustment, and random occlusion are performed. Finally, the enhanced retinal image is pixel-normalized to the range of [0, 1] and then standardized using the mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225] to generate a preprocessed image of size 512×512×3 as the input data of the image classification model.

[0060] Among them, random affine transformation includes random rotation, translation and scaling; random flipping includes random horizontal flipping and random vertical flipping; random color adjustment includes adjustment of brightness, contrast and saturation; random occlusion is to probabilistically add 3-5 elliptical occlusion areas; noise injection uses Gaussian noise injection, where the standard deviation of the noise σ∈[0.01,0.05].

[0061] In one embodiment, the normalization process includes:

[0062] The data augmented image after pixel value normalization is converted into a three-channel image according to the channel processing rules;

[0063] Convert the three-channel image to a grayscale image;

[0064] Based on a predetermined central reference point, the grayscale image is cropped and filled to generate a preprocessed image.

[0065] In one embodiment, the channel processing rules are:

[0066] If the data augmented image is a single-channel image, copy the single-channel image to generate a three-channel image;

[0067] If the data augmented image is a four-channel image, the data of the first three channels are retained to generate a three-channel image;

[0068] If the data augmented image is a three-channel image, the data augmented image is directly output.

[0069] Furthermore, the normalization processing includes: converting the input image into a grayscale image, and cropping and padding it with the retina as the center to ensure the normalized input of the image; wherein the channel processing rules include: for a single-channel grayscale image, repeating the channel to generate a three-channel image; for a four-channel RGBA image, retaining only the first three channels to obtain a three-channel image.

[0070] In one embodiment, the data processing steps of the image classification model are:

[0071] Extract multi-scale feature maps of preprocessed images based on EfficientNet network and feature pyramid network;

[0072] Perform cross-scale fusion on the multi-scale feature maps to generate multi-scale fusion features, which are then input into the lesion attention module to generate an attention map of the lesion area.

[0073] The multi-scale fusion features are enhanced using the lesion area attention map to obtain the lesion enhanced feature map;

[0074] The classification module performs category recognition on the lesion enhancement feature map and outputs the classification results.

[0075] In one embodiment, the feature pyramid network includes:

[0076] Feature alignment unit, used to perform feature mapping and feature alignment on multi-scale feature maps to generate pre-fused features;

[0077] The fusion unit concatenates the pre-fused features along the channel dimension to generate multi-scale fused features.

[0078] Furthermore, the image classification model includes a feature extraction module, a feature fusion module, and a classification module. The feature extraction module includes a multi-scale feature extraction network, which uses the pre-trained EfficientNet-B5 model combined with a feature pyramid network (FPN) to extract multi-scale features from the pre-processed image. The feature fusion module includes a lesion attention module, which learns a lesion area attention map to highlight lesion features and improve classification accuracy. The lesion attention module includes a convolutional layer for extracting feature maps, a sigmoid activation function for generating the lesion area attention map, and an element-wise multiplication operation for multiplying the lesion area attention map with the original image to be classified to enhance the features of the lesion area.

[0079] The input preprocessed image is first subjected to feature extraction by the EfficientNet-B5 model to generate the first channel feature map (low-level feature map) P1, the second channel feature map (mid-level feature map) P2, and the third channel feature map (high-level feature map) P3; these feature maps are fused after feature mapping and upsampling to generate multi-scale fusion features to better capture the multi-scale information of the lesion.

[0080] Furthermore, in the process of feature extraction and fusion, as Figure 3 As shown in the figure, the input image size is 512×512×3. After being processed by the EfficientNet-B5 model, low-level, intermediate and high-level feature maps are obtained from the 2nd, 3rd and 4th output layers. The features at each level are mapped by 2D convolution (1×1) + batch normalization + ReLU, and then the low-level feature map is upsampled by 4 times to obtain the upsampled low-level features; the intermediate feature map is upsampled by 2 times to obtain the upsampled intermediate features; the high-level feature map maintains the original size to obtain the high-level features, and the obtained low-level features, intermediate features and high-level features are spliced ​​and fused to generate multi-scale fusion features, which is expressed as follows:

[0081] ;

[0082] in, Splicing for channel dimension; is the first channel feature map; is the upsampling operation; is the second channel feature map; is the third channel feature map; , is the feature space; For batches; is the height of the feature map; is the width of the feature map.

[0083] In one embodiment, if Figure 4 As shown in the figure, the multi-scale fused features are first processed through a 3×3 convolution and a ReLU activation function, and then through a 1×1 convolution and a Sigmoid activation function to generate attention weights. The attention weights are element-wise multiplied with the original features (features of the input image) to generate weighted features that highlight the lesion area. Finally, the weighted features are input into a fully connected classifier for classification.

[0084] Among them, the lesion area attention map of the lesion attention module is generated by the following formula:

[0085] ;

[0086] Where, is the activation function; for convolution; for Activation function; for convolution; is the input feature.

[0087] And the enhanced feature map is calculated by the following formula:

[0088] ;

[0089] Where, It is a multi-scale fusion feature; is the attention map of the lesion area; This is element-wise multiplication.

[0090] In one embodiment, the weighted features are input into a fully connected classifier (i.e., a classification module), and the classification result is output through a fully connected layer FC(256×3→512)→batch normalization (BatchNorm)→ReLU activation function→random inactivation (Dropout)→fully connected layer FC(512→5).

[0091] In one embodiment, during the model training phase, a hybrid loss function is used to optimize the model by combining cross entropy loss and category weighting strategy. The formula is:

[0092] ;

[0093] Among them, the cross entropy loss is expressed as:

[0094] ;

[0095] Where, is the number of samples in a training batch; The actual retinal lesion diagnosis results annotated in the image dataset; is the retinal disease diagnosis result predicted by the model; the loss function also includes the category weighting strategy:

[0096] ;

[0097] in, For category The number of samples.

[0098] In one embodiment, a mixed precision training system is adopted, and the calculation of the initial image classification model is converted to half precision through the autocast context manager, and the loss value is gradient scaled using GradScaler; and the kappa coefficient of the initial image classification model is obtained in real time, and the optimal parameters are saved based on the kappa coefficient comparison trigger mechanism to obtain the image classification model.

[0099] Furthermore, the training process adopts a mixed-precision training system, including GradScaler gradient scaling and autocast half-precision context, to reduce video memory usage and increase training speed.

[0100] In one embodiment, a dynamic model saving strategy is adopted during the model training process, and the best model is saved based on the kappa coefficient comparison trigger mechanism. The file naming logic is a double identification of timestamp_index value (for example: 20230815_1530_kappa0.8923.pth).

[0101] In one embodiment, the classification module compresses the feature map into a fixed-size feature vector through adaptive average pooling, and then maps the feature vector into a classification space through a fully connected layer and a Softmax classifier to output a multi-classification result.

[0102] In one embodiment, a multi-index evaluation system is used to evaluate model performance:

[0103] A variety of evaluation indicators, including accuracy, precision, recall, F1 value and Cohen's Kappa coefficient, are used to comprehensively evaluate the performance of the model.

[0104] In one embodiment, an automated report generation mechanism is also included:

[0105] Report generation: Generates detailed classification reports and confusion matrices based on the model's diagnostic results, providing clinicians with comprehensive diagnostic information.

[0106] On the other hand, this embodiment also discloses a deep learning-based diabetic retinal image classification system, including:

[0107] A training set construction module is used to collect and preprocess retinal image datasets to construct training datasets;

[0108] The model training module is used to build an initial image classification model based on a deep neural network and iteratively train the initial image classification model using a training dataset to obtain an image classification model. The initial image classification model uses the pre-trained EfficientNet network as the backbone network and integrates a feature pyramid network, a lesion attention module, and a classification module.

[0109] Image processing module, used to collect images to be classified and perform preprocessing on them;

[0110] The diagnosis module is used to input the preprocessed image to be classified into the image classification model and output the classification result.

[0111] In one embodiment, multi-threaded data loading is adopted in the system implementation, and the number of data loading threads is 4.

[0112] In one embodiment, the system implementation includes a model interpretation module for generating a heat map, where the heat map is generated by a Grad-CAM++ algorithm, and the visualization rule includes an α-blending coefficient, where 0.35≤α≤0.65.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0114] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A deep learning-based diabetic retinal image classification method, characterized in that: The specific steps are: Collect and preprocess retinal image datasets to construct training datasets; An initial image classification model is constructed based on a deep neural network, and the initial image classification model is iteratively trained using the training data set to obtain an image classification model; the initial image classification model uses the pre-trained EfficientNet network as the backbone network, and is constructed by integrating a feature pyramid network, a lesion attention module, and a classification module; the data processing steps of the image classification model are as follows: a multi-scale feature map of the pre-processed image is extracted based on the EfficientNet network and the feature pyramid network; the multi-scale feature map is subjected to feature mapping by 2D convolution, batch normalization, and ReLU in sequence to obtain multiple mapping features; the multiple mapping features are up-sampled to obtain multiple up-sampled features; the multiple up-sampled features are cross-scale fused to generate multi-scale fused features, and the multi-scale fused features are input into the lesion attention module to generate a lesion area attention map; the multi-scale fused features are feature enhanced using the lesion area attention map to obtain a lesion enhancement feature map; the classification module performs category recognition on the lesion enhancement feature map and outputs a classification result; Obtain an image to be classified, pre-process the image to be classified, input it into the image classification model, and output a classification result.

2. The deep learning-based diabetic retinal image classification method according to claim 1, characterized in that: The pretreatment includes: performing contrast enhancement processing on the retinal images in the retinal image dataset using a contrast-limited adaptive histogram equalization technique to generate a contrast-enhanced image; Performing a data enhancement operation on the contrast-enhanced image to obtain a data-enhanced image; the data enhancement operation includes: random affine transformation, random flipping, random color adjustment, random occlusion, and noise injection; The data augmented image is subjected to pixel value normalization and standardization processing to generate a preprocessed image.

3. The deep learning-based diabetic retinal image classification method according to claim 2, characterized in that: The specific steps of standardization are: Generating a three-channel image from the data-augmented image after pixel value normalization according to a channel processing rule; Converting the three-channel image into a grayscale image; Based on a predetermined central reference point, cropping and filling operations are performed on the grayscale image to generate a preprocessed image.

4. The deep learning-based diabetic retinal image classification method according to claim 3, characterized in that: The channel processing rules are: If the data-augmented image is a single-channel image, copying the single-channel image to generate a three-channel image; If the data-augmented image is a four-channel image, retain the data of the first three channels to generate a three-channel image; If the data-enhanced image is a three-channel image, the data-enhanced image is directly output.

5. The deep learning-based diabetic retinal image classification method according to claim 1, characterized in that: The lesion attention module includes: a first convolutional layer, a ReLU activation layer, a second convolutional layer, a Sigmoid activation layer and a weighted fusion layer.

6. The deep learning-based diabetic retinal image classification method according to claim 1, characterized in that: The calculation expression of the lesion enhancement feature map is: ; in, ; Where, It is a multi-scale fusion feature; is the attention map of the lesion area; is element-wise multiplication; is the Sigmoid activation function; for convolution; for Activation function; for convolution; is the input feature.

7. The deep learning-based diabetic retinal image classification method according to claim 1, characterized in that: The classification module includes a first fully connected layer, a batch normalization layer, Activation layer, Dropout layer and second fully connected layer.

8. The deep learning-based diabetic retinal image classification method according to claim 1, characterized in that: The training method of the initial image classification model is: A mixed precision training system is used to convert the calculation of the initial image classification model to half precision through the autocast context manager, and the loss value is gradient scaled using GradScaler; The kappa coefficient of the initial image classification model is obtained in real time, and the optimal parameters are saved based on the kappa coefficient comparison trigger mechanism to obtain the image classification model.

9. A deep learning-based diabetic retinal image classification system, characterized in that: include: A training set construction module is used to collect and preprocess retinal image datasets to construct training datasets; A model training module is used to construct an initial image classification model based on a deep neural network, and iteratively train the initial image classification model using the training data set to obtain an image classification model; the initial image classification model uses the pre-trained EfficientNet network as the backbone network, and integrates a feature pyramid network, a lesion attention module and a classification module to construct; the data processing steps of the image classification model are as follows: extracting a multi-scale feature map of the pre-processed image based on the EfficientNet network and the feature pyramid network; performing feature mapping on the multi-scale feature map through 2D convolution, batch normalization and ReLU in sequence to obtain multiple mapping features; performing upsampling processing on the multiple mapping features to obtain multiple up-sampled features; performing cross-scale fusion on the multiple up-sampled features to generate multi-scale fusion features, and inputting the multi-scale fusion features into the lesion attention module to generate a lesion area attention map; using the lesion area attention map to perform feature enhancement on the multi-scale fusion features to obtain a lesion enhancement feature map; the classification module performs category recognition on the lesion enhancement feature map and outputs a classification result; Image processing module, used to collect images to be classified and perform preprocessing on them; The diagnosis module is used to input the pre-processed image to be classified into the image classification model and output the classification result.

Citation Information

Patent Citations

  • Attention mechanism-based in-depth learning diabetic retinopathy classification method

    CN108021916A

  • Grading device for diabetic retinopathy

    CN115937194A

  • Medical image classification method based on feature fusion and attention mechanism

    CN116543197A

  • Lightweight classification model and classification method for fundus multi-type lesions

    CN117237693A

  • Diabetic retinopathy grading model system and method using frequency domain attention mechanism

    CN119724541A

Cited By

  • Method and system for predicting harvesting loss of corn ear harvester based on deep learning

    CN121074689A

  • A Deep Learning-Based Method and System for Predicting Harvest Loss in Corn Ear Harvesters

    CN121074689B

  • Intelligent periodontal disease classification method, device, medium, equipment and product

    CN121211348A

  • Age prediction method and system based on facial fine granularity

    CN121686535A