Diabetic retinopathy grading method based on vision-language pre-training model and ranking perception prompt strategy
By applying the RankCLIP method based on the CLIP model in diabetic retinopathy rating, the ranking perceptual prompt and similarity matrix smoothing module are used to solve the problems of inconsistent grading results and lack of robustness in the prior art, and achieve higher accuracy and robustness.
Patent Information
- Application Number
- CN202510231929.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems with inconsistent diagnostic results and lack of robustness in the automated grading of diabetic retinopathy, especially in the case of low treatment levels and insufficient correct treatment rates in primary areas.
A diabetic retinopathy grading method based on the visual-language pre-training model CLIP and ranking-aware cue strategy is proposed. By converting DR grading problems into image-text matching tasks, and introducing ranking-aware cue and similarity matrix smoothing modules, the natural ordinal information of DR rating is fully explored.
It improves the accuracy and robustness of diabetic retinopathy grading, can better deal with data imbalance and long-tail problems in category, and provides more reliable diagnostic support.
Smart Images

Figure QLYQS_1 
Figure BDA0005291730700000031 
Figure HDA0005291730710000011
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and medical image processing, and particularly relates to a diabetic retinopathy grading method (RankCLIP) based on a vision-language pre-training model (CLIP) and a ranking-aware prompting strategy. Background Art
[0002] Diabetic Retinopathy (DR) is one of the ocular complications caused by diabetes and is one of the main causes of blindness in adults. With the increase in the number of diabetic patients, the incidence of DR is also continuously rising. The diagnosis of DR mainly relies on the analysis of fundus images, especially through the changes in retinal blood vessels, bleeding points, and the structural changes of the retina. Traditional DR diagnosis methods rely on doctors' visual experience, but due to the differences between different doctors and the subjectivity in the diagnosis process, the diagnostic results may be inconsistent. Therefore, automated fundus image analysis methods based on deep learning have become a current research hotspot.
[0003] Fundus image-assisted diagnosis has become the mainstream method for ophthalmic disease screening due to its non-invasive, low-risk, and easy-to-operate characteristics. Fundus images capture information such as the retina, optic nerve, and blood vessels through imaging technology, which can help doctors judge the ocular health status and early identify common eye diseases such as diabetic retinopathy and glaucoma. Especially in the early diagnosis of chronic eye diseases such as Diabetic Retinopathy (DR), fundus images have become an indispensable important tool. Automated fundus image analysis technology, especially image segmentation and classification methods based on deep learning, is gradually replacing manual diagnosis and becoming one of the core technologies in ophthalmic medicine.
[0004] According to statistics, there are currently only 44,800 ophthalmologists in the country. On average, there is only 1.6 ophthalmologists for every 50,000 people, and the gap of ophthalmologists is huge. At the same time, as a complex eye disease, the treatment level of fundus diseases in grass-roots areas is far from that in first-tier cities. In such a background, it is difficult for patients to receive timely and correct treatment. According to the previous research data of IQVIA, currently, more than 3 million new fundus disease patients enter the hospital in China every year, but the correct treatment rate is less than 10%. The phenomenon that patients miss the best treatment opportunity and cause the aggravation of eye diseases occurs frequently. In this situation, computer-aided detection and computer-aided diagnosis have become the needs of the times. Summary of the Invention
[0005] The object of the present invention is to provide a diabetic retinopathy grading method using text knowledge guidance and ranking-aware prompting strategies (hereinafter referred to as the "RankCLIP method"). By transforming the DR grading problem into an image-text matching task and introducing a ranking-aware prompt and a similarity matrix smoothing module on this basis, the natural ordinal information of the DR grades is fully exploited to improve the accuracy and robustness of grading.
[0006] To achieve the above object, the present invention adopts the following technical solutions. The flow chart is as Figure 1 including the following steps:
[0007] S1. Preprocess the fundus color photographs, including denoising, standardization, and data augmentation;
[0008] S2. Divide the preprocessed data set into a training set and a test set according to a ratio of 7:3;
[0009] S3. A new network model: RankCLIP is proposed. It combines the visual and language contrast learning techniques of the CLIP model and is improved by introducing a similarity matrix smoothing module and a ranking loss to enhance the order between categories for the diabetic retinopathy grading task.
[0010] S4. Load the divided test set data set into the optimal model for testing, obtain the final grading results, and compare them with the corresponding labels in the public data set.
[0011] Further, S1 generally includes the following four steps. The data changes are as Figure 2 shown:
[0012] The first step: For each original fundus image, first calculate the central position of the image. Then, cut out a region from the center of the image according to a preset fixed size (e.g., (512, 512)). This step helps to focus on the key information of the fundus image, especially the retinal region.
[0013] The second step: Apply a Gaussian filter to smooth the image, reduce high-frequency noise, and at the same time maintain the structural features of the image. This method is suitable for removing minor random noise.
[0014] The third step: Perform contrast-limited adaptive histogram equalization (CLAHE) on the image data to enhance the details of low-contrast images and remove noise;
[0015] The fourth step: To increase the diversity of the data set and help improve the generalization ability of the model, perform data augmentation on the original image, including: mirror inversion, central rotation, horizontal flipping, and randomly adjusting saturation, contrast, and brightness;
[0016] Step S2 is further summarized as including the following step:
[0017] The first step: First, list all images and their corresponding label paths, and assign a unique label to each pair of images and labels, with the label range from 1 to 10. Then, distribute the data into the test set and the training set according to the labels, where the data with labels from 1 to 7 is used for the test set, and the data with labels from 8 to 10 is used for the training set. In this way, not only is the randomness of the data ensured, but also the proportional balance between the training set and the test set is guaranteed. Finally, store the data of the test set and the training set into two Excel files respectively.
[0018] Step S3 is further summarized as including the following 3 steps:
[0019] The first step: In the training stage of the proposed RankCLIP model, first preprocess and augment the data of fundus color images, crop the images into the same size of 512×512, and then send the processed fundus color images into the RankCLIP model for training. After training, special optimal weights for diabetic retinopathy grading detection will be generated.
[0020] The second step: Use the pre-trained CLIP model to extract high-dimensional features from the preprocessed images through the image encoder, and at the same time construct a DR grade text prompt containing shared context tokens and convert it into a text embedding through the text encoder. Among them, the learnable context tokens are kept consistent among all DR grades to provide a unified task context, while different DR grades are determined by the corresponding keywords.
[0021] The third step: Construct a similarity matrix and a ranking-aware prompt, calculate the inner product between the image features and the text embedding to obtain the similarity matrix; design a ranking loss to ensure that for any image, the similarity of its correct DR grade is higher than that of the adjacent grades, so that the model can better capture the ordinal relationship between DR grades; at the same time, introduce the SMS module to smooth and correct the similarity matrix to reduce the impact of data imbalance.
[0022] Step S4 is further summarized as including the following 1 step:
[0023] The first step: Import the training set sorted in step 2 into the model for test verification, and finally obtain the optimal diabetic retinopathy grading detection result.
[0024] Experimental parameters:
[0025] In the diabetic retinopathy grading experiment, the hardware environment is: NVIDIA GeForce RTX2080Ti Graphics card with 11G of video memory; operating system: Windows 11; Pytorch deep learning framework. The Adam optimizer is used to update the parameters. Among them, the initial learning rate (lr) is set to (1e-4)*3, the betas parameter is set to (0.9, 0.999), the batch size is set to 14, and the total number of training epochs is set to 500. During training, the cosine annealing decay strategy for the learning rate is adopted, and the learning rate at the t-th epoch can be expressed as:
[0026] In the prediction stage, the performance of the model is evaluated by the area under the ROC curve (AUC) and the macro F1-score. Brief Description of the Drawings
[0027] To clearly illustrate the technical solution of the present invention, the drawings required for the description of the existing technology will be introduced below. Those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0028] Figure 1 is a hierarchical flowchart.
[0029] Figure 2 is a flowchart for dataset production.
[0030] Figure 3 is the result graph of image enhancement in the embodiment of the present invention.
[0031] Figure 4 is an overview of the RankCLIP model proposed by the present invention for training and inference. Detailed Embodiments
[0032] The following is a clear and complete description of the technical solution of the present invention in combination with the drawings in the embodiments of the present invention. The described embodiments are part of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.
[0033] A method for grading diabetic retinopathy based on a vision-language pre-training model and a ranking-aware prompting strategy is as Figure 1 shown and includes the following steps:
[0034] I. Dataset Preparation Stage
[0035] As Figure 2 shown
[0036] (1), Image acquisition
[0037] : First, collect fundus color images through a clinical fundus camera. These images should include different grades of diabetic retinopathy (DR), including but not limited to "no lesion", "mild lesion", "moderate lesion", "severe lesion", and "proliferative lesion", etc. All images should be accompanied by accurate lesion grade labels for subsequent training and testing.
[0038] (2), Image preprocessing
[0039] The collected fundus color images need to be preprocessed for subsequent feature extraction and model training. The preprocessing steps include: 1. Image size adjustment: Since the collected images have different resolutions, the image size needs to be uniformly adjusted to a fixed size, such as 512×512 or 256×256, to ensure that the network can process images from different sources consistently. 2. Denoising: To reduce interference, apply an image denoising algorithm (such as median filtering or Gaussian filtering) to remove background noise and improve the clarity of the fundus lesion area. 3. Image enhancement: To improve the robustness of the model, enhance the images. Enhance the image contrast through methods such as histogram equalization (CLAHE) to make the distinction between the lesion area and the normal area higher.
[0040] II. Image and text encoding
[0041] After the data preprocessing is completed, the image and text encoding operations are carried out next, which are specifically divided into the following steps:
[0042] (1), Image encoding
[0043] Use a pre-trained CLIP image encoder (such as Transformer) to convert the fundus images into high-dimensional feature vectors. These feature vectors can effectively capture the visual information in the images, including blood vessel morphology, lesion areas, etc., and become the basis for subsequent tasks.
[0044] (2), Text encoding
[0045] For each DR grade (such as "no lesion", "mild lesion", "moderate lesion", etc.), use the corresponding natural language description (such as "the image shows no DR"). Convert these text descriptions into text embeddings through the CLIP text encoder. Each text embedding represents the semantic information corresponding to the DR grading.
[0046] (3), Image-text matching
[0047] Calculate the similarity between the image features and the text embeddings through the inner product to obtain the matching degree between the image and the text. This matching degree will be used as the basis for subsequent training, and the model is trained by optimizing the similarity between the image and the text.
[0048] III. Ranking-Aware Prompting and Feature Matching
[0049] To improve the performance of the model, the present invention introduces Ranking-aware Prompting, whose main purpose is to strengthen the model's learning of the natural order between different DR levels. The specific implementation steps are as follows:
[0050] (1) Ranking Loss
[0051] For the matching degree between the image and the DR level text, we design a ranking loss function L rank . This loss function requires the model to learn the following relationship: for the same image, the correct DR level text should have a higher similarity. For example, the similarity between the image and the "mild lesion" label should be higher than that with the "no lesion" label, thus ensuring the order relationship between different lesion levels. In this way, the model not only learns the grading information of the image but also can understand the relationship between these gradings.
[0052] (2) Similarity Matrix Generation and Smoothing
[0053] After calculating the similarity between the image features and the text embeddings, a similarity matrix is generated, where each row represents the matching degree of an image with each DR level text. To mitigate the impact of data imbalance and class tail, the present invention introduces a Similarity Matrix Smoothing (SMS) module to smooth and correct the similarity matrix, enabling the model to better handle rare classes or classes with less data volume.
[0054] IV. Training the Model
[0055] After data preprocessing and feature extraction are completed, the image and text embeddings will be fed into the RankCLIP model for training. The training steps are as follows:
[0056] (1) Network Construction
[0057] The RankCLIP method proposed by the present invention is an end-to-end deep neural network based on the CLIP model and the ranking-aware prompting strategy, and the overall structure is as Figure 4As shown. During the construction of the network, combining the powerful functions of the vision-language model and the characteristics of the DR grading task, the model design includes image feature extraction, text embedding, similarity calculation, and ranking-aware prompting module. Each module cooperates with each other to form an efficient and accurate grading system. Specifically, the network consists of the following main parts: 1. Image feature extraction module: This module uses the pre-trained ResNet50 as the backbone network and is responsible for extracting high-level visual features from the input fundus images. To improve the model's expressive ability, the network expands the dimension of the input image and performs feature mapping through a layer of convolution. This process converts high-dimensional sparse data into low-dimensional dense data, thereby improving the model's ability to capture image features in the latent space. 2. Text feature extraction module: The text encoder of the Transformer architecture is used to encode the text labels related to each DR grading (such as "no lesion", "mild lesion", etc.) and convert them into text embeddings. This embedding is generated by the pre-trained model of CLIP and can provide rich semantic representations for each grading label. 3. Similarity calculation module: To achieve the matching between images and texts, the inner product is used to calculate the similarity between image features and text embeddings. The calculated similarity matrix is used for subsequent training and guides the model to learn the correspondence between images and texts. 4. Ranking-aware prompting module: This module introduces a ranking loss to ensure that the similarity between an image and its correct grading text is higher than that with other non-adjacent grades. By explicitly strengthening the ordinal relationship between adjacent grades, the ranking-aware prompting helps the model better understand the sequentiality between each grade. 5. Similarity matrix smoothing (SMS) module: To reduce the negative impact of long-tailed distribution or class imbalance on model training, the present invention introduces similarity matrix smoothing technology. This module performs smoothing correction on the similarity matrix to ensure that the model does not overly bias towards common classes when dealing with rare classes.
[0058] In the training process of the model, the cross-entropy loss function is used to optimize the image classification task. At the same time, the ranking loss part ensures the sequential relationship between classes. The optimizer adopts the Adam optimization algorithm, combined with dynamic learning rate adjustment and early stopping strategy, to ensure the efficient convergence of the model and avoid overfitting.
[0059] The present invention uses a comprehensive loss function for optimization, combining the main loss (KL divergence) and the ranking loss. The two work together to enable the model to not only accurately classify but also maintain the sequential relationship between classes.
[0060] During the training process, the parameters are adjusted through multiple iterations to finally achieve the optimal effect. The evaluation metrics include F1 score, AUC, and accuracy, etc., to ensure the classification accuracy and robustness of the model at different DR grades.
[0061] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate understanding of the present invention by those skilled in the art, and it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
Claims
1. A diabetic retinopathy grading method based on a visual-linguistic pre-training model (CLIP) and a ranking-aware prompting strategy, characterized in that: include: S11: Preprocessing of fundus color photography, including denoising, standardization and data enhancement; S12: Divide the preprocessed data set into a training set and a test set in a ratio of 7:3; S13: A new network model, RankClip, is proposed. It combines the visual and language contrast learning techniques of the CLIP model and improves it by introducing a similarity matrix smoothing module and ranking loss to enhance the order between categories for the diabetic retinopathy grading task. S14: Load the divided test set data set into the optimal model for testing, obtain the final classification results and compare them with the corresponding labels in the public data set.
2. The method for grading diabetic retinopathy based on visual-linguistic pre-training model (CLIP) and ranking-aware prompting strategy (hereinafter referred to as "RankCLIP method") according to claim 1, characterized in that: The S1 includes: S21: For each original fundus image, first calculate the center position of the image. Then cut out a region from the center of the image according to a preset fixed size (e.g., (512, 512)). This step helps focus on the key information of the fundus image, especially the retinal area. S22: Apply a Gaussian filter to smooth the image, reduce high-frequency noise, and maintain the structural characteristics of the image. This method is suitable for removing slight random noise. S23: performing contrast-limited adaptive histogram equalization (CLAHE) processing on the image data to enhance low-contrast image details and remove noise; S24: In order to increase the diversity of the dataset and help improve the generalization ability of the model, data enhancement is performed on the original images, including: mirror inversion, center rotation, horizontal flipping, and random adjustment of saturation, contrast, and brightness.
3. The method for grading diabetic retinopathy based on a visual-linguistic pre-training model (CLIP) and a ranking-aware prompting strategy according to claim 1, characterized in that: The S2 includes: S31: Network Construction RankCLIP combines image encoding and text encoding, and uses contrastive learning techniques to align images with category labels to improve classification accuracy. The method has made structural improvements, especially in image feature extraction and classification performance optimization. To enhance the performance of the image encoder, RankCLIP introduces a multi-scale feature fusion module. By fusing features of different scales, the model can extract more hierarchical visual information. Low-resolution and high-resolution features are fused through skip connections to retain more detail information. The text encoder uses a pre-trained Transformer model to process text data and convert category labels into vector representations. These text embedding vectors can be compared in the same latent space as the image feature vectors to establish the association between images and categories. The cosine similarity or inner product between the image feature vectors and the text feature vectors is then calculated to evaluate their matching. Images and their corresponding category labels should have a high similarity in the latent space, while unrelated images and category labels should have a low similarity. Contrastive loss is one of the key components in the RankCLIP framework, which is responsible for optimizing the similarity between images and texts, making images and their corresponding category labels as close as possible, while unrelated category labels are far away. The InfoNCE loss function is used to optimize the similarity. As shown in formula (3.1): After the contrast loss is optimized, the SMS module generates a smoothed similarity matrix by smoothing the original similarity matrix to ensure that the similarity between categories is more stable. The ranking loss is introduced to optimize the order relationship so that the model can learn the relative positions between different categories and ensure that the order relationship is preserved. Finally, RankCLIP can calculate the similarity between the image and the text of each category and output a category label to indicate the severity of diabetic retinopathy corresponding to the image (such as mild, moderate or severe) to perform grading prediction of diabetic retinopathy. S32: Network model data training According to the network built by S31, the preprocessed and data augmented data sets were selected for training on the RankCLIP network. The contrast loss and ranking loss were used to guide the network learning and obtain the diabetic retinopathy grading results.
4. The diabetic retinopathy grading method based on CLIP framework and contrastive learning according to claim 1, characterized in that: Improvements to the RankCLIP network include: S41: In the traditional CLIP framework, contrastive learning between images and text is usually achieved by calculating the similarity matrix, while RankCLIP optimizes this process through an improved similarity matrix smoothing module (SMS), which reduces the impact of data imbalance and enhances the order relationship between categories. This improvement enables the model to effectively learn the corresponding features when dealing with a small number of sample categories; S42: Ranking loss is introduced to optimize the order relationship between categories. In the diabetic retinopathy grading task, there is a clear order relationship between categories (such as "mild" and "moderate"). RankCLIP optimizes the similarity difference between adjacent categories through ranking loss, thereby further improving the accuracy of model grading.
Citation Information
Cited By
Congenital heart disease perioperative period risk early warning method and system based on fundus color photo
CN120895240A
Hip multi-disease automatic classification method and system
CN121278474A
Hip multi-disease automatic classification method and system
CN121278474B
Eye fundus image few-sample classification method and system based on context learning
CN121482446A