Self-adaptive multi-scale model for fine-grained grading of diabetic retinopathy
Through the adaptive multi-scale model, combined with multi-scale feature extraction and adaptive attention enhancement, the problem of difficult to identify multi-scale and multi-type lesion characteristics in the prior art is solved, and a higher-precision grading of diabetic retinopathy is achieved.
Patent Information
- Application Number
- CN202510319235.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing CNN-based diabetic retinopathy (DR) grading method is difficult to take into account subtle differences and complex lesion associations when facing multi-scale and multi-type lesion characteristics, resulting in insufficient diagnostic accuracy, especially in the identification and classification of adjacent-level lesions.
Adaptive multi-scale model is adopted, including a hierarchical global context module, a multi-scale adaptive attention module and a relational multi-head attention module. Combined with a dynamic weighted fusion strategy, multi-scale feature extraction and adaptive attention enhancement, the complex dependence between different spatial locations and features is captured, and the model's ability to identify fine-grained lesion features is enhanced.
It significantly improved the performance of the model in the grading task of diabetic retinopathy, improved the identification and classification accuracy of multi-scale lesions, enhanced the robustness and generalization ability of the model, and could more accurately analyze the diverse characteristics of DR lesions.
Smart Images

Figure CN120340099A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and specifically discloses an adaptive multi-scale model for fine-grained grading of diabetic retinopathy. Background Art
[0002] Diabetic retinopathy (DR) is one of the common complications among diabetic patients and is also one of the main causes of blindness. With the continuous increase in the global prevalence of diabetes, the incidence of DR has also shown a significant growth trend. Diabetic retinopathy occurs due to retinal vascular diseases and abnormal retinal blood flow. DR is usually characterized by four different levels of the disease, including mild, moderate, severe, and proliferative diabetic retinopathy.
[0003] Traditional DR diagnosis mainly relies on ophthalmologists to evaluate through fundus photos. This method is not only time-consuming and laborious but also overly dependent on the clinical experience accumulated by ophthalmologists over the years; the quality of retinal images is affected by aspects such as light, equipment, and operation techniques; there are some tiny features in the images that are difficult to distinguish with the naked eye, etc. Especially in resource-scarce areas, the phenomena of diagnostic delays and misdiagnoses are more common. To improve the efficiency and accuracy of diagnosis, researchers have been committed to developing automatic grading methods based on image processing and pattern recognition. Before the popularization of deep learning, the detection and grading of diabetic retinopathy (DR) mainly relied on traditional machine learning. Roychowdhury S et al. proposed a computer-aided screening system at that time. This system analyzes fundus images with different illuminations and fields of view and uses a set of different classifiers, such as Gaussian mixture model (GMM), k-nearest neighbor (KNN), support vector machine (SVM), and AdaBoost classifiers, to classify retinal lesions and non-lesions. However, these methods rely on manual feature extraction, such as color, texture, and shape features, and can handle small-scale datasets well, but when faced with large-scale datasets, their generalization ability is significantly insufficient, and their performance is severely limited. In recent years, deep learning has developed rapidly in the field of medical imaging, especially the application of convolutional neural network (CNN) has made remarkable progress. CNN has been widely concerned and has been experimentally proven to have a strong driving force for the development of automatic classification of diabetic retinopathy (DR).
[0004] With the development of deep learning, single-branch convolutional neural networks, such as ResNet and Inception, have been widely used in DR automatic grading. These networks usually improve the grading accuracy by pre-training on large datasets such as ImageNet and fine-tuning on the DR dataset. For example, Pan Y et al. used Inception V3 and ResNet-50 deep learning models to classify fundus images into three categories: normal, macular degeneration, and tessellated fundus, in order to identify and treat fundus diseases in a timely manner, providing a reference for the clinical diagnosis or screening of diabetic retinopathy and other eye diseases. However, they used a relatively balanced dataset collected by themselves, but their network structure mainly relied on a single receptive field or fixed scale, which was difficult to capture multi-scale and multi-type lesion information simultaneously in actual clinical practice. Although it has good adaptability to small-scale and balanced data scenarios, when faced with a more complex real environment, the model may not be able to fully identify or distinguish large lesions from extremely subtle lesions, resulting in insufficient recognition ability for severe or border lesions, thus affecting the overall diagnostic performance. In addition, most of these studies ignored the intricate spatial correlations between different levels of DR lesions: lesions at adjacent levels often have obvious overlaps or continuities in visual features, and the boundaries of the lesion areas are not always clearly distinguishable, which greatly increases the difficulty of accurately distinguishing different levels of lesions. To address this issue, Wang X et al. proposed a joint learning multi-level grading method based on convolutional neural network (CNN), which successfully alleviated the decline in grading performance under low-resolution conditions. However, this method did not deeply explore the extremely subtle lesion differences between adjacent levels, especially lacking refined modeling means for levels with highly similar visual features and blurred lesion ranges, resulting in a significant risk of confusion in multi-level lesion scenarios. Although the joint learning strategy can share deep features to a certain extent, thereby improving the overall generalization ability of the model, when the morphology and distribution of lesions show continuous or gradual changes between adjacent levels, the model still seems unable to handle these details. Due to the lack of targeted processing for similar lesion areas and the full exploration of multi-scale features, the model is ultimately difficult to accommodate all possible lesion morphologies in actual applications, resulting in insufficient discrimination and impaired discrimination accuracy. Especially in the context of clinical diagnosis, this misclassification of lesions at approximate levels is extremely likely to lead to deviations in subsequent treatment and intervention strategies, highlighting the importance of deeply modeling the spatial correlations and subtle visual differences of lesions in multi-level lesion recognition tasks. He A et al. proposed a new Class Attention Block (CAB) for the unbalanced DR data distribution, and constructed CABNet to perform DR grading by deeply mining the differential regional features of each DR level and achieving "equal treatment" for different classes in the network.In terms of enhancing the attention to the minority class, this method does alleviate the model bias caused by data imbalance to a certain extent. However, CABNet is still insufficient in capturing the transitions and correlations between levels, making it difficult to fully reflect the gradual changes from mild to severe lesions. At the same time, the network structure with a single receptive field or fixed scale also shows obvious limitations when facing multi-scale and multi-type lesion features, and it is difficult to take into account the subtle differences of different lesions. Due to the lack of refined modeling of complex or gradual lesion morphologies, the discrimination difficulty at the boundaries of this method remains high, and the lesion features at adjacent levels are prone to confusion, which further leads to non-negligible classification errors in actual grading.
[0005] Although existing CNN-based DR grading methods have improved the diagnostic efficiency and accuracy to a certain extent, they still face many challenges when dealing with the highly diverse and interrelated lesion features of diabetic retinopathy (DR). DR lesions often appear simultaneously in different scales and forms in the images, such as microaneurysms, hemorrhages, exudates, etc. These lesions may have blurred edges or dense distributions, making it difficult for a network structure with a single receptive field or fixed scale to take into account all the lesion information. Moreover, the spatial correlations between different levels of DR lesions are extremely complex: the lesions at adjacent levels often have a certain degree of overlap or continuity in visual features, and the boundaries of the lesion areas are not necessarily clear, thus increasing the risk of grading confusion. Taking grade 1 DR as an example, its early manifestation is only subtle microaneurysms or mild hemorrhages. If the image quality is poor or the resolution is insufficient, it is very easy to be confused with normal fundus (grade 0). The number and complexity of lesions in grade 2 DR are significantly increased, but there is a certain visual similarity with the small lesions in grade 1 DR, further leading to misjudgments between adjacent classes. The course of DR development essentially has a progressive and continuous feature, but the clinical diagnosis uses a relatively discrete grading standard, ignoring the transitions and correlations between levels, making it more difficult for the model to make judgments at the boundaries. Therefore, to further improve the accuracy and stability of DR grading in aspects such as multi-scale lesion detection, complex lesion correlation modeling, and disease continuity characterization, more flexible and refined network structures and learning strategies are still needed. Therefore, in view of this, the inventors provide an adaptive multi-scale model for fine-grained grading of diabetic retinopathy to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide an adaptive multi-scale model for fine-grained grading of diabetic retinopathy to comprehensively and delicately analyze the diverse features of DR lesions.
[0007] To achieve the above purpose, the basic solution of the present invention provides an adaptive multi-scale model for fine-grained grading of diabetic retinopathy, including:
[0008] Hierarchical Global Context Module: Integrating multi-scale feature extraction, adaptive attention enhancement, and dynamic weighted fusion to construct a multi-scale representation mechanism that complements global and local information, generating more expressive feature representations;
[0009] Multi-scale Adaptive Attention Module: Capturing complex dependencies between different spatial positions in the feature map through the synergistic effect of multi-scale feature extraction and adaptive attention enhancement;
[0010] Relational Multi-head Attention Module: Deeply mining and capturing complex associations between different features in the feature map to enhance the model's ability to identify fine-grained lesion features and classification performance;
[0011] Attention Mechanism Module: Combining channel and spatial attention mechanisms to capture key channels and important spatial regions in the feature map.
[0012] Furthermore, in the hierarchical global context module, for a given feature map, the number of channels of the input feature map is mapped to a lower dimension through convolution;
[0013] The feature map mapped to a lower dimension is downsampled through adaptive average pooling of different sizes to obtain corresponding pooled features;
[0014] The pooled features are subjected to convolution, normalization, and PreLU activation processing;
[0015] The processed pooled features are sampled to the size of the original feature map;
[0016] An attention mechanism is introduced into the upsampled feature map, and an attention map is generated through convolution and activation functions;
[0017] Feature enhancement is achieved by element-wise multiplication of the upsampled feature map and the attention map, and weighted by learnable weights;
[0018] The enhanced features of all scales are concatenated in the channel dimension, and the number of channels is compressed and fused through convolution to output the final multi-scale feature representation.
[0019] Furthermore, in the hierarchical global context module, the expression for mapping the number of channels of the input feature map to a lower dimension is as follows:
[0020]
[0021] where X down is the feature map mapped to a lower dimension;
[0022] Conv2d 3×3 represents a 3×3 convolution operation;
[0023] BN is batch normalization;
[0024] φ is the PreLU activation function;
[0025] H and W are the height and width of the feature map respectively;
[0026] The expression for obtaining the pooled feature is as follows:
[0027]
[0028] where F i is the pooled feature;
[0029] The expressions for performing convolution, normalization, and PreLU activation on the pooled feature are as follows:
[0030] F′ i = φ(BN(Conv2d 1×1 (F i )))
[0031] where F' i is the processed pooled feature;
[0032] The expression for sampling the processed pooled feature to the size of the original feature map is as follows:
[0033]
[0034] The expression for generating the attention map is as follows:
[0035]
[0036] where σ is the Sigmoid function;
[0037] The weights for weighting are normalized by the Softmax function:
[0038]
[0039] where p i is the learned parameter vector;
[0040] ⊙ represents element-wise multiplication;
[0041] The final multi-scale feature representation is as follows:
[0042]
[0043] where X HGCM is the final multi-scale feature.
[0044] Furthermore, in the multi-scale adaptive attention module, the feature map is divided into multiple equal groups to obtain multiple sub-feature maps;
[0045] Convolution operations are performed on each sub-feature map through a group of dynamically generated convolutional kernels to obtain each convolutional output feature map;
[0046] Each convolutional output feature map is enhanced through a channel attention mechanism, and the channel importance of the convolutional output feature map is adjusted through channel weights to obtain multi-scale feature maps;
[0047] The multi-scale feature maps are stacked in a new dimension through a dynamic weight generation mechanism;
[0048] The dynamically generated weights are multiplied element-wise with the stacked feature maps to perform weighted enhancement on each scale feature;
[0049] The weighted feature maps are reorganized back to the original number of channels.
[0050] Furthermore, in the multi-scale adaptive attention module, the expression for dividing the feature map into multiple equal groups is as follows:
[0051]
[0052] where S is the number of equal groups, X is the feature map to be divided, is the number of channels of each sub-feature map, C is the number of input channels, and H and W are the height and width of the feature map respectively;
[0053] The dynamic convolutional kernels for performing convolution operations on each sub-feature map consist of multiple convolutional layers, and the size of its convolutional kernel varies with the number of groups S. The expression is as follows:
[0054]
[0055] where Conv2d represents the standard two-dimensional convolution operation;
[0056] The groups parameter implements grouped convolution;
[0057] The channel attention mechanism adopts a squeeze-and-excitation module to generate channel weights through adaptive pooling and fully connected layers. The expression is as follows:
[0058]
[0059] where φ represents the ReLU activation function;
[0060] are the weights of the fully connected layer;
[0061] σ is the Sigmoid activation function;
[0062] The expression for adjusting the channel importance of the convolutional output feature map through channel weights is as follows:
[0063]
[0064] Among them, ⊙ represents element-wise multiplication;
[0065] The expression for stacking multi-scale feature maps on a new dimension is as follows:
[0066]
[0067] The expression for the dynamic weight generation mechanism is as follows:
[0068] W = Softmax(Conv2d(AdaptiveAvgPool2d(X), kernel size = 1))
[0069] The expression for weighted enhancement of each scale feature is as follows:
[0070] conv_outs s,c ' ,h,w = conv_outs s,c ' ,h,w ⊙W s
[0071] The expression for generating the final output feature map is as follows:
[0072]
[0073] Among them, X MSAM is the finally generated output feature map.
[0074] Furthermore, in the relational multi-head attention module, the feature map is mapped to the k×classes dimensional space through convolution, and the expression is as follows:
[0075]
[0076] Among them, k is the feature dimension of each class;
[0077] classes is the number of classes;
[0078] H and W are the height and width of the feature map respectively;
[0079] Batch normalization and ReLU activation are performed on the convolution output, and the expression is as follows:
[0080]
[0081] Among them, BN represents the batch normalization operation, and ReLU is the activation function;
[0082] The feature map is weighted through the channel attention mechanism, and the expression is as follows:
[0083]
[0084] F1' = F1 ⊙ ChannelAttentionMap
[0085] where σ is the Sigmoid activation function;
[0086] ⊙ represents element-wise multiplication;
[0087] The multi-head attention mechanism is introduced into the relational multi-head attention module, and the channel dimension k×classesk is flattened and transposed to adapt to the multi-head attention mechanism. The expression is as follows:
[0088]
[0089] where H and W are the height and width of the feature map respectively;
[0090] The core formula of the multi-head attention mechanism is:
[0091]
[0092] where Q, K, and V are the query, key, and value vectors respectively;
[0093] d k is the dimension of the key vector;
[0094] The output of the multi-head attention is obtained by concatenating the attention results of multiple heads and fusing them through a linear transformation. The expression is as follows:
[0095]
[0096] Generate an attention map through a 1×1 convolution and apply Softmax normalization to obtain weights. The expression is as follows:
[0097]
[0098] where Conv2d 1×1 is used to generate the class attention map, and the Softmax operation is normalized in the class dimension;
[0099] Rearrange the feature map F1' into the class dimension and expand the class attention map to match the dimension. The expression is as follows:
[0100]
[0101] Apply the attention map to weight the feature map. The expression is as follows:
[0102]
[0103] Average in the k dimension, and the expression is as follows:
[0104]
[0105] Fuse the feature map in the category dimension back to the original number of channels through 1×1 convolution, and perform element-wise multiplication with the input feature map to achieve semantic weighting. The expression is as follows:
[0106]
[0107] Among them, Semantic is the result of semantic weighting.
[0108] Furthermore, the adaptive multi-scale model adopts a dynamic weight adjustment mechanism based on performance feedback, and its comprehensive loss function is defined as:
[0109] L = α(t)·L cls + β(t)·L reg + L r
[0110] Among them, α(t) and β(t) are weight coefficients that are dynamically adjusted during the training process;
[0111] L cls is the cross-entropy loss function;
[0112] L reg is the mean square error loss function.
[0113] Furthermore, α(t) and β(t) are adaptively adjusted based on the real-time performance of the model in classification accuracy and regression performance. When the relative improvement in classification accuracy is greater, increase the weight α(t) of the classification loss; when the relative improvement in regression performance is greater, increase the weight β(t) of the regression loss.
[0114] Furthermore, in the attention mechanism module, channel shuffling operations are introduced to promote cross-group information exchange, thereby enhancing the diversity of feature representations and the generalization ability of the model.
[0115] Based on the same inventive concept, the present invention discloses a method for processing retinal images for diabetic retinopathy, including using the above-mentioned adaptive multi-scale model to perform fine-grained grading processing on retinal images.
[0116] The principle and effect of this solution are as follows:
[0117] 1. The hierarchical global context module combines multi-scale context aggregation, attention guidance mechanism and dynamic feature fusion strategy. It can not only efficiently capture global semantic information through multi-scale pooling, but also retain subtle local structural features in the spatial dimension, ensuring the balance between large-scale context and precise local response. By introducing an attention weighting mechanism, the hierarchical global context module further endows the network with the ability to adaptively perceive key lesion areas, highlighting significant features while suppressing redundant information. In addition, the dynamic multi-scale feature fusion strategy enables collaborative optimization of features at different scales in the channel dimension, significantly enhancing the robustness and diversity of feature representation. The hierarchical global context module provides an efficient and flexible solution for fine-grained lesion detection, greatly improving the performance of the network in the diabetic retinopathy grading task.
[0118] 2. The multi-scale adaptive attention module realizes the adaptive enhancement of key features and the effective suppression of non-key features by introducing dynamic convolution kernels and an advanced channel attention mechanism. The design of this module not only significantly improves the spatial sensitivity of features, enabling the model to capture local lesion details more meticulously, but also enhances the robustness and generalization ability of the model through diverse feature expressions. The multi-scale adaptive attention module ensures efficient feature extraction and expression in various complex lesion scenarios by dynamically adjusting feature weights, thus significantly improving the overall performance and accuracy of the model in the fine-grained lesion recognition task. This innovative design provides an efficient and flexible feature processing method for medical image analysis, significantly promoting the development of DR fine-grained grading technology.
[0119] 3. The relational multi-head attention module ensures that the model can understand and analyze lesion features in DR images from multiple perspectives and scales by processing feature relationships in different subspaces in parallel. The RMA module aims to deeply mine and capture the complex associations between different features in the feature map, thereby significantly enhancing the model's ability to identify fine-grained lesion features and classification performance.
[0120] 4. Using the cross-entropy loss function can effectively measure the difference between the predicted probability distribution and the true label distribution, and the mean squared error loss function is used to measure the gap between the predicted value and the true value. Weighted fusion of the cross-entropy loss function and the mean squared error loss function can significantly improve the model performance. The model can adaptively balance the requirements of the two tasks in the same optimization process, ensuring that the regression task can provide accurate continuous score predictions, while ensuring that these predictions can accurately map to the corresponding categories. This design not only improves the prediction accuracy of the model, but also enhances the model's ability to understand the internal order relationship of the data. The core advantage of multi-task learning lies in the feature sharing and information complementarity between tasks, and the dynamic weighting mechanism ensures that this sharing and complementarity can be optimally balanced during training.
[0121] 5. The Shuffle Attention mechanism module effectively captures important channels and key spatial regions in the feature map by combining channel attention and spatial attention mechanisms. It introduces a channel shuffle operation to promote cross-group information exchange, thereby enhancing the diversity of feature representation and the generalization ability of the model. This module has a simple structure and high computational efficiency, is applicable to the feature extraction stage of various deep learning models, and demonstrates a significant performance improvement in complex medical image analysis tasks, achieving a better balance between computational efficiency and performance improvement, making it show significant performance advantages in the DR grading task.
[0122] 6. Through the collaborative work of the above multiple modules in this embodiment, it can comprehensively and delicately analyze the diverse features of DR lesions, significantly enhancing the generalization ability and robustness of the model. Compared with traditional single-branch attention models and other mainstream algorithms, this embodiment has achieved a significant improvement in the grading performance on the public dataset, further verifying the effectiveness and superiority of this method in the DR fine-grained grading task. BRIEF DESCRIPTION OF THE DRAWINGS
[0123] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0124] Figure 1 shows a schematic structural diagram of an adaptive multi-scale model for fine-grained grading of diabetic retinopathy proposed in the embodiments of the present application;
[0125] Figure 2 shows a schematic diagram of a hierarchical global context module in an adaptive multi-scale model for fine-grained grading of diabetic retinopathy proposed in the embodiments of the present application;
[0126] Figure 3 shows a schematic diagram of a multi-scale adaptive attention module in an adaptive multi-scale model for fine-grained grading of diabetic retinopathy proposed in the embodiments of the present application;
[0127] Figure 4 shows a schematic diagram of a relational multi-head attention module in an adaptive multi-scale model for fine-grained grading of diabetic retinopathy proposed in the embodiments of the present application;
[0128] Figure 5The figure shows a schematic diagram of the Shuffle Attention mechanism module in an adaptive multi-scale model for fine-grained grading of diabetic retinopathy proposed in an embodiment of the present application. Detailed implementation manner
[0129] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manner, structure, features and their effects of the present invention as follows.
[0130] An adaptive multi-scale model for fine-grained grading of diabetic retinopathy - MAFNet (Multi-scale Adaptive Fine-grained Network), as shown in the embodiments Figure 1 is shown, where Figure 1The description of the character parameters involved is as follows: Shuffle Attention: Shuffle attention module; HGCM: Hierarchical Global Context Module; MSAM: Multi-Scale Adaptive Attention Module; RMA: Relational Multihead Attention; BaseNet: Basic network; Concat: Concatenation; GAP: Global Average Pooling; CE loss: Cross-entropy loss function; MSE loss: Mean Squared Error loss function; Mapping: Mapping. To address the problem that DR lesions may have pathological features of different sizes (such as microaneurysms, hemorrhages, exudates, etc.) simultaneously in images, it is difficult for traditional methods with a single receptive field to simultaneously focus on lesion features at different scales. In this embodiment, a Hierarchical Global Context Module (HGCM) is proposed, which can capture global semantic information from tiny lesions to large-scale lesions simultaneously. Due to the complex spatial correlation between DR lesions, the distribution location and combination pattern of different lesions have an important impact on the judgment of disease severity. In this embodiment, a Multi-Scale Adaptive Attention Module (MSAM) is proposed, which can accurately capture the dependence relationship between lesions at different spatial positions. To further enhance the model's ability to understand and model the deep relationship between lesion features, a Relational Multihead Attention module (RMA) is proposed in this embodiment. RMA parallelly captures and fuses the complex relationships between features in multiple feature subspaces through the multihead attention mechanism, and is specifically optimized for DR lesions at the feature channel and high-order interaction levels. Through the synergistic effect of the Hierarchical Global Context Module (HGCM), the Multi-Scale Adaptive Attention Module (MSAM), and the Relational Multihead Attention module (RMA), MAFNet can comprehensively and meticulously analyze diverse lesion features in DR images and achieve more accurate and reliable disease grading. The development of DR is a progressive and continuous process, but the clinical grading is a discrete category division, ignoring the continuity between adjacent categories. By transforming the problem into a dual-task learning with a regression task as the core and a classification task as an assistant, the continuity of disease progression is maintained while ensuring the accuracy of the classification result.
[0131] The ConvNeXt backbone network serves as the basis for feature extraction. It adopts multi-level depthwise separable convolutions and layer normalization (LayerNorm) modules to gradually extract high-level features of the input retinal images. To further optimize the feature representation, the model introduces a Shuffle Attention module at specific layers to enhance the features.
[0132] Among them, the Hierarchical Global Context Module (HGCM) is an innovative multi-scale feature extraction component proposed in this embodiment, aiming to break through the limitations of traditional convolutional neural networks in capturing fine-grained pathological features under a single receptive field. As Figure 2 shown, HGCM combines multi-scale context aggregation, attention guidance mechanism and dynamic feature fusion strategy to systematically solve the problem of insufficient information expression of lesion regions at different scales, where Figure 2 the descriptions of each character parameter involved are as follows: HGCM: Hierarchical Global Context Module; F: input feature map; F’: output feature map; Conv: convolution; Pool: pooling; BN: batch normalization; PReLU: activation function; Upsampled: upsampling; Sigmoid: activation function; Concat: concatenation. Compared with existing methods, HGCM can not only efficiently capture global semantic information through multi-scale pooling, but also retain fine local structural features in the spatial dimension, ensuring the balance between large-scale context and precise local response. By introducing an attention weighting mechanism, HGCM further endows the network with the ability to adaptively perceive key lesion regions, highlighting significant features while suppressing redundant information. In addition, the dynamic multi-scale feature fusion strategy enables features at different scales to achieve collaborative optimization in the channel dimension, significantly enhancing the robustness and diversity of feature expression. The HGCM module provides an efficient and flexible solution for fine-grained lesion detection, greatly improving the performance of the network in the task of diabetic retinopathy grading.
[0133] In HGCM, for a given where C in is the number of input channels, and H and W are the height and width of the feature map respectively. To reduce the computational overhead and adapt to subsequent operations, HGCM maps the number of channels C in of the input feature map to a lower dimension C down through a 3×3 convolution:
[0134]
[0135] where Conv2d 3×3 represents a 3×3 convolution operation, BN is batch normalization, and φ is the PreLU activation function.
[0136] In the multi-scale pooling stage, dimensionality reduction operations are performed on X down through adaptive average pooling AdaptiveAvgPool2d with different sizes. Specifically, for the i-th pooling scale s i ∈{1,2,3,4}, the pooled feature F i is expressed as:
[0137]
[0138] To further refine the pooled features, perform 1×1 convolution, batch normalization, and PreLU activation on F i :
[0139] F' i = φ(BN(Conv2d 1×1 (F i )))
[0140] These pooled features F' i are upsampled to the size of the original feature map for subsequent fusion, and the expression is as follows:
[0141]
[0142] On the upsampled feature map, an attention mechanism is introduced to generate an attention map through a 3×3 convolution and a Sigmoid activation function:
[0143]
[0144] where A i is the generated attention map, and σ is the Sigmoid function, which is used to map the values of the attention map to the range [0,1].
[0145] Feature enhancement is achieved by element-wise multiplication of the feature map obtained in the above steps and the generated attention map, and a learnable weight w i is introduced for weighting, and the weight is normalized by the Softmax function:
[0146]
[0147] where p i is the learned parameter vector, and ⊙ represents element-wise multiplication.
[0148] Finally, the enhanced features of all scales are concatenated in the channel dimension and fused through 1×1 convolution to compress the number of channels and output the final multi-scale feature representation:
[0149]
[0150] The Hierarchical Global Context Module (HGCM) effectively constructs a multi-scale representation mechanism that complements global and local information through the fusion of multi-scale feature extraction, adaptive attention enhancement, and dynamic weighted fusion. Adaptive average pooling at different scales expands the receptive field of the model, ensuring the collaborative capture of global context and local details; the attention mechanism highlights key regions through saliency weighting, suppresses redundant information, and enhances the semantic integrity and spatial sensitivity of feature expressions. In addition, the dynamic weight mechanism adaptively balances the importance of information at each scale, enabling the model to intelligently adjust the contributions of multi-scale features according to the input feature distribution. Finally, PPM efficiently integrates multi-scale information through channel fusion, generates more expressive feature representations, provides accurate and robust feature support for the fine-grained grading task of diabetic retinopathy (DR), and effectively breaks through the limitations of single-scale methods.
[0151] The Multi-Scale Adaptive Attention Module (MSAM) aims to accurately capture the complex dependencies between different spatial positions in the feature map. By introducing dynamic convolutional kernels and an advanced channel attention mechanism, it achieves the adaptive enhancement of key features and the effective suppression of non-key features. The design of this module not only significantly improves the spatial sensitivity of features, enabling the model to capture local lesion details more precisely, but also enhances the robustness and generalization ability of the model through diverse feature expressions. The MSAM module ensures efficient feature extraction and expression in various complex lesion scenarios by dynamically adjusting feature weights, thus significantly improving the overall performance and accuracy of the model in the fine-grained lesion recognition task. This innovative design provides an efficient and flexible feature processing method for medical image analysis, significantly promoting the development of DR fine-grained grading technology.
[0152] As Figure 3 shown, Figure 3 the descriptions of each character parameter in Where C is the number of input channels, and H and W are the height and width of the feature map respectively. The MSAM module first divides the input feature map into S equal groups to achieve multi-scale and multi-angle feature extraction. Specifically, the feature map X is rearranged into S sub-feature maps, and the number of channels of each sub-feature map is
[0153]
[0154] Each sub-feature map X i (i = 1, 2, …, S) is convolved through a group of dynamically generated convolutional kernels to capture feature representations at different spatial scales. The dynamic convolutional kernels consist of multiple convolutional layers, and the size of the convolutional kernels varies with the number of groups S:
[0155]
[0156] where Conv2d represents the standard two-dimensional convolution operation, and the groups parameter implements grouped convolution to enhance the diversity and fine-grained expression of features.
[0157] By dividing the feature map and using dynamic convolutional kernels of different sizes, the MSAM module extracts rich feature information at multiple spatial scales. This design effectively expands the receptive field of the model, enabling it to simultaneously focus on the global context and local details.
[0158] After the dynamic convolution operation, each convolutional output feature map is further enhanced through a channel attention mechanism. The channel attention mechanism adopts a Squeeze-and-Excitation (SE) module to generate channel weights through adaptive pooling and fully connected layers:
[0159]
[0160] where φ represents the ReLU activation function, are the weights of the fully connected layer, σ is the Sigmoid activation function, ensuring that the weights are in the range of [0, 1]. The channel weights are used to adjust the channel importance of the feature map :
[0161]
[0162] where ⊙ represents element-wise multiplication. By introducing the SE module and generating channel weights through adaptive pooling and fully connected layers, the PSA module can adaptively adjust the importance of each channel. This mechanism not only improves the semantic expression ability of the feature map but also enhances the sensitivity of the model to key lesion regions.
[0163] Stacking multi-scale feature maps in a new dimension makes it easier to apply corresponding dynamic weights to each scale feature later:
[0164]
[0165] In order to further optimize the fusion of features at different scales, a dynamic weight generation mechanism is introduced:
[0166] W=Softmax(Conv2d(AdaptiveAvgPool2d(X),kernel size=1))
[0167] The dynamically generated weight W is multiplied element by element with the stacked feature map to achieve weighted enhancement of each scale feature:
[0168] conv_outs s,c ' ,h,w =conv_outs s,c ' ,h,w ⊙W s
[0169] Finally, the weighted feature map is reorganized back to the original number of channels C through .view(C,H,W). Channel compression and information integration are performed through 1x1 convolution and batch normalization to generate the final output feature map X MSAM :
[0170]
[0171] The Multiscale Adaptive Attention Module (MSAM module) significantly improves the performance of the model in the fine-grained classification task of diabetic retinopathy through the synergy of multi-scale feature extraction and adaptive attention enhancement. The MSAM module not only enhances the model's perception of complex lesion structures, but also improves the flexibility and robustness of feature expression, providing an efficient and innovative feature processing solution for the field of medical image analysis. The MSAM module solves the limitations of the traditional attention mechanism in processing multi-scale lesion features by dynamically adjusting feature weights and flexibly integrating multi-scale features, significantly improving the adaptability and accuracy of the model in complex lesion scenarios.
[0172] The Relational Multihead Attention (RMA) module processes feature relationships in parallel in different subspaces, ensuring that the model can understand and parse lesion features in DR images from multiple angles and scales. The RMA module is designed to deeply explore and capture the complex associations between different features in the feature map, thereby significantly enhancing the model's recognition and classification performance of fine-grained lesion features.
[0173] like Figure 4As shown Figure 4 The descriptions of each character parameter in Figure 4 are as follows: RMA: Relational Multi-Head Attention; ReLU: Activation function; Conv: Convolution; BN: Batch Normalization; Softmax: Softmax function; Key: Key; Query: Query; Value: Value. In the RMA module, given the input feature map First, it is mapped to the k×classes dimensional space through a 1×1 convolution to extract class-related features:
[0174]
[0175] where k is the feature dimension of each class and classes is the number of classes. Batch Normalization and ReLU activation are performed on the convolution output C:
[0176]
[0177] where BN represents the batch normalization operation and ReLU is the activation function.
[0178] The feature map is weighted through the channel attention mechanism to enhance the responses of important channels:
[0179]
[0180] F1' = F1 ⊙ ChannelAttentionMap
[0181] where σ is the Sigmoid activation function and ⊙ represents element-wise multiplication. The channel attention module generates channel weights through adaptive pooling and convolution operations, and then adjusts the response intensity of each channel. Through the channel attention mechanism, RMA can adaptively adjust the importance of different channels, strengthen the responses of key feature channels, and suppress the interference of unimportant channels, thereby improving the model's ability to identify fine-grained lesion features.
[0182] To model the complex relationships between different features in the feature map, the RMA module introduces the multi-head attention mechanism. The channel dimension k×classes is flattened and transposed to adapt to the multi-head attention mechanism:
[0183]
[0184] where H and W are the height and width of the feature map respectively. The core formula of the multi-head attention mechanism is:
[0185]
[0186] where Q, K, and V are the Query, Key, and Value vectors respectively, and dk is the dimension of the key vector. The attention weights are calculated by scaling the dot product and the value vectors are weighted and summed to obtain the attention output. The output of multi-head attention is obtained by concatenating the attention results of multiple heads and fusing them through a linear transformation:
[0187]
[0188] The multi-head attention mechanism allows the model to capture diverse feature relationships in parallel in different subspaces, enhancing the ability to understand and represent complex lesion features. This is particularly important for the DR grading task because there are subtle differences in lesion features and semantic associations between different DR grades.
[0189] Generate an attention map through a 1×1 convolution and apply Softmax normalization to obtain the weights:
[0190]
[0191] where Conv2d 1×1 is used to generate the class attention map, and the Softmax operation normalizes in the class dimension to ensure that the sum of the weights of each pixel point in the class dimension is 1.
[0192] Rearrange the feature map F1' to the class dimension and expand the class attention map to match the dimension:
[0193]
[0194] Apply the attention map to weight the feature map:
[0195]
[0196] Take the average in the k dimension:
[0197]
[0198] Fuse the feature map in the class dimension back to the original number of channels through a 1×1 convolution and multiply it element-wise with the input feature map to achieve semantic weighting:
[0199]
[0200] The semantic weighting operation enables the model to adaptively strengthen the important regions in the input feature map, improve the response ability to key lesion regions, while maintaining the spatial structure integrity of the feature map, enhancing the discriminative ability and accuracy of the model.
[0201] The Relational Multi-Head Attention module (RMA) effectively enhances the model's ability to understand and represent complex feature relationships in DR images. This module not only enhances the model's ability to identify fine-grained lesion features but also improves the overall model performance and robustness through a flexible feature fusion strategy.
[0202] To focus on the continuity between different categories in Diabetic Retinopathy (DR) images, this embodiment regards the overall problem as a continuous regression task and then converts the regression output into a discrete classification result through a mapping mechanism. The regression task aims to predict a continuous numerical score related to the image, while the classification task achieves the final classification prediction by mapping these continuous values into five predefined categories. To achieve the collaborative optimization of these two tasks, this embodiment discloses a weighted integrated loss function that combines Cross-Entropy Loss (CE) and Mean Squared Error Loss (MSE) and introduces a regularization term to prevent the model from overfitting.
[0203] The goal of the image classification task is to accurately classify the input image into five predefined categories. For this purpose, this embodiment adopts the Cross-Entropy Loss function (CE), which is one of the most commonly used and effective loss functions in classification tasks. The Cross-Entropy Loss function can effectively measure the difference between the predicted probability distribution and the true label distribution, and its mathematical expression is as follows:
[0204]
[0205] m represents the number of inputs, k is the number of classification categories, t j represents that the true label is the j-th category, and prob i represents the predicted label probability after passing through the activation function. Through the Softmax activation function, the model can map the output of the classification head into a probability distribution, ensuring that the sum of the predicted probabilities for each sample is 1. The advantage of the Cross-Entropy Loss function is that it imposes a large penalty on prediction errors, especially when the predicted probability is far from the true label probability, and the loss value increases rapidly. However, using the Cross-Entropy Loss function alone in multi-task learning may ignore the optimization requirements of the regression task, leading to optimization conflicts between tasks.
[0206] The regression task is the core of the research on Diabetic Retinopathy image processing, and its goal is to predict the continuous numerical score corresponding to the input image. For this purpose, this embodiment adopts the Mean Squared Error Loss function (MSE) to measure the gap between the predicted value and the true value, and its mathematical expression is as follows:
[0207]
[0208] y is the output score. The MSE loss mainly considers the gap between the predicted label and the true label, but there is a problem with the optimization efficiency of the MSE loss function: when the difference between the predicted value and the true value is large, its gradient is relatively small, which makes it difficult for the model to quickly correct large prediction biases; when the difference is small, its gradient is relatively large, which affects the precise convergence of the model when approaching the optimal solution. This characteristic does not conform to the ideal learning process because researchers in this field hope that the greater the prediction error, the greater the adjustment strength of the model, so as to achieve more efficient optimization.
[0209] To overcome the above problems, this paper proposes an innovative dynamic weighted loss function integration strategy. Different from the traditional fixed-weight method, this embodiment discloses a dynamic weight adjustment mechanism based on performance feedback, and its combined loss function is defined as:
[0210] L = α(t)·L cls + β(t)·L reg + L r
[0211] where α(t) and β(t) are weight coefficients that are dynamically adjusted during the training process. These weight coefficients are adaptively adjusted based on the real-time performance of the model in terms of classification accuracy and regression performance (Kappa coefficient). When the relative improvement in classification accuracy is greater, the weight α(t) of the classification loss is increased; when the relative improvement in the Kappa coefficient is greater, the weight β(t) of the regression loss is increased. To ensure the stability of training, a momentum mechanism (momentum = 0.9) is used to smooth the change of weights, and the individual weights are restricted within the range of [0.2, 0.8] to avoid a certain task completely dominating the optimization process. L r is the regularization loss, and its purpose is to prevent the model from overfitting. Introducing the regularization term finds a balance between the complexity of the model and the degree of fitting of the training data.
[0212] Weighted fusion of the CE loss and the MSE loss can significantly improve the model performance, which stems from the complementarity of the two loss functions. The CE loss focuses on the accurate division of discrete categories, ensuring that the model can make clear category judgments; while the MSE loss focuses on the continuity of numerical values, helping the model understand the transitional relationship between categories. The MSE loss can provide corresponding penalties according to the distance between the predicted value and the true value, while the CE loss ensures that the model can maintain high discriminative ability in each category. This complementarity enables the model to not only accurately distinguish different categories but also maintain sensitivity to the order relationship between categories.
[0213] There are significant differences in objectives and loss functions between image classification and regression tasks. With the regression task as the main focus, the prediction of continuous scores is optimized through the MSE loss, while the CE loss is used to ensure the accuracy of discrete class predictions. Through a dynamic weighted loss function fusion mechanism, the model can adaptively balance the requirements of the two tasks in the same optimization process, ensuring that the regression task can provide accurate continuous score predictions and that these predictions can be accurately mapped to the corresponding classes. This design not only improves the prediction accuracy of the model but also enhances the model's ability to understand the internal order relationship of the data. The core advantage of multi-task learning lies in the feature sharing and information complementarity between tasks, and the dynamic weighting mechanism ensures that this sharing and complementarity can achieve an optimal balance during training.
[0214] To improve the feature expression ability of deep neural networks in the Diabetic Retinopathy (DR) grading task, the ShuffleAttention attention mechanism module is introduced, as Figure 5 shown, Figure 5 The descriptions of each character parameter in
[0215] are as follows: Group: grouping; Split: splitting; Channel Shuffle: channel shuffling; aggregate: summarizing; Concat: concatenating; Fuse: fusing. The ShuffleAttention attention mechanism module effectively captures important channels and key spatial regions in the feature map by combining channel attention and spatial attention mechanisms. In addition, the channel shuffle operation is introduced to promote cross-group information exchange, thereby enhancing the diversity of feature representations and the generalization ability of the model. This module has a simple structure and high computational efficiency, is suitable for the feature extraction stage of various deep learning models, and shows significant performance improvement in complex medical image analysis tasks.
[0216] This embodiment has been experimentally verified on three public datasets, namely DDR, Messdior-2, and APTOS.
[0217] The DDR dataset is currently the largest retinal image dataset in China, containing 13,673 color retinal images from 147 hospitals in 23 provinces. The retinal images in this dataset were scored by multiple professional scorers according to the international DR scoring standard, and the final scoring results were determined by voting. In addition, the dataset includes more than 1,000 non-gradable low-quality images. This embodiment only focuses on the gradable images, a total of 12,522. The sample distribution of each grade is shown in the following table. In the experiment, each category was randomly divided in a ratio of 7:1.5:1.5 to obtain the training set, validation set, and test set.
[0218] The Messdior-2 dataset consists of 874 pairs of retinal images, each pair containing two retinal images centered on the macula. This dataset contains 1,748 images, and the sample distribution of each grade is shown in the following table. It is an extension of the Messdior dataset, maintaining a consistent graphic style, color, and visual integrity, and providing higher-quality pictures than other DR datasets. In the experiment, each category was randomly divided in a ratio of 7:1.5:1.5 to obtain the training set, validation set, and test set.
[0219] The ATPOS dataset originated from the Kaggle competition "ATPOS2019 Blindness Detection". It was collected by doctors from a hospital in India using various devices in rural areas and was annotated by professional doctors. However, only the annotation information of the training set is publicly available, consisting of 3,662 images, and the sample distribution of each grade is shown in the following table. In the experiment, each category was randomly divided in a ratio of 7:1.5:1.5 to obtain the training set, validation set, and test set.
[0220] The sample distribution of each grade in the DDR, Messdior-2, and APTOS datasets is shown in the following table:
[0221]
[0222] In terms of data preprocessing, since most of the images have large black edges around the eyeballs and the positions of the eyeballs in the images are not the same. To remove these unnecessary factors and interfere with the model's extraction of effective features, the redundant black edges of the retinal images are cropped to ensure that the eyeballs are centered in the images. In terms of training, in this embodiment, each round of images will be randomly flipped, cropped, and the contrast will be increased to increase the diversity of the images. Finally, the image size is uniformly adjusted to 512×512. In the experiment, the initial learning rate is set to 5e-4 and the learning rate scheduler is combined to train the network. The epoch is set to 50 in the experiment. The Adam optimizer and the cross-entropy loss function are used during the training process. The batch_size is set to 8 during the training process. Pytorch is used as the experimental framework, and all experiments are run on an NVIDIA GeForce RTX 3090 GPU with 24G of memory.
[0223] In the experiment, quadratic weighted Kappa κ is used to measure the performance of diabetic retinopathy grading. This is the official metric in the Kaggle diabetic retinopathy grading competition and is especially applicable to imbalanced datasets. The calculation formula of quadratic weighted Kappa is as follows:
[0224]
[0225] where C represents the total number of classes, w is the quadratic weight matrix, and the subscripts i and j represent the row and column indices of the matrix respectively. The weight w ij is defined as The range of κ is from -1 to 1, and -1 and 1 represent complete disagreement and complete agreement. o ij represents the number of samples in which the observed (actually occurring) class i is predicted as class j, that is, the element in the i-th row and j-th column of the confusion matrix. e ij represents the expected number of samples in which class i is predicted as class j under random prediction.
[0226] To analyze the effects of the Hierarchical Global Context Module (HGCM), Multi-Scale Adaptive Attention Module (MSAM), Relational Multihead Attention (RMA), ShuffleAttention attention mechanism module, and multi-task learning, ablation experiments are conducted on the DDR dataset, and the experimental results are shown in the following table:
[0227]
[0228] First, the first two rows in the table show the experimental results exploring the effectiveness of the Hierarchical Global Context Module (HGCM). The experiments show that, compared with the basic network, after adding the HGCM, the model improves the grading ability of DR, and its quadratic weighted kappa value increases by 0.7%. Since the HGCM breaks through the limitation of the traditional convolutional neural network in capturing fine-grained pathological features under a single receptive field, it can capture more lesion information, extract more discriminative features, and thus improve the grading ability.
[0229] Next, the first and third rows in the table show the experimental results verifying the effectiveness of the Multi-Scale Adaptive Attention Module (MSAM). The experiments show that, compared with the basic network, after adding the HGCM, the model improves the grading ability of DR, and its quadratic weighted kappa value increases by 2.0%.
[0230] It can be seen that through the collaborative work of the above multiple modules in this embodiment, the diverse features of DR lesions can be comprehensively and finely analyzed, significantly improving the generalization ability and robustness of the model. The experiments show that, compared with the traditional single-branch attention model and other mainstream algorithms, MAFNet has achieved a significant improvement in the grading performance on the public dataset, further verifying the effectiveness and superiority of this method in the DR fine-grained grading task.
[0231] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to obtain equivalent embodiments with equivalent changes, but as long as the technical content of the present invention is not departed from, any indirect modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy, characterized in that Including: Hierarchical global context module: Integrating multi-scale feature extraction, adaptive attention enhancement, and dynamic weighted fusion to construct a multi-scale representation mechanism that complements global and local information, generating more expressive feature representations. Multi-scale adaptive attention module: Capturing complex dependencies between different spatial positions in the feature map through the synergistic effect of multi-scale feature extraction and adaptive attention enhancement. Relational multi-head attention module: Deeply mining and capturing complex associations between different features in the feature map to enhance the model's ability to identify fine-grained lesion features and classification performance. Attention mechanism module: Combining channel and spatial attention mechanisms to capture key channels and important spatial regions in the feature map.
2. The adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 1, wherein In the hierarchical global context module, for a given feature map, the number of channels of the input feature map is mapped to a lower dimension through convolution. Dimensionality reduction operations are performed on the feature map mapped to a lower dimension through adaptive average pooling of different sizes to obtain corresponding pooled features. The pooled features are subjected to convolution, normalization, and PreLU activation processing. The processed pooled features are sampled to the size of the original feature map. An attention mechanism is introduced into the upsampled feature map, and an attention map is generated through convolution and activation functions. Feature enhancement is achieved by element-wise multiplication of the upsampled feature map and the attention map, and weighted by learnable weights. The enhanced features of all scales are concatenated in the channel dimension, and the number of channels is compressed and fused through convolution to output the final multi-scale feature representation.
3. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 2, characterized in that, In the hierarchical global context module, the expression for mapping the number of channels of the input feature map to a lower dimension is as follows: Among them, X down is the feature map mapped to a lower dimension; Conv2d 3×3 Represents a convolution operation of 3×3; BN is batch normalization. φ is the PreLU activation function. H and W are the height and width of the feature map respectively. The expression for obtaining the pooled features is as follows: Among them, F i is the pooled feature; The expression for performing convolution, normalization, and PreLU activation processing on the pooled features is as follows: F’ i = φ(BN(Conv2d 1×1 (F i ))) Among them, F' i is the processed pooling feature; The expression for sampling the processed pooled features to the size of the original feature map is as follows: The expression for generating the attention map is as follows: where σ is the Sigmoid function. The weights for weighting are normalized by the Softmax function: where p i is the parameter vector for learning; ⊙ represents element-wise multiplication. The final multi-scale feature representation is as follows: Among them, X HGCM is the final multi-scale feature.
4. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 1, characterized in that, In the multi-scale adaptive attention module, the feature map is divided into multiple equal groups to obtain multiple sub-feature maps. Each sub-feature map is convolved through a group of dynamically generated convolutional kernels to obtain each convolutional output feature map. Each convolutional output feature map is enhanced through a channel attention mechanism, and the channel importance of the convolutional output feature map is adjusted through channel weights to obtain a multi-scale feature map. The multi-scale feature maps are stacked in a new dimension through a dynamic weight generation mechanism. The dynamically generated weights are multiplied element-wise with the stacked feature maps to enhance the weighting of each scale feature. The weighted feature maps are reorganized back to the original number of channels.
5. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 4, characterized in that, In the multi-scale adaptive attention module, the expression for dividing the feature map into multiple equal groups is as follows: Among them, S is the number of equal groups, X is the feature map to be divided, is the number of channels of each sub-feature map, C is the number of input channels, and H and W are the height and width of the feature map respectively; The dynamic convolutional kernels for performing convolution on each sub-feature map consist of multiple convolutional layers, and the size of the convolutional kernels varies with the number of groups S, and the expression is as follows: Among them, Conv2d represents the standard two-dimensional convolution operation; The groups parameter implements grouped convolution; The channel attention mechanism adopts the squeeze-and-excitation module, and generates channel weights through adaptive pooling and fully connected layers. The expression is as follows: Among them, φ represents the ReLU activation function; is the weight of the fully connected layer; σ is the Sigmoid activation function; The expression for adjusting the channel importance of the convolution output feature map through the channel weight is as follows: Among them, ⊙ represents element-wise multiplication; The expression for stacking multi-scale feature maps on a new dimension is as follows: The expression for the dynamic weight generation mechanism is as follows: W = Softmax(Conv2d(AdaptiveAvgPool2d(X), kernel size = 1)) The expression for the weighted enhancement of each scale feature is as follows: conv_outs s,c',h,w = conv_outs s,c',h,w ⊙W s The expression for generating the final output feature map is as follows: Among them, X MSAM is the finally generated output feature map.
6. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 1, characterized in that In the relational multi-head attention module, the feature map is mapped to the k×classes dimensional space through convolution. The expression is as follows: Among them, k is the feature dimension of each class; classes is the number of classes; H and W are the height and width of the feature map respectively; Batch normalization and ReLU activation are performed on the convolution output. The expression is as follows: Among them, BN represents the batch normalization operation, and ReLU is the activation function; The feature map is weighted through the channel attention mechanism. The expression is as follows: F1' = F1⊙ChannelAttentionMap Among them, σ is the Sigmoid activation function; ⊙ represents element-wise multiplication; In the relational multi-head attention module, the multi-head attention mechanism is introduced, and the channel dimension k×classesk is flattened and transposed to adapt to the multi-head attention mechanism. The expression is as follows: Among them, H and W are the height and width of the feature map respectively; The core formula of the multi-head attention mechanism is: Among them, Q, K, and V are the query (Query), key (Key), and value (Value) vectors respectively; d k is the dimension of the key vector; The output of the multi-head attention is obtained by concatenating the attention results of multiple heads and fusing them through a linear transformation. The expression is as follows: An attention map is generated through a 1×1 convolution, and Softmax normalization is applied to obtain the weights. The expression is as follows: Among them, Conv2d 1×1 is used to generate class attention maps, and the Softmax operation normalizes in the class dimension; The feature map F1' is rearranged into the class dimension, and the class attention map is extended to match the dimension. The expression is as follows: The attention map is applied to weight the feature map. The expression is as follows: An average is taken in the k dimension. The expression is as follows: The feature map in the class dimension is fused back to the original number of channels through a 1×1 convolution and multiplied element-wise with the input feature map to achieve semantic weighting. The expression is as follows: Among them, Semantic is the result of semantic weighting.
7. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 1, characterized in that, The adaptive multi-scale model adopts a dynamic weight adjustment mechanism based on performance feedback, and its comprehensive loss function is defined as: L = α(t)·L cls + β(t)·L reg + L r Among them, α(t) and β(t) are weight coefficients that are dynamically adjusted during the training process; L cls is the cross-entropy loss function; L reg is the mean squared error loss function.
8. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 7, characterized in that The α(t) and β(t) are adaptively adjusted based on the real-time performance of the model in terms of classification accuracy and regression performance. When the relative improvement in classification accuracy is greater, the weight α(t) of the classification loss is increased; when the relative improvement in regression performance is greater, the weight β(t) of the regression loss is increased.
9. An adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to claim 1, characterized in that, In the attention mechanism module, a channel shuffle operation is introduced to promote cross-group information exchange, thereby enhancing the diversity of feature representations and the generalization ability of the model.
10. A method for processing retinal images for diabetic retinopathy, characterized in that, It includes performing fine-grained grading processing on retinal images using the adaptive multi-scale model for fine-grained grading of diabetic retinopathy according to any one of claims 1-9.
Citation Information
Cited By
Abrasive hardening detection method and system of deep network structure
CN121213571A
Zero-sample blood glucose prediction method based on dual-scale Transform framework
CN121964155A