Heterogeneous double-flow fusion method and system for grading diabetic retinopathy

By employing a heterogeneous dual-stream fusion method, combining a lightweight ViT encoder and an EfficientNet-B5 CNN encoder, and utilizing a symmetrical bidirectional cross-attention module to achieve adaptive deep interaction between global contextual information and local lesion details, this approach addresses the limitations of model representation capabilities and high computational costs in existing technologies, thereby achieving efficient and accurate grading of diabetic retinopathy.

CN121033041AActive Publication Date: 2025-11-28HUNAN NORMAL UNIVERSITY

Patent Information

Application Number
CN202511559234.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-11-28
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

Existing methods for grading diabetic retinopathy suffer from limitations in the representation capabilities of single models, high computational costs of large models, high barriers to clinical deployment, and insufficient fusion of heterogeneous features in hybrid architectures.

Method used

A heterogeneous dual-stream fusion approach is adopted, which transfers the knowledge of the large-scale basic model to the lightweight ViT encoder through composite knowledge distillation technology. Combined with the EfficientNet-B5 CNN encoder, the adaptive deep interaction between global context information and local lesion details is achieved by using a symmetrical bidirectional cross attention module, and the model is optimized by cross-entropy loss.

Benefits of technology

It achieves improved accuracy and robustness in grading diabetic retinopathy while reducing computational costs and the number of parameters. It can be easily deployed in resource-constrained medical devices and has the ability to capture both global contextual information and local lesion details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033041A_ABST
    Figure CN121033041A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous double-flow fusion method and system for diabetic retinopathy grading. The method comprises the following steps: obtaining an output result of diabetic retinopathy grading by utilizing a heterogeneous double-flow architecture; processing an input fundus image into images with different resolutions; extracting global context features from the low-resolution image by using a lightweight visual Transform model distilled by composite knowledge, and extracting local focus features from the high-resolution image by using a convolutional neural network model; performing interactive fusion on the global context features and the local focus features of the double-branch architecture through a symmetric bidirectional cross attention fusion module to obtain enhanced fusion feature representation; and finally, inputting the fusion features into a classifier, and outputting a severity grading result of the lesion. The method aims at improving the accuracy and robustness of hierarchical diagnosis through deep analysis of global information and local details, and can be applied to the medical fields of clinical computer-aided diagnosis, eye image analysis and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of medical image processing, and particularly relates to a heterogeneous dual-flow fusion method and system for diabetic retinopathy grading. BACKGROUND

[0002] Diabetic retinopathy (DR) is an ocular fundus complication caused by long-term high blood sugar, which causes progressive damage to retinal blood vessels, and is one of the main causes of blindness in working-age adults. The development of DR can be divided into different stages such as non-proliferative (NPDR) and proliferative (PDR), and the symptoms of early stage such as microaneurysm and hemorrhagic spots are mild or even asymptomatic. If not discovered and intervened in time, it will cause damage to vision, and even lead to blindness.

[0003] In practice, the screening and severity grading diagnosis of DR basically rely on manual reading of ophthalmologists. With the gradual increase in the number of DR patient population, manual analysis of cases consumes a lot of manpower and time, and the subtle differences between different degrees of DR lesions and the increasingly complex lesions are easily affected by the subjective experience of doctors, resulting in challenges to the consistency and accuracy of diagnosis. Therefore, it has become an urgent medical need to develop a computer-aided diagnosis (CAD) tool that can perform large-scale and efficient screening.

[0004] In recent years, existing research has proposed various deep learning-based DR automatic grading schemes, which can be mainly summarized as follows: methods using convolutional neural network (CNN) as the core architecture, which make good use of the inherent local inductive bias of CNN in extracting local lesion features (such as microaneurysms, exudates, etc.) in fundus images. However, the architecture design of CNN is essentially limited by the local receptive field of convolution operation, making it difficult to effectively establish the dependency between distant pixels in the image, such as understanding the overall trend of retinal blood vessels or the distribution of lesions; methods using visual Transformer (ViT) architecture, especially using large pre-trained base models such as RETFound model on a large amount of image data. This kind of method can capture the global dependency and context information of the image by virtue of its self-attention mechanism, and has shown strong performance in the field of image processing. However, the large number of parameters results in high computational cost and high clinical deployment threshold, and in addition, due to the lack of CNN's inductive bias, its accuracy is sometimes not as good as CNN model when analyzing and identifying fine local lesions with small pixel ratio but important texture. Today's research focus tries to fuse the hybrid architecture of CNN and ViT, and the existing fusion methods usually adopt a serial structure, and for parallel structures, the feature fusion method is mostly simple splicing or element-wise addition. This kind of simple static fusion strategy has shortcomings, which easily ignores the heterogeneity between local spatial hierarchical features and global context features, and cannot realize the deep fusion of heterogeneous features, thereby limiting the improvement of model performance.

[0005] There is currently a lack of a diabetic retinopathy severity automatic grading method in the technical field that can simultaneously consider the capturing ability of global context information and local lesion details, achieve a good balance between model performance and computational efficiency, and realize deep adaptive fusion of heterogeneous features. SUMMARY

[0006] The purpose of the present application is to provide a heterogeneous dual-flow fusion method for diabetic retinopathy grading. To solve the problems of existing diabetic retinopathy grading methods, such as limited representation ability of single model, high computational cost of large model, high clinical deployment threshold, and insufficient fusion of heterogeneous features in hybrid architecture,

[0007] To achieve the above purpose, the present application provides the following scheme:

[0008] The present application provides a heterogeneous dual-flow fusion method for diabetic retinopathy grading, comprising the following steps:

[0009] S1: Global feature extraction module: This module uses a composite knowledge distillation technology to transfer the knowledge of a large base model to a lightweight ViT encoder for processing 224x224 low-resolution images, efficiently capturing global context information while significantly reducing computational cost and parameter size to meet the needs of efficient lightweight deployment.

[0010] S2: This branch uses a CNN encoder (EfficientNet-B5) pre-trained on ImageNet. It is used to process 456x456 high-resolution images, and uses the inherent local inductive bias of CNN to extract fine local texture and spatial hierarchical features of small lesions, effectively compensating for the shortcomings of ViT in analyzing small lesions.

[0011] S3: Symmetric bidirectional cross-attention module: This module is designed to achieve adaptive and symmetric bidirectional cross-attention depth interaction and enhancement between two heterogeneous branches (global ViT features and local CNN features). Through this dynamic fusion mechanism, the method ensures that both the microscopic details of DR lesions and the macroscopic structure of the retinal image are fully considered, avoiding information conflicts caused by simple static fusion.

[0012] S4: To map the fused high-dimensional feature vector to the final class prediction, a cross-entropy loss is used to optimize the model. The fused features are fed into a multi-layer perceptron classification head or a DR grading diagnosis result.

[0013] Further, in S1, a composite knowledge distillation scheme is designed, which includes soft target distillation, feature distillation, and contrast distillation, to transfer knowledge from the teacher model to the lightweight student model in multiple dimensions (including prediction distribution, feature hierarchy, and data relationship). During the experiment, the RETFound encoder is used as the teacher model, and a ViT-Base model with reduced depth and width is used as the student model. Under the premise of freezing the teacher model parameters, the student model is optimized by a composite loss function , which specifically includes:

[0014] S1.1, soft target distillation technology:

[0015] Simply using hard labels for supervision ignores the implicit knowledge in class probability prediction, i.e., the similarity information between classes. The present invention uses a Softmax function with a temperature coefficient to soften the logits output of the teacher model. By minimizing the KL divergence between the soft targets of the student model and the teacher model, the student model is forced to fit the fine distribution of the teacher model for class probability. This part of the loss function is defined as: In the above formula, and These are the logits outputs of the student and teacher models, respectively. It is the Softmax function. It is the temperature coefficient. When When the value is greater than 1, the output distribution becomes flatter, which helps the student model learn richer inter-category structural information. The factor is used to scale the gradient during backpropagation, ensuring that it is applied at different temperatures. Under these conditions, the contribution magnitudes of the loss terms for soft and hard targets remain relatively stable.

[0016] S1.2, Characteristic distillation technology:

[0017] To guide student models in simulating more robust intermediate feature representations of teacher models, this invention introduces feature distillation. This method encourages student models to utilize pre-defined intermediate layers... The feature map of the student model should be matched with the feature map of the teacher model. Considering that the differences in network structure between the student and teacher models may lead to inconsistent feature map dimensions, this invention introduces an adaptation layer φ to align dimensions. This loss term is defined as the square of the L2 norm between the adapted student features and teacher features at a specified intermediate layer, as follows: In the above formula, It is the selected set of intermediate layers. and The student and teacher models respectively represent the model in the 19th century. Layer feature representation, It is for the first Learnable adaptation layers of layer feature maps. Minimize this loss by directly supervising the intermediate representations of the student model in the feature space.

[0018] S1.3, Comparative learning of distillation techniques:

[0019] Higher-level knowledge is embedded in the structural relationships between data samples. To transfer the relational knowledge learned by the teacher model, this study constructs a contrastive learning distillation loss. Its core idea is that if two samples are similar (positive sample pairs) in the teacher model's representation space, then they should also be brought closer in the student model's representation space; otherwise, they should be pushed apart. This loss can be formalized as: In the above formula, This refers to the batch size. It is the student model for anchor point samples Feature representation, It is the feature representation of the teacher model for the same positive sample, while This represents the teacher feature representation of N negative samples. The vector inner product is used to calculate similarity. It is a temperature hyperparameter used to adjust the sharpness of the similarity distribution.

[0020] The above three distillation losses are compared with the standard supervised learning loss (i.e., the standard cross-entropy loss). The weighted combinations are then used to form the final distillation objective function: in, It is the cross-entropy loss between the student model and the real label, used to ensure the model's basic prediction performance on the real label. These are hyperparameters used to balance the relative importance of the three losses: soft target distillation, characteristic distillation, and comparative distillation. The composite objective function is jointly optimized. This invention successfully reduced the number of parameters in the student model from 3.7 billion to 327 million, while minimizing performance loss in key downstream tasks. The resulting ViT architecture includes image patch embedding, location embedding, class labeling, and a Transformer encoder layer, successfully providing a powerful and efficient global feature encoder for hybrid models, while inheriting the representation capabilities of the RETFound base model.

[0021] S2 selects EfficientNet-B5 as the local feature CNN branch, focusing on extracting fine local texture and spatial hierarchy features from high-resolution images:

[0022] Among the various variants of EfficientNet, the B5 variant offers an ideal balance between parameter count and performance, and its powerful texture processing capabilities are well-suited for medical image analysis. This invention uses weights pre-trained on the large natural image dataset ImageNet to... Perform initialization, As an initial feature extractor, it can capture local spatial patterns and hierarchical features in images. The main motivation for transfer learning lies in the limited number of high-resolution fundus images available for model training. Utilizing knowledge from diverse datasets such as ImageNet enhances the model's generalization ability and maintains good performance even with a limited number of training images. In practical processing, this branch receives... High-resolution preprocessed image (e.g., 456×456×3) As input, the image is encoded through convolutional layers of EfficientNet-B5, and after a series of MBConv blocks, hierarchical features from low to high levels are extracted step by step. This invention takes the final feature map before the global average pooling layer and flattens it to obtain a... Local eigenvectors of dimension .

[0023] Furthermore, in step S3, this invention designs a symmetrical bidirectional cross-attention fusion module. The core mechanism of this module is: to promote... and Deep interaction allows two features to act as queries, actively seeking the desired information from each other's features (as keys and values). The specific method is as follows:

[0024] Global-to-local information augmentation: The module uses three independent learnable linear projection layers (i.e., fully connected layers, with weight matrices respectively) , , This is used to generate the query, key, and value matrix required for the attention mechanism. Global features. As a query, from local features The process extracts relevant local details to enrich the global view with fine texture. The attention output generated in this process is added to the original global features via residual connections, then stabilized through a normalization layer, ultimately resulting in an enhanced global feature that absorbs local details. This process can be formally represented as:

[0025] in, Depend on Generate, and and Depend on Linear projection generation. Output. This is the enhanced global feature.

[0026] Local-to-global information enhancement: Symmetrically, local features As a query, from global features The process involves identifying the context within the lesion to understand its importance from a global perspective. This process also includes residual connections and layer normalization to obtain an enhanced local feature that incorporates global context awareness. This process can be formally represented as:

[0027] in, Depend on Generate, and and Depend on generate.

[0028] After the attention output from the two branches described above, a unified residual fusion function and layer normalization strategy are applied after each branch. The extracted fusion information is then fused with the original features again through residual connections (i.e., element-wise addition). The fused result is then input into a layer normalization unit for processing to avoid missing key information. This process ultimately generates features that absorb key local details and fuse global information, ensuring the model's training stability. Finally, these two feature vectors, which have undergone bidirectional interaction and enhancement, are... and By splicing together, a fusion feature is formed. The data is then fed into a classifier for decision-making. This process can be formally represented as:

[0029] The pre-trained, powerful EfficientNet-B5 branch inherits the precise ability of CNNs to capture local lesions (such as microaneurysms and hemorrhages), while the lightweight ViT branch, after knowledge distillation, constructs global context dependencies, fusing the features of both to generate more comprehensive features. Through this design, HDNet overcomes the limitation of traditional static fusion, which tends to ignore deep heterogeneous relationships between features, achieving a dynamic, context-aware feature enhancement that enables the model to simultaneously perceive the microscopic details of DR lesions and the macroscopic structure of retinal images.

[0030] To fuse the high-dimensional feature vector To map to the final class prediction, a classification head is designed, and the model is optimized using cross-entropy loss. The data is fed into a classification head consisting of a multilayer perceptron (MLP) containing two hidden layers with ReLU activation (dimensions 512 and 256, respectively), designed for learning. A non-linear mapping between the score and the category label. Outputs an M-dimensional original score vector (logits). The Softmax function is used to transform z into a class probability distribution. The entire model is trained in an end-to-end manner, and its optimization objective is to minimize the predicted probability distribution. With real labels Classification cross-entropy loss For those containing The dataset of a sample, the total loss on the entire dataset It is the mean of the losses of all samples, and the total loss is defined as:

[0031] in It is a one-hot encoded real tag. The model predicts the sample. Category The probability of this is determined. Finally, the classification head outputs the diagnostic result of the diabetic lesion grading based on the input fusion enhancement features.

[0032] Furthermore, the present invention also provides a heterogeneous dual-stream fusion system for grading diabetic retinopathy, comprising a processor and a memory interconnected thereon, wherein the memory stores computer-executable instructions, and when the computer-executable instructions are executed by the at least one processor, the microprocessor is programmed or configured to any of the aforementioned heterogeneous dual-stream fusion methods for grading diabetic retinopathy.

[0033] Furthermore, the present invention also provides a computer-readable storage medium having program code loaded thereon, the program code comprising a set of instructions which, when executed on a computing device, cause the computing device to perform the steps defined in any of the aforementioned heterogeneous dual-stream fusion methods for grading diabetic retinopathy.

[0034] Furthermore, the present invention also provides a computer program product comprising program instructions stored on a computer-readable medium, which, when loaded and executed by a processor, are intended to implement any of the aforementioned heterogeneous dual-stream fusion methods for grading diabetic retinopathy.

[0035] Compared with existing technologies, this invention has the following main advantages: it can automatically obtain accurate grading results for diabetic lesions by comprehensively analyzing the input fundus images. Specifically, this invention first processes the fundus images input by the device into different resolutions and inputs them into two parallel feature extraction branches: the global feature extraction branch designs a composite distillation scheme, using a lightweight visual Transformer model that is significantly compressed using techniques such as feature distillation and contrastive learning distillation, reducing the number of parameters of the feature extractor from 3.7G to 327MB, which is only 8.8% of its original size. This efficiently captures global contextual information such as retinal vessel trends and lesion distribution on low-resolution images with low computational cost; the local lesion feature extraction branch uses the convolutional neural network EfficientNet-b5 to accurately focus on key local lesion details such as microaneurysms and hemorrhages on high-resolution images, capturing local spatial patterns and hierarchical features in the image. Traditional hybrid models use simple feature splicing or element-wise addition fusion methods, which easily ignore deep heterogeneous relationships between features. In contrast, this method introduces a symmetrical bidirectional cross-attention fusion module. The symmetrical bidirectional design of this module ensures adaptive deep interaction between the global features extracted by ViT and the local features extracted by CNN. This allows global features to actively "query" and absorb the most relevant lesion details from local features, while also enabling local lesion features to "perceive" their context within the global retinal structure. This overcomes the performance bottleneck of simple addition or splicing, generating highly complementary and representative fusion features that balance global contextual information with high-resolution local information. Therefore, this invention effectively solves the technical challenges of single-model representation being one-sided (e.g., CNN lacking a global perspective), the high clinical deployment threshold of large models (e.g., the excessive computational cost of large ViT models), and the insufficient handling of heterogeneous features by existing fusion methods. While ensuring or even surpassing the accuracy of large models in grading diabetic retinopathy, this invention significantly reduces the number of model parameters and computational complexity, making it easily deployable in resource-constrained medical devices or primary healthcare institutions to improve the accuracy of diabetic retinopathy diagnosis. It can be applied to fields such as ocular image analysis and precision medicine. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the basic process of the method in an embodiment of the present invention;

[0037] Figure 2 This is an algorithm framework diagram of a heterogeneous two-stream fusion method for grading diabetic retinopathy provided in an embodiment of the present invention;

[0038] Figure 3 This is a framework diagram of the composite knowledge distillation scheme in an embodiment of the present invention;

[0039] Figure 4 This is a diagram showing the internal structure of the symmetrical bidirectional cross-attention fusion module in an embodiment of the present invention.

[0040] Figure 5 This is a comparison of model performance on the MESSIDOR-2 dataset in the embodiments of the present invention;

[0041] Figure 6 This is a comparison of model performance on the IDRID dataset in the embodiments of the present invention;

[0042] Figure 7 This is the ROC curve for multi-class classification on the MESSIDOR-2 dataset in an embodiment of the present invention;

[0043] Figure 8 This is the ROC curve for multi-class classification on the IDRID dataset in an embodiment of the present invention;

[0044] Figure 9 This is a diagram illustrating the heatmap effect in an actual sample according to an embodiment of the present invention.

[0045] Figure 10 This is a heatmap illustration of the ViT branch in an actual sample in an embodiment of the present invention;

[0046] Figure 11 This is a heatmap illustration of the CNN branch effect in actual samples in an embodiment of the present invention. Detailed Implementation

[0047] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. The described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the protection scope of the present invention.

[0048] The core of this invention is to provide a heterogeneous dual-stream fusion method for grading diabetic retinopathy, in order to solve the problems existing in the prior art.

[0049] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] Figure 1 This is a basic flowchart illustrating a heterogeneous two-stream fusion method for grading diabetic retinopathy provided in an embodiment of the present invention, as shown below. Figure 1 As shown, it includes the following steps:

[0051] S1: Receives the input image and performs parallel resizing to generate images at different resolutions. Specifically, the 224×224 pixel image is fed into the ViT branch, while the 456×456 pixel image is fed into the CNN branch.

[0052] S2: Two branches work in parallel to extract features from different dimensions. The ViT branch extracts feature vectors representing global contextual information from its input. The CNN branch (EfficientNet) extracts feature vectors representing local texture information from its input.

[0053] S3: The feature vectors of ViT and CNN are fed into the vector cross-attention module. This module implements bidirectional information exchange: ViT features act as queries to extract information from CNN features, and CNN features act as queries to extract information from ViT features. Through a standard multi-head attention mechanism and residual connections, deeply fused and enhanced fused features are generated.

[0054] S4: Integrate the two enhanced feature vectors output by the cross-attention module. These two vectors absorb each other's advantageous information from the ViT and CNN features. By directly concatenating them along the feature dimension, a final fused feature with higher information density and stronger expressive power is formed;

[0055] S5: Input the fused feature vector into a multilayer perceptron (MLP) classification head. This classification head maps the high-dimensional features to the final category space through fully connected layers, normalization layers, and activation functions, completing the classification of image content and outputting the final prediction result.

[0056] Figure 2 This invention provides an algorithmic framework diagram for a heterogeneous two-stream fusion method for grading diabetic retinopathy (DR). To effectively address the challenges in DR grading (i.e., subtle lesion identification, global context understanding and multi-scale integration, model efficiency and deployment limitations, and the difficulty of heterogeneous feature fusion), this invention proposes a heterogeneous two-stream fusion method called HDNet. Its core design is its innovative parallel dual-branch heterogeneous architecture. By processing input images at different scales, the heterogeneous branches leverage their efficient feature extraction capabilities, and finally, through deep fusion, achieve the goal of DR grading. This design can solve the problems of accurate identification of subtle lesions, comprehensive understanding of global context, model efficiency and deployability, and dynamic and effective fusion of heterogeneous features, ultimately achieving a more accurate and robust grading of DR severity. The overall framework of HDNet consists of three core components: Global Feature Extraction Module: This module uses composite knowledge distillation technology to transfer the knowledge of a large base model to a lightweight ViT encoder for processing low-resolution 224×224 images. It efficiently captures global contextual information while significantly reducing computational costs and the number of parameters to meet the needs of efficient and lightweight deployment.

[0057] Lesion Focusing Module: This branch employs a CNN encoder (EfficientNet-B5) pre-trained on ImageNet. It is used to process high-resolution 456×456 images, leveraging the inherent local inductive bias of CNNs to extract fine local textures and spatial hierarchy features of minute lesions, effectively compensating for the shortcomings of ViT in local detail analysis.

[0058] Symmetrical Bidirectional Cross-Attention Module (BCAF): This module is designed to achieve adaptive, symmetrical bidirectional cross-attention deep interaction and enhancement between two heterogeneous branches (global ViT features and local CNN features). Through this dynamic fusion mechanism, HDNet ensures that it can comprehensively consider both the microscopic details of DR lesions and the macroscopic structure of retinal images, avoiding information conflicts caused by simple static fusion.

[0059] like Figure 3 As shown, this invention designs a composite distillation scheme. It comprehensively transfers knowledge from the teacher model to a lightweight student model from multiple dimensions, including prediction distribution, feature hierarchy, and data relationships. In the experiment, the RETFound encoder serves as the teacher model, and a ViT-Base model with reduced depth and width serves as the student model. With the teacher model parameters frozen, a composite loss function is used to jointly optimize the student model, specifically including soft target distillation, feature distillation, and contrastive learning distillation. Simply using hard labels for supervision ignores the hidden knowledge inherent in class probability prediction, i.e., the similarity information between classes. This invention uses a temperature coefficient... The Softmax function is used to soften the logits output of the teacher model. By minimizing the KL divergence between the soft objectives of the student and teacher models, the student model is forced to fit the finer distribution of the class probabilities of the teacher model. To guide the student model to simulate a more robust intermediate feature representation of the teacher model, this invention introduces feature distillation. This method prompts the student model to perform feature distillation at a predefined intermediate layer. The feature map of the student model should be matched with the feature map of the teacher model. Considering that the differences in network structure between the student and teacher models may lead to inconsistent feature map dimensions, this invention introduces an adaptation layer φ to align dimensions. Higher-level knowledge is implied in the structural relationships between data samples. To transfer the relational knowledge learned by the teacher model, this invention constructs a contrastive learning distillation loss. Its core idea is that if two samples are similar (positive sample pairs) in the representation space of the teacher model, then they should also be brought closer in the representation space of the student model; otherwise, they should be pushed apart.

[0060] like Figure 4 As shown, this invention designs a symmetrical bidirectional cross-attention fusion module. This module performs cross-attention calculations in two directions in parallel: global-to-local information enhancement and global feature fusion. As a query, from local features The process extracts relevant local details to enrich the global view with fine texture. The attention output generated in this process is added to the original global features via residual connections, then stabilized through a normalization layer, ultimately resulting in an enhanced global feature that absorbs local details. This process can be formally represented as: .

[0061] in, Depend on Generate, and and Depend on Linear projection generation; Local-to-global information enhancement: Symmetrically, local features As a query, from global features The process involves identifying the context within the lesion to understand its importance from a global perspective. This process also includes residual connections and layer normalization to obtain an enhanced local feature that incorporates global context awareness. This process can be formally represented as: .

[0062] in, Depend on Generate, and and Depend on generate; After the attention output from the bi-branch approach described above, a unified residual fusion function and layer normalization strategy are applied after each branch. The extracted fusion information is then fused again with the original features through residual connections (i.e., element-wise addition). The fused result is then input into a layer normalization unit for further processing to avoid missing key information. This process ultimately generates features that absorb key local details and fuse global information, ensuring the model's training stability. These two feature vectors, after bi-directional interaction and enhancement, are then... and By splicing together, a fusion feature is formed. The data is then fed into a classifier for decision-making. This process can be formally represented as: .

[0063] See Figure 5 and Figure 6 It can be seen that the heterogeneous dual-stream fusion method for grading diabetic retinopathy in this invention is superior to existing diabetic retinopathy diagnostic methods in all aspects.

[0064] To further illustrate the effectiveness of the heterogeneous two-stream fusion method for grading diabetic retinopathy in the examples of this invention, Figure 7 and Figure 8 The ROC curves further demonstrate the model's discriminative ability. Different colored curves represent the performance of different models, while the diagonal dashed line represents the random guessing baseline (AUC=0.5), indicating the model has no classification ability. The visualization shows that the ROC curves representing all categories significantly deviate from the diagonal dashed line and strongly bulge towards the upper left corner of the graph. This clearly indicates that the invention achieves a very high true positive rate while maintaining an extremely low false positive rate when identifying each category. On the MESSIDOR2 dataset, its micro-average ROC area reaches 95%, demonstrating excellent performance across all categories. The macro-average ROC area also reaches 92%. The results demonstrate the model's powerful ability in multi-class classification tasks on the MESSIDOR-2 dataset. On the IDRiD dataset, the micro-average ROC area reaches 88%, demonstrating excellent performance across all categories. The macro-average ROC area also reaches 88%.

[0065] Furthermore, to illustrate the effectiveness of the heterogeneous two-stream fusion method for grading diabetic retinopathy according to the embodiments of the present invention, Figure 9 , Figure 10 , Figure 11Heatmaps are used to explain and interpret the model's decision-making mechanism. By overlaying a color layer onto the original input image, it visually demonstrates which regions in the image the model focuses on during classification decisions. The internal working principle of the proposed dual-branch parallel model is analyzed in depth. Figure 10 and Figure 11 The activation heatmaps of the ViT branch and the CNN branch in the model are shown separately when they operate independently. The figures reveal that the decision-making processes of the two independent branches each have their own emphasis but also significant limitations: the ViT branch ( Figure 10 CNNs, with their self-attention mechanism, demonstrate the ability to capture global context and long-range dependencies. Their activation regions are typically more macroscopic and complete, leaning more towards exploring global information. (CNN branch) Figure 11 The focus is more on the local texture and edge details of the image, and it can accurately capture some subtle features, such as the distribution of retinal blood vessels and local information of the eyeball.

[0066] Figure 9 The heatmap effect of the method shown in the right image illustrates that after integrating the information from both methods, it absorbs the precise localization capability of CNN for key local details and integrates the macroscopic grasp of overall structure and context of ViT. Ultimately, the hotspot region is accurately located on the most distinctive core target, making the identification of lesion areas more targeted. This intuitive comparison powerfully demonstrates that the fusion architecture proposed in this invention is not a simple feature concatenation, but rather, through effective bidirectional information interaction, it enables the model to learn how to collaboratively utilize information from different dimensions to form a more comprehensive and accurate decision.

[0067] In summary, the heterogeneous two-stream fusion method for grading diabetic retinopathy in this invention processes the input fundus image using a heterogeneous two-stream parallel architecture to obtain accurate grading results for diabetic retinopathy. First, the input fundus image is acquired and processed into images of different resolutions. Then, a lightweight visual Transformer model is used in parallel to extract global contextual features from the low-resolution image, and a convolutional neural network, EfficientNet-B5, is used to extract local lesion features from the high-resolution image. A symmetrical bidirectional cross-attention fusion module deeply fuses the global contextual features and local lesion features of the two-branch architecture to obtain an enhanced fused feature representation. Finally, this fused feature representation is input into a classifier to output the severity grading result of diabetic retinopathy. This invention aims to improve the accuracy and robustness of grading diagnosis through deep collaborative analysis of global information and local details, and can be applied to fields such as clinical computer-aided diagnosis.

[0068] Furthermore, the present invention also provides a heterogeneous two-stream fusion method for grading diabetic retinopathy, comprising a processor and a memory interconnected thereon, wherein the memory stores computer-executable instructions, and when the computer-executable instructions are executed by the at least one processor, the microprocessor is programmed or configured to any of the aforementioned heterogeneous two-stream fusion methods for grading diabetic retinopathy.

[0069] Furthermore, the present invention also provides a computer-readable storage medium having program code loaded thereon, the program code comprising a set of instructions which, when executed on a computing device, cause the computing device to perform the steps defined in any of the aforementioned heterogeneous dual-stream fusion methods for grading diabetic retinopathy.

[0070] Furthermore, the present invention also provides a computer program product comprising program instructions stored on a computer-readable medium, which, when loaded and executed by a processor, are intended to implement any of the aforementioned heterogeneous dual-stream fusion methods for grading diabetic retinopathy.

[0071] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus (systems), or computer program products. Therefore, the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of computer program products embodied on one or more computer-readable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-executable program code. The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flowchart illustration and / or block, and combinations of flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to create a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce an implementation for the process. Figure 1 One or more processes and / or frames Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or frames Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or frames Figure 1 Figure 1 The steps of the function specified in one or more boxes.

[0072] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A heterogeneous dual-stream fusion method for grading diabetic retinopathy, characterized in that, This involves using a pre-trained neural network model and a lightweight feature extractor obtained through a composite knowledge distillation scheme to process a large amount of eye image information. A symmetrical bidirectional cross-attention fusion module then fuses the outputs of these two different model architectures to obtain the diagnostic results for retinal disease grading. The heterogeneous dual-stream fusion model generates the retinal disease grading diagnostic results through the following steps: S1: Obtain the input eye image and preprocess the image to different resolutions; the 224×224 pixel image is fed into the ViT branch through the composite knowledge distillation scheme, and the 456×456 pixel image is fed into the CNN branch; S2: Parallel branches extract information from different dimensions; the ViT branch extracts global features from low-resolution images; the CNN branch extracts local lesion features from high-resolution images. S3: Feed the feature vectors of ViT and CNN into the vector cross-attention module; perform bidirectional information interaction and enhancement within this module to generate deeply fused features; S4: The fused feature vectors are fed into the multilayer perceptron classification head, which integrates multiple feature vectors and maps the high-dimensional features to the final category space to output the final prediction result.

2. The heterogeneous dual-stream fusion method for grading diabetic retinopathy according to claim 1, characterized in that, Different branches have corresponding optimal resolution requirements. Preprocessing the input image information into different resolutions can improve the efficiency of feature extraction by the model. The ViT branch mentioned in S1 is obtained through a composite knowledge distillation scheme, specifically including: S1.1 utilizes soft target distillation technology to enable the student model to learn the teacher model's accurate final prediction probability distribution and capture the hidden knowledge between categories; S1.2 utilizes feature distillation technology to guide the student model to simulate the robust feature representation learned by the teacher model in the intermediate layer, and performs dimension alignment and matching through the adaptation layer; S1.3 utilizes contrastive learning distillation techniques to transfer the teacher model's understanding of high-level structural relationships between data samples, namely similarity and dissimilarity; Finally, the losses from soft target distillation, feature distillation, and contrastive learning distillation are combined with the standard supervised learning loss, i.e., the standard cross-entropy loss, through a joint loss function. We perform weighted combinations to transfer the deep knowledge of the RETFound encoder from multiple dimensions to a lightweight student model.

3. The heterogeneous dual-stream fusion method for grading diabetic retinopathy according to claim 1, characterized in that, In S2, the parallel branch extracts information from different dimensions, while the ViT branch extracts global features from the low-resolution image. In terms of extracting local lesion features from high-resolution images, the CNN branch specifically includes: S2.1, the ViT branch extracts global features from low-resolution images: the large RETFound model is migrated to the lightweight ViT model through composite knowledge distillation, and its number of parameters is greatly compressed to 327 million, which is 8.8% of the original RETFound parameters. It efficiently extracts global contextual features from 224×224 low-resolution images, meeting the requirements of lightweight deployment. S2.2, the CNN branch extracts local lesion features from high-resolution images: The EfficientNet-B5 CNN encoder, pre-trained on ImageNet, is used to process 456×456 high-resolution images. It leverages the inherent local induction bias of CNNs to focus on fine local lesion features, thus compensating for the shortcomings of ViT in detail analysis.

4. The heterogeneous dual-stream fusion method for grading diabetic retinopathy according to claim 1, characterized in that, In S3, features extracted from different branches are encouraged to interact deeply, allowing two feature paths to act as queries, actively seeking the required information from each other's features, thereby generating enhanced features with higher density and stronger discriminative power. Specifically, it includes: S3.1, Global Features As a query, from local features The process extracts relevant local details; the attention output generated in this process is added to the original global features through residual connections, and then stabilized through a normalization layer, finally yielding a global feature that absorbs local details and is enhanced. This process can be formally represented as: ; S3.2, symmetrically, local features As a query, from global features The process of finding the context environment involves residual connections and layer normalization, resulting in an enhanced local feature that integrates global context awareness. This process can be formally represented as: 。 5. The heterogeneous dual-stream fusion method for grading diabetic retinopathy according to claim 1, characterized in that, In S4, the fused feature vectors are fed into the multilayer perceptron classification head, which integrates multiple feature vectors and maps the high-dimensional features to the final class space, outputting the final prediction result. Specifically, this includes: The data is fed into a classification head consisting of a multilayer perceptron (MLP) containing two hidden layers with ReLU activation, having dimensions of 512 and 256 respectively, for learning purposes. A non-linear mapping between the score and the category label; outputs an M-dimensional original score vector. The Softmax function is used to transform z into a class probability distribution. The entire model is trained in an end-to-end manner, and its optimization objective is to minimize the predicted probability distribution. With real labels Classification cross-entropy loss ; The final diagnosis of diabetic retinopathy is obtained by converting the data into a class probability distribution using the Softmax function.

6. A heterogeneous dual-stream fusion system for grading diabetic retinopathy, comprising an interconnected microprocessor and a memory, characterized in that, The microprocessor is programmed or configured to perform the heterogeneous dual-stream fusion method for grading diabetic retinopathy as described in any one of claims 1 to 5.

7. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the heterogeneous dual-stream fusion method for grading diabetic retinopathy as described in any one of claims 1 to 5.

8. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the heterogeneous dual-stream fusion method for grading diabetic retinopathy as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Transform-based fine-grained image classification method

    CN114676776A

  • Interactive attention-based eye fundus image processing method, system and terminal

    CN119963879A

  • Dual-mode guided interactive diffusion medical image segmentation method

    CN120219735A

  • Abdominal aortic aneurysm progress prediction method and system based on cross-modal knowledge distillation

    CN120809238A

  • Deep fake face image detection method based on double-flow CNN and ViT hybrid model

    CN120833637A

Cited By

  • Pulmonary arterial hypertension detection method fusing image features and tricuspid regurgitation velocity

    CN121304657A

  • Electric power safety supervision image detection method and system

    CN121640375A

  • Automatic screening system for fundus color illumination sugar net lesions based on image discrimination

    CN122023950A

  • Image recognition-based fundus color photograph retinopathy screening system

    CN122023950B

  • Diabetic retinopathy ultra-widefield fundus image grading method

    CN122391766A