Eye fundus image multi-disease combined diagnosis system based on deep learning

By designing a multi-disease joint diagnosis system for fundus images based on deep learning, the problem of joint diagnosis of multiple diseases in the prior art is solved, efficient and accurate diagnosis of diabetic retinopathy and diabetic macular edema is achieved, and the interpretability of the system is enhanced.

CN120072271APending Publication Date: 2025-05-30SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510221869.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing deep learning methods mainly diagnose a single disease, making it difficult to effectively diagnose multiple diseases of diabetic retinopathy and diabetic macular edema. Inadequate information sharing, insufficient extraction of fine-grained lesions features is insufficient, and low interpretability is low.

Method used

A joint diagnosis system for fundus image multidiseases based on deep learning is designed, and a multi-task learning framework is adopted to extract disease-related features through a shared backbone network module, a unique lesion feature is dynamically captured using an adaptive lesion attention module, and an intrinsic correlation between diseases is learned through the deep feature fusion module, and finally the disease level is output through the disease rating module.

Benefits of technology

It improves the accuracy and efficiency of the combined diagnosis of multiple diseases of diabetic retinopathy and diabetic macular edema, reduces diagnostic errors, enhances the interpretability of the system, and can more accurately extract the fine-grained characteristics of the lesion area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072271A_ABST
    Figure CN120072271A_ABST
Patent Text Reader

Abstract

The invention discloses an eye fundus image multi-disease combined diagnosis system based on deep learning, and the system comprises a data importing module which is used for loading eye fundus image data and carrying out the preprocessing of the data; the combined diagnosis training module is used for training a fundus image multi-disease combined diagnosis model, extracting disease-related feature maps through the shared backbone network module, dynamically capturing unique lesion features of diabetic retinopathy and diabetic macular edema by using the two adaptive lesion attention modules, and outputting the unique lesion features of the diabetic retinopathy and the diabetic macular edema. Learning the correlation of the two through a deep feature fusion module, and performing optimization training by using a four-branch weighted composite loss function to obtain an optimal multi-disease combined diagnosis model; and the combined diagnosis and prediction module is used for outputting the disease levels of multiple diseases and providing a visual thermodynamic diagram. According to the method, combined diagnosis of diabetic retinopathy and diabetic macular edema of complications of the diabetic retinopathy is realized by using deep learning, the screening efficiency is improved, the diagnosis error is reduced, and the interpretability of the system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and medical image analysis, and in particular to a multi-disease joint diagnosis system for fundus images based on deep learning. Background Art

[0002] With the rapid development of modern medical imaging technology, fundus image analysis has become an important auxiliary tool in ophthalmic diagnosis. Diabetic retinopathy is a very common diabetic eye complication and has been widely recognized as a major factor seriously threatening human vision. At any stage of the development of diabetic retinopathy, diabetic macular edema may occur, which is a common complication of diabetic retinopathy. Approximately one-third of patients with diabetic retinopathy show symptoms of diabetic macular edema. Because the early disease has no obvious symptoms, it is easily ignored by patients. By the time the patient feels obvious abnormalities and seeks medical treatment, irreversible vision damage often has occurred. Therefore, it is very important to perform intervention treatment before substantial and irreversible vision loss occurs in the fundus. Due to the large number of patients with eye diseases, the uneven distribution of existing medical resources, and the limited number of professional ophthalmologists, large-scale manual screening cannot be achieved. In recent years, automated fundus image analysis methods based on deep learning have gradually become the mainstream technology to solve this problem. Convolutional neural networks have achieved remarkable results in medical image analysis due to their strong feature learning ability. However, current deep learning methods usually diagnose single diseases, and there is still relatively little research on multi-disease joint diagnosis. Since diabetic macular edema is a complication of diabetic retinopathy and often coexists, and their progress affects each other, there is a strong internal correlation between the image features of the two. It is necessary to develop a multi-disease joint diagnosis system that can simultaneously diagnose the degree of diabetic retinopathy and diabetic macular edema. In terms of multi-disease joint diagnosis, although some studies have explored the possibility of joint diagnosis, the following problems still exist: 1. The information sharing between the two diseases is insufficient; 2. The extraction of fine-grained lesion features is not precise enough; 3. The interpretability is relatively low. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a multi-disease joint diagnosis system for fundus images based on deep learning, aiming to accurately identify and diagnose the degree of diabetic retinopathy and diabetic macular edema by automatically processing fundus images, improve the screening efficiency, reduce the diagnostic error, and enhance the interpretability of the system.

[0004] To achieve the above purpose, the technical solution provided by the present invention is: A multi-disease joint diagnosis system for fundus images based on deep learning, comprising:

[0005] A data import module for loading fundus image data, including fundus color digital photos and corresponding labels of the degree of diabetic retinopathy and diabetic macular edema, and preprocessing the fundus image data to obtain enhanced and standardized fundus image data;

[0006] A joint diagnosis training module for training a multi-disease joint diagnosis model for fundus images. This model is based on a multi-task learning framework. It extracts disease-related feature maps of fundus images through a shared backbone network module. Two adaptive lesion attention modules are used to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema respectively. The adaptive lesion attention module adopts an adaptive weighted fusion mechanism, combines channel attention and spatial attention, and accurately extracts the fine-grained features of the lesion area, thereby improving the effectiveness of lesion feature extraction. And a deep feature fusion module is used to learn the internal correlation between diabetic retinopathy and diabetic macular edema. The deep feature fusion module uses a local attention mechanism with shared parameters to learn disease-related deep features, and effectively captures the internal correlation between diseases through feature fusion. Finally, two disease grading modules output the disease grades of diabetic retinopathy and diabetic macular edema, realizing the multi-disease joint diagnosis of fundus images; To further optimize the model's ability to handle multiple tasks, a four-branch weighted composite loss function is used, including cross-entropy losses for learning the unique lesion features of the two diseases and cross-entropy loss after deep feature fusion, guiding the optimizer to adjust the model parameters, and finally training the optimal multi-disease joint diagnosis model;

[0007] A joint diagnosis prediction module for applying the optimal multi-disease joint diagnosis model obtained by the joint diagnosis training module to actual prediction tasks. Input the fundus image data to be predicted into the optimal multi-disease joint diagnosis model, output the disease grades of diabetic retinopathy and diabetic macular edema, and visualize the lesion areas concerned by the model through a heat map to assist clinical diagnosis.

[0008] Furthermore, the data import module includes a data loading module and a data preprocessing module, where:

[0009] The data loading module reads fundus image data from local, including fundus images in tif format and annotation files in xls format. The fundus images are posterior pole fundus color digital photos taken by a color camera with a 45-degree field of view. The annotation file contains the names of fundus images, labels of the disease grades of diabetic retinopathy, and labels of the disease grades of diabetic macular edema;

[0010] The data preprocessing module uniformly adjusts the size of fundus images to 256×256, standardizes the input size, which helps improve the efficiency of batch processing and reduces the computational amount; then obtains an image size of 224×224 through central cropping, effectively reducing the interference of peripheral irrelevant regions on the model and ensuring the integrity of important feature regions in the image, providing a more accurate input; adopts data augmentation methods, including random horizontal flipping, random vertical flipping, color jittering, and random rotation, to increase the diversity of fundus image data, alleviate the overfitting problem, and improve the generalization ability of the model; finally, performs normalization processing so that each pixel value is within a unified numerical range, which helps improve the stability of the training process and accelerates the convergence speed of the model.

[0011] Furthermore, the joint diagnosis training module includes a backbone network module, an adaptive lesion attention module, a deep feature fusion module, a disease grading module, and a four-branch weighted composite loss function, where:

[0012] The backbone network module is used to initially extract disease-related feature maps, which is constructed based on the convolutional neural network DenseNet-121. The convolutional neural network DenseNet-121 establishes the connection relationship between different layers through a dense connection structure, fully utilizes the feature information, alleviates the gradient disappearance phenomenon, and can effectively extract low-level to high-level features in fundus images; to reduce the complexity of the neural network and alleviate the overfitting phenomenon, the Dropout regularization technique is incorporated into the backbone network module. Finally, a preliminary disease-related feature map F is generated.

[0013] The adaptive lesion attention module is used to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema, including the following steps:

[0014] 1) Strengthen the attention to important channel features:

[0015] The preliminary disease-related feature map F extracted by the backbone network module is used as the input of the adaptive lesion attention module. The global maximum pooling layer and the global average pooling layer are used to capture the global information of the feature map F, and its spatial dimension is reduced to 1×1; for the outputs of the two poolings, a multi-layer perceptron is used to increase the non-linear mapping, enabling the model to learn more complex mapping relationships, thereby enhancing the expression ability of the features, and respectively obtaining the global feature maps F Max and F Avg ; an adaptive weighted fusion mechanism is adopted to obtain the channel attention weight A C_att , this mechanism dynamically adjusts the contributions of the max-pooling and average-pooling strategies by learning a trainable parameter α, with an initial value set to 0.5, and is optimized through the backpropagation algorithm during the training process, so that the model can adaptively adjust the weighted ratio of the pooled features according to the input data. This strategy is expressed by the following formula:

[0016]

[0017] In the formula, σ represents the Sigmoid activation function, which is used to generate the channel attention weight in the range of [0, 1]. represents the element-wise addition of two feature maps at the same position, and the generated channel attention weight A C_att is multiplied element-wise by the preliminary disease-related feature map F along the channels to obtain the weighted feature map F';

[0018] 2) Enhance the attention to important spatial features:

[0019] The weighted feature map F' is subjected to global max pooling and global average pooling operations along the width and height respectively, and the information fusion degree of the two pooling strategies is dynamically adjusted through the trainable parameter β with an initial value of 0.5 to obtain the weighted feature maps F h ' and F w ', as shown in the following formula:

[0020]

[0021] In the formula, GAP represents global average pooling, GMP represents global max pooling, and the subscripts w and h represent along the width and height directions; the weighted results are concatenated to jointly model the spatial attention information in the height and width directions, and after 1×1 convolution and batch normalization to accelerate the model convergence and stabilize the training process, and then through a non-linear activation function to generate the spatial feature map F S ':

[0022] F' S = δ(BN(Conv(Concat(F h ', F w '))))

[0023] In the formula, δ represents the non-linear activation function, BN represents the batch normalization layer, Conv represents the 1×1 convolution layer, and Concat represents the concatenation operation; the spatial feature map F S ' is split to obtain two tensors in the height and width directions, and respectively pass through 1×1 convolution and Sigmoid activation function to obtain the spatial attention weights A' S_h and A' S_w in the height and width directions; finally, the feature map F' is multiplied element-wise by the spatial attention weights A' S_h and A' S_w in the height and width directions to finely control the importance of features in different spatial dimensions of the image, and respectively obtain the unique lesion feature map F DR of diabetic retinopathy and the unique lesion feature map F of diabetic macular edemaDME ;

[0024] The depth feature fusion module is used to learn the intrinsic correlation between diabetic retinopathy and diabetic macular edema, including the following steps:

[0025] 1) Obtain a general feature representation across tasks:

[0026] The unique lesion feature map F of diabetic retinopathy obtained by the adaptive lesion attention module DR and the unique lesion feature map F of diabetic macular edema DME pass through the global average pooling layer and the fully connected layer to obtain the input feature map F D ′ R and F D ′ ME of the depth feature fusion module, and then obtain their respective attention weights A′ DR_att and A′ DME_att through the local attention mechanism with shared parameters, which helps the model learn the regions jointly concerned by diabetic retinopathy and diabetic macular edema, and assigns higher weights to these regions, so that the model can obtain a general feature representation across tasks, especially the regions jointly affected by these diseases:

[0027] A′ DR_att = σ(ConB(ReLU(ConB(F D ′ R ))))

[0028] A′ DME_att = σ(ConB(ReLU(ConB(F D ′ ME ))))

[0029] In the formula, ConB means performing 1×1 convolution first and then batch normalization. The first convolution is used for dimensionality reduction, reducing the number of channels from 1024 to 512, reducing the computational amount and memory consumption, and at the same time compressing the feature information to extract more expressive features. The second convolution increases the number of channels from 512 to 1024 for subsequent operations. ReLU represents a non-linear activation function, which is used to alleviate the problem of gradient disappearance and accelerate the training process;

[0030] 2) Perform feature fusion operations:

[0031] Multiply the obtained attention weights A′ DR_att and A′ DME_att with the input feature map F′ DR and F D ′ MEMultiply them so that the model can focus on the features useful for the task, thereby strengthening the information of the relevant features. Then, respectively correspond the strengthened feature maps to the feature maps F of the input depth feature fusion module D ′ R and F D ′ ME Perform an addition operation to obtain the fused feature map F D ′ R_fus and F D ′ ME_fus , enhancing the model's learning of the common features of diabetic retinopathy and diabetic macular edema lesions, and capturing the internal correlation between the two:

[0032]

[0033] In the formula, represents element-wise multiplication;

[0034] The disease grading module is used to output the disease grades of diabetic retinopathy and diabetic macular edema, realizing the joint diagnosis of multiple diseases in fundus images. It contains two parallel fully connected layers, which are respectively used for the grading tasks of different diseases. The feature maps F D ′ R_fus and F D ′ ME_fus are used as the inputs of the fully connected layers, and the grading results of different diseases are output respectively;

[0035] The four-branch weighted composite loss function is used to improve the model's ability to handle multiple tasks, so as to obtain the optimal multi-disease joint diagnosis model:

[0036] WCLoss = λ(L DR + L DME ) + (1 - λ)(L' DR + L' DME )

[0037] In the formula, WCLoss represents the four-branch weighted composite loss function, L DR and L DME respectively represent the cross-entropy losses after the depth feature fusion of diabetic retinopathy and diabetic macular edema. L' DR and L' DME respectively represent the cross-entropy losses for learning the unique lesion features of diabetic retinopathy and diabetic macular edema. The coefficient λ controls the weight distribution of each loss term, and its value range is [0, 1], adjusting the degree of attention of the model to the losses of each task during the multi-task learning process. The aim is to balance the independent learning of the unique features of diabetic retinopathy and diabetic macular edema and the correlation learning of the fused features, so as to improve the overall performance of the model. The model is trained using the backpropagation algorithm to obtain the optimal multi-disease joint diagnosis model.

[0038] Furthermore, the combined diagnosis prediction module includes the following steps:

[0039] 1) Use the data loading module in the data import module to read the fundus image data to be predicted from the local, and use the data preprocessing module to process the fundus image data to be predicted;

[0040] 2) Input the fundus image data processed in step 1) into the optimal multi-disease combined diagnosis model trained by the combined diagnosis training module to diagnose the disease levels of diabetic retinopathy and diabetic macular edema, and use the Grad-CAM method to generate a heat map of the model's attention area, intuitively showing the decision-making basis of the multi-disease combined diagnosis model, and enhancing the interpretability and credibility in medical diagnosis.

[0041] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0042] 1. A new multi-disease combined diagnosis model for fundus images is designed. Based on the multi-task learning framework, it realizes the multi-disease combined grading diagnosis of diabetic retinopathy and diabetic macular edema.

[0043] 2. The designed combined diagnosis model extracts disease-related feature maps through a shared backbone network module, uses two adaptive lesion attention modules to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema respectively, and learns the internal correlation between diabetic retinopathy and diabetic macular edema through a deep feature fusion module.

[0044] 3. A new four-branch weighted composite loss function adapted to multi-task learning is designed for training to improve the model's ability to handle multi-tasks;

[0045] 4. Provide a heat map to visualize the areas concerned by the model, intuitively show the decision-making basis of the multi-disease combined diagnosis model, and enhance the interpretability and credibility of the multi-disease combined diagnosis system in medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the relationship between the various modules of the system of the present invention; in the figure, DR represents diabetic retinopathy, and DME represents diabetic macular edema.

[0047] Figure 2 It is a flowchart of the training and prediction of the system of the present invention.

[0048] Figure 3It is the overall network structure diagram used in the system of the present invention; in the figure, GAP represents global average pooling, FC represents fully connected layer, Conv represents convolutional layer, BN represents batch normalization layer, ReLU represents non-linear activation function, MaxPool represents max pooling, Dense Block represents dense connection block, Transition represents transition layer, L DR and L DME respectively represent the cross-entropy losses after the depth feature fusion of diabetic retinopathy and diabetic macular edema. L′ DR and L′ DME respectively represent the cross-entropy losses for learning the unique lesion features of diabetic retinopathy and diabetic macular edema.

[0049] Figure 4 It is the structure diagram of the adaptive lesion attention module used in the system of the present invention; in the figure, GMP represents global max pooling, GAP represents global average pooling, MLP represents multi-layer perceptron, represents the element-wise addition of two feature maps at the same position. Sigmoid represents the activation function, which is used to generate the attention weight in the range of [0,1]. represents element-wise multiplication. The subscripts w and h represent along the width and height directions. Concat represents the concatenation operation. Conv represents the convolutional layer. BN represents the batch normalization layer. Non-linear represents the non-linear activation function.

[0050] Figure 5 It is the structure diagram of the depth feature fusion module used in the system of the present invention; in the figure, Conv represents the convolutional layer, BN represents the batch normalization layer, ReLU represents the non-linear activation function, Sigmoid represents the activation function, which is used to generate the attention weight in the range of [0,1]. represents element-wise multiplication. represents the element-wise addition of two feature maps at the same position. Detailed implementation manners

[0051] The present invention will be further described below in conjunction with specific embodiments.

[0052] This embodiment discloses a multi-disease joint diagnosis system for fundus images based on deep learning, which is a multi-disease joint diagnosis system developed using the Python language and can run on Windows devices. The relationship between the modules of the system is as Figure 1 shown, and the flowcharts of system training and prediction are as Figure 2 shown. It includes:

[0053] A data import module, used to load fundus image data, including fundus color digital photos and corresponding labels of the degree of diabetic retinopathy and diabetic macular edema, and preprocess the fundus image data to obtain enhanced and standardized fundus image data;

[0054] A joint diagnosis training module, used to train a multi-disease joint diagnosis model for fundus images. As Figure 3 shown, this model is based on a multi-task learning framework. It extracts disease-related feature maps of fundus images through a shared backbone network module. Two adaptive lesion attention modules are respectively used to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema. The adaptive lesion attention module adopts an adaptive weighted fusion mechanism, combines channel attention and spatial attention, and accurately extracts the fine-grained features of the lesion area, thereby improving the effectiveness of lesion feature extraction. And a deep feature fusion module is used to learn the internal correlation between diabetic retinopathy and diabetic macular edema. The deep feature fusion module uses a local attention mechanism with shared parameters to learn disease-related deep features, and effectively captures the internal correlation between diseases through feature fusion. Finally, two disease grading modules output the disease grades of diabetic retinopathy and diabetic macular edema, realizing the multi-disease joint diagnosis of fundus images. To further optimize the model's ability to handle multiple tasks, a four-branch weighted composite loss function is used, including the cross-entropy loss for learning the unique lesion features of the two diseases and the cross-entropy loss after deep feature fusion, guiding the optimizer to adjust the model parameters, and finally training the optimal multi-disease joint diagnosis model;

[0055] A joint diagnosis prediction module, used to apply the optimal multi-disease joint diagnosis model obtained by the joint diagnosis training module to actual prediction tasks, input the fundus image data to be predicted into the optimal multi-disease joint diagnosis model, output the disease grades of diabetic retinopathy and diabetic macular edema, and visualize the lesion areas concerned by the model through a heat map to assist clinical diagnosis.

[0056] Specifically, the data import module includes a data loading module and a data preprocessing module, where:

[0057] The data loading module reads fundus image data from local, including fundus images in tif format and annotation files in xls format. The fundus images are posterior pole fundus color digital photos taken by a color camera with a 45-degree field of view. The annotation files contain the names of fundus images, the labels of the disease grades of diabetic retinopathy, and the labels of the disease grades of diabetic macular edema;

[0058] The data preprocessing module uniformly adjusts the fundus image size to 256×256, normalizes the input size, helps to improve the efficiency of batch processing and reduce the amount of calculation; then obtains an image size of 224×224 through center cropping, effectively reduces the interference of peripheral irrelevant areas on the model, and ensures the integrity of important feature areas in the image, providing more accurate input; adopts data enhancement methods, including random horizontal flipping, random vertical flipping, color jittering, random rotation and other operations, to increase the diversity of fundus image data, alleviate overfitting problems, and improve the generalization ability of the model; finally, normalization processing is performed so that each pixel value is within a uniform numerical range, which helps to improve the stability of the training process and accelerate the convergence speed of the model.

[0059] Specifically, the joint diagnosis training module includes a backbone network module, an adaptive lesion attention module, a deep feature fusion module, a disease classification module, and a four-branch weighted composite loss function, wherein:

[0060] The backbone network module is used to preliminarily extract disease-related feature maps, and is built based on the convolutional neural network DenseNet-121. The convolutional neural network DenseNet-121 establishes connection relationships between different layers through a dense connection structure, fully utilizes feature information, alleviates the gradient vanishing phenomenon, and can effectively extract low-level to high-level features in fundus images. In order to reduce the complexity of the neural network and alleviate the overfitting phenomenon, the Dropout regularization technology is integrated into the backbone network module, and finally, a preliminary disease-related feature map F is generated;

[0061] like Figure 4 As shown, the adaptive lesion attention module is used to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema, including the following steps:

[0062] 1) Strengthen attention to important channel characteristics:

[0063] The preliminary disease-related feature map F extracted by the backbone network module is used as the input of the adaptive lesion attention module. After the global maximum pooling layer and the global average pooling layer, the global information of the feature map F is captured, and its spatial dimension is reduced to 1×1. For the two pooled outputs, a multi-layer perceptron is used to increase nonlinear mapping, so that the model can learn more complex mapping relationships, thereby enhancing the expression ability of features, and obtaining the global feature maps F and F. Max and F Avg ; An adaptive weighted fusion mechanism is used to obtain the channel attention weight A C_att, this mechanism dynamically adjusts the contributions of the max pooling and average pooling strategies by learning a trainable parameter α with an initial value of 0.5, and optimizes it through the backpropagation algorithm during training, enabling the model to adaptively adjust the weighted ratio of the pooled features according to the input data. This strategy can be expressed by the following formula:

[0064]

[0065] In the formula, σ represents the Sigmoid activation function, which is used to generate the channel attention weights in the range of [0, 1]. represents the element-wise addition of two feature maps at the same position, and the generated channel attention weights A C_att are multiplied element-wise by the preliminary disease-related feature map F along the channels to obtain the weighted feature map F';

[0066] 2) Enhance the attention to important spatial features:

[0067] The weighted feature map F' is subjected to global max pooling and global average pooling operations along the width and height respectively, and the information fusion degree of the two pooling strategies is dynamically adjusted through a trainable parameter β with an initial value of 0.5 to obtain the weighted feature maps F h ' and F w ', as shown in the following formula:

[0068]

[0069] In the formula, GAP represents global average pooling, GMP represents global max pooling, and the subscripts w and h represent along the width and height directions. represents the element-wise addition of two feature maps at the same position; the weighted results are concatenated to jointly model the spatial attention information in the height and width directions, and after 1×1 convolution and batch normalization to accelerate the model convergence and stabilize the training process, and then through a non-linear activation function to generate the spatial feature map F S :

[0070] F' S = δ(BN(Conv(Concat(F' h , F' w ))))

[0071] In the formula, δ represents the non-linear activation function, BN represents the batch normalization layer, Conv represents the 1×1 convolution layer, and Concat represents the concatenation operation; the spatial feature map F' S is segmented to obtain two tensors in the height and width directions, and respectively passes through 1×1 convolution and Sigmoid activation function to obtain the spatial attention weights A' S_h and A' S_w; Finally, multiply the feature map F′ element - by - element with the spatial attention weights A′ S_h and A′ S_w in the height and width directions, which finely controls the importance of features in different spatial dimensions of the image, and respectively obtains the unique lesion feature map F DR of diabetic retinopathy and the unique lesion feature map F DME of diabetic macular edema;

[0072] As Figure 5 shown, the depth feature fusion module is used to learn the internal correlation between diabetic retinopathy and diabetic macular edema, including the following steps:

[0073] 1) Obtain a general - purpose feature representation across tasks:

[0074] The unique lesion feature map F DR of diabetic retinopathy and the unique lesion feature map F DME of diabetic macular edema obtained by the adaptive lesion attention module pass through the global average pooling layer and the fully - connected layer to obtain the input feature maps F D ′ R and F D ′ ME of the depth feature fusion module, and then pass through the local attention mechanism with shared parameters to obtain their respective attention weights A′ DR_att and A′ DME_att , which helps the model learn the regions that diabetic retinopathy and diabetic macular edema jointly focus on, and assigns higher weights to these regions, so that the model can obtain a general - purpose feature representation across tasks, especially the regions jointly affected by these diseases:

[0075] A′ DR_att =σ(ConB(ReLU(ConB(F D ′ R ))))

[0076] A′ DME_att =σ(ConB(ReLU(ConB(F D ′ ME ))))

[0077] In the formula, σ represents the Sigmoid activation function, which is used to generate attention weights in the range of [0,1]. ConB represents performing 1×1 convolution first and then batch normalization. The first convolution is used for dimensionality reduction, reducing the number of channels from 1024 to 512, reducing the computational amount and memory consumption, while compressing feature information and extracting more expressive features. The second convolution increases the number of channels from 512 to 1024 for subsequent operations. ReLU represents the non - linear activation function, which is used to alleviate the vanishing gradient problem and accelerate the training process;

[0078] 2) Perform feature fusion operation:

[0079] Multiply the obtained attention weights A′ DR_att and A′ DME_att with the input feature maps F D ′ R and F D ′ ME of the deep feature fusion module, so that the model can focus on the features useful for the task, thereby strengthening the information of the relevant features, and then respectively correspond the strengthened feature maps to the feature maps F D ′ R and F D ′ ME of the input deep feature fusion module for addition operation to obtain the fused feature maps F D ′ R_fus and F D ′ ME_fus , enhancing the model's learning of the common features of diabetic retinopathy and diabetic macular edema lesions, and capturing the internal correlation between the two:

[0080]

[0081] In the formula, represents element-wise multiplication, represents the addition of elements at the same position of two feature maps;

[0082] The disease grading module is used to output the disease grades of diabetic retinopathy and diabetic macular edema, realizing the multi-disease joint diagnosis of fundus images. It contains two parallel fully connected layers, which are respectively used for the grading tasks of different diseases. The feature maps F D ′ R_fus and F D ′ ME_fus are used as the inputs of the fully connected layers, and the grading results of different diseases are output respectively;

[0083] The four-branch weighted composite loss function is used to improve the model's ability to handle multiple tasks, so as to obtain the optimal multi-disease joint diagnosis model:

[0084] WCLoss = λ(L DR + L DME ) + (1 - λ)(L′ DR + L′ DME )

[0085] In the formula, WCLoss represents the four-branch weighted composite loss function, L DR and L DMErespectively represent the cross-entropy loss after the depth feature fusion of diabetic retinopathy and diabetic macular edema, \(L'\) DR and \(L'\) DME respectively represent the cross-entropy loss for learning the unique lesion features of diabetic retinopathy and diabetic macular edema. The coefficient \(\lambda\) controls the weight distribution of each loss term, and its value range is \([0, 1]\), which adjusts the attention degree of the model to the losses of each task during the multi-task learning process. This design aims to balance the independent learning of the unique features of diabetic retinopathy and diabetic macular edema and the correlation learning of the fusion features, so as to improve the overall performance of the model. The model is trained using the backpropagation algorithm to obtain the optimal multi-disease joint diagnosis model.

[0086] Specifically, the joint diagnosis prediction module includes the following steps:

[0087] 1) Use the data loading module in the data import module to read the fundus image data to be predicted from the local, and use the data preprocessing module to process the fundus image data to be predicted;

[0088] 2) Input the fundus image data processed in step 1) into the optimal multi-disease joint diagnosis model trained by the joint diagnosis training module to diagnose the disease levels of diabetic retinopathy and diabetic macular edema, and use the Grad-CAM method to generate a heat map of the model's attention area, visually displaying the decision-making basis of the multi-disease joint diagnosis model and enhancing the interpretability and credibility in medical diagnosis.

[0089] The above embodiments are only the preferred embodiments of the present invention, and do not limit the implementation scope of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A multi-disease joint diagnosis system for fundus images based on deep learning, characterized in that: include: A data import module is used to load fundus image data, including color digital fundus photos and corresponding diabetic retinopathy and diabetic macular edema disease severity labels, and preprocess the fundus image data to obtain enhanced standardized fundus image data; The joint diagnosis training module is used to train the fundus image multi-disease joint diagnosis model. The model is based on a multi-task learning framework. The disease-related feature map of the fundus image is extracted through a shared backbone network module. Two adaptive lesion attention modules are used to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema respectively. The adaptive lesion attention module adopts an adaptive weighted fusion mechanism, combined with channel attention and spatial attention, to accurately extract fine-grained features of the lesion area, thereby improving the effectiveness of lesion feature extraction, and learns the intrinsic correlation between diabetic retinopathy and diabetic macular edema through a deep feature fusion module. The deep feature fusion module uses a local attention mechanism with shared parameters to learn disease-related deep features and effectively captures the intrinsic correlation between diseases through feature fusion. Finally, two disease grading modules output the disease grades of diabetic retinopathy and diabetic macular edema to achieve fundus image multi-disease joint diagnosis. In order to further optimize the model's ability to handle multiple tasks, a four-branch weighted composite loss function is used, including the cross entropy loss learned for the unique lesion features of the two diseases and the cross entropy loss after deep feature fusion, to guide the optimizer to adjust the model parameters, and finally train the optimal multi-disease joint diagnosis model. The joint diagnosis prediction module is used to apply the optimal multi-disease joint diagnosis model obtained by the joint diagnosis training module to actual prediction tasks, input the fundus image data to be predicted into the optimal multi-disease joint diagnosis model, output the disease levels of diabetic retinopathy and diabetic macular edema, and visualize the lesion area of ​​concern of the model through a heat map to assist clinical diagnosis.

2. The deep learning-based fundus image multi-disease joint diagnosis system according to claim 1, characterized in that: The data import module includes a data loading module and a data preprocessing module, wherein: The data loading module reads fundus image data from the local computer, including fundus images in tif format and annotation files in xls format, wherein the fundus image is a color digital photograph of the posterior pole fundus taken with a color camera having a 45-degree field of view, and the annotation file includes the fundus image name, diabetic retinopathy disease grade label and diabetic macular edema disease grade label; The data preprocessing module uniformly adjusts the fundus image size to 256×256, normalizes the input size, helps to improve the efficiency of batch processing and reduce the amount of calculation; then obtains an image size of 224×224 through center cropping, effectively reduces the interference of peripheral irrelevant areas on the model, and ensures the integrity of important feature areas in the image, providing more accurate input; adopts data enhancement methods, including random horizontal flipping, random vertical flipping, color jittering and random rotation, to increase the diversity of fundus image data, alleviate overfitting problems, and improve the generalization ability of the model; finally, normalization is performed so that each pixel value is within a uniform numerical range, which helps to improve the stability of the training process and accelerate the convergence speed of the model.

3. The deep learning-based fundus image multi-disease joint diagnosis system according to claim 2, characterized in that: The joint diagnosis training module includes a backbone network module, an adaptive lesion attention module, a deep feature fusion module, a disease classification module and a four-branch weighted composite loss function, wherein: The backbone network module is used to preliminarily extract disease-related feature maps, and is built based on the convolutional neural network DenseNet-121. The convolutional neural network DenseNet-121 establishes connection relationships between different layers through a dense connection structure, fully utilizes feature information, alleviates the gradient vanishing phenomenon, and can effectively extract low-level to high-level features in fundus images. In order to reduce the complexity of the neural network and alleviate the overfitting phenomenon, the Dropout regularization technology is integrated into the backbone network module, and finally, a preliminary disease-related feature map F is generated; The adaptive lesion attention module is used to dynamically capture the unique lesion features of diabetic retinopathy and diabetic macular edema, including the following steps: 1) Strengthen attention to important channel characteristics: The preliminary disease-related feature map F extracted by the backbone network module is used as the input of the adaptive lesion attention module. After the global maximum pooling layer and the global average pooling layer, the global information of the feature map F is captured, and its spatial dimension is reduced to 1×1. For the two pooled outputs, a multi-layer perceptron is used to increase nonlinear mapping, so that the model can learn more complex mapping relationships, thereby enhancing the expression ability of features, and obtaining the global feature maps F and F. Max and F Avg ; An adaptive weighted fusion mechanism is used to obtain the channel attention weight A C_att , this mechanism dynamically adjusts the contribution of the maximum pooling and average pooling strategies by learning a trainable parameter α, with an initial value of 0.5, and optimizes it through the back-propagation algorithm during the training process, so that the model can adaptively adjust the weighted ratio of the pooling features according to the input data. The strategy is expressed as follows: In the formula, σ represents the Sigmoid activation function, which is used to generate the channel attention weight of [0,1]. It means that the elements at the same position of the two feature maps are added together, and the generated channel attention weight A C_att Multiply the initial disease-related feature map F element by channel to obtain the weighted feature map F′; 2) Strengthen attention to important spatial features: The weighted feature map F′ is subjected to global maximum pooling and global average pooling operations along the width and height respectively, and the information fusion degree of the two pooling strategies is dynamically adjusted through the trainable parameter β, with the initial value set to 0.5 to obtain the weighted feature map F h ′ and F w ′, as shown in the following formula: In the formula, GAP represents global average pooling, GMP represents global maximum pooling, and the subscripts w and h represent the width and height directions. The weighted results are concatenated to jointly model the spatial attention information in the height and width directions. After 1×1 convolution and batch normalization, the model convergence is accelerated and the training process is stabilized. Then, the spatial feature map F′ is generated through a nonlinear activation function. S : F′ S =δ(BN(Conv(Concat(F′ h ,F′ w )))) In the formula, δ represents a nonlinear activation function, BN represents a batch normalization layer, Conv represents a 1×1 convolutional layer, and Concat represents a concatenation operation; the spatial feature map F′ S The segmentation obtains two tensors in the height and width directions, and the spatial attention weights A′ in the height and width directions are obtained by 1×1 convolution and Sigmoid activation function respectively. S_h and A′ S_w ; Finally, the feature map F′ is combined with the spatial attention weights A′ in the height and width directions S_h and A′ S_w By multiplying each element, the importance of features in different spatial dimensions of the image is finely controlled, and the unique lesion feature map F of diabetic retinopathy is obtained respectively. DR Figure F shows the unique pathological features of diabetic macular edema DME ; The deep feature fusion module is used to learn the intrinsic correlation between diabetic retinopathy and diabetic macular edema, including the following steps: 1) Obtain universal feature representation across tasks: Unique lesion features of diabetic retinopathy obtained by adaptive lesion attention module F DR Figure F shows the unique pathological features of diabetic macular edema DME After the global average pooling layer and the fully connected layer, the input feature map F′ of the deep feature fusion module is obtained DR and F′ DME , and then obtain their respective attention weights A′ through the local attention mechanism of shared parameters DR_att and A′ DME_att , helping the model learn the areas of common interest for diabetic retinopathy and diabetic macular edema, assigning higher weights to these areas so that the model can obtain a universal feature representation across tasks, especially the areas commonly affected by these diseases: TO' DR_att =σ(ConB(ReLU(ConB(F D ′ R )))) TO' DME_att =σ(ConB(ReLU(ConB(F′ DME )))) In the formula, ConB means that 1×1 convolution is performed first and then batch normalization is performed. The first convolution is used for dimensionality reduction, reducing the number of channels from 1024 to 512, reducing the amount of calculation and memory consumption, while compressing feature information and extracting more expressive features. The second convolution increases the number of channels from 512 to 1024 for easy subsequent operations. ReLU represents a nonlinear activation function, which is used to alleviate the gradient vanishing problem and accelerate the training process. 2) Perform feature fusion operation: The acquired attention weight A′ DR_att and A′ DME_att The input feature map F of the deep feature fusion module D ' R and F D ' ME Multiplying them enables the model to focus on the features that are useful for the task, thereby strengthening the information of the relevant features, and then corresponding the enhanced feature maps to the feature maps F of the input deep feature fusion module D ' R and F D ' ME Perform the addition operation to obtain the fused feature map F D ' R_fus and F D ' ME_fus , enhance the model's common feature learning of diabetic retinopathy and diabetic macular edema features, and capture the intrinsic correlation between the two: In the formula, Represents element-wise multiplication; The disease classification module is used to output the disease grades of diabetic retinopathy and diabetic macular edema, and realize the joint diagnosis of multiple diseases in fundus images. It contains two parallel fully connected layers, which are used for the classification tasks of different diseases respectively. The feature map F′ DR_fus and F′ DME_fus As the input of the fully connected layer, the classification results of different diseases are output respectively; The four-branch weighted composite loss function is used to improve the model's ability to handle multiple tasks, thereby obtaining the optimal multi-disease joint diagnosis model: WCLoss=λ(L DR +L DME )+(1-λ)(L′ DR +L′ DME ) Where WCLoss represents the four-branch weighted composite loss function, L DR and L DME denote the cross entropy loss after deep feature fusion of diabetic retinopathy and diabetic macular edema, respectively, and L′ DR and L′ DME They represent the cross entropy loss for learning the unique lesion features of diabetic retinopathy and diabetic macular edema respectively. The coefficient λ controls the weight distribution of each loss term, and its value range is [0, 1]. It adjusts the degree of emphasis on the loss of each task in the multi-task learning process, aiming to balance the independent learning of the unique features of diabetic retinopathy and diabetic macular edema with the associative learning of the fusion features to improve the overall performance of the model. The model is trained using the back propagation algorithm to obtain the optimal multi-disease joint diagnosis model.

4. The deep learning-based fundus image multi-disease joint diagnosis system according to claim 3, characterized in that: The combined diagnosis prediction module comprises the following steps: 1) using the data loading module in the data import module to read the fundus image data to be predicted from the local, and using the data preprocessing module to process the fundus image data to be predicted; 2) The fundus image data processed in step 1) is input into the optimal multi-disease joint diagnosis model trained by the joint diagnosis training module to diagnose the disease grade of diabetic retinopathy and diabetic macular edema, and the Grad-CAM method is used to generate a heat map of the model's focus area to intuitively display the decision-making basis of the multi-disease joint diagnosis model and enhance the interpretability and credibility of medical diagnosis.

Citation Information

Cited By

  • Color eye fundus image classification model and classification method based on virtual multi-mode technology

    CN121280815A

  • Diabetic complication risk prediction system based on disease prior mask

    CN122314422A

  • A diabetes complication risk prediction system based on disease prior masks

    CN122314422B