Intelligent glaucoma diagnosis method based on multi-task learning

By adding VGG16 network to the U-Net network, a multi-task learning network is formed, which solves the problem of low accuracy in glaucoma classification in the existing technology, and achieves more efficient glaucoma diagnosis.

CN120014688APending Publication Date: 2025-05-16TIANJIN UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311515532.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing medical image segmentation algorithm based on deep learning has a low classification accuracy when judging glaucoma, and the cup-dish ratio information obtained by image segmentation alone is not enough to accurately judge glaucoma.

Method used

Adding VGG16 network on the basis of the U-Net network is formed to form a multi-task learning network, realizing the visual cup, disc segmentation and glaucoma classification of fundus images, and sharing image features to improve classification accuracy.

Benefits of technology

Through the multi-task learning network, the accuracy of glaucoma classification is improved, the model training time and data calculation amount are reduced, and more efficient diagnostic results are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014688A_ABST
    Figure CN120014688A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent glaucoma diagnosis method based on multi-task learning. According to the method, a multi-task learning network model is constructed, and the model is composed of a U-Net network and a VGG network. The U-Net network is responsible for extracting features from the eye fundus image to obtain a cup-to-disk ratio, and inputting the cup-to-disk ratio as one of the features of the eye fundus image into the VGG network, and the VGG network performs glaucoma classification on the eye fundus image by using the features. By sharing the encoder portion of the U-Net network, the two networks can share the learned feature information. After being trained by a fundus image data set, the model obtains very high glaucoma classification accuracy. According to the invention, an accurate and efficient glaucoma diagnosis function is realized, the method can be widely applied to large-scale glaucoma screening tasks, the work of ophthalmologists is greatly facilitated, and the harm of glaucoma to people is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image classification and deep learning, and specifically relates to an intelligent diagnosis method for glaucoma based on multi-task learning. Background Art

[0002] Glaucoma is a chronic eye disease that causes irreversible vision loss. There is currently no cure for glaucoma, but early detection of glaucoma can significantly reduce vision loss. The most common screening methods for glaucoma include measuring intraocular pressure, cup-to-disc ratio, optical coherence tomography, and visual field testing. CDR is closely related to intraocular pressure. Usually, increased intraocular pressure can lead to damage to the optic nerve head and may be accompanied by an increase in CDR. Fundus images are more economical while being able to obtain cup-to-disc ratio information. Therefore, by observing the size of the optic disc and optic cup in fundus images in clinical practice, cup-to-disc ratio information can be obtained, which can be used for preliminary screening of glaucoma.

[0003] Since there are many latent glaucoma patients and a shortage of ophthalmologists, the use of deep learning to detect glaucoma has become a mainstream solution. In order to obtain the cup-to-disc ratio information of the fundus image, it is necessary to segment the optic cup and optic disc of the fundus image to obtain the cup-to-disc ratio. At present, most of the medical image segmentation algorithms based on deep learning are improved based on U-Net. Compared with convolutional neural networks, U-Net networks are more suitable for medical image segmentation. However, the classification accuracy is too low when only relying on image segmentation to obtain the cup-to-disc ratio to judge glaucoma. Summary of the invention

[0004] In order to solve the above problems, the present invention proposes a glaucoma intelligent diagnosis method based on multi-task learning. The network model does not change the basic network structure of U-Net, but adds a VGG16 network on the basis of the original U-Net network structure, and realizes the glaucoma classification of fundus images while segmenting the optic cup and optic disc. The classification and segmentation tasks are performed simultaneously and share the features of the image, forming multi-task learning to jointly classify fundus images.

[0005] The technical solutions adopted by the present invention are as follows:

[0006] A multi-task learning network is used to classify glaucoma in fundus images;

[0007] The multi-task learning network consists of a U-Net network and a VGG network;

[0008] The cross entropy function is used as the loss function of the U-Net network and the VGG network to optimize the parameter weights of the network model;

[0009] The optimal network model is obtained through the fundus images and label information in the training set;

[0010] The fundus image of the test set is input into the optimized network model to obtain the probability that the fundus image is glaucoma.

[0011] The VGG network and the U-Net network share the image information extracted by the U-Net network encoder.

[0012] The multi-task learning network model consists of 15 convolutional layers, 3 maximum pooling layers, 3 deconvolutional layers and 2 fully connected layers, wherein all convolutional layer filters are 3×3 convolutional kernels with a step size of 1, and the filters of the pooling layer and the deconvolution layer are 2×2 convolutional kernels with a step size of 2;

[0013] The 14 convolutional layers, 3 maximum pooling layers, and 3 deconvolutional layers in the multi-task learning network constitute the U-Net network, which is used to segment the optic cup and optic disc of the fundus image;

[0014] The 8 convolutional layers, 3 maximum pooling layers, and 3 fully connected layers in the multi-task learning network constitute the VGG network, which is used to classify fundus images for glaucoma;

[0015] The U-Net network performs the segmentation task to obtain the optic cup and optic disc information of the fundus image, and then obtains the cup-disc ratio CDR=D of the fundus image. cup / D disc , D cup and D disc They refer to the longitudinal lengths of the optic cup and optic disc of the fundus image, respectively;

[0016] The cup-to-disc ratio is input into the fully connected layer of the VGG network according to a certain weight η, and finally the VGG network obtains the probability that the image is glaucoma.

[0017] In the multi-task learning network, the encoder of the U-Net network extracts fundus image features, the VGG network shares the encoder part of the U-Net network, and the image information extracted by the encoder is input into the decoder of the U-Net network and the VGG network using skip connections and fully connected layers respectively. The U-Net network implements the segmentation of the image optic cup and optic disc, thereby obtaining CDR; the CDR is input into the fully connected layer of the VGG network as one of the features of the image, thereby obtaining the image classification result;

[0018] The U-Net network completes the segmentation task. The training uses the cross entropy function as the loss function. The loss of the segmentation task is defined as:

[0019]

[0020] Where: L s is the loss of the U-Net network, p s is the predicted value of the U-Net network, ys is the corresponding true value, M represents the number of pixels in the image;

[0021] The diagnosis of glaucoma is a binary classification task, which is completed by the VGG network. The training uses cross entropy as the loss function. The loss of the classification task is defined as:

[0022] L C =-y C log(p C )-(1-y C )log(1-p C )

[0023] Where: L c is the loss of the VGG network, p c is the predicted value of the VGG network, y c is the corresponding true value;

[0024] The U-Net network obtains the CDR of the fundus image and inputs the CDR into the fully connected layer of the VGG network. The output of the fully connected layer is:

[0025]

[0026] Where: y is the output of the current layer, f is the activation function, x i is the output of the i-th input neuron, b is the bias term, n is the number of neurons in the previous layer, CDR is the value predicted by the network, η is the weight of CDR in the fully connected layer, ω i is the weight of the ith input neuron.

[0027] Since the fundus images used for training are 3-channel images with pixels of 1634×1634 and 2124×2056, and the network model input is a 3-channel image with a size of 224×224, the images in the dataset are reduced in size;

[0028] The training set consists of 1,200 annotated fundus images, of which 121 are from glaucoma patients. The proportion of positive training samples is only 1 / 10, so data enhancement is required, including flipping, rotation, cropping, translation and other methods to enhance the data.

[0029] The preprocessed training set images are input into the network model. The downsampling part of the U-Net network extracts features from the input images, and then they are input into the upsampling part of the U-Net network and the fully connected layer of the VGG network. The upsampling part of the U-Net network first performs segmentation of the optic cup and optic disc to obtain the cup-disc ratio, and then the cup-disc ratio is input into the VGG network according to the weight η. Finally, the SoftMax layer of the VGG network outputs the probability p that the image is glaucoma. c , and the true value y cIn contrast, the model weights are optimized through the same loss function.

[0030] After completing the training of the network model weights, the fundus images of the test set are used to evaluate the performance of the network model. The fundus images of the test set are first subjected to image scaling preprocessing operations and then input into the network model. The network model outputs the probability p that the fundus image is glaucoma. When p>0.6, the image is determined to be glaucoma. When p<=0.6, the image is determined not to be glaucoma. After testing all images, the accuracy of the model prediction is obtained.

[0031] Compared with other model methods, the present invention has the following characteristics:

[0032] The U-Net network and the VGG network are trained simultaneously, which reduces the time required for model training;

[0033] The U-Net network and the VGG network share the downsampling part of the U-Net network, which reduces the amount of data calculation during model training;

[0034] The cup-to-disc ratio obtained by the U-Net network was input into the fully connected layer of the VGG network as one of the features of the fundus image, resulting in a higher glaucoma classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a structural diagram of the multi-task learning network model proposed in the present invention;

[0036] Figure 2 This is a structural diagram of a network model for performing optic cup and optic disc segmentation tasks in the multi-task learning network model proposed in the present invention;

[0037] Figure 3 A structural diagram of a network model for performing glaucoma classification tasks in the multi-task learning network model proposed in the present invention;

[0038] Figure 4 The experimental results of verifying the multi-task learning network model proposed in the present invention in the task of segmenting the optic cup and optic disc;

[0039] Figure 5 The figure shows the experimental results for verifying the multi-task learning network model proposed in the present invention on the glaucoma classification task. Specific implementation plan

[0040] The present invention will be further described below in conjunction with the accompanying drawings and specific implementations.

[0041] Embodiments of the present invention are as follows:

[0042] A multi-task learning network model consisting of a U-Net network and a VGG network was established. The VGG network shared the encoder part of the U-Net network. The U-Net network performed the optic cup and optic disc segmentation task, and the cup-disc ratio information was obtained and input into the VGG network. The VGG network then performed the glaucoma classification task.

[0043] The network model is specifically shown in the attached Figure 1 .

[0044] Depend on Figure 1 It can be seen that the multi-task learning network model consists of 15 convolutional layers, 3 maximum pooling layers, 3 deconvolutional layers and 2 fully connected layers. All convolutional layer filters are 3×3 convolution kernels with a stride of 1, and the filters of the pooling layer and deconvolution layer are 2×2 convolution kernels with a stride of 2.

[0045] The 14 convolutional layers, 3 maximum pooling layers, and 3 deconvolutional layers in the multi-task learning network constitute the U-Net network. For details, see the attached Figure 2 ,The U-Net network is used to segment the optic cup and optic disc of fundus images.

[0046] The 8 convolutional layers, 3 maximum pooling layers, and 3 fully connected layers in the multi-task learning network constitute the VGG network. See the attached Figure 3 , VGG network is used to classify glaucoma in fundus images.

[0047] After the preprocessed fundus image training set is input into the multi-task learning network, the image features of the fundus images are obtained by the downsampling part of the U-Net network and input into the upsampling layer of the U-Net network and the fully connected layer of the VGG network to perform segmentation and classification tasks respectively.

[0048] The loss function of the U-Net network is defined as:

[0049]

[0050] Where: L s is the loss of the U-Net network, p s is the predicted value of the U-Net network, y s is the corresponding true value, and M represents the number of pixels in the image.

[0051] After obtaining the segmentation results, calculate the cup-disc ratio CDR = D cup / D disc , D cup and D disc They refer to the longitudinal lengths of the optic cup and optic disc of the fundus image, respectively.

[0052] The VGG network is completed, and the training uses cross entropy as the loss function. The loss of the classification task is defined as:

[0053] L C =-y C log(p C )-(1-y C )log(1-p C )

[0054] Where: L c is the loss of the VGG network, p c is the predicted value of the VGG network, y c is the corresponding true value.

[0055] The U-Net network obtains the CDR of the fundus image and inputs the CDR into the fully connected layer of the VGG network. The output of the fully connected layer is:

[0056]

[0057] Where: y is the output of the current layer, f is the activation function, x i is the output of the i-th input neuron, b is the bias term, n is the number of neurons in the previous layer, CDR is the value predicted by the network, η is the weight of CDR in the fully connected layer, ω i is the weight of the ith input neuron.

[0058] The present invention uses the REFUGE challenge dataset for experimental verification. The REFUGE challenge dataset consists of 1,200 annotated retinal images, 121 of which are from glaucoma patients. The retinal images in this dataset are centered on the macula, and the image size is 1634×1634 or 2124×2056 3-channel images.

[0059] After preprocessing, two groups of experiments were set up, one for multi-task learning (MTL), that is, the U-Net network and the VGG16 network were trained simultaneously; the other for single-task learning (STL), the U-Net network was trained first and the VGG16 network was trained later. Both groups of experiments selected the Adam optimizer, the learning rate was set to 0.0001, the batch_size was set to 16, and the epoch was set to 300.

[0060] The experimental results are attached. Figure 4 and attached Figure 5 .

[0061] By the attached Figure 4 It can be seen that the results obtained by multi-task learning (MTL) and single-task learning (STL) are basically the same when performing the optic cup and optic disc segmentation task.

[0062] By the attached Figure 5It can be seen that the AUC of multi-task learning (MTL) is 0.9788, and the AUC of single-task learning (STL) is 0.9563. The network model based on multi-task learning has a higher glaucoma classification accuracy.

[0063] It can be seen from the experiment that compared with other model methods, the present invention has the following characteristics:

[0064] The U-Net network and the VGG network are trained simultaneously, which reduces the time required for model training;

[0065] The U-Net network and the VGG network share the downsampling part of the U-Net network, which reduces the amount of data calculation during model training;

[0066] The cup-to-disc ratio obtained by the U-Net network was input into the fully connected layer of the VGG network as one of the features of the fundus image, resulting in a higher glaucoma classification accuracy.

Claims

1. An intelligent glaucoma diagnosis method based on multi-task learning, characterized in that: The following steps are involved: A multi-task learning network is used to classify glaucoma in fundus images; The multi-task learning network consists of a U-Net network and a VGG network; The cross entropy function is used as the loss function of the U-Net network and the VGG network to optimize the weight parameters of the network model; The optimal network model is obtained through the fundus images and label information in the training set; The fundus image of the test set is input into the optimized network model to obtain the probability that the fundus image is glaucoma.

2. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The multi-task learning network is a neural network that combines multiple convolutional neural networks to achieve multiple task requirements, and it completes multiple learning tasks simultaneously in a learning process. The convolutional neural network is a deep learning model, which is mainly used to process data with a grid structure such as images, and can automatically learn and extract features such as texture, shape, and edges in images. The VGG network and the U-Net network are both deep convolutional neural network models. The VGG network is a convolutional neural network that includes multiple convolutions. The number of convolutions can be reasonably set according to task requirements. It is used to classify fundus features with or without glaucoma. The U-Net network is a convolutional neural network that processes images at the pixel level. It is used to extract optic cup and disc features from fundus images to obtain the cup-disc ratio. The VGG network and the U-Net network share the feature information of the input image.

3. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The cross entropy function is a kind of loss function. The loss function is a function used to measure the difference between the true value and the predicted value of the network model. In the U-Net network, it is used to describe the difference between the pixel-level model predicted value and the true value, and in the VGG network, it is used to describe the difference between the image-level model predicted value and the true value. The parameter weight is used to measure the importance of the feature information contained in the input image of the network model. The prediction accuracy of the network model is improved by adjusting the parameter weight.

4. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The label information refers to the information annotation of the collected fundus images, which are divided into two categories: those with glaucoma and those without glaucoma. Inputting the data set with label information into the network model for pre-training can improve the prediction accuracy of the network model.

5. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The U-Net network is composed of 14 convolutional layers, 3 maximum pooling layers, and 3 deconvolution layers. Both the convolution and deconvolution are mathematical methods based on integral transformation. The convolution layer extracts different features of the fundus image, and the deconvolution layer restores the low-resolution input image to a high-resolution image. The pooling is a form of downsampling that reduces the number of parameters and the amount of calculation in the network model.

6. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The VGG network is composed of 8 convolutional layers, 3 maximum pooling layers, and 3 fully connected layers. In the fully connected layer, each neuron node is connected to all the neuron nodes in the previous layer, which is used to integrate the features extracted previously and play the role of a "classifier" to classify the fundus images.

7. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The U-Net network performs image segmentation tasks, uses segmentation mask technology to accurately separate the optic cup and optic disc in the fundus image, obtains a segmentation mask image of the fundus image, performs edge detection on the segmentation mask image, divides the boundary between the optic disc and the optic cup, and can respectively obtain the longitudinal height of the optic cup and the longitudinal height of the optic disc, thereby obtaining the cup-to-disc ratio of the fundus image.

8. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: The VGG network performs the glaucoma classification task, and the CDR (cup-to-disc ratio) mentioned in claim 7 is input into the VGG network according to a certain proportional coefficient. Finally, the VGG network obtains the probability that the fundus image is glaucoma.

9. The intelligent glaucoma diagnosis method based on multi-task learning according to claim 1, characterized in that: When the probability that the image is glaucoma obtained by the VGG network described in claim 8 is greater than 0.6, the fundus image is determined to be glaucoma; otherwise, the fundus image is determined not to be glaucoma.