Fundus Image Classification Algorithm Based on Fusion Decision Tree and Improved UNet++
By adopting a fusion decision tree and an improved UNet++ algorithm in glaucoma fundus image classification, the problem of poor classification of glaucoma fundus image severity detection in the prior art is solved, and higher accuracy and sensitivity are achieved.
Patent Information
- Application Number
- CN202211134603.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The prior art has poor classification effect and low accuracy in detecting the severity of glaucoma fundus images.
The fundus image classification algorithm based on the fusion decision tree and improved UNet++ is adopted to enhance the texture information and contrast of glaucoma fundus images through the pre-processing stage. The feature extraction stage uses the UNet++ model based on the residual module and attention mechanism to improve. The image classification stage uses the decision tree C4.5 for multi-classification.
Compared with traditional algorithms, this algorithm has improved the average accuracy, average specificity and average sensitivity by 9.2%, 6.4%, and 6.5% respectively in the classification of glaucoma fundus images, which significantly improves the classification effect.
Smart Images

Figure CN115601822B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image classification, and particularly relates to a fundus image classification algorithm based on a fusion decision tree and an improved UNet++. Background Art
[0002] Glaucoma, as the second most common cause of blindness in the world, is the main cause of blindness caused by optic nerve damage. It is the focus of research on the classification detection of fundus images and has become the focus of attention of experts at home and abroad.
[0003] Among them, He Xiaoyun et al. proposed an improved U-Net network model, which incorporated residual blocks, cascaded dilated convolutions, and an embedded attention mechanism into the U-Net model to achieve retinal vessel segmentation; Sabri Deari et al. proposed a retinal vessel segmentation network model based on a transfer learning strategy. This model enhanced the dataset through pixel-level transformation and reflection transformation, and then used the U-Net model to train retinal features to achieve retinal vessel segmentation; Yuan Zhou et al. proposed a network model that fuses the attention mechanism and UNet++. This model extracts image features based on the UNet++ model and incorporates the attention mechanism into the convolutional unit to enhance features, thereby completing end-to-end image detection; Ali Serener et al. proposed an image classification algorithm based on a single CNN convolutional neural network model. This method achieved the classification detection of glaucoma lesion images by creating multiple fusions of CNNs; Guo Fan et al. proposed a glaucoma image detection method that combines MobileNet v2 and the VGG classification network. This method used the MobileNet v2 segmentation model to segment and locate the optic disc image, and combined the VGG classification network and the attention module to screen for glaucoma; Gupta et al. proposed a method for detecting retinal vessels by random forest classification. This method segmented retinal images and extracted the texture features and gray-scale features of image patches in units of blocks to classify retinal images; Ke Shiyuan et al. used a multi-view ensemble learning method of support vector machines and logistic regression to predict glaucoma; DAS et al. proposed a glaucoma detection method based on the CDR and ISNT rules. This method used region growing and watershed transformation to segment the OC and OD, and then achieved glaucoma image classification.
[0004] Although the above algorithms can screen and judge glaucoma fundus lesions, the accuracy of detecting the severity of glaucoma fundus lesions is relatively low, and the classification effect is not good. Summary of the Invention
[0005] The purpose of the present invention is to use an improved UNet++ algorithm that fuses a decision tree to achieve the classification of the severity of glaucoma for the problem of poor image classification effect caused by low contrast of glaucoma images.
[0006] The specific solution of the present invention is as follows:
[0007] A fundus image classification algorithm based on a fusion decision tree and improved UNet++ includes:
[0008] In the preprocessing stage, the green component image of the fundus image is extracted, and an improved Butterworth transfer function based on a power function is used to enhance the texture information and contrast of the glaucomatous fundus image;
[0009] In the feature extraction stage, an improved UNet++ model based on a residual module and an attention mechanism is used to extract image features;
[0010] In the image classification stage, decision tree C4.5 is used for multi-classification of the image to obtain the classification detection result of glaucomatous lesions.
[0011] Furthermore, the preprocessing stage specifically includes:
[0012] Separate the RGB image and extract the green component image;
[0013] Use the improved Butterworth transfer function for frequency division processing to obtain high-frequency information P h and low-frequency information P l , and its calculation formula is
[0014]
[0015]
[0016] where R h represents the high-frequency gain coefficient of the glaucomatous fundus image, R l represents the low-frequency gain coefficient of the glaucomatous fundus image. When R h >1, it means that the enhanced fundus image is high-frequency information. When R l <1, it means that the low-frequency information of the fundus image is weakened. a represents the sharpening coefficient, D 0 represents the cut-off frequency, n represents the filter order, D(x,y) represents the distance from the frequency (x,y) to the filter center (x 0 ,y 0 ), and the calculation uses the Euclidean distance formula
[0017]
[0018] Use the inverse Fourier transform to convert the frequency domain information into a spatial domain image, convert the high and low frequency information into high and low frequency images, and the inverse Fourier transform is
[0019]
[0020] Among them, F(t) represents a function in the time domain, F(w) represents a function in the frequency domain, F(t) is the original function of F(w), and after processing, the high-frequency image F h (x, y) and the low-frequency image F l (x, y);
[0021] For the high-frequency image F h (x, y) and the low-frequency image F l (x, y), after local enhancement respectively, weighted fusion is carried out to obtain the enhanced fundus image, and the fusion formula is
[0022] G(x, y) = aF′ h (x, y) + bF′ l (x, y)
[0023] Among them, a and b respectively represent the weighting constants, and G(x, y) represents the enhanced green component map of the fundus.
[0024] Furthermore, the preprocessing stage specifically further includes
[0025] For the fused enhanced fundus image, noise reduction processing is carried out in combination with the power function curve method. The power function adjusts the image contrast mode through parameters and is adjusted using the image mapping relationship. Its calculation formula is
[0026] G′ = ax t + bx (t-1) + …… + cx + d
[0027] Among them, t is the power, which is a controllable parameter. After processing, the preprocessed enhanced image G′ is obtained.
[0028] Specifically, the local enhancement of the high-frequency image F h (x, y) is specifically as follows: Using the SMQT algorithm to perform gray-level region expansion processing on the high-frequency image F h (x, y) to achieve non-linear stretching of the image gray level.
[0029] Specifically, the local enhancement of the low-frequency image F l (x, y) respectively is specifically as follows: The low-frequency image is converted to the Lab space, and the histogram equalization method is used for the L channel to process the contrast. Specifically, the image is divided into blocks, each image block is classified respectively, and the fat equalization method is used to perform interpolation operations on each pixel to obtain the processed gray-scale image F′ l .
[0030] Specifically, the SMQT algorithm includes:
[0031] The image reading points are hierarchically processed using a binary tree, and the outputs of each layer are linearly superimposed to obtain a locally enhanced high-frequency image. The calculation formula is
[0032]
[0033] where m represents a certain pixel in the image D(m), F′ h (m) is the output of SMQT, v(m) represents the gray value of the pixel, U(m) is the result of gray value quantization, L represents the number of layers of the binary tree, and n represents the output number of MQN with layer number l.
[0034] Furthermore, in the improved UNet++ model based on the residual module and attention mechanism, a residual module is introduced between the upsampling and downsampling convolutional layers of the UNet++ network, and a hybrid-domain attention mechanism is added before each residual convolutional module;
[0035] The hybrid attention mechanism includes a channel attention mechanism and a spatial attention mechanism. First, the input fundus feature map is sent into the channel attention mechanism to perceive the global texture information, and the extracted information is fused with the original image to obtain the global feature processing result. Then, the global enhanced feature processing result is sent into the spatial attention mechanism for local texture feature enhancement, and after processing, it is weighted and summed with the global enhanced feature processing result to obtain the local and global feature enhancement result. The calculation formula is
[0036] F M =CBAM(F i )=SAM(CAM(F i ))×F i ×(CAM(F i )×F i )
[0037] where CBAM(F i ) represents the operation result of the hybrid-domain attention mechanism, F i represents the input fundus image, CAM(F i ) represents the operation of the channel attention mechanism, SAM represents the spatial attention mechanism, and × represents matrix convolution operation.
[0038] Specifically, the channel attention mechanism uses average pooling and max pooling to aggregate the spatial information of the feature map, obtaining max pooling and average pooling respectively. Then, the max pooling and average pooling are forwarded to a shared hidden layer MLP network. Next, the dimensions of the two channel attention mechanism maps obtained by max pooling and average pooling respectively are set to C×1×1. The result after average pooling is processed by the sigmoid function, and finally, the element-wise addition of the two is performed to obtain the processing result of the channel attention mechanism. The calculation formula is
[0039] CAM(F i ) = sigmod(MLP(AvgPool(F i )) + MLP(MaxPool(F i )))
[0040] Among them, sigmod represents the activation function, AvgPool represents average pooling processing, MaxPool represents max pooling processing, MLP represents the MLP neural network, that is, multi-layer perceptron processing, and the number of neurons in the hidden layer is set to r is a hyperparameter;
[0041] The spatial attention mechanism performs average pooling and max pooling processing along the channel axis, and after processing, the two obtained feature maps are concatenated for convolution operation, and finally the spatial attention mechanism's processing result is obtained by using sigmoid activation. Its calculation formula is
[0042] SAM(CAM(F i )) = sigmod(conv([AvgPool(M c ) + MaxPool(M c )]))
[0043] Among them, SAM represents the operation of the spatial domain attention mechanism, and conv represents the convolution operation.
[0044] Specifically, the improved UNet++ model based on the residual module and the attention mechanism trains the model in a deep supervision mode, and the loss function uses the combination of binary cross entropy and DICE coefficient. Its calculation formula is
[0045]
[0046] Among them, and Y b respectively represent the flattened predicted probability and the flattened ground truth of the b-th image, and N represents the batch size.
[0047] Specifically, the decision tree C4.5 algorithm searches for splitting attributes from all texture information extracted from features for segmentation, generates texture information and non-texture information, continuously splits the nodes with texture information, and then classifies the lesions of glaucoma fundus images into four categories: normal images, mild glaucoma, moderate glaucoma, and severe glaucoma.
[0048] After adopting the above solution, the beneficial effects of the present invention are as follows: Compared with traditional algorithms, the present invention has improvements in terms of accuracy, average specificity, and average sensitivity. Specifically, its average accuracy, average specificity, and average sensitivity are increased by 9.2%, 6.4%, and 6.5% respectively. It can be seen that the improved algorithm has good results in the classification of glaucoma fundus images. For specific effects, please refer to the specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the overall flowchart of the algorithm of the present invention;
[0050] Figure 2 is the improved UNet++ model diagram of the present invention;
[0051] Figure 3 Structural diagram of the improved residual module;
[0052] Figure 4 is the hybrid-domain attention mechanism module diagram of the present invention;
[0053] Figure 5 is the dataset sample diagram in the specific implementation of the present invention, where (a) is a normal glaucoma image, (b) is a mild glaucoma image, (c) is a moderate glaucoma image, and (d) is a severe glaucoma image;
[0054] Figure 6 is the analysis diagram of the average accuracy of the model under different iteration times in the specific implementation of the present invention.
[0055] SPECIFIC IMPLEMENTATION MANNER
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0057] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] The following will elaborate on each step of the present invention based on what is known to those skilled in the art.
[0059] This embodiment will elaborate in detail on the fundus image classification algorithm of the present invention based on the fusion decision tree and the improved UNet++. The overall flowchart of the algorithm of the present invention is as Figure 1 shown, including:
[0060] S1. In the preprocessing stage, extract the green component image of the fundus image, and use the Butterworth transfer function based on the power function to enhance the texture information and contrast of the glaucoma fundus image;
[0061] The preprocessing stage specifically includes:
[0062] S101. Separate the RGB image and extract the green component image;
[0063] S102. Use the improved Butterworth transfer function for frequency division processing to obtain the high-frequency information P h and the low-frequency information P l , and its calculation formula is
[0064]
[0065]
[0066] where, R h represents the high-frequency gain coefficient of the glaucoma fundus image, R l represents the low-frequency gain coefficient of the glaucoma fundus image. When R h > 1, it means that the enhanced fundus image is high-frequency information. When R l < 1, it means that the low-frequency information of the fundus image is weakened. a represents the sharpening coefficient, D 0 represents the cut-off frequency, n represents the filter order, D(x, y) represents the distance from the frequency (x, y) to the filter center (x 0 , y 0 ), and the calculation uses the Euclidean distance formula
[0067]
[0068] S103. Use the inverse Fourier transform to convert the frequency-domain information into a spatial domain image, and convert the high- and low-frequency information into high- and low-frequency images. The inverse Fourier transform is
[0069]
[0070] where, F(t) represents the function in the time domain, F(w) represents the function in the frequency domain, F(t) is the original function of F(w), and after processing, the high-frequency image F h (x, y) and the low-frequency image F l (x, y) are obtained;
[0071] S104. For the high-frequency image F hLocal enhancement of (x, y) specifically means using the SMQT algorithm to process the high-frequency image F h Perform gray-level region expansion processing on (x, y) to achieve non-linear stretching of the image gray level. The SMQT algorithm improves the local contrast, enhances the texture details of the image, and enhances the pixel points. The SMQT algorithm includes:
[0072] Use a binary tree to hierarchically process the image reading points, and linearly superimpose the outputs of each layer to obtain a locally enhanced high-frequency image. The calculation formula is
[0073]
[0074] Among them, m represents a certain pixel in the image D(m), F′ h (m) is the output of SMQT, v(m) represents the gray value of the pixel, U(m) is the gray value quantization result, L represents the number of layers of the binary tree, and n represents the MQN output number of layer l.
[0075] The local enhancement of the low-frequency image F l (x, y) is specifically as follows. The low-frequency image is converted to the Lab space, and the histogram equalization method is used for the L channel to process the contrast, reducing the influence of the image color components on the detection. Specifically, the image is divided into blocks, each image block is classified, and the fat equalization method is used to perform interpolation operations on each pixel to obtain the processed gray-scale image F′ l . For the high-frequency image F h (x, y) and the low-frequency image F l (x, y), after local enhancement respectively, weighted fusion is performed to obtain an enhanced fundus image. The fusion formula is
[0076] G(x, y) = aF′ h (x, y) + bF′ l (x, y)
[0077] Among them, a and b respectively represent the weighting constants, and G(x, y) represents the enhanced fundus green component map.
[0078] S105. For the fused enhanced fundus image, noise reduction is performed in combination with the power function curve method. The power function adjusts the image contrast mode through parameters and is adjusted using the image mapping relationship. The calculation formula is
[0079] G′ = ax t + bx (t-1) + …… + cx + d
[0080] Among them, t is the power, which is a controllable parameter. After processing, the preprocessed enhanced image G′ is obtained.
[0081] S2. Feature extraction stage, use the improved UNet++ model based on the residual module and attention mechanism to extract image features. The model structure diagram is as Figure 2 shown. The UNet++ network consists of an encoder and a decoder. x i,j represents the output of node x i,j . Among them, i represents the layer number, and j represents the jth convolutional layer of the current layer. The skip path is used to change the connectivity between the encoder and decoder sub-networks. In UNet, the decoder directly receives the feature map of the encoder; while in UNet++, it goes through a dense convolutional block, and all convolutional layers on the skip path use a kernel of size 3×3. The skip path formula is
[0082]
[0083] where, X i,j represents the output of node X i,j . Among them, i indexes the downsampling layer along the encoder, j indexes the convolutional layer of the dense block along the skip path, H(·) represents the convolutional operation and activation function, μ(·) represents the upsampling layer, and [] represents the concatenation layer. The node at level j = 0 only receives one input from the previous layer of the encoder; the node at level j = 1 receives two inputs, both from the encoder sub-network, but at two consecutive levels; and the node at level j > 1 receives j + 1 inputs, where j inputs are the outputs of the previous j nodes in the same skip path, and the last input is the upsampling output from the lower skip path.
[0084] In the model, to solve the problem of gradient disappearance, a residual module is introduced between the upsampling and downsampling convolutional layers of the UNet++ network, and a hybrid-domain attention mechanism is added before each residual convolutional module to obtain more local and global texture information. The improved residual block is as Figure 3 shown. The implementation principle of the residual module is to add the input feature map to the feature extraction module to obtain feature information, so that the network includes the feature information of the input feature map during forward propagation, effectively solving the degradation problem of network model convolution processing. The residual block formula is
[0085] H(x) = F(x) + x
[0086] where, x is the input of the network, F(x) represents the feature extraction module, and H(x) represents the output result of fundus image feature extraction.
[0087] Such as Figure 4As shown in the figure, the hybrid attention mechanism includes a channel attention mechanism and a spatial attention mechanism. First, the input fundus feature map is fed into the channel attention mechanism to perceive the global texture information, and the extracted information is fused with the original image to obtain the global feature processing result. Then, the global enhanced feature processing result is fed into the spatial attention mechanism for local texture feature enhancement. After processing, it is weighted and summed with the global enhanced feature processing result to obtain the local and global feature enhancement result. The calculation formula is
[0088] F M =CBAM(F i )=SAM(CAM(F i ))×F i ×(CAM(F i )×F i )
[0089] Among them, CBAM(F i ) represents the operation result of the hybrid domain attention mechanism, F i represents the input fundus image, CAM(F i ) represents the operation of the channel attention mechanism, SAM represents the spatial attention mechanism, and × represents matrix convolution operation.
[0090] Specifically, the channel attention mechanism uses average pooling and max pooling to aggregate the spatial information of the feature map, respectively obtaining max pooling and average pooling. Then, the max pooling and average pooling are forwarded to a shared hidden layer MLP network. Next, the dimensions of the two channel attention mechanism maps obtained by max pooling and average pooling respectively are set to C×1×1. The result after average pooling is processed by the sigmoid function, and finally, the element-wise addition of the two is performed to obtain the processing result of the channel attention mechanism. The calculation formula is
[0091] CAM(F i )=sigmod(MLP(AvgPool(F i ))+MLP(MaxPool(F i )))
[0092] Among them, sigmod represents the activation function, AvgPool represents average pooling processing, MaxPool represents max pooling processing, MLP represents the MLP neural network, that is, multi-layer perceptron processing. The number of neurons in the hidden layer is set to r is a hyperparameter.
[0093] The spatial attention mechanism performs average pooling and max pooling processing along the channel axis. After processing, the two obtained feature maps are concatenated for convolution operation, and finally, the sigmoid activation is used to obtain the processing result of the spatial attention mechanism. The calculation formula is
[0094] SAM(CAM(F i )) = sigmod(conv([AvgPool(M c ) + MaxPool(M c )]))
[0095] Among them, SAM represents the operation of the spatial domain attention mechanism, and conv represents the convolution operation.
[0096] During model training, deep supervision is adopted to enable the UNet++ model to run in an accurate mode and a fast mode. In the accurate mode, the output results of all segmentation branches are averaged. In the fast mode, only one segmentation branch is selected, and the others are pruned. The selection result is used to determine the degree of model pruning and the speed gain.
[0097] The combination of binary cross-entropy and DICE coefficient is used as the loss function for the four semantic levels of {X 0,j , j ∈ {1, 2, 3, 4}} as
[0098]
[0099] Among them, and Y b respectively represent the flattened predicted probability and the flattened ground truth of the b-th image, and N represents the batch size.
[0100] S3. In the image classification stage, the decision tree C4.5 is used for image multi-classification to obtain the classification detection results of glaucoma lesions. The decision tree C4.5 algorithm searches for splitting attributes from all texture information extracted by features for segmentation, generating textured information and non-textured information, and continuously splitting the textured information nodes, thereby classifying the glaucoma fundus image lesions into four categories: normal images, mild glaucoma, moderate glaucoma, and severe glaucoma. The implementation of the decision tree C4.5 algorithm is divided into two stages: the generation of the initial decision tree and the pruning of the decision tree. The algorithm flow is as follows:
[0101] Input: Training set decision table: Training set D = {(d1, k1), (d2, k2),..., (dn, kn)} and attribute set A = {a1, a2,..., am}
[0102] Output: Decision tree with Node as the root node
[0103] 1: function Build_DT(D, A) Tree-building function
[0104] 2: Generate node node;
[0105] 3: if all samples in D belong to the same category C then
[0106] 4: Mark the node as a leaf node of class C; return
[0107] 5: end if
[0108] 6: if the samples in D have the same value on A then then
[0109] 7: Mark the node as a leaf node of the class with the largest number of samples in D; return
[0110] 8: end if
[0111] 9: Select the optimal attribute from A, that is, a* = arg max a∈AGR(D,a) the attribute with the highest gain ratio;
[0112] 10: For each attribute value av* of a* do
[0113] 11: Generate a branch for the node; let Dv be the subset of samples in D that take the value av* on a*;
[0114] 12: if Dv is empty then
[0115] 13: Mark the branch node as a leaf node of the class with the largest number of samples in D; return
[0116] 14: else
[0117] 15: Use Build DT(Dv,A\{a*}) as the branch node;
[0118] 16: end if
[0119] 17: end for
[0120] 18: end function
[0121] After decision tree classification, it is detected whether the fundus image of glaucoma belongs to a normal image, mild glaucoma, moderate glaucoma or severe glaucoma.
[0122] In this specific implementation, the dataset provided by Paddle Paddle is used, and 480 glaucoma datasets are selected for training, with 120 for normal glaucoma, mild glaucoma, moderate glaucoma and severe glaucoma respectively, as Figure 5 shown.
[0123] Using the Intel i7-7800 CPU, NVIDIA Ge Force GTX1080i graphics card, Paddle Paddle 2G GPU computing power, and deep learning frameworks Keras, OpenCV, and TensorFlow. Since the input layer of the UNet++ network requires 1024×1024 pixels, the crop operation in the pillow library of Python is used to set a fixed cropping area to crop the size of all images to 1024×1024 and train them in a 7:3 ratio.
[0124] The research uses accuracy Acc, specificity S p , and sensitivity S n to objectively evaluate the classification of glaucomatous fundus lesions. The calculation formula is
[0125]
[0126]
[0127]
[0128] Among them, TP represents the number of normal fundus images correctly classified, TN represents the number of glaucomatous lesion images correctly classified, FN represents the number of images misclassified as normal fundus images, FP represents the number of glaucomatous lesion fundus images misclassified, and TN and FP respectively represent the sum of the total number of three degrees of glaucomatous lesion images for correct and incorrect judgments. The calculation formula is
[0129] TN = TN 1 + TN 2 + TN 3
[0130] FP = FP 1 + FP 2 + FP 3
[0131] Among them, TN 1 represents the number of mildly diseased fundus images correctly judged, TN 2 represents the number of moderately diseased fundus images correctly judged, TN 3 represents the number of severely diseased fundus images correctly judged, FP 1 represents the number of mildly diseased fundus images misjudged, TFP 2 represents the number of moderately diseased fundus images misjudged, FP 3 represents the number of severely diseased fundus images correctly judged.
[0132] To make the gradient of the loss function reach the global optimum, through continuous experiments to adjust the network weight hyperparameters, the best learning rate of 0.001 is finally selected for implementation. During the model training process, the accuracies of experiments with different iteration times are analyzed, and the analysis results are as Figure 6 shown. As can be seen from the figure, when the learning rate of the research algorithm is 0.001, the algorithm has the best average accuracy effect for classifying glaucoma fundus images at about 12,000 iterations, and the average accuracy is 94.46%.
[0133] To verify the effects of different algorithms on classifying glaucoma fundus images in the same experimental environment, the research uses accuracy, specificity, and sensitivity to analyze CNN, the improved UNet algorithm, the multi-fusion algorithm of the CNN model, and the algorithm of the present invention. The analysis results are shown in Table 1.
[0134] Table 1 Comparison of different neural networks (%)
[0135]
[0136] As can be seen from Table 1, the lowest average accuracy, average specificity, and average sensitivity for glaucoma detection are all for the classic CNN algorithm, and the best effect is the algorithm of this article, reaching 94.46%, 91.74%, and 95.89% respectively. Compared with the traditional network model, the average accuracy, average specificity, and average sensitivity are increased by 9.2%, 6.4%, and 6.5% respectively. The improved algorithm has a good effect on classifying glaucoma fundus lesions.
[0137] To verify the effects of different algorithms on classifying glaucoma fundus images in the same experimental environment, performance analysis is carried out on the classic support vector machine, random forest method, UNet++ algorithm with attention mechanism, image-level recognition algorithm with local variation microscopic inspection mode, multi-view ensemble learning method of Dempster-Shafer (DS) evidence inference, image detection method of CDR and ISNT rules, and the algorithm of the present invention. The analysis results are shown in Table 2.
[0138] Table 2 Comparison of different classifiers
[0139]
[0140] As can be seen from Table 2, the best classification effect is the research algorithm of this article, and its accuracy, specificity, and sensitivity are 94.46%, 91.74%, and 95.89% respectively. Compared with the traditional algorithm, they are increased by 3.6%, 4.5%, and 3.5% on average respectively. The improved algorithm has certain advantages in detecting glaucoma fundus images.
[0141] It should be understood that the algorithm of the present invention can be applied not only to the classification and detection of glaucoma fundus lesions, but also to the classification of other medical images and traffic images.
[0142] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0143] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A fundus image classification algorithm based on a fusion decision tree and improved UNet++. It is characterized in that it includes: In the preprocessing stage, the green component image of the fundus image is extracted, and an improved Butterworth transfer function based on a power function is used to enhance the texture information and contrast of the glaucoma fundus image; In the feature extraction stage, an improved UNet++ model based on a residual module and an attention mechanism is used to extract image features; In the image classification stage, decision tree C4.5 is used for multi-classification of the image to obtain the classification detection result of glaucoma lesions; In the improved UNet++ model based on the residual module and the attention mechanism, a residual module is introduced between the upsampling and downsampling convolutional layers of the UNet++ network, and a hybrid-domain attention mechanism is added before each residual convolutional module; The hybrid-domain attention mechanism includes a channel attention mechanism and a spatial attention mechanism. First, the input fundus feature map is sent into the channel attention mechanism to perceive the global texture information, and the extracted information is fused with the original image to obtain the global feature processing result. The global enhanced feature processing result is sent into the spatial attention mechanism for local texture feature enhancement, and after processing, it is weighted and summed with the global enhanced feature processing result to obtain the local and global feature enhancement result. Its calculation formula is F M = CBAM(F i ) = SAM(CAM(F i )) × F i × (CAM(F i ) × F i ) Among them, CBAM(F i ) represents the operation result of the hybrid domain attention mechanism, F i represents the input fundus image, CAM(F i ) represents the operation of the channel attention mechanism, SAM represents the spatial attention mechanism, and × represents the matrix convolution operation.
2. The fundus image classification algorithm based on a fusion decision tree and improved UNet++ according to claim 1, it is characterized in that the preprocessing stage specifically includes: Separate the RGB image and extract the green component image; Frequency division processing is performed using an improved Butterworth transfer function to obtain high-frequency information P h and low-frequency information P l , Among them, R h represents the high-frequency gain coefficient of the glaucoma fundus image, and R l represents the low-frequency gain coefficient of the glaucoma fundus image. When R h > 1, it indicates that the enhanced fundus image is high-frequency information. When R l < 1, it indicates that the low-frequency information of the fundus image is weakened. a represents the sharpening coefficient, D 0 represents the cut-off frequency, n represents the filter order, and D(x, y) represents the distance from the frequency (x, y) to the filter center (x 0 , y 0 ). The calculation uses the Euclidean distance formula Use the inverse Fourier transform to convert the frequency-domain information into a spatial-domain image, convert the high-frequency and low-frequency information into high-frequency and low-frequency images, and the inverse Fourier transform is Among them, F(t) represents a function in the time domain, F(ω) represents a function in the frequency domain, F(t) is the original function of the image of F(ω), and after processing, the high-frequency image F h (x, y) and the low-frequency image F l (x, y); For the high-frequency image F h (x, y) and the low-frequency image F l (x, y), after local enhancement respectively, weighted fusion is performed to obtain the enhanced fundus image, and the fusion formula is G(x,y) = aF′ h (x,y) + bF′ l (x,y) where a and b respectively represent weighted constants, and G(x, y) represents the enhanced fundus green component map.
3. The fundus image classification algorithm based on a fusion decision tree and improved UNet++ according to claim 2, it is characterized in that the preprocessing stage also specifically includes: For the fused enhanced fundus image, noise reduction is performed by combining the power function curve method. The power function adjusts the image contrast mode through parameters and is adjusted using the image mapping relationship. Its calculation formula is G′ = ax t + bx (t-1) + …… + cx + d where t is the power, which is a controllable parameter, and after processing, the preprocessing enhanced image G′ is obtained.
4. The fundus image classification algorithm based on a fusion decision tree and improved UNet++ according to claim 2, it is characterized in that The local enhancement of the high-frequency image F h (x, y) is specifically as follows: Using the SMQT algorithm to perform gray-level region expansion processing on the high-frequency image F h (x, y) to achieve non-linear stretching of the image gray level.
5. The fundus image classification algorithm based on a fusion decision tree and improved UNet++ according to claim 2, it is characterized in that The local enhancement of the low-frequency image F l (x, y) is specifically as follows: The low-frequency image is converted to the Lab color space, and the histogram equalization method is used to process the contrast of the L channel. Specifically, the image is divided into blocks, each image block is classified, and the fat equalization method is used to perform interpolation operations on each pixel to obtain the processed grayscale image F'. l .
6. The fundus image classification algorithm based on a fusion decision tree and improved UNet++ according to claim 4, it is characterized in that the SMQT algorithm includes: Use a binary tree to hierarchically process the image reading points, and linearly superimpose the outputs of each layer to obtain a locally enhanced high-frequency image. The calculation formula is Among them, m represents a certain pixel in the image D(m), F′ h (m) is the output of SMQT, v(m) represents the gray value of the pixel, U(m) is the result of gray value quantization, L represents the number of layers of the binary tree, and n represents the output number of MQN with the layer number l.
7. The fundus image classification algorithm based on a fusion decision tree and improved UNet++ according to claim 1, it is characterized in that The channel attention mechanism uses average pooling and max pooling to aggregate the spatial information of the feature map, obtaining the max pooling and average pooling respectively. Then, the max pooling and average pooling are forwarded to a shared hidden layer MLP network. Next, the dimensions of the two channel attention mechanism maps obtained by the max pooling and average pooling respectively are set to C×1×1. The result after average pooling is processed by the sigmoid function. Finally, the element-wise addition of the two is performed to obtain the processing result of the channel attention mechanism, and its calculation formula is CAM(F i ) = sigmod(MLP(AvgPool(F i )) + MLP(MaxPool(F i ))) Among them, sigmod represents the activation function, AvgPool represents the average pooling process, MaxPool represents the max pooling process, MLP represents the MLP neural network, that is, the multi-layer perceptron process, and the number of neurons in the hidden layer is set to r is a hyperparameter; The spatial attention mechanism performs average pooling and max pooling along the channel axis. After processing, the two obtained feature maps are concatenated and then subjected to a convolution operation. Finally, the spatial attention mechanism's processing result is obtained using the sigmoid activation, and its calculation formula is SAM(CAM(F i )) = sigmod(conv([AvgPool(M c ) + MaxPool(M c ))) Among them, SAM represents the operation of the spatial attention mechanism, and conv represents the convolution operation.
8. An algorithm for classifying fundus images based on a fusion decision tree and improved UNet++ according to claim 1, characterized in that the UNet++ model improved based on the residual module and the attention mechanism trains the model in a deep supervision mode, and the loss function uses the combination of binary cross-entropy and DICE coefficient. Among them, and Y b respectively represent the flattened predicted probability and the flattened ground truth of the b-th picture, and N represents the batch size.
9. An algorithm for classifying fundus images based on a fusion decision tree and improved UNet++ according to claim 1, characterized in that The decision tree C4.5 algorithm searches for splitting attributes from all texture information extracted from features for segmentation, generating texture information and non-texture information, and continuously splitting the nodes with texture information, thereby classifying the lesions of glaucomatous fundus images into four categories: normal images, mild glaucoma, moderate glaucoma, and severe glaucoma.
Citation Information
Patent Citations
Unet network brain tumor MRI image segmentation method for improving attention module
CN113554669A
Genetic fuzzy tree-based retinal diabetes mellitus variable depth network detection method
CN114494196A
Ultrasound image denoising model establishing method and ultrasound image denoising method
WO2022083026A1