Fundus Lesion Exudate Segmentation Method and Device Based on Deep Learning

By using the Encoder-Decoder structure and the improved loss function in the exudation segmentation of fundus lesions, the problems of small receptive fields, high model complexity and uneven exudation size distribution in the existing methods are solved, and higher detection accuracy and efficiency are achieved.

CN115082491BActive Publication Date: 2025-06-24XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210164617.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-06-24
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

The existing fundus lesions exudation segmentation method based on deep learning has problems such as small receptive field, high model complexity, and uneven exudation size distribution, resulting in low detection accuracy.

Method used

The Encoder-Decoder structure is used to build a fully convolutional deep learning network, extract multi-scale features through multiple downsampling and upsampling modules, and optimize model performance using the improved loss function EdgeLoss.

Benefits of technology

It improves the accuracy of detection results of exudate segmentation, reduces the number of model parameters and algorithm complexity, and improves the detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082491B_ABST
    Figure CN115082491B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for exudate segmentation of fundus lesions based on deep learning. The method includes: obtaining a fundus image to be processed; constructing a deep learning network for exudate segmentation of fundus lesions based on the Encoder-Decoder structure; inputting the fundus image to be processed into the deep learning network to obtain an exudate segmentation result; wherein, the encoding end of the deep learning network includes a plurality of downsampling modules, which perform multiple downsamplings on the fundus image to be processed and then output a feature map to the decoding end; correspondingly, the decoding end includes a plurality of upsampling modules, which perform multiple upsamplings on the feature map to finally obtain an exudate segmentation result. The present invention constructs a fully convolutional deep learning network for exudate segmentation of fundus lesions by using the Encoder-Decoder structure. The network model not only has a small number of parameters and a simple algorithm, but also has the ability to quickly extract multi-scale features of images, thereby improving the accuracy of detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a method and device for exudate segmentation of fundus lesions based on deep learning. Background Art

[0002] Diabetic Retinopathy (DR) is one of the main complications of diabetes. The global incidence of diabetes is extremely high, and DR is the leading cause of blindness in people over 18 years old. DR can effectively control vision decline when detected early and intervened in a timely manner. Therefore, the early diagnosis of DR is particularly important.

[0003] The traditional method of manually judging fundus color photos requires many years of experience of doctors and a large amount of time and energy, and has low efficiency. At the same time, patients often cannot obtain the results in real time. In recent years, scholars at home and abroad have mainly solved this problem from two aspects: traditional machine learning methods and deep learning methods.

[0004] Currently, the methods based on deep learning are mainly divided into two categories. The first category is the method of image segmentation based on segmented patches or sliding windows. Although this method can achieve certain results, due to the size limitation of the patches, the receptive field that the model can obtain is small, and it is impossible to obtain global information. Therefore, it is easy to misjudge drusen as exudates. And the patch-based model cannot share operations, and multiple model predictions are required for a single image to obtain the result of the entire image, increasing the algorithm complexity.

[0005] The second category of methods is image segmentation based on full-image input. However, when performing exudate segmentation, this method is limited by the number of model parameters and feature extraction capabilities, and there will be a problem of uneven distribution of exudate sizes. The small exudate areas are very small but numerous, while the large exudate areas are large but few in number, thus affecting the detection accuracy. Summary of the Invention

[0006] In order to solve the above problems existing in the prior art, the present invention provides a method and device for exudate segmentation of fundus lesions based on deep learning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] In a first aspect, the present invention provides a method for exudate segmentation of fundus lesions based on deep learning, including:

[0008] Obtain the fundus image to be processed;

[0009] Construct a deep learning network for exudate segmentation of fundus lesions based on the Encoder-Decoder structure;

[0010] Input the fundus image to be processed into the deep learning network to obtain the exudate segmentation result;

[0011] Among them, the encoding end of the deep learning network includes multiple downsampling modules, which perform multiple downsamplings on the fundus image to be processed and then output the feature map to the decoding end; correspondingly, the decoding end includes multiple upsampling modules, which perform multiple upsamplings on the feature map to finally obtain the exudate segmentation result.

[0012] In an embodiment of the present invention, each of the downsampling modules includes a convolutional layer for halving the image height and width and doubling the number of channels, and a downsampling layer for performing multi-scale feature extraction on the image; wherein, the downsampling layer includes convolutions of various different scales.

[0013] In an embodiment of the present invention, the downsampling layer uses three different scales of convolutions: namely, a 1*1 convolution, a 3*3 convolution, and two consecutive 3*3 convolutions to perform feature extraction on the image and splice to obtain feature maps with different receptive fields; then, a 3*3 convolution is used to fuse the features to obtain a fused feature map.

[0014] In an embodiment of the present invention, each of the upsampling modules includes a transposed convolutional layer for doubling the image height and width.

[0015] In an embodiment of the present invention, after each upsampling module performs upsampling, it further includes: using two consecutive 3*3 convolutions to perform feature fusion on the feature map with the same size in the downsampling module and the feature map obtained by the current-level upsampling module, and using the result as the input of the next-level upsampling module.

[0016] In an embodiment of the present invention, the decoding end further includes a segmentation output head connected to the last-level upsampling module, wherein the segmentation output head includes two consecutive 3*3 convolutions.

[0017] In an embodiment of the present invention, the loss function of the deep learning network is:

[0018] EdgeLoss = α·avg∑(gt i -pred i ) 2 ·ifcnt(i)

[0019]

[0020] Among them, α represents the weight parameter of this loss term, avg represents taking the average value, gt i is a point on the label result gt, and pred iThe point corresponding to the prediction result pred, ξ (i) It represents the result obtained by convolving the i-th point on the graph with the edge extraction operator ξ.

[0021] In a second aspect, the present invention provides a fundus lesion exudate segmentation device based on deep learning, including:

[0022] A data acquisition module for acquiring fundus images to be processed;

[0023] A model construction module for constructing a deep learning network for fundus lesion exudate segmentation based on the Encoder-Decoder structure;

[0024] An image segmentation module for inputting the fundus image to be processed into the deep learning network to obtain an exudate segmentation result;

[0025] Among them, the encoding end of the deep learning network includes a plurality of downsampling modules to perform multiple downsamplings on the fundus image to be processed and then output a feature map to the decoding end; correspondingly, the decoding end includes a plurality of upsampling modules to perform multiple upsamplings on the feature map to finally obtain an exudate segmentation result.

[0026] In a third aspect, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0027] The memory is used to store a computer program;

[0028] The processor is used to implement the method steps described in the above embodiments when executing the program stored in the memory.

[0029] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps described in the above embodiments are implemented.

[0030] Advantages of the present invention:

[0031] The fundus lesion exudate segmentation method based on deep learning provided by the present invention uses the Encoder-Decoder structure to construct a fully convolutional deep learning network for fundus lesion exudate segmentation. This network model not only has a small number of parameters and a simple algorithm, but also has the ability to quickly extract multi-scale features of images, thereby improving the accuracy of detection results.

[0032] The following will further describe the present invention in detail with reference to the accompanying drawings and embodiments. Description of the Drawings

[0033] Figure 1It is a schematic flowchart of a fundus lesion exudate segmentation method based on deep learning provided by an embodiment of the present invention;

[0034] Figure 2 It is a schematic diagram of a model structure of a deep learning network provided by an embodiment of the present invention;

[0035] Figure 3 It is a schematic diagram of the process of a deep learning network model processing pictures provided by an embodiment of the present invention;

[0036] Figure 4 It is an example diagram of an edge extraction operator provided by an embodiment of the present invention;

[0037] Figure 5 It is a schematic diagram of the structure of a fundus lesion exudate segmentation device based on deep learning provided by an embodiment of the present invention;

[0038] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;

[0039] Figure 7 It is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present invention. Specific embodiments

[0040] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0041] Embodiment 1

[0042] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a fundus lesion exudate segmentation method based on deep learning provided by an embodiment of the present invention, and it includes:

[0043] S1: Obtain the fundus picture to be processed;

[0044] Specifically, the fundus picture to be detected can be obtained through an imaging device and processed to a certain size, such as 1440x960 pixels. This size can not only retain most of the picture information but also prevent the video memory from being insufficient to support model inference due to the overly large picture size.

[0045] S2: Construct a deep learning network for fundus lesion exudate segmentation based on the Encoder-Decoder structure;

[0046] Among them, the encoding end of the deep learning network includes multiple downsampling modules to perform multiple downsamplings on the fundus picture to be processed and then output a feature map to the decoding end; correspondingly, the decoding end includes multiple upsampling modules to perform multiple upsamplings on the feature map to finally obtain the exudate segmentation result.

[0047] In this embodiment, the encoding end of the deep learning network model includes three levels of downsampling modules, and the decoding end includes two levels of upsampling modules. Please refer to Figure 2 , Figure 2 which is a schematic diagram of a model structure of the deep learning network provided by the embodiment of the present invention.

[0048] Among them, each of the downsampling modules includes a convolutional layer for halving the height and width of the image and doubling the number of channels, and a downsampling layer for performing multi-scale feature extraction on the image; wherein, the downsampling layer includes convolutions of multiple different scales.

[0049] Specifically, the downsampling layer uses three different scales of convolutions: namely, a 1*1 convolution, a 3*3 convolution, and two consecutive 3*3 convolutions to extract features from the image and splice them to obtain feature maps with different receptive fields; then, the features are fused through a 3*3 convolution to obtain a fused feature map.

[0050] Optionally, as an implementation manner, the first-level downsampling module includes a first convolutional layer with parameters c7 and s2 to halve the height and width of the image and double the number of channels. Wherein, c represents the convolutional kernel size and s represents the sliding step. The downsampling layer in the first-level downsampling module includes a convolutional layer with parameters c1 and s1, a convolutional layer with parameters c3 and s1, and two consecutive convolutional layers with parameters c3 and s1 to extract features from the image at different scales and splice them to obtain feature maps with different receptive fields; then, the features are fused through a convolutional layer with parameters c3 and s1 to obtain a first fused feature map.

[0051] The second-level downsampling module includes a second convolutional layer with parameters c3 and s2 to halve the height and width of the first fused feature map and double the number of channels. The structure of the downsampling layer in the second-level downsampling module is the same as that of the first level. The convolutional layer with parameters c1 and s1, the convolutional layer with parameters c3 and s1, and two convolutional layers with parameters c3 and s1 are continuously used to extract features from the image at different scales and splice them to obtain feature maps with different receptive fields; then, the features are fused through a convolutional layer with parameters c3 and s1 to obtain a second fused feature map.

[0052] The third-level downsampling module includes a third convolutional layer with parameters c3 and s2 to halve the height and width of the second fused feature map and double the number of channels. The downsampling layer in the third-level downsampling module also includes a convolutional layer with parameters c1 and s1, a convolutional layer with parameters c3 and s1, and two consecutive convolutional layers with parameters c3 and s1 to extract features from the image at different scales respectively, and splice the obtained feature maps with different receptive fields; then, the features are fused through a convolutional layer with parameters c3 and s1 to obtain the third fused feature map, which is transmitted to the decoding end as the output of the encoding end.

[0053] Since most of the exudative lesion points in the segmentation task are tiny, in this embodiment, a smaller convolutional kernel is used in the upsampling module, so that the part with a small receptive field in the feature map can better complete the fine segmentation task; at the same time, multiple downsamplings and stacked 3*3 convolutions are used to achieve the purpose of extracting features at a larger scale.

[0054] At the decoding end, each of the upsampling modules includes a deconvolution layer for doubling the height and width of the image.

[0055] Optionally, in this embodiment, all three upsampling modules use a deconvolution layer with parameters Dc3 and s2 for upsampling. Among them, Dc is the deconvolution kernel size and s is the stride.

[0056] In this embodiment, after upsampling in each of the upsampling modules, it further includes: using two consecutive 3*3 convolutions to fuse the feature maps of the same size in the downsampling module with the feature maps obtained from the current-level upsampling module, and taking the result as the input of the next-level upsampling module, as Figure 2 shown.

[0057] Specifically, the third fused feature map obtained by the third-level downsampling module at the encoding end is used as the input of upsampling 1 at the decoding end, the second fused feature map obtained by the second-level downsampling module at the encoding end is fused with the output image of upsampling 1 at the decoding end as the input of upsampling 2, and the first fused feature map obtained by the first-level downsampling module at the encoding end is fused with the output image of upsampling 2 at the decoding end.

[0058] In addition, the decoding end further includes a segmentation output head connected to the last-level upsampling module. Among them, the segmentation output head includes two consecutive 3*3 convolutions. The first fused feature map obtained by the first-level downsampling module at the encoding end and the output image of the second-level upsampling module at the decoding end are fused through two convolutions with parameters c3 and s1, and then output a feature map with the same size as the original image through the segmentation output head.

[0059] In this embodiment, deconvolution is used for upsampling, doubling the height and width of the feature map and halving the number of channels. Compared with the traditional interpolation method that needs to double the entire feature map first and then perform convolution operations, the deconvolution method saves the memory space required for operations. In addition, in order to maximally retain the detailed information extracted by downsampling at the encoding end, that is, the features extracted by convolution with a small receptive field, the feature map in the upsampling module is concatenated with the upsampled feature map.

[0060] S3: Input the to-be-processed fundus image into the deep learning network to obtain an exudate segmentation result.

[0061] Specifically, input the to-be-processed fundus image obtained in step S1 into the deep learning network constructed in step S2 to obtain an image with a pixel channel number of 2 and the same size as the input original image (i.e., the to-be-processed fundus image). Among them, each channel of the output image represents the classification result of each class (i.e., the 0th channel represents the 0th class, and the 1st channel represents the 1st class). After performing Softmax operation on the output image, extract the result of the 1st channel, and consider those with a confidence level > 0.99 as the 1st class, and others as the 0th class. Thus, an exudate segmentation result is obtained.

[0062] To more clearly describe the process of the deep learning network provided by the present invention for processing images, the following combines Figure 3 , taking an image with an input size of 3 * 960 * 1440 as an example, to illustrate the processing process of the deep learning network.

[0063] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the process of the deep learning network model provided by the embodiment of the present invention for processing images.

[0064] First, input the image 3 * 960 * 1440. After being processed by the first convolutional layer c7 and s2, a feature map of 64 * 480 * 720 is obtained. Then, through c1 and s1, c3 and s1, and 2 * (c3 and s1) in the first-level downsampling module respectively, three feature maps with different scales are obtained, namely 64 * 480 * 720, 32 * 480 * 720, and 32 * 480 * 720. Concatenate these three feature maps and fuse them through the convolutional layer c3 and s1 to obtain a first fused feature map of 64 * 480 * 720. Thus, the first-level downsampling is completed.

[0065] Then, the second-level downsampling is performed. Specifically, the obtained first fusion feature map of 64*480*720 is processed by the second convolutional layer c3, s2 to obtain a feature map of 128*240*360, and then through c1, s1, c3, s1 and 2*(c3, s1) in the second-level downsampling module respectively, three feature maps of different scales are obtained, namely 128*240*360, 64*240*360 and 64*240*360; these three feature maps are concatenated and fused through the convolutional layer c3, s1 to obtain a second fusion feature map of 128*240*360. Thus, the second-level downsampling is completed.

[0066] Next, the third-level downsampling is carried out. Specifically, the obtained second fusion feature map of 128*240*360 is processed by the third convolutional layer c3, s2 to obtain a feature map of 256*120*180, and then through c1, s1, c3, s1 and 2*(c3, s1) in the third-level downsampling module respectively, three feature maps of different scales are obtained, namely 256*120*180, 64*120*180 and 64*120*180; these three feature maps are concatenated and fused through the convolutional layer c3, s1 to obtain a third fusion feature map of 256*120*180, which is used as the input of the decoding end and input into the first upsampling module of the decoding end.

[0067] The first upsampling module at the decoding end processes the third fusion feature map of 256*120*180 through a deconvolutional layer of Dc3, s2 to obtain a fourth feature map of 256*240*360. It is convolved with the second feature map of 128*240*360 through 2*(c3, s1) to obtain a fifth feature map of 128*240*360 as the input of the second upsampling.

[0068] The second upsampling module processes the fifth feature map of 128*240*360 through a deconvolutional layer of Dc3, s2 to obtain a sixth feature map of 128*480*720. It is convolved with the first feature map of 64*480*720 through 2*(c3, s1) to obtain a seventh feature map of 64*480*720.

[0069] The seventh feature map of 64*480*720 is input into the segmentation output head, and after deconvolutional processing of Dc3, s2 and convolutional processing of 2*(c3, s1), an output image of 2*960*1440 is finally obtained.

[0070] The exudate segmentation method for fundus lesions based on deep learning provided in this embodiment constructs a fully convolutional deep learning network for fundus lesion exudate segmentation using an Encoder-Decoder structure. This network model not only has a small number of parameters and a simple algorithm, but also has the ability to quickly extract multi-scale features of images, thereby improving the accuracy of detection results.

[0071] In another embodiment of the present invention, an improved loss function is also proposed for the above deep learning network. It is mainly optimized for the segmentation of complex object boundaries, and an EdgeLoss loss term is added.

[0072] First, use the edge extraction operator ξ to extract the edges predicted by the model and the edges of the label annotations. The definition of ξ is as Figure 4 shown.

[0073] Since exudates are labeled as the first class, the exudate parts should also be predicted as class 1 during prediction. Use ξ to perform convolution operations on the model prediction map. Ideally, when the center of ξ is located at the exudate edge, the convolution operation result should be positive. When ξ is located in the background area or the middle area of the exudate, the convolution operation result is 0. When the center of ξ is located 1 pixel outside the edge area, the convolution operation result is negative. The points with the absolute value of the result greater than 0.1 are used to calculate the loss and increase the penalty parameter. ξ (i) represents the result obtained by performing convolution on point i on the graph using ξ. Let the ground truth label result be gt and the prediction result be pred. gt i is a point on gt, and pred i is the corresponding point on pred. Then the EdgeLoss is defined as follows:

[0074] EdgeLoss = α·avg∑(gt i - pred i ) 2 ·ifcnt(i)

[0075]

[0076] By adding the EdgeLoss loss term in this embodiment, the performance of the network model can be comprehensively improved.

[0077] Embodiment 2

[0078] To verify the beneficial effects of the present invention, the method of the present invention will be compared and described below through three data sets.

[0079] This embodiment uses the OIA-DDR data set, the E_ophtha_EX data set, and the IDRiD data set to verify the effectiveness of the present invention.

[0080] In this embodiment, to measure the performance of the model, pixel-level Intersection over union (IoU), Recall, Precision, and F1 score are used for measurement. In this segmentation task, the background (non-exudative lesions) is defined as class 0, and the exudative lesions (foreground) are defined as class 1. When the i-th pixel in a fundus image is of class 1: when the model correctly classifies it as class 1, this pixel is called a true positive pixel (TP); if the model misclassifies it as class 0, this pixel is called a false negative (FN). When the i-th pixel in a fundus image is of class 0: when the model classifies it as class 1, this pixel is called a false positive (FP); if the model correctly classifies it as class 0, it is a true negative (TN). TPs represent all true positive pixels in the predictions. Similarly, FNs, FPs, and TNs represent all false negatives, false positives, and true negatives, respectively.

[0081] Among them, Intersection over union (IoU), Recall, Precision, and F1 are defined as follows:

[0082]

[0083]

[0084]

[0085]

[0086] IoU is the ratio of the intersection to the union of the predicted foreground and the true foreground shapes. Simply put, the larger the intersection over union ratio, the more consistent the prediction is with the actual situation.

[0087] Recall is the ratio of true positives in the predictions to all actual positive samples, reflecting the proportion of positive samples recalled. The larger the Recall value, the stronger the model's ability to detect lesions, and vice versa.

[0088] Precision is the ratio of true positives in the predictions to the sum of true positives and false positives, reflecting the proportion of true positives among all predicted positive samples. The larger the Precision value, the fewer false positives the model predicts, and vice versa.

[0089] Since Recall and Precision cannot comprehensively reflect the performance of the model, the F1 score is introduced. The F1 score assumes that Recall and Precision are equally important, and the larger the F1 score when both Recall and Precision values are larger.

[0090] 1. The existing HED method, DeepLab-v3+, and the method of the present invention were compared on the OIA-DDR dataset, and the results are shown in Table 1 as follows:

[0091] Table 1 Comparison of various methods on the OIA-DDR dataset

[0092]

[0093] As can be seen from Table 1, on the OIA-DDR dataset, the IoU of the present method on the validation set is 0.4941, far exceeding the performance of the DeepLab-v3+ and HED algorithms. At the same time, the IoU on the test set is 0.4140, and this performance also greatly exceeds the existing methods.

[0094] 2. The existing HED method, FCRN method, DeepLab-v3+, LWEnet, and the method of the present invention were compared on the E_ophtha_EX dataset, and the results are shown in Table 2 as follows:

[0095] Table 2 Comparison of various methods on the E_ophtha_EX dataset

[0096]

[0097]

[0098] Among them, the above methods all use five-fold cross-validation to verify the model performance.

[0099] As can be seen from Table 2, on the E_ophtha_EX dataset, except that the Precision of the present method is slightly less than that of the HED algorithm, the Recall and F1 values far exceed other methods.

[0100] 3. The existing HED method, FCRN method, DeepLab-v3+, LWEnet, and the method of the present invention were compared on the IDRiD dataset, and the results are shown in Table 3 as follows:

[0101] Table 3 Comparison of various methods on the IDRiD dataset

[0102]

[0103] As can be seen from Table 3, on the IDRiD dataset, except that the Precision of the present method is slightly less than that of LWEnet, the Recall and F1 values are the best among all methods.

[0104] 4. On the IDRiD dataset, the HED method, FCRN method, DeepLab-v3+, LWEnet, and the method of the present invention after adding the EdgeLoss term were compared, and the results are shown in Table 4:

[0105] Table 4

[0106]

[0107]

[0108] As can be seen from Table 4, after adding the EdgeLoss loss term, the performance of the model has been comprehensively improved. The Recall value reached 0.7945, the Precision reached 0.7976, and the F1 value was 0.7960, and the results all exceeded the performance of other methods.

[0109] In summary, the fundus lesion exudate segmentation method based on deep learning provided by the present invention has surpassed the performance of existing methods on multiple datasets with fewer model parameters and faster inference time, and has important application value.

[0110] Example 3

[0111] Based on the above Example 1, this example provides a fundus lesion exudate segmentation device based on deep learning. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a fundus lesion exudate segmentation device provided by an embodiment of the present invention, and it includes:

[0112] A data acquisition module for acquiring fundus images to be processed;

[0113] A model construction module for constructing a deep learning network for fundus lesion exudate segmentation based on the Encoder-Decoder structure;

[0114] An image segmentation module for inputting the fundus image to be processed into the deep learning network to obtain an exudate segmentation result;

[0115] Among them, the encoding end of the deep learning network includes a plurality of downsampling modules to perform multiple downsamplings on the fundus image to be processed and output a feature map to the decoding end; correspondingly, the decoding end includes a plurality of upsampling modules to perform multiple upsamplings on the feature map to finally obtain an exudate segmentation result.

[0116] The device provided in this example can implement the method provided in the above Example 1, and the detailed process will not be repeated here.

[0117] Example 4

[0118] Based on the above-mentioned First Embodiment, this embodiment provides an electronic device. Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0119] The memory is used to store computer programs;

[0120] When the processor is used to execute the program stored on the memory, it implements the method steps provided in the above-mentioned First Embodiment. The implementation principle and technical effects are similar and will not be elaborated here.

[0121] Embodiment Five

[0122] Based on the above-mentioned First Embodiment, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer-readable storage medium provided by an embodiment of the present invention. A computer-readable storage medium provided by this embodiment stores a computer program. When the above-mentioned computer program is executed by a processor, it implements the method steps provided in the above-mentioned First Embodiment. The implementation principle and technical effects are similar and will not be elaborated here.

[0123] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for exudate segmentation of fundus lesions based on deep learning, characterized in that Including: Obtain the fundus image to be processed; Construct a deep learning network for fundus lesion exudate segmentation based on the Encoder-Decoder structure; Input the fundus image to be processed into the deep learning network to obtain the exudate segmentation result; Among them, the encoding end of the deep learning network includes multiple downsampling modules, which perform multiple downsamplings on the fundus image to be processed and then output the feature map to the decoding end; each downsampling module includes a convolutional layer for halving the image height and width and doubling the number of channels, and a downsampling layer for multi-scale feature extraction of the image; among them, the downsampling layer includes convolutions of multiple different scales; the downsampling layer uses three different scales of convolutions: namely, a 1*1 convolution, a 3*3 convolution, and two consecutive 3*3 convolutions to perform feature extraction on the image and splice to obtain feature maps with different receptive fields; then, the features are fused through a 3*3 convolution to obtain a fused feature map; Correspondingly, the decoding end includes multiple upsampling modules, which perform multiple upsamplings on the feature map to finally obtain the exudate segmentation result; each upsampling module includes a transposed convolutional layer for doubling the image height and width; After each upsampling module performs upsampling, it further includes: using two consecutive 3*3 convolutions to fuse the feature maps of the same size in the downsampling module with the feature map obtained by the current upsampling module at the current level, and using the result as the input of the next upsampling module; The decoding end further includes a segmentation output head connected to the last upsampling module, where the segmentation output head includes two consecutive 3*3 convolutions; The loss function of the deep learning network is: EdgeLoss = α·avg∑(gt i - pred i ) 2 ·ifcnt(i) Among them, α represents the weight parameter, avg represents taking the average value, and gt i is a point on the label result gt, and pred i is the corresponding point of the prediction result pred, and ξ( i ) represents the result obtained by performing a convolution operation on the i-th point on the graph using the edge extraction operator ξ.

2. A fundus lesion exudate segmentation device based on deep learning, which is used to implement the method described in claim 1, and is characterized in that, Including: A data acquisition module for obtaining the fundus image to be processed; A model construction module for constructing a deep learning network for fundus lesion exudate segmentation based on the Encoder-Decoder structure; An image segmentation module for inputting the fundus image to be processed into the deep learning network to obtain the exudate segmentation result; Among them, the encoding end of the deep learning network includes multiple downsampling modules, which perform multiple downsamplings on the fundus image to be processed and then output the feature map to the decoding end; correspondingly, the decoding end includes multiple upsampling modules, which perform multiple upsamplings on the feature map to finally obtain the exudate segmentation result.

3. An electronic device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in claim 1 when executing the programs stored on the memory.

4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method steps described in claim 1.

Citation Information

Patent Citations

  • Sugar net analysis method and system and electronic equipment

    CN113576399A