An Optic Disc and Cup Segmentation Method Based on MLP and CNN

By combining MLP and CNN technical means, the lightweight segmentation method is adopted to solve the problems of low accuracy and slow inference speed in visual disc cup segmentation, achieving high precision, low complexity and fast inference effects, which are suitable for real-time requirements of industrial applications.

CN115578562BActive Publication Date: 2025-06-27SHANXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211258174.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2025-06-27
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

The prior art has problems of low accuracy and slow inference speed in the visual disc cup segmentation task, which cannot meet the real-time requirements of industrial applications.

Method used

A lightweight segmentation method based on MLP and CNN is adopted, and through technical means such as convolution module, SE-Block, SR-MLP module and pyramid average pooling, image features are extracted and improved, receptive fields are expanded, and details are restored through cross-layer connections to achieve fast and accurate visual disc and cup segmentation.

Benefits of technology

It realizes high-precision visual disc and cup segmentation, has low model complexity and fast inference speed, can greatly reduce hardware costs, and is suitable for real-time requirements of industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003889761080000011
    Figure HDA0003889761080000011
  • Figure HDA0003889761080000012
    Figure HDA0003889761080000012
  • Figure HDA0003889761080000013
    Figure HDA0003889761080000013
Patent Text Reader

Abstract

A method for optic disc and cup segmentation based on MLP and CNN, belonging to the field of medical image segmentation; it solves the problems of low segmentation accuracy and slow inference speed of optic disc and cup. The technical solution is as follows: (1) In the first stage, a convolutional neural network and a channel attention mechanism are used to initially extract image features; (2) In the second and third stages, a convolutional neural network and an RB-MLP module are used to induce the network to focus on location information and mark features to improve the feature quality; (3) In the fourth and fifth stages, a convolutional neural network and an R-MLP module are used to mark features to improve the feature quality; (4) In the sixth stage, the feature maps of the second to fifth stages are fused, and pyramid average pooling is used to expand the receptive field of the features; (5) In the seventh stage, the feature map is upsampled and cross-layer connected with the feature map output by the first layer to restore the lost details; (6) In the eighth stage, the feature map is upsampled to the original size, and the final prediction result is output through 1×1 convolution. The present invention has the advantages of high segmentation accuracy, real results, low network complexity, and fast inference speed, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and particularly relates to a lightweight segmentation method based on the fusion of MLP and CNN. Background Art

[0002] Glaucoma, as an irreversible chronic neurodegenerative disease, is one of the three major blinding diseases in the world. Since the optic nerve damage and vision loss caused by glaucoma are irreversible, early screening and diagnosis of glaucoma play a crucial role in maintaining vision. Experienced ophthalmologists mainly evaluate through indicators such as the diameter ratio of the optic cup to the optic disc, i.e., the cup to disk ratio (CDR), in the early screening and clinical diagnosis stages. Generally speaking, the cup to disk ratio is proportional to the probability of having glaucoma. However, manually dividing the OD and OC is a time-consuming task and requires experienced experts to operate personally, with a high cost. Therefore, using computer algorithms to assist in dividing the OD and OC has become a research hotspot. Currently, the mainstream segmentation methods mainly include traditional methods, deep learning methods based on convolution, and deep learning methods based on Transformer.

[0003] Among them, the research on traditional OD and OC segmentation is based on handcrafted features, such as handcrafted features like color, texture, contrast, and gradient. Although classical methods have achieved certain success in the optic disc and optic cup segmentation tasks. However, traditional methods have the problem of low accuracy and cannot meet the requirements of industrial applications;

[0004] Deep learning methods based on convolution have achieved great success in natural image segmentation and medical image segmentation, and they have the advantages of high accuracy and strong generalization. For example, Zhao et al. added an attention gate between the encoder and decoder of U-Net to focus on the target area. However, the above methods are all restricted by the receptive field and cannot fully consider global information;

[0005] Deep learning methods based on Transformer can consider the global features of images, which helps to improve the accuracy of the network. For example, Junde Wu et al. associated each single low-level feature with multi-scale features, and then used the segmentation features to interact with the diagnostic features for modeling, and achieved good results in the OD and OC segmentation tasks and glaucoma screening. However, the above methods have high complexity, slow inference speed, and high requirements for hardware, and cannot meet the real-time requirements of the industrial application field. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems of low accuracy and slow inference speed existing in the above-mentioned prior art, and provide a faster inference speed optic disc and optic cup segmentation method based on MLP and CNN.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a method for segmenting optic disc and optic cup based on MLP and CNN, wherein: the following steps are included:

[0008] (1) Use the convolutional module and SE-Block to initially extract image features and output a feature map with 32 output channels;

[0009] The convolutional module consists of a single 3x3 convolution operation, a single BatchNormal operation, and an H-Swish activation function, mainly used for feature extraction. Among them, a single convolution operation can ensure efficient inference speed. H-Swish has the characteristics of no upper bound, having a lower bound, being smooth, and non-monotonic, and can also avoid calculating the complex sigmod function, having great advantages in deep models. The lightweight SE-Block channel attention module can not only improve the quality of the feature map but also not increase the complexity of the network.

[0010] (2) The second and third stages: the convolutional module + SR-MLP module to improve the feature quality;

[0011] First, use the convolutional module to further extract features. Subsequently, use the SR-MLP module to label the feature map to improve the quality of the feature map. The SR-MLP module has MLP as the backbone and has a faster speed than the fully convolutional network. Specifically: the first step: divide the feature map into 5 groups on average for translation: 2 units to the upper right; 1 unit to the upper right; no movement; 1 unit to the lower left; 2 units to the lower left. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into MLP and DW-Conv in sequence, where MLP quadruples the number of channels to obtain richer features, and DW-Conv is used to encode the position information of the feature map. The second step: divide the feature map into 5 groups on average for translation: 2 units to the lower left; 1 unit to the lower left; no movement; 1 unit to the upper right; 2 units to the lower right. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into MLP and DW-Conv in sequence, where MLP reduces the number of channels by four times. The third step: perform a residual connection between the current feature map and the feature map after the first translation, which can fuse adjacent feature information and prevent performance degradation in the deep network. Finally, output a feature map with 64 output channels. See formulas (1)(2)(3) for details;

[0012] X shift = Shift w,h (X) (1)

[0013] X stage1 = DWConv(MLP(X shift )) (2)

[0014] Xoutput = DWConv(MLP(Shift w,h (X stage1 )))+X shift (3)

[0015] where X represents the feature map, Shift w,h () represents the translation splitting function, DWConv represents the DW-Conv operation, and MLP represents the multi-layer perceptron.

[0016] (3) Fourth and fifth stages: The convolution module + S-MLP module improve the feature quality;

[0017] First, use the convolution module to further extract features. Subsequently, use the S-MLP module to label the feature map to improve the quality of the feature map. Since the deep features contain sufficiently broad information and considering the downstream design, the faster S-MLP module is used. Specifically: Step 1: Divide the feature map into 5 groups on average for translation: 2 units to the upper right; 1 unit to the upper right; no movement; 1 unit to the lower left; 2 units to the lower left. Then, crop and merge the excess parts of each group into a feature map. Input this feature map into MLP and DW-Conv in sequence, where MLP quadruples the number of channels. Then, perform Step 2: Input this feature map into MLP and DW-Conv in sequence, where MLP reduces the number of channels by four times. Step 3: Perform a residual connection between the current feature map and the feature map after the first translation, which can prevent performance degradation in the deep network. Finally, output a feature map with 96 channels.

[0018] See formulas (4)(5)(6) for details;

[0019] X shift = Shift w,h (X) (4)

[0020] X stage1 = DWConv(MLP(X shift )) (5)

[0021] X output = DWConv(MLP(X stage1 ))+X shift (6)

[0022] where X represents the feature map, Shift w,h () represents the translation splitting function, DWConv represents the DW-Conv operation, and MLP represents the multi-layer perceptron.

[0023] (4) Sixth stage: Use pyramid average pooling to expand the receptive field;

[0024] After upsampling the feature maps of the second and third stages to the same size, perform residual connection to obtain X1. After upsampling the feature maps of the fourth and fifth stages to the same size, perform residual connection to obtain X2. Subsequently, connect X1 and X2, and then use pyramid average pooling to further expand the receptive field of the feature map. The sizes of the average pooling are 1, 2, 3, and 6 respectively. Larger-sized average pooling can achieve the purpose of expanding the receptive field, while smaller-sized average pooling can supplement the detailed information lost by larger-sized average pooling. Finally, output a feature map with 64 output channels.

[0025] (5) The seventh stage: After upsampling, perform cross-layer connection with the feature map output in the first stage to restore the lost details, and output a feature map with 32 output channels;

[0026] (6) The eighth stage: Upsample to the original size and output the final prediction result via 1×1 convolution.

[0027] This solution has good operability and portability, can train excellent accuracy in a smaller sample set, and has high segmentation accuracy, low model complexity, and fast inference speed, which can greatly reduce the hardware cost. Brief Description of the Drawings

[0028] Figure 1 is a flowchart of a method for optic disc and cup segmentation based on MLP and CNN according to the present invention.

[0029] Figure 2 is a flowchart of RB-MLP of the present invention.

[0030] Figure 3 is a flowchart of R-MLP of the present invention. Detailed Embodiment

[0031] As Figures 1 to 3 shown, the method for optic disc and cup segmentation based on MLP and CNN described in this embodiment, wherein: includes the following steps:

[0032] (1) The first stage: Use a convolution module and an SE-Block to initially extract image features, and output a feature map with 32 output channels;

[0033] (2) The second and third stages: First, use the convolutional module to further extract features. Subsequently, perform the first feature marking: Divide the feature map into 5 groups on average for translation: 2 units to the upper right; 1 unit to the upper right; no movement; 1 unit to the lower left; 2 units to the lower left. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into the MLP and DW-Conv in sequence, where the MLP quadruples the number of channels. Then, perform the second feature marking operation: Divide the feature map into 5 groups on average for translation: 2 units to the lower left; 1 unit to the lower left; no movement; 1 unit to the upper right; 2 units to the lower right. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into the MLP and DW-Conv in sequence, where the MLP reduces the number of channels by four times. Finally, perform a residual connection between the current feature map and the feature map after the first translation, and output a feature map with 64 channels. See Formulas (1), (2), and (3) for details;

[0034] X shift = Shift w,h (X) (1)

[0035] X stage1 = DWConv(MLP(X shift )) (2)

[0036] X output = DWConv(MLP(Shift w,h (X stage1 )))+X shift (3)

[0037] where X represents the feature map, Shift w,h () represents the translation splitting function, DWConv represents the DW-Conv operation, and MLP represents the multi-layer perceptron.

[0038] (3) The fourth and fifth stages: First, use the convolutional module to further extract features. Subsequently, perform the first feature marking: Divide the feature map into 5 groups on average for translation: 2 units to the upper right; 1 unit to the upper right; no movement; 1 unit to the lower left; 2 units to the lower left. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into the MLP and DW-Conv in sequence, where the MLP quadruples the number of channels. Then, perform the second feature marking operation: Input this feature map into the MLP and DW-Conv in sequence, where the MLP reduces the number of channels by four times. Finally, perform a residual connection between the current feature map and the feature map after the first translation, and output a feature map with 96 channels. See Formulas (4), (5), and (6) for details;

[0039] X shift = Shift w,h (X) (4)

[0040] Xstage1 = DWConv(MLP(X shift )) (5)

[0041] X output = DWConv(MLP(X stage1 )) + X shift (6)

[0042] where X represents the feature map, Shift w,h () represents the translation splitting function, DWConv represents the DW-Conv operation, and MLP represents the multi-layer perceptron.

[0043] (4) The sixth stage: The feature maps of the second and third stages are upsampled to the same size and then residual connections are made to obtain X1. The feature maps of the fourth and fifth stages are upsampled to the same size and then residual connections are made to obtain X2. Subsequently, X1 and X2 are connected, and after pyramid average pooling with sizes of 1, 2, 3, and 6, a feature map with 64 output channels is output;

[0044] (5) The seventh stage: After upsampling, cross-layer connection is made with the feature map output in the first stage to restore the lost details, and a feature map with 32 output channels is output;

[0045] (6) The eighth stage: Upsample to the original size and output the final prediction result through a 1×1 convolution.

[0046] Among them, the convolution module consists of one 3x3 convolution operation, one BatchNormal operation, and one H-Swish activation function.

Claims

1. A method for optic disc and cup segmentation based on MLP and CNN, characterized in that: It includes the following steps: (1) The first stage: Use the convolutional module and SE-Block to initially extract image features and output a feature map with 32 output channels; (2) The second and third stages: First, use the convolutional module to further extract features; Subsequently, perform the first feature marking: Divide the feature map into 5 groups on average for translation: 2 units to the upper right; 1 unit to the upper right; no movement; 1 unit to the lower left; 2 units to the lower left. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into MLP and DW-Conv in sequence, where MLP quadruples the number of channels. Then, perform the second feature marking operation: Divide the feature map into 5 groups on average for translation: 2 units to the lower left; 1 unit to the lower left; no movement; 1 unit to the upper right; 2 units to the lower right. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into MLP and DW-Conv in sequence, where MLP reduces the number of channels by four times. Finally, perform a residual connection between the current feature map and the feature map after the first translation to output a feature map with 64 output channels; see Formulas (1), (2), and (3) for details; X shift = Shift w,h (X) (1) X stage1 = DWConv(MLP(X shift )) (2) X output = DWConv(MLP(Shift w,h (X stage1 )))+X shift (3) where X represents the feature map, DWConv represents the DW-Conv operation, and MLP represents the multi-layer perceptron; (3) The fourth and fifth stages: First, use the convolutional module to further extract features; Subsequently, perform the first feature marking: Divide the feature map into 5 groups on average for translation: 2 units to the upper right; 1 unit to the upper right; no movement; 1 unit to the lower left; 2 units to the lower left. Then, crop and merge the excess parts of each group into a single feature map. Input this feature map into MLP and DW-Conv in sequence, where MLP quadruples the number of channels. Then, perform the second feature marking operation: Input this feature map into MLP and DW-Conv in sequence, where MLP reduces the number of channels by four times. Finally, perform a residual connection between the current feature map and the feature map after the first translation to output a feature map with 96 output channels; see Formulas (4), (5), and (6) for details; X shift = Shift w,h (X) (4) X stage1 = DWConv(MLP(X shift )) (5) X output = DWConv(MLP(X stage1 )) + X shift (6) (4) The sixth stage: Upsample the feature maps of the second and third stages to the same size and then perform a residual connection to obtain X1. Upsample the feature maps of the fourth and fifth stages to the same size and then perform a residual connection to obtain X2. Subsequently, connect X1 and X2, and perform pyramid average pooling with sizes of 1, 2, 3, and 6 to output a feature map with 64 output channels; (5) The seventh stage: After upsampling, perform a cross-layer connection with the feature map output in the first stage to restore the lost details and output a feature map with 32 output channels; (6) The eighth stage: Upsample to the original size and output the final prediction result through a 1×1 convolution.

2. The method for optic disc and cup segmentation based on MLP and CNN according to claim 1, wherein: The convolutional module consists of a 3x3 convolution operation, a BatchNormal operation, and an H-Swish activation function.

Citation Information

Patent Citations

  • Image noise reduction method and device, electronic equipment and storage medium

    CN113850741A

  • Eye fundus image optic cup and optic disc segmentation method under unified framework

    CN113870270A