Automatic retinopathy segmentation system for fundus image
By combining the hybrid architecture of CNN and Transformer, the fundus image retinopathy automatic segmentation system using MSFormer Block and downsampling units, the segmentation problem of small target areas is solved, and retinopathy segmentation is achieved with higher accuracy and robust retinopathy segmentation.
Patent Information
- Application Number
- CN202510483147.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
When the existing retinopathy segmentation model is difficult to accurately segment and is prone to lose key information when dealing with small target areas, resulting in inaccurate segmentation results.
A fundus image retinopathy automatic segmentation system is adopted, combined with the hybrid architecture of CNN and Transformer, and small target areas are extracted through MSFormer Block and downsampling units, and decoded using a jump connection mechanism to avoid the loss of target information caused by traditional convolution, and decompose features through Haar wavelet transformation to reduce the calculation amount.
Effectively capture small target areas of retinopathy, improve segmentation accuracy, reduce calculation amount, and enhance the accuracy and robustness of the model when processing fine-grained features.
Smart Images

Figure BDA0005363363150000071 
Figure BDA0005363363150000091 
Figure BDA0005363363150000111
Abstract
Description
Technical Field
[0001] The field of the present invention relates to the fields of artificial intelligence, deep learning, and medical image processing, and particularly relates to an automatic segmentation system for retinal lesions in fundus images. Background Art
[0002] As important medical images, fundus images can provide important information about eye diseases. Retinal lesions, especially diseases such as diabetic retinopathy (DR) and age-related macular degeneration (AMD), cause severe visual impairment globally. Therefore, quickly and accurately extracting the retinal lesion area from fundus images is of great significance for the comprehensive analysis and treatment of diseases. Ophthalmologists usually rely on clinical and geometric features to manually identify and label these lesion areas, but this process is not only time-consuming and laborious but also vulnerable to subjective factors, which may limit the accuracy and consistency of diagnosis. Therefore, studying automatic segmentation methods in fundus image analysis is particularly important.
[0003] In recent years, with the rapid development of deep learning technology, many advanced automatic segmentation methods have emerged, especially in the field of medical image segmentation, where significant progress has been made. The automatic segmentation method based on convolutional neural network (CNN), especially the U-Net [1] architecture, has become the mainstream technology in the field of medical image segmentation. However, CNN extracts features in the local receptive field through convolutional operations, which limits its ability to capture global information. Especially when dealing with complex structures or long-range dependencies, it usually needs to expand the receptive field by increasing the network depth or using a larger convolutional kernel, but this will lead to a sharp increase in computational complexity.
[0004] To make up for the deficiencies of CNN in global information and long-range dependence modeling, the Transformer [2] architecture, as an emerging deep learning model, has gradually been introduced into the field of image processing. The great success of Transformer in the field of natural language processing (NLP) has also promoted its application in image understanding tasks. However, although Transformer has powerful global modeling capabilities, it is not as efficient as CNN in processing fine-grained local features (such as edges, textures, etc.).
[0005] To combine the local feature extraction ability of CNN and the global modeling ability of Transformer, many researchers have proposed hybrid architectures that combine CNN and Transformer, such as TransFuse [3] and TransUNe t[4]However, in some specific tasks, especially when dealing with small object segmentation tasks such as retinopathy, there are still many challenges. The target areas of retinopathy are usually very small, often only a few pixels, which makes it difficult to accurately identify them in fundus images. Traditional convolutional and pooling downsampling operations will lead to a reduction in image resolution, thus causing the loss of these small target areas and affecting the segmentation accuracy. Therefore, existing models often cannot effectively capture these subtle areas, resulting in inaccurate segmentation results and easy loss of key information.
[0006] [1] O.Ronneberger, P.Fischer, T.Brox, et al. U-Net: Convolutional networks for biomedical image segmentation[C]. International Conference on Medical Image Computing and Computer-Assisted Intervention, 2015: 234-241.
[0007] [2] Vaswani, A., Shazeer, N., Parmar, N., et al. Attention is all you need. In Advances in Neural Information Processing Systems, 2017: 5998-6008.
[0008] [3] Zhuang, J., Dong, X., Liang, Z., et al. TransFuse: Fusing Transformer and CNNs for Medical Image Segmentation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021: 15617-15627.
[0009] [4] Chen, J., Dou, Q., Yu, L., et al. TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation. In Medical Image Analysis, 2021: 101525. Summary of the Invention
[0010] The present invention provides an automatic segmentation system for fundus image retinopathy to solve the problem of insufficient processing of small targets in the existing retinopathy segmentation.
[0011] To solve the above problems, the present invention adopts the following technical solutions:
[0012] An automatic segmentation system for fundus image retinopathy includes an encoder and a decoder. The encoder includes four serially connected encoding modules: the first encoding module to the fourth encoding module. For a given fundus image, it passes through four serially connected encoding modules in turn. The spatial resolution of the feature map processed by each encoding module is halved, and at the same time, the channel dimension is expanded to twice the original. The first encoding module contains an MSFormer Block, and the second to fourth encoding modules each contain an MSFormer Block and a downsampling unit, combining the characteristics of CNN and Transformer to extract small target regions and prevent the loss of small targets caused by downsampling.
[0013] The decoder contains four serially connected decoding modules: the first decoding module to the fourth decoding module. The first to third decoding modules each contain an MSFormer Block and an upsampling unit, while the fourth decoding module only contains an MSFormer Block. The final output of the encoder is used as the initial input of the first decoding module. The decoding process adopts a skip connection mechanism: First, the output of the first decoding module is concatenated with the features of the third encoding module; then, the concatenated result is input to the second decoding module, and its output is concatenated with the features of the second encoding module; then, the concatenated result is input to the third decoding module and concatenated with the features of the first encoding module; finally, after being processed by the fourth decoding module, the features are restored to the target resolution.
[0014] For the MSFormer Block, first, the input feature map is normalized by LayerNorm. Then, the module uses three different sizes of convolutional kernels (3x3, 5x5, and 7x7) to perform convolutional operations on the input respectively, and the convolutional results of the three scales are averaged and fused. After that, through a 1×1 convolution, a GELU activation function, and a 11×11 convolution and other operations to obtain A, and a 1×1 convolution to obtain V. Next, multiply A and V and then perform a 1×1 convolution to the final output feature: The calculation formula is as follows:
[0015] f′=(Conv 3×3 (LN(X i ))+Conv 3×3 (LN(X i ))+Conv3×3 (LN(X i ))) / 3,
[0016] A = Conv 11×11 (GELU(Conv 1×1 (f′))),
[0017] V(Conv 1×1 (f′), f = Conv 1×1 (A⊙V)
[0018] Wherein, X i (i ∈ [1~4]) represents the input features of each MSFmormer Block, ⊙ represents the Hadamard product, LN represents the normalization operation, and f represents the output features of each MSFmormer Block.
[0019] For the described downsampling unit, first, perform 1×1 convolution on the input features, and then decompose them into low-frequency features and high-frequency features through wavelet transform. Secondly, perform 1×1 convolution block operation on the low-frequency features, and then obtain a gating coefficient through the Sigmoid activation function; perform 1×1 convolution block operation on the high-frequency features and then multiply them by the gating coefficient obtained from the low-frequency features to get the final output: The calculation formula is as follows:
[0020] f L , f H = Conv 1×1 (DWT(f)),
[0021] f = Conv 1×1 (BN(ReLU(f L )))⊙Sigmoid(Conv 1×1 (BN(ReLU(f H )))),
[0022] Wherein, f L , f H represent the low-frequency features and high-frequency features respectively, DWT(·) represents the haar wavelet convolution, ⊙ represents the Hadamard product, and BN represents the normalization operation.
[0023] The wavelet transform in the described downsampling unit is the Haar wavelet transform.
[0024] The described upsampling unit is composed of 2-fold bilinear interpolation, 3×3 convolution, and the ReLU activation function.
[0025] Advantages of the present invention:
[0026] The MSFormer Block proposed by the present invention effectively obtains global information through convolutional operations, not only avoiding the loss of target information caused by the limited receptive field of traditional convolutions, but also reducing the high computational cost problem brought by the attention mechanism; the designed downsampling unit solves the problem of small target loss that may occur during pooling and convolutional downsampling. Brief Description of the Drawings
[0027] Figure 1 : Schematic diagram of an automatic segmentation system for fundus image retinopathy proposed;
[0028] Figure 2 : Schematic diagram of the structure of the proposed MSFmormer Block;
[0029] Figure 3 : Schematic diagram of the structure of the proposed downsampling module;
[0030] Figure 4 : Schematic diagram of the structure of Haar wavelet transform;
[0031] Figure 5 : Schematic diagram of the implementation process of the proposed MSFmormer Block;
[0032] Figure 6 : Schematic diagram of the implementation process of the proposed downsampling module;
[0033] Figure 7 : Visualization diagram of the proposed MSFNet and other methods on different eye disease datasets; Detailed Description of the Invention
[0034] The present invention will be described in detail below in conjunction with specific embodiments.
[0035] An automatic segmentation system for fundus image retinopathy includes an encoder and a decoder, as Figure 1 shown.
[0036] The encoder includes four serially connected encoding modules: the first encoding module to the fourth encoding module. For a given fundus image, it sequentially passes through four serially connected encoding modules. The spatial resolution of the feature map processed by each encoding module is halved, and at the same time, the channel dimension is expanded to twice the original. The first encoding module contains an MSFormer Block, and the second to fourth encoding modules each contain an MSFormer Block and a downsampling unit, which combines the characteristics of CNN and Transformer to extract small target regions and prevent the loss of small targets caused by downsampling.
[0037] The described decoder includes four serial decoding modules: the first to the fourth decoding modules. The first to the third decoding modules each contain an MSFormer Block and an upsampling unit, while the fourth decoding module only contains an MSFormer Block. The final output of the encoder is used as the initial input of the first decoding module. The decoding process adopts a skip connection mechanism: First, the output of the first decoding module is concatenated with the features of the third encoding module; then, the concatenated result is input into the second decoding module, and its output is concatenated with the features of the second encoding module; then, the concatenated result is input into the third decoding module and concatenated with the features of the first encoding module; finally, after being processed by the fourth decoding module, the features are restored to the target resolution.
[0038] The described MSFormer Block is as Figure 2 shown. First, the input feature map is normalized by LayerNorm. Then, the module performs convolution operations on the input using three different sizes of convolutional kernels (3x3, 5x5, and 7x7) respectively, averages and fuses the convolution results of the three scales, and then obtains A through operations such as a 1×1 convolution, a GELU activation function, and an 11×11 convolution, obtains V through a 1×1 convolution. Next, after multiplying A and V, a 1×1 convolution is performed to the final output feature: The calculation formula is as follows:
[0039] f′ = (Conv 3×3 (LN(X i )) + Conv 3×3 (LN(X i )) + Conv 3×3 (LN(X i ))) / 3,
[0040] A = Conv 11×11 (GELU(Conv 1×1 (f′))),
[0041] V(Conv 1×1 (f′), f = Conv 1×1 (A⊙V)
[0042] where X i (i ∈ [1~4]) represents the input features of each MSFmormer Block, ⊙ represents the Hadamard product, LN represents the normalization operation, and f represents the output features of each MSFmormer Block.
[0043] The described downsampling unit is as Figure 3As shown below. First, perform a 1×1 convolution on the input features, and then decompose them into low-frequency features and high-frequency features through wavelet transform. Second, perform a 1×1 convolution block operation on the low-frequency features, and then obtain a gating coefficient through the Sigmoid activation function; perform a 1×1 convolution block operation on the high-frequency features and multiply them by the gating coefficient obtained from the low-frequency features to get the final output. The calculation formula is as follows:
[0044] f L ,f H =Conv 1×1 (DWT(f)),
[0045] f=Conv 1×1 (BN(ReLU(f L )))⊙Sigmoid(Conv 1×1 (BN(ReLU(f H )))),
[0046] Among them, f L ,f H represent the low-frequency features and high-frequency features respectively, DWT(·) represents the haar wavelet convolution, ⊙ represents the Hadamard product, and BN represents the normalization operation.
[0047] The wavelet transform in the downsampling unit is the Haar wavelet transform, as Figure 4 shown, where A(·) represents the low-pass filter, D(·) represents the high-pass filter, A, H, V, D represent the low-frequency component, horizontal high-frequency component, vertical high-frequency component, and diagonal high-frequency component respectively, A represents the low-frequency features, and H, V, D are concatenated as the high-frequency features after splicing.
[0048] The upsampling unit consists of 2× bilinear interpolation, 3×3 convolution, and the ReLU activation function.
[0049] Example 1: Applied to diabetic retinopathy (DR) lesions in the IDRiD dataset
[0050] The specific steps are as follows:
[0051] First step, divide the collected IDRiD dataset into a training set and a test set according to a ratio of 7:3.
[0052] Second step, input the image, construct and train the proposed automatic segmentation method for fundus image retinopathy MSFNet. The specific steps are as follows:
[0053] First, construct an encoder module. Four serial encoding modules are used for encoding. The spatial resolution of the feature map processed by each encoding module is halved, while the channel dimension is expanded to twice the original. The first encoding module contains an MSFormerBlock, and the second to fourth encoding modules each contain an MSFormer Block and a downsampling unit to extract features. Among them, the MSFormer Block contains a LayerNorm layer, three 1×1 convolutions, a 3×3 convolution, a 5×5 convolution, a 7×7 convolution, an 11×11 convolution, and a GELU activation function. The specific implementation process is as Figure 5 shown; the downsampling module contains three 1×1 convolutions, two ReLU activation functions, and two BatchNorms. The specific implementation process is as Figure 6 shown.
[0054] Secondly, construct a decoder, which also consists of four stages. Among them, the first three modules each contain an MSFormerBlock and an upsampling unit, while the fourth module only contains an MSFormer Block. The final output of the encoder is used as the initial input of the decoder. The decoding process adopts a skip connection mechanism: First, the output of the first decoder module is concatenated with the features of the third encoder module; then, the concatenated result is input into the second decoder module, and its output is concatenated with the features of the second encoder module; then, the concatenated result is input into the third decoder module and concatenated with the features of the first encoder module; finally, after being processed by the final decoder module, the features are restored to the target resolution. Among them, the MSFormer Block contains a LayerNorm layer, three 1×1 convolutions, a 3×3 convolution, a 5×5 convolution, a 7×7 convolution, an 11×11 convolution, and a GELU activation function as in the previous step; the upsampling module consists of 2x bilinear interpolation, a 3×3 convolution, and a ReLU activation function.
[0055] Finally, construct a 1×1 convolution layer as the segmentation head, and decode based on the feature map finally output by the decoder to obtain the final segmentation mask.
[0056] Thirdly, use the divided training set and test set to train the model. The specific steps are as follows:
[0057] First, input the original images in the training set into the network, and obtain the output results of the current iteration through calculation. Then, compare the model output results with the corresponding segmentation results in the labels, and use the loss function to calculate the loss value. In this embodiment, a weighted combination of two loss functions L s = α·Loss BCE + β·Loss DICE, with weights α and β, where Loss BCE represents the binary cross-entropy loss function, and Loss DICE represents the Dice loss function, and the weight values are 0.5 and 0.5 respectively. Next, the Adam optimizer is used to calculate the gradients, and the weights in the network are updated through backpropagation. This process is iterated until the loss value reaches the predetermined error requirement, and finally a trained network model is obtained. Finally, the original images of the test set are input into the trained model to obtain the final segmentation results, which are compared with the true labels of the test set to verify the performance of the model. The experimental results of the MSFNet network in this embodiment are shown in Table 1.
[0058] In Table 1, ED refers to Euclidean distance, and the smaller the value, the better the effect. It can be observed that the effect of MSFNet on IDRiD dataset for ED F (fovea detection) reaches the best, which is 55 and 18.3 lower than UNet and UNe++ respectively. For ED OD (optic disc detection), the effect also reaches the best, which is 32.8 and 18.4 lower than UNet and UNe++ respectively.
[0059] Table 1 Comparative experimental results of MSFNet on IDRiD dataset
[0060]
[0061] Example 2: Application to age-related macular degeneration (AMD) lesions in the ADAM dataset
[0062] The implementation method is the same as that in Example 1, and the specific steps are as follows:
[0063] First step, the collected ADAM dataset is divided into a training set and a test set according to a ratio of 7:3.
[0064] Second step, input the images, construct and train the proposed MSFNet, an automatic segmentation method for fundus image retinopathy, and the specific steps are as follows:
[0065] First, construct the encoder module, which performs encoding in four serial stages. The spatial resolution of the feature map processed in each stage is halved, while the channel dimension is expanded to twice the original. The first-stage encoding module contains an MSFormer Block, and the second to fourth-stage encoding modules each contain an MSFormer Block and a downsampling unit for feature extraction. Among them, the MSFormer Block contains a LayerNorm layer, three 1×1 convolutions, a 3×3 convolution, a 5×5 convolution, a 7×7 convolution, an 11×11 convolution, and a GELU activation function. The specific implementation process is as Figure 5 shown; the downsampling module contains three 1×1 convolutions, two ReLU activation functions, and two BatchNorms. The specific implementation process is as Figure 6 shown.
[0066] Secondly, construct the decoder, which also consists of four stages. Among them, the first three modules each contain an MSFormer Block and an upsampling unit, while the fourth module only contains an MSFormer Block. The final output of the encoder is used as the initial input of the decoder. The decoding process adopts a skip connection mechanism: First, the output of the first decoder module is concatenated with the features of the third encoder module; then, the concatenated result is input into the second decoder module, and its output is concatenated with the features of the second encoder module; then, the concatenated result is input into the third decoder module and concatenated with the features of the first encoder module; finally, after being processed by the final decoder module, the features are restored to the target resolution. Among them, the MSFormer Block is the same as in the previous step and contains a LayerNorm layer, three 1×1 convolutions, a 3×3 convolution, a 5×5 convolution, a 7×7 convolution, an 11×11 convolution, and a GELU activation function; the upsampling module consists of 2x bilinear interpolation, a 3×3 convolution, and a ReLU activation function.
[0067] Finally, construct a 1×1 convolution layer as the segmentation head, which decodes based on the feature map finally output by the decoder to obtain the final segmentation mask.
[0068] Thirdly, use the divided training set and test set to train the model. The specific steps are as follows:
[0069] First, input the original images in the training set into the network, and obtain the output result of the current iteration through calculation. Then, compare the model output result with the corresponding segmentation result in the label, and use the loss function to calculate the loss value. In this embodiment, a weighted combination L of two loss functions is adopted s = α·Loss BCE + β·Loss DICE, with weights α and β, where Loss BCE represents the binary cross-entropy loss function, and Loss DICE represents the Dice loss function, and the weight values are 0.5 and 0.5 respectively. Next, the Adam optimizer is used to calculate the gradient, and the weights in the network are updated through backpropagation. This process is iterated until the loss value reaches the predetermined error requirement, and finally a trained network model is obtained. Finally, the original images in the test set are input into the trained model to obtain the final segmentation result, which is compared with the true labels in the test set to verify the performance of the model. The experimental results of the MSFNet network in this embodiment are shown in Table 2.
[0070] In Table 2, ED refers to Euclidean distance, and the smaller the value, the better the effect. Dice is a segmentation metric, and the larger the value, the better the effect. It can be observed that the effect of MSFNet on ED F (fovea detection) reaches the best, which is 45.3 and 1.5 lower than UNet and UNe++ respectively. For Dice OD (optic disc segmentation), the effect also reaches the best, which is 0.206 and 0.080 higher than UNet and UNe++ respectively.
[0071] Table 2 Comparative experimental results of MSFNet on the ADAM dataset
[0072]
[0073] Example 3: Application to glaucoma lesions in the REFUGE dataset
[0074] The implementation method is the same as that in Example 1, and the specific steps are as follows:
[0075] First step, the collected REFUGE dataset is divided into a training set and a test set according to a ratio of 7:3.
[0076] Second step, input the images, construct and train the proposed MSFNet, an automatic segmentation method for retinal lesions in fundus images. The specific steps are as follows:
[0077] First, construct an encoder module that performs encoding in four sequential stages. The spatial resolution of the feature map after each stage is halved, while the channel dimension is expanded to twice the original. The first-stage encoding module contains an MSFormer Block, and the second to fourth-stage encoding modules each contain an MSFormer Block and a downsampling unit for feature extraction. Among them, the MSFormer Block contains a LayerNorm layer, three 1×1 convolutions, a 3×3 convolution, a 5×5 convolution, a 7×7 convolution, an 11×11 convolution, and a GELU activation function. The specific implementation process is as Figure 5 shown; the downsampling module contains three 1×1 convolutions, two ReLU activation functions, and two BatchNorms. The specific implementation process is as Figure 6 shown.
[0078] Secondly, construct a decoder, which also consists of four stages. Among them, the first three modules each contain an MSFormerBlock and an upsampling unit, while the fourth module only contains an MSFormer Block. The final output of the encoder is used as the initial input of the decoder. The decoding process adopts a skip connection mechanism: First, the output of the first decoder module is concatenated with the features of the third encoder module; then, the concatenated result is input into the second decoder module, and its output is concatenated with the features of the second encoder module; then, the concatenated result is input into the third decoder module and concatenated with the features of the first encoder module; finally, after being processed by the final decoder module, the features are restored to the target resolution. Among them, the MSFormer Block contains a LayerNorm layer, three 1×1 convolutions, a 3×3 convolution, a 5×5 convolution, a 7×7 convolution, an 11×11 convolution, and a GELU activation function as in the previous step; the upsampling module consists of 2x bilinear interpolation, a 3×3 convolution, and a ReLU activation function.
[0079] Finally, construct a 1×1 convolutional layer as the segmentation head, which decodes based on the feature map output by the decoder finally to obtain the final segmentation mask.
[0080] Thirdly, use the divided training set and test set to train the model. The specific steps are as follows:
[0081] First, input the original images in the training set into the network and obtain the output results of the current iteration through calculation. Then, compare the model output results with the corresponding segmentation results in the labels, and use a loss function to calculate the loss value. In this embodiment, a weighted combination of two loss functions L s = α·Loss BCE + β·Loss DICE, with weights α and β, where Loss BCE represents the binary cross - entropy loss function, and Loss DICE represents the Dice loss function, and the weight values are 0.5 and 0.5 respectively. Next, the Adam optimizer is used to calculate the gradients, and the weights in the network are updated through backpropagation. This process is iterated until the loss value reaches the predetermined error requirement, and finally a trained network model is obtained. Finally, the original images in the test set are input into the trained model to obtain the final segmentation results, which are compared with the true labels in the test set to verify the performance of the model. The experimental results of the MSFNet network in this embodiment are shown in Table 3.
[0082] In Table 3, ED refers to Euclidean distance, and the smaller the value, the better the effect. Dice is a segmentation metric, and the larger the value, the better the effect. It can be observed that MSFNet has the best effect on ED F (fovea detection), which is 32.7 and 5.1 lower than UNet and UNet++ respectively. For Dice OD (optic disc segmentation), the effect is also the best, which is 0.141 and 0.012 higher than UNet and UNet++ respectively.
[0083] Table 3 Comparative experimental results of MSFNet on the REFUGE dataset
[0084]
[0085] Figure 7 Shows the visual comparison results of the proposed MSFNet with UNet and UNet++ on different eye disease datasets. The experiments cover four typical fundus conditions: the first row represents healthy fundus images, the second row is age - related macular degeneration (AMD) lesions, the third row is glaucoma (Glaucoma) lesions, and the fourth row is diabetic retinopathy (DR) lesions. In healthy fundus images, both UNet and UNet++ show problems of missing segmentation in some areas; for AMD cases, the segmentation results of the two comparison methods are incomplete; particularly noteworthy is that in Glaucoma cases, UNet misjudges some vascular structures as fundus background; and in DR cases, there are obvious area losses in the segmentation results of UNet. In contrast, MSFNet shows the optimal segmentation performance in all four fundus types, and its segmentation results are highly consistent with the Ground Truth, fully demonstrating the effectiveness and robustness of the proposed method.
[0086] It should be understood that those of ordinary skill in the art can make improvements or transformations based on the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An automatic segmentation system for fundus image retinopathy, characterized in that, It includes an encoder and a decoder. The encoder includes four serially connected encoding modules: the first encoding module to the fourth encoding module. For a given fundus image, it sequentially passes through the four serially connected encoding modules. The spatial resolution of the feature map processed by each encoding module is halved, while the channel dimension is expanded to twice the original. The first encoding module contains an MSFormer Block, and the second to the fourth encoding modules each contain an MSFormer Block and a downsampling unit, combining the characteristics of CNN and Transformer to extract small target regions and prevent the loss of small targets caused by downsampling. The decoder contains four serially connected decoding modules: the first decoding module to the fourth decoding module. The first to the third decoding modules each contain an MSFormer Block and an upsampling unit, while the fourth decoding module only contains an MSFormer Block. The final output of the encoder is used as the initial input of the first decoding module. The decoding process uses a skip connection mechanism: First, the output of the first decoding module is concatenated with the features of the third encoding module; then, the concatenated result is input into the second decoding module, and its output is concatenated with the features of the second encoding module; then, the concatenated result is input into the third decoding module and concatenated with the features of the first encoding module; finally, after being processed by the fourth decoding module, the features are restored to the target resolution.
2. The automatic segmentation system for fundus image retinopathy according to claim 1, characterized in that, In the MSFormer Block, first, the input feature map is normalized by LayerNorm. Then, the module uses convolutional kernels of three different sizes, 3x3, 5x5, and 7x7, to perform convolutional operations on the input respectively, and the convolutional results of the three scales are averaged and fused. After that, through a 1×1 convolution, a GELU activation function, and a 11×11 convolution and other operations to obtain A, and a 1×1 convolution to obtain V. Next, A and V are multiplied and then a 1×1 convolution is performed to obtain the final output feature.
3. The automatic segmentation system for fundus image retinopathy according to claim 2, characterized in that, The calculation formula of the MSFormer Block is as follows: f ′ = (Conv 3×3 (LN(X i )) + Conv 3×3 (LN(X i )) + Conv 3×3 (LN(X i ))) / 3, A = Conv 11×11 (GELU(Conv 1×1 (f ′ ))), V(Conv 1×1 (f ′ ), f = Conv 1×1 (A⊙V) Among them, X i (i ∈ [1~4]) represents the input feature of each MSFmormer Block, ⊙ represents the Hadamard product, LN represents the normalization operation, and f represents the output feature of each MSFmormer Block.
4. The automatic segmentation system for fundus image retinopathy according to claim 1, characterized in that, For the downsampling unit, first, a 1×1 convolution is performed on the input feature, and then it is decomposed into low-frequency features and high-frequency features through wavelet transform; secondly, a 1×1 convolutional block operation is performed on the low-frequency features, and then a gating coefficient is obtained through the Sigmoid activation function; a 1×1 convolutional block operation is performed on the high-frequency features and then multiplied by the gating coefficient obtained from the low-frequency features to obtain the final output.
5. The automatic segmentation system for fundus image retinopathy according to claim 4, characterized in that, The calculation formula of the downsampling unit is as follows: f L , f H = Conv 1×1 (DWT(f)), f = Conv 1×1 (BN(ReLU(f L ))) ⊙ Sigmoid(Conv 1×1 (BN(ReLU(f H )))), where, f L , f H represent low-frequency features and high-frequency features respectively, DWT(·) represents Haar wavelet convolution, ⊙ represents Hadamard product, and BN represents the normalization operation.
6. The automatic segmentation system for fundus image retinopathy according to claim 1, wherein, The wavelet transform in the downsampling unit is the Haar wavelet transform.
7. The automatic segmentation system for fundus image retinopathy according to claim 1, characterized in that, The upsampling unit consists of 2x bilinear interpolation, 3×3 convolution, and ReLU activation function.