Retina microaneurysm region segmentation-classification method and imaging method for fundus image
By adopting U-Net network and residual multi-scale attention network in fundus image processing, combined with preprocessing and multi-attention tract fusion module, the reliability and accuracy of microarse aneurysm region segmentation and classification in fundus images are solved, and efficient retinal microarse aneurysm region segmentation and classification are achieved.
Patent Information
- Application Number
- CN202510047841.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
When used in actual application, the existing fundus image retinal microarseism region segmentation and classification schemes are affected by low contrast of small red dots and complex background noise, resulting in poor reliability and accuracy of segmentation and classification results.
The retinal microarseizure region segmentation-classification method based on U-Net network and residual multi-scale attention network is adopted. The training data set is constructed through pre-processing steps such as background cropping, green channel extraction and contrast enhancement, and the accuracy of feature extraction and segmentation is improved using the multi-attention gate fusion module and RMS layer.
High reliability and high accuracy segmentation-classification of retinal microarsem areas in fundus images are achieved, false positive areas are eliminated, deep and shallow semantic feature information in the image are fully utilized, and the accuracy of segmentation and classification results is improved.
Smart Images

Figure CN119992172A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image processing, and in particular relates to a retinal microaneurysm region segmentation-classification method and an imaging method for fundus images. Background Art
[0002] Fundus images are important image data that reflect the fundus status of the person being tested, and they are of great significance in both clinical medicine and basic medical research.
[0003] Accurate and reliable segmentation results of retinal microaneurysm areas in fundus images can provide more reliable data support in clinical and basic research. Therefore, the identification of retinal microaneurysm areas in fundus images has always been one of the research focuses of researchers.
[0004] At present, the commonly used retinal microaneurysm region segmentation and classification schemes for fundus images are mainly based on the research of deep learning algorithms. This type of deep learning method can extract high-level semantic features from retinal images without relying on traditional manual feature extraction, and can improve the accuracy of microaneurysm segmentation and classification through a specific network architecture. However, when applied to the segmentation and classification of microaneurysm regions in fundus images, this type of scheme still has some defects: the microaneurysm region of the fundus image usually appears as a small red dot with low contrast with the background image, and the retinal image contains complex elements, such as fine blood vessel fragments, small hemorrhages and other background noise and other interference factors; therefore, the existing schemes will be interfered by such factors in actual application, resulting in poor reliability and accuracy of the segmentation and classification results. Summary of the invention
[0005] One of the objectives of the present invention is to provide a retinal microaneurysm region segmentation-classification method for fundus images with high reliability and good accuracy.
[0006] A second object of the present invention is to provide an imaging method including the retinal microaneurysm region segmentation-classification method for fundus images.
[0007] The retinal microaneurysm region segmentation-classification method for fundus images provided by the present invention comprises the following steps:
[0008] S1. Acquire existing fundus image data;
[0009] S2. Preprocessing the fundus image data obtained in step S1 to construct a training data set;
[0010] S3. Based on the U-Net network and residual multi-scale attention network, an initial model for retinal microaneurysm region segmentation and classification was constructed;
[0011] S4. Using the training data set obtained in step S2, the retinal microaneurysm region segmentation-classification initial model constructed in step S3 is trained to obtain a retinal microaneurysm region segmentation-classification model;
[0012] S5. Use the retinal microaneurysm region segmentation-classification model obtained in step S4 to perform segmentation-classification of the retinal microaneurysm region on the actual fundus image.
[0013] The preprocessing described in step S2 specifically includes the following steps:
[0014] Perform background cropping on the acquired fundus image to obtain a core image area;
[0015] Extract the green channel of the obtained core image area;
[0016] For the extracted image data, contrast-limited adaptive histogram equalization is used to enhance the image contrast.
[0017] The initial model for retinal microaneurysm region segmentation and classification based on the U-Net network and the residual multi-scale attention network described in step S3 specifically includes the following steps:
[0018] The constructed initial model for retinal microaneurysm region segmentation and classification includes a segmentation model, a clipper and a classification model;
[0019] The segmentation model includes an encoder and a decoder; the encoder is used to process the input data layer by layer to extract key features and information, and transmit the obtained feature data to the decoder; the encoder includes a first encoding module, a second encoding module, a third encoding module, a fourth encoding module and a fifth encoding module connected in series; the first encoding module includes a 2d convolution layer, a ReLU activation function layer and a BN normalization layer connected in series; the second encoding module includes three ResRF convolution layers connected in series; the third encoding module includes four ResRF convolution layers connected in series; the fourth encoding module includes six ResRF convolution layers connected in series; the fifth encoding module includes The decoder is used to map the output of the encoder back to the original data space and restore the details and structure of the input data as much as possible to generate a reconstructed output that is as similar as possible to the original input data, and transmit the obtained image features to the decoder; the decoder includes a first multi-attention gate fusion module, a first processing layer, a second multi-attention gate fusion module, a second processing layer, a third multi-attention gate fusion module, a third processing layer, a fourth multi-attention gate fusion module and a Four processing layers; the structures of the first to fourth multi-attention gate fusion modules are the same; the multi-attention gate fusion module is used to improve the selectivity of feature fusion, so that the segmentation model can focus more on learning the features in the target area and reduce the interference of noise; the structures of the first to fourth processing layers are the same; the processing layer is used to process the output of the multi-attention gate fusion module; the output of the fourth encoding module and the output of the fifth encoding module are both used as the input of the first multi-attention gate fusion module; the output of the first multi-attention gate fusion module and the output of the upsampled fifth encoding module are used as the input of the first processing layer; the output of the first processing layer and the output of the third encoding module All of them are used as the input of the second multi-attention gate fusion module; the output of the second multi-attention gate fusion module and the output of the first processing layer after upsampling are used as the input of the second processing layer; the output of the second processing layer and the output of the second encoding module are both used as the input of the third multi-attention gate fusion module; the output of the third multi-attention gate fusion module and the output of the second processing layer after upsampling are used as the input of the third processing layer; the output of the third processing layer and the output of the first encoding module are both used as the input of the fourth multi-attention gate fusion module; the output of the fourth multi-attention gate fusion module and the output of the third processing layer after upsampling are used as the input of the fourth processing layer; the output of the fourth processing layer is used as the output of the decoder;
[0020] The cropper is used to crop the image features output by the decoder by center cropping to obtain candidate area blocks, and transmit the obtained candidate area blocks to the classification model;
[0021] The classification model is used to perform feature extraction and region classification based on the feature data of the candidate region blocks to complete the classification of the retinal microaneurysm region of the input image; the classification model includes a 7×7 convolutional layer, a BN normalization layer, a ReLU activation function layer, a first RMS layer, a second RMS layer, a third RMS layer, a fourth RMS layer and a fully connected layer connected in series; the structures of the first to fourth RMS layers are the same, and the RMS layer is used to learn multi-scale information, refine the texture information between the target area and the background, and perform feature response calibration on the extracted feature map to highlight the area of interest.
[0022] The processing of the ResRF convolutional layer includes the following steps:
[0023] The input feature data is processed by the first 1×1 convolution sublayer to obtain the first feature; the input feature data is processed by the second 1×1 convolution sublayer, the RFCBAM sublayer and the third 1×1 convolution sublayer in sequence to obtain the second feature; the first feature and the second feature are connected to obtain the output of the ResRF convolution layer;
[0024] The processing of the RFCBAM sublayer includes the following steps:
[0025] The input data of the RFCBAM sublayer is divided into two paths:
[0026] The first path is processed in sequence through the global average pooling layer, the fully connected layer, the ReLU activation function layer, the fully connected layer, and the Sigmoid activation function layer to obtain the channel feature F channel ; F channel The calculation formula is F channel =σ(FC(ReLU(FC(GAP(F))))), where F is the input data of the RFCBAM sublayer, GAP() is the global average pooling layer processing function; FC() is the fully connected layer processing function; ReLU() is the ReLU activation function; σ() is the Sigmoid activation function;
[0027] The second path is processed by the grouped convolution layer, ReLU activation function layer, BN normalization layer and Reshape layer in turn to obtain the main feature F main ; F main The calculation formula is F main =Reshape(BN(ReLU(GroupConv(F)))), where GroupConv() is the group convolution layer processing function, BN() is the BN normalization layer processing function, and Reshape() is the Reshape layer processing function; the Reshape layer is used to adjust the dimension of the data;
[0028] The main feature F mainAfter being processed through the maximum pooling layer and the average pooling layer respectively, the output of the maximum pooling layer and the output of the average pooling layer are cascaded, and then processed through the 1×1 convolution layer and the Sigmoid activation function layer in turn to obtain the spatial feature F spatial ; F spatial The calculation formula is F spatial =σ(Conv 1×1 ([Avg(F main ),Max(F main )])), where Max() is the maximum pooling layer processing function, Avg() is the average pooling layer processing function, Conv 1×1 () is the 1×1 convolution layer processing function; [] is the cascade operation;
[0029] The channel feature F channel Main features F main and spatial feature F spatial Multiply them and then process them through a 3×3 convolutional layer to obtain the output data of the RFCBAM sub-layer.
[0030] The processing of the multi-attention gate fusion module includes the following steps:
[0031] The input of the multi-attention gate fusion module includes the corresponding encoder data and upsampled data;
[0032] The upsampled data is processed by the grouped convolution layer, BN normalization layer, ReLU activation function layer, channel attention mechanism layer and spatial attention mechanism layer in turn to obtain the first fusion feature;
[0033] The encoder data passes through the grouped convolution layer, BN normalization layer, ReLU activation function layer, channel attention mechanism layer and spatial attention mechanism layer in turn to obtain the second fusion feature;
[0034] The first fusion feature and the second fusion feature are added to obtain a third fusion feature;
[0035] The third fused feature is processed in sequence through the ReLU activation function layer, the 1×1 convolution layer, and the Sigmoid activation function layer to obtain the fourth fused feature;
[0036] The fourth fusion feature is multiplied by the encoder data and then added to the encoder data to obtain the output of the multi-attention gate fusion module. MAGF ;
[0037] Output of the multi-attention gate fusion module MAGF The calculation formula is
[0038] Output MAGF =x×W4+x
[0039] W4=σ(Conv 1×1 (ReLU(W3)))
[0040] W3=W1+W2
[0041] W1=f SE (f ECA (ReLU(BN(GroupConv(g)))))
[0042] W2=f SE (f ECA (ReLU(BN(GroupConv(x)))))
[0043] Where x is the encoder data; W4 is the fourth fusion feature; W3 is the third fusion feature; W1 is the first fusion feature; W2 is the second fusion feature; f SE () is the spatial attention mechanism layer processing function; f ECA () is the channel attention mechanism layer processing function; g is the upsampled data.
[0044] The processing process of the processing layer includes the following steps:
[0045] The processing layer includes a first 2d convolution layer, a first ReLU activation function layer, a second 2d convolution layer, and a second ReLU activation function layer connected in series in sequence;
[0046] The data input to the processing layer is processed in sequence by the first 2D convolution layer, the first ReLU activation function layer, the second 2D convolution layer, and the second ReLU activation function layer to obtain the output of the processing layer.
[0047] The processing of the RMS layer includes the following steps:
[0048] The input of the RMS layer is divided into two paths:
[0049] The first path is processed by a 1×1 convolution layer, a 3×3 convolution layer, and a channel attention layer in sequence to obtain the first feature;
[0050] The second path is processed by 1×1 convolution layer, RFB layer, 3×3 convolution layer and channel attention layer in sequence to obtain the second feature;
[0051] After the first feature and the second feature are cascaded, they are processed through the spatial attention layer and the 1×1 convolution layer to obtain the cascaded processing feature;
[0052] The cascade processing features are added to the input of the RMS layer and used as the output of the RMS layer. RMS ;
[0053] Output of the RMS layer RMS The calculation formula is
[0054] Output RMS =Conv 1×1 (SA(Concat(F1,F2)))+X
[0055] F1=CA(Conv 3×3 (Conv 1×1 (X)))
[0056] F2=CA(Conv 3×3 (RFB(Conv 1×1 (X))))
[0057] Where SA is the processing function of the spatial attention layer; Concat() is the cascade function; X is the input of the RMS layer; F1 is the first feature; F2 is the second feature; CA is the processing function of the channel attention layer; RFB() is the processing function of the RFB layer;
[0058] The processing of the RFB layer comprises the following steps:
[0059] The input of the RFB layer is divided into 4 channels:
[0060] The first path is processed by a 1×1 convolutional layer and a 3×3 dilated convolutional layer with a dilation rate of 1 to obtain the first sub-feature;
[0061] The second path is processed in sequence by a 1×1 convolutional layer, a 3×3 convolutional layer, and a 3×3 dilated convolutional layer with a dilation rate of 3 to obtain the second sub-feature;
[0062] The third path is processed in sequence by a 1×1 convolutional layer, a 5×5 convolutional layer, and a 3×3 dilated convolutional layer with a dilation rate of 5 to obtain the third sub-feature;
[0063] The fourth path directly serves as the fourth sub-feature;
[0064] After cascading the first sub-feature, the second sub-feature, and the third sub-feature, they are processed through a 1×1 convolutional layer to obtain the cascaded sub-feature;
[0065] The cascaded sub-features are added to the fourth sub-feature to obtain the output of the RFB layer.
[0066] The present invention also provides an imaging method including the retinal microaneurysm region segmentation-classification method for fundus images, further comprising the following steps:
[0067] S6. The segmentation and classification results obtained in step S5 are marked and re-imaged on the actual fundus image to obtain a fundus image with the segmentation and classification results of the retinal microaneurysm area.
[0068] The retinal microaneurysm region segmentation-classification method and imaging method for fundus images provided by the present invention acquires and processes fundus image data of retinal microaneurysms to construct a training data set, and trains and uses a segmentation-classification model constructed based on a U-Net network and a residual multi-scale attention network. This not only realizes the segmentation-classification of retinal microaneurysm regions for fundus images, but also has higher reliability and better accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 Schematic diagram of the process flow of the segmentation-classification method of the present invention.
[0070] Figure 2 Schematic diagram of comparison curve of E-Ophtha-MA data set in an embodiment of the method of the present invention.
[0071] Figure 3 Schematic diagram of comparison curve of IDRiD data set of the method embodiment of the present invention.
[0072] Figure 4 The figure is a schematic diagram of the method flow of the imaging method of the present invention. DETAILED DESCRIPTION
[0073] like Figure 1 The figure shows a flow chart of the segmentation-classification method of the present invention: the retinal microaneurysm region segmentation-classification method for fundus images disclosed in the present invention comprises the following steps:
[0074] S1. Acquire existing fundus image data;
[0075] S2. Preprocess the fundus image data obtained in step S1 to construct a training data set; specifically, the steps include:
[0076] Perform background cropping on the acquired fundus image to obtain a core image area;
[0077] Since microaneurysms have a higher contrast in the green channel of the fundus image than in other channels, the green channel is extracted from the obtained core image area;
[0078] For the extracted image data, contrast-limited adaptive histogram equalization is used to enhance the image contrast;
[0079] S3. Based on the U-Net network and the residual multi-scale attention network, an initial model for retinal microaneurysm region segmentation and classification is constructed; specifically, the following steps are included:
[0080] The constructed initial model for retinal microaneurysm region segmentation and classification includes a segmentation model, a clipper and a classification model;
[0081] The segmentation model includes an encoder and a decoder; the encoder is used to process the input data layer by layer to extract key features and information, and transmit the obtained feature data to the decoder; the encoder includes a first encoding module, a second encoding module, a third encoding module, a fourth encoding module and a fifth encoding module connected in series; the first encoding module includes a 2d convolution layer, a ReLU activation function layer and a BN normalization layer connected in series; the second encoding module includes three ResRF convolution layers connected in series; the third encoding module includes four ResRF convolution layers connected in series; the fourth encoding module includes six ResRF convolution layers connected in series; the fifth encoding module includes The decoder is used to map the output of the encoder back to the original data space and restore the details and structure of the input data as much as possible to generate a reconstructed output that is as similar as possible to the original input data, and transmit the obtained image features to the decoder; the decoder includes a first multi-attention gate fusion module, a first processing layer, a second multi-attention gate fusion module, a second processing layer, a third multi-attention gate fusion module, a third processing layer, a fourth multi-attention gate fusion module and a Four processing layers; the structures of the first to fourth multi-attention gate fusion modules are the same; the multi-attention gate fusion module is used to improve the selectivity of feature fusion, so that the segmentation model can focus more on learning the features in the target area and reduce the interference of noise; the structures of the first to fourth processing layers are the same; the processing layer is used to process the output of the multi-attention gate fusion module; the output of the fourth encoding module and the output of the fifth encoding module are both used as the input of the first multi-attention gate fusion module; the output of the first multi-attention gate fusion module and the output of the upsampled fifth encoding module are used as the input of the first processing layer; the output of the first processing layer and the output of the third encoding module All of them are used as the input of the second multi-attention gate fusion module; the output of the second multi-attention gate fusion module and the output of the first processing layer after upsampling are used as the input of the second processing layer; the output of the second processing layer and the output of the second encoding module are both used as the input of the third multi-attention gate fusion module; the output of the third multi-attention gate fusion module and the output of the second processing layer after upsampling are used as the input of the third processing layer; the output of the third processing layer and the output of the first encoding module are both used as the input of the fourth multi-attention gate fusion module; the output of the fourth multi-attention gate fusion module and the output of the third processing layer after upsampling are used as the input of the fourth processing layer; the output of the fourth processing layer is used as the output of the decoder;
[0082] The cropper is used to crop the image features output by the decoder in a central cropping manner to obtain candidate region blocks, and transmit the obtained candidate region blocks to the classification model; the output of the decoder still retains some non-microaneurysm areas that are difficult to distinguish (such as the intersection of small blood vessels), so the cropper is used for cropping, and the cropped image is further classified;
[0083] The classification model is used to perform feature extraction and regional classification based on the feature data of the candidate region blocks to complete the classification of the retinal microaneurysm region of the input image; the classification model includes a 7×7 convolution layer, a BN normalization layer, a ReLU activation function layer, a first RMS layer, a second RMS layer, a third RMS layer, a fourth RMS layer and a fully connected layer connected in series; the structures of the first to fourth RMS layers are the same, and the RMS layer is used to learn multi-scale information, refine the texture information between the target area and the background, and perform feature response calibration on the extracted feature map to highlight the area of interest; at the same time, the classification model can also reduce the impact of changes in the size of the microaneurysm structure on the classification results;
[0084] When implementing:
[0085] The processing of the ResRF convolutional layer includes the following steps:
[0086] Compared with traditional encoders, ResRF convolution can provide more detailed feature information for microaneurysm segmentation tasks and can better optimize the performance of the model;
[0087] The input feature data is processed by the first 1×1 convolution sublayer to obtain the first feature; the input feature data is processed by the second 1×1 convolution sublayer, the RFCBAM sublayer and the third 1×1 convolution sublayer in sequence to obtain the second feature; the first feature and the second feature are connected to obtain the output of the ResRF convolution layer;
[0088] The processing of the RFCBAM sublayer includes the following steps:
[0089] The RFCBAM sublayer focuses on the spatial features of the receptive field, which is enhanced by improving the receptive field attention and also enhances the attention to channel features, thereby improving the feature extraction of spatial and channel dimensions, and capturing the global context information of the image through global pooling technology;
[0090] The input data of the RFCBAM sublayer is divided into two paths:
[0091] The first path is processed in sequence through the global average pooling layer, the fully connected layer, the ReLU activation function layer, the fully connected layer, and the Sigmoid activation function layer to obtain the channel feature F channel ; F channel The calculation formula is Fchannel =σ(FC(ReLU(FC(GAP(F))))), where F is the input data of the RFCBAM sublayer, GAP() is the global average pooling layer processing function; FC() is the fully connected layer processing function; ReLU() is the ReLU activation function; σ() is the Sigmoid activation function;
[0092] The second path is processed by the grouped convolution layer, ReLU activation function layer, BN normalization layer and Reshape layer in turn to obtain the main feature F main ; F main The calculation formula is F main =Reshape(BN(ReLU(GroupConv(F)))), where GroupConv() is the group convolution layer processing function, BN() is the BN normalization layer processing function, and Reshape() is the Reshape layer processing function; the Reshape layer is used to adjust the dimension of the data;
[0093] The main feature F main After being processed through the maximum pooling layer and the average pooling layer respectively, the output of the maximum pooling layer and the output of the average pooling layer are cascaded, and then processed through the 1×1 convolution layer and the Sigmoid activation function layer in turn to obtain the spatial feature F spatial ; F spatial The calculation formula is F spatial =σ(Conv 1×1 ([Avg(F main ),Max(F main )])), where Max() is the maximum pooling layer processing function, Avg() is the average pooling layer processing function, Conv 1×1 () is the 1×1 convolution layer processing function; [] is the cascade operation;
[0094] The channel feature F channel Main features F main and spatial feature F spatial Multiply and then process through a 3×3 convolutional layer to obtain the output data of the RFCBAM sub-layer;
[0095] The processing of the multi-attention gate fusion module includes the following steps:
[0096] The main function of the multi-attention gate fusion module is to improve the selectivity of feature fusion, so that the model can focus on learning important features in the target area and reduce the interference of irrelevant noise;
[0097] The input of the multi-attention gate fusion module includes the corresponding encoder data and upsampled data;
[0098] The upsampled data is processed by the grouped convolution layer, BN normalization layer, ReLU activation function layer, channel attention mechanism layer and spatial attention mechanism layer in turn to obtain the first fusion feature;
[0099] The encoder data passes through the grouped convolution layer, BN normalization layer, ReLU activation function layer, channel attention mechanism layer and spatial attention mechanism layer in turn to obtain the second fusion feature;
[0100] The first fusion feature and the second fusion feature are added to obtain a third fusion feature;
[0101] The third fused feature is processed in sequence through the ReLU activation function layer, the 1×1 convolution layer, and the Sigmoid activation function layer to obtain the fourth fused feature;
[0102] The fourth fusion feature is multiplied by the encoder data and then added to the encoder data to obtain the output of the multi-attention gate fusion module. MAGF ;
[0103] The use of grouped convolution can effectively reduce computational complexity; the processing of spatial attention mechanism and efficient channel attention mechanism can be more conducive to information segmentation; when the correlation between features is weak, the connection and processing method of the multi-attention gate fusion module can reduce the impact of high-level semantic features on low-level semantic features;
[0104] Output of the multi-attention gate fusion module MAGF The calculation formula is
[0105] Output MAGF =x×W4+x
[0106] W4=σ(Conv 1×1 (ReLU(W3)))
[0107] W3=W1+W2
[0108] W1=f SE (f ECA (ReLU(BN(GroupConv(g)))))
[0109] W2=f SE (f ECA (ReLU(BN(GroupConv(x)))))
[0110] Where x is the encoder data; W4 is the fourth fusion feature; W3 is the third fusion feature; W1 is the first fusion feature; W2 is the second fusion feature; f SE () is the spatial attention mechanism layer processing function; f ECA() is the channel attention mechanism layer processing function; g is the upsampled data;
[0111] The processing process of the processing layer includes the following steps:
[0112] The processing layer includes a first 2d convolution layer, a first ReLU activation function layer, a second 2d convolution layer, and a second ReLU activation function layer connected in series in sequence;
[0113] The data input to the processing layer is processed in sequence by the first 2d convolution layer, the first ReLU activation function layer, the second 2d convolution layer, and the second ReLU activation function layer to obtain the output of the processing layer;
[0114] The processing of the RMS layer includes the following steps:
[0115] The input of the RMS layer is divided into two paths:
[0116] The first path is processed by a 1×1 convolution layer, a 3×3 convolution layer, and a channel attention layer in sequence to obtain the first feature;
[0117] The second path is processed by 1×1 convolution layer, RFB layer, 3×3 convolution layer and channel attention layer in sequence to obtain the second feature;
[0118] After the first feature and the second feature are cascaded, they are processed through the spatial attention layer and the 1×1 convolution layer to obtain the cascaded processing feature;
[0119] The cascade processing features are added to the input of the RMS layer and used as the output of the RMS layer. RMS ;
[0120] Output of the RMS layer RMS The calculation formula is
[0121] Output RMS =Conv 1×1 (SA(Concat(F1,F2)))+X
[0122] F1=CA(Conv 3×3 (Conv 1×1 (X)))
[0123] F2=CA(Conv 3×3 (RFB(Conv 1×1 (X))))
[0124] Where SA is the processing function of the spatial attention layer; Concat() is the cascade function; X is the input of the RMS layer; F1 is the first feature; F2 is the second feature; CA is the processing function of the channel attention layer; RFB() is the processing function of the RFB layer;
[0125] The processing of the RFB layer comprises the following steps:
[0126] The input of the RFB layer is divided into 4 channels:
[0127] The first path is processed by a 1×1 convolutional layer and a 3×3 dilated convolutional layer with a dilation rate of 1 to obtain the first sub-feature;
[0128] The second path is processed in sequence by a 1×1 convolutional layer, a 3×3 convolutional layer, and a 3×3 dilated convolutional layer with a dilation rate of 3 to obtain the second sub-feature;
[0129] The third path is processed in sequence by a 1×1 convolutional layer, a 5×5 convolutional layer, and a 3×3 dilated convolutional layer with a dilation rate of 5 to obtain the third sub-feature;
[0130] The fourth path directly serves as the fourth sub-feature;
[0131] After cascading the first sub-feature, the second sub-feature, and the third sub-feature, they are processed through a 1×1 convolutional layer to obtain the cascaded sub-feature;
[0132] Add the cascaded sub-features to the fourth sub-feature to get the output of the RFB layer;
[0133] The receptive field is expanded through dilated convolutions with different dilation rates, multi-scale information is learned, and the texture information between the target area and the background is refined. Then, the feature response is recalibrated through spatial attention and channel attention, allowing the model to highlight the area of interest during learning.
[0134] S4. Using the training data set obtained in step S2, the retinal microaneurysm region segmentation-classification initial model constructed in step S3 is trained to obtain a retinal microaneurysm region segmentation-classification model;
[0135] S5. Use the retinal microaneurysm region segmentation-classification model obtained in step S4 to perform segmentation-classification of the retinal microaneurysm region on the actual fundus image.
[0136] The retinal microaneurysm region segmentation-classification model proposed in the present invention can eliminate a large number of false positive areas in the segmentation-classification results, can make full use of the deep semantic feature information and shallow semantic feature information in the image, can enable the model to achieve segmentation-classification of small target areas with low contrast such as microaneurysms, and improve the accuracy of the segmentation results and classification results.
[0137] The advantages of the method of the present invention are described below in conjunction with a comparative example.
[0138] The free response receiving operating characteristic curve (FROC) of the proposed method and the existing methods were experimentally evaluated on the E-Ophtha-MA and IDRiD datasets. The evaluation indicators also included the competitive performance measure (CPM) and the area under the curve (F AUC ), CPM is the average of the 7 sensitivity values. The comparison results are shown in Table 1. For a more intuitive comparison, the FROC curves of the two data sets are plotted as follows Figure 2 and Figure 3 shown.
[0139] Table 1 Comparative data diagram
[0140]
[0141] For the E-Ophtha-MA dataset, the CPM score of the proposed method is 0.622, and the F AUC The score is 0.775. The method of the present invention achieves the AUC The state-of-the-art results evaluated surpass most of the state-of-the-art methods in CPM scores. The proposed method achieves a CPM score of 0.494 on the IDRiD dataset, with an F AUC The score was 0.691. Compared with other methods, the sensitivity of the method of the present invention was the highest when the FPI level was 4 and 8, which were 0.626 and 0.712 respectively.
[0142] Through Table 1, Figure 2 and Figure 3 It can be seen that compared with the existing solutions, the method of the present invention has obvious performance advantages, higher reliability and better accuracy.
[0143] like Figure 4 The method flow diagram of the imaging method of the present invention is shown as follows: the imaging method disclosed in the present invention, which includes the retinal microaneurysm region segmentation-classification method for fundus images, comprises the following steps:
[0144] S1. Acquire existing fundus image data;
[0145] S2. Preprocessing the fundus image data obtained in step S1 to construct a training data set;
[0146] S3. Based on the U-Net network and residual multi-scale attention network, an initial model for retinal microaneurysm region segmentation and classification was constructed;
[0147] S4. Using the training data set obtained in step S2, the retinal microaneurysm region segmentation-classification initial model constructed in step S3 is trained to obtain a retinal microaneurysm region segmentation-classification model;
[0148] S5. Using the retinal microaneurysm region segmentation-classification model obtained in step S4, segmentation-classification of the retinal microaneurysm region is performed on the actual fundus image;
[0149] S6. The segmentation and classification results obtained in step S5 are marked and re-imaged on the actual fundus image to obtain a fundus image with the segmentation and classification results of the retinal microaneurysm area.
[0150] The imaging method provided by the present invention can be directly applied to an existing fundus image acquisition device (such as a fundus camera, or an OCTA device, etc.), or directly applied to a terminal (such as a computer); in specific application, the existing scheme is used to acquire the actual fundus image, and then the acquired data is input into the corresponding machine equipment (such as a fundus camera, or an OCTA device, etc.) or terminal. At this time, the machine equipment or terminal can obtain the actual segmentation-classification result of the retinal microaneurysm area according to the imaging method disclosed in the present invention, and display the segmentation-classification result on the original fundus image through different types of representation (such as color), and then perform secondary imaging and output; at this time, the output fundus image is a fundus image with the segmentation-classification result of the retinal microaneurysm area, and the image can reflect the actual segmentation-classification result of the retinal microaneurysm area, thereby greatly facilitating clinical medical personnel and laboratory experimenters to carry out subsequent work.
Claims
1. A retinal microaneurysm region segmentation-classification method for fundus images, comprising the following steps: S1. Acquire existing fundus image data; S2. Preprocessing the fundus image data obtained in step S1 to construct a training data set; S3. Based on the U-Net network and residual multi-scale attention network, the initial model of retinal microaneurysm region segmentation and classification was constructed; S4. Using the training data set obtained in step S2, the retinal microaneurysm region segmentation-classification initial model constructed in step S3 is trained to obtain a retinal microaneurysm region segmentation-classification model; S5. Use the retinal microaneurysm region segmentation-classification model obtained in step S4 to perform segmentation-classification of the retinal microaneurysm region on the actual fundus image.
2. The retinal microaneurysm region segmentation-classification method for fundus images according to claim 1, characterized in that The preprocessing described in step S2 specifically includes the following steps: Perform background cropping on the acquired fundus image to obtain a core image area; Extract the green channel of the obtained core image area; For the extracted image data, contrast-limited adaptive histogram equalization is used to enhance the image contrast.
3. The retinal microaneurysm region segmentation-classification method for fundus images according to claim 2, characterized in that The initial model for retinal microaneurysm region segmentation based on the U-Net network and the residual multi-scale attention network described in step S3 specifically includes the following steps: The constructed initial model for retinal microaneurysm region segmentation and classification includes a segmentation model, a clipper and a classification model; The segmentation model includes an encoder and a decoder; the encoder is used to process the input data layer by layer to extract key features and information, and transmit the obtained feature data to the decoder; the encoder includes a first encoding module, a second encoding module, a third encoding module, a fourth encoding module and a fifth encoding module connected in series; the first encoding module includes a 2d convolution layer, a ReLU activation function layer and a BN normalization layer connected in series; the second encoding module includes three ResRF convolution layers connected in series; the third encoding module includes four ResRF convolution layers connected in series; the fourth encoding module includes six ResRF convolution layers connected in series; the fifth encoding module includes three ResRF convolution layers connected in series; the ResRF convolution layer is used to focus on the spatial features of the receptive field, enhance the attention to spatial features and channel features, and improve the feature extraction of spatial and channel dimensions; the decoder is used to map the output of the encoder back to the original data space, and restore the details and structure of the input data as much as possible to generate a reconstructed output as similar as possible to the original input data, and transmit the obtained image features to the decoder; The decoder includes a first multi-attention gate fusion module, a first processing layer, a second multi-attention gate fusion module, a second processing layer, a third multi-attention gate fusion module, a third processing layer, a fourth multi-attention gate fusion module and a fourth processing layer in sequence; the first multi-attention gate fusion module to the fourth multi-attention gate fusion module have the same structure; the multi-attention gate fusion module is used to improve the selectivity of feature fusion, so that the segmentation model is more focused on learning features in the target area and reducing noise interference; the first processing layer to the fourth processing layer have the same structure; the processing layer is used to process the output of the multi-attention gate fusion module; the output of the fourth encoding module and the output of the fifth encoding module are both used as inputs of the first multi-attention gate fusion module; the output of the first multi-attention gate fusion module and the output of the upsampled fifth encoding module are used as inputs of the first processing layer; the output of the first processing layer and the output of the third encoding module are both used as inputs of the second multi-attention gate fusion module; The output of the second multi-attention gate fusion module and the output of the first processing layer after upsampling are used as the input of the second processing layer; the output of the second processing layer and the output of the second encoding module are both used as the input of the third multi-attention gate fusion module; The output of the third multi-attention gate fusion module and the output of the second processing layer after upsampling are used as the input of the third processing layer; the output of the third processing layer and the output of the first encoding module are both used as the input of the fourth multi-attention gate fusion module; The output of the fourth multi-attention gate fusion module and the output of the third processing layer after upsampling are used as the input of the fourth processing layer; the output of the fourth processing layer is used as the output of the decoder; The cropper is used to crop the image features output by the decoder by center cropping to obtain candidate area blocks, and transmit the obtained candidate area blocks to the classification model; The classification model is used to perform feature extraction and region classification based on the feature data of the candidate region blocks to complete the classification of the retinal microaneurysm region of the input image; the classification model includes a 7×7 convolutional layer, a BN normalization layer, a ReLU activation function layer, a first RMS layer, a second RMS layer, a third RMS layer, a fourth RMS layer and a fully connected layer connected in series; the structures of the first to fourth RMS layers are the same, and the RMS layer is used to learn multi-scale information, refine the texture information between the target area and the background, and perform feature response calibration on the extracted feature map to highlight the area of interest.
4. The retinal microaneurysm region segmentation-classification method for fundus images according to claim 3, characterized in that The processing of the ResRF convolutional layer includes the following steps: The input feature data is processed by the first 1×1 convolution sublayer to obtain the first feature; the input feature data is processed by the second 1×1 convolution sublayer, the RFCBAM sublayer and the third 1×1 convolution sublayer in sequence to obtain the second feature; the first feature and the second feature are connected to obtain the output of the ResRF convolution layer; The processing of the RFCBAM sublayer includes the following steps: The input data of the RFCBAM sublayer is divided into two paths: The first path is processed in sequence through the global average pooling layer, the fully connected layer, the ReLU activation function layer, the fully connected layer, and the Sigmoid activation function layer to obtain the channel feature F channel ; F channel The calculation formula is F channel =σ(FC(ReLU(FC(GAP(F))))), where F is the input data of the RFCBAM sublayer, GAP() is the global average pooling layer processing function; FC() is the fully connected layer processing function; ReLU() is the ReLU activation function; σ() is the Sigmoid activation function; The second path is processed by the grouped convolution layer, ReLU activation function layer, BN normalization layer and Reshape layer in turn to obtain the main feature F main ; F main The calculation formula is F main =Reshape(BN(ReLU(GroupConv(F)))), where GroupConv() is the group convolution layer processing function, BN() is the BN normalization layer processing function, and Reshape() is the Reshape layer processing function; the Reshape layer is used to adjust the dimension of the data; The main feature F main After being processed through the maximum pooling layer and the average pooling layer respectively, the output of the maximum pooling layer and the output of the average pooling layer are cascaded, and then processed through the 1×1 convolution layer and the Sigmoid activation function layer in turn to obtain the spatial feature F spatial ; F spatial The calculation formula is F spatial =σ(Conv 1×1 ([Avg(F main ),Max(F main )])), where Max() is the maximum pooling layer processing function, Avg() is the average pooling layer processing function, Conv 1×1 () is the 1×1 convolution layer processing function; [] is the cascade operation; The channel feature F channel Main features F main and spatial feature F spatial Multiply them and then process them through a 3×3 convolutional layer to obtain the output data of the RFCBAM sub-layer.
5. The retinal microaneurysm region segmentation-classification method for fundus images according to claim 4, characterized in that The processing of the multi-attention gate fusion module includes the following steps: The input of the multi-attention gate fusion module includes the corresponding encoder data and upsampled data; The upsampled data is processed by the grouped convolution layer, BN normalization layer, ReLU activation function layer, channel attention mechanism layer and spatial attention mechanism layer in turn to obtain the first fusion feature; The encoder data passes through the grouped convolution layer, BN normalization layer, ReLU activation function layer, channel attention mechanism layer and spatial attention mechanism layer in turn to obtain the second fusion feature; The first fusion feature and the second fusion feature are added to obtain a third fusion feature; The third fused feature is processed in sequence through the ReLU activation function layer, the 1×1 convolution layer, and the Sigmoid activation function layer to obtain the fourth fused feature; The fourth fusion feature is multiplied by the encoder data and then added to the encoder data to obtain the output of the multi-attention gate fusion module. MAGF ; Output of the multi-attention gate fusion module MAGF The calculation formula is Output MAGF =x×W4+x W4=σ(Conv 1×1 (ReLU(W3))) W3=W1+W2 W1=f SE (f ECA (ReLU(BN(GroupConv(g))))) W2=f SE (f ECA (ReLU(BN(GroupConv(x))))) Where x is the encoder data; W4 is the fourth fusion feature; W3 is the third fusion feature; W1 is the first fusion feature; W2 is the second fusion feature; f SE () is the spatial attention mechanism layer processing function; f ECA () is the channel attention mechanism layer processing function; g is the upsampled data.
6. The retinal microaneurysm region segmentation-classification method for fundus images according to claim 5, characterized in that The processing process of the processing layer includes the following steps: The processing layer includes a first 2d convolution layer, a first ReLU activation function layer, a second 2d convolution layer, and a second ReLU activation function layer connected in series in sequence; The data input to the processing layer is processed in sequence by the first 2D convolution layer, the first ReLU activation function layer, the second 2D convolution layer, and the second ReLU activation function layer to obtain the output of the processing layer.
7. The retinal microaneurysm region segmentation-classification method for fundus images according to claim 6, characterized in that The processing of the RMS layer includes the following steps: The input of the RMS layer is divided into two channels: The first path is processed by a 1×1 convolution layer, a 3×3 convolution layer, and a channel attention layer in sequence to obtain the first feature; The second path is processed by 1×1 convolution layer, RFB layer, 3×3 convolution layer and channel attention layer in sequence to obtain the second feature; After the first feature and the second feature are cascaded, they are processed through the spatial attention layer and the 1×1 convolution layer to obtain the cascaded processing feature; The cascade processing features are added to the input of the RMS layer and used as the output of the RMS layer. RMS ; Output of the RMS layer RMS The calculation formula is Output RMS =Conv 1×1 (SA(Concat(F1,F2)))+X F1=CA(Conv 3×3 (Conv 1×1 (X))) F2=CA(Conv 3×3 (RFB(Conv 1×1 (X)))) Where SA is the processing function of the spatial attention layer; Concat() is the cascade function; X is the input of the RMS layer; F1 is the first feature; F2 is the second feature; CA is the processing function of the channel attention layer; RFB() is the processing function of the RFB layer; The processing of the RFB layer comprises the following steps: The input of the RFB layer is divided into 4 channels: The first path is processed by a 1×1 convolutional layer and a 3×3 dilated convolutional layer with a dilation rate of 1 to obtain the first sub-feature; The second path is processed in sequence by a 1×1 convolutional layer, a 3×3 convolutional layer, and a 3×3 dilated convolutional layer with a dilation rate of 3 to obtain the second sub-feature; The third path is processed in sequence by a 1×1 convolutional layer, a 5×5 convolutional layer, and a 3×3 dilated convolutional layer with a dilation rate of 5 to obtain the third sub-feature; The fourth path directly serves as the fourth sub-feature; After cascading the first sub-feature, the second sub-feature, and the third sub-feature, they are processed through a 1×1 convolutional layer to obtain the cascaded sub-feature; The cascaded sub-features are added to the fourth sub-feature to obtain the output of the RFB layer.
8. An imaging method comprising the retinal microaneurysm region segmentation-classification method for fundus images according to any one of claims 1 to 7, characterized in that The following steps are also included: S6. The segmentation and classification results obtained in step S5 are marked and re-imaged on the actual fundus image to obtain a fundus image with the segmentation and classification results of the retinal microaneurysm area.