Semantic Segmentation Method for Sentinel Lymph Node Ultrasound Images Based on Deep Learning
By constructing the RA-U-Net++ network, combining the residual cavity pyramid module and deep supervision training, the problem of many manual interventions in sentinel lymph node ultrasonic image segmentation is solved, and efficient automatic segmentation and high-precision semantic segmentation effects are achieved.
Patent Information
- Application Number
- CN202411071788.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-08-06
AI Technical Summary
The prior art requires a lot of manual intervention in semantic segmentation of sentinel lymph node ultrasound images, and manual segmentation is time-consuming and labor-intensive and the segmentation results are poor.
The semantic segmentation of sentinel lymph node ultrasound images is used based on deep learning. By constructing a U-shaped structure RA-U-Net++ network, combining residual cavity pyramid module and deep supervision training, automatic segmentation of sentinel lymph node ultrasound images is achieved.
Automatic segmentation of sentinel lymph node ultrasound images is realized, and it gets rid of tedious and time-consuming manual segmentation, and has good segmentation effect, which can achieve higher intercombination and comparison.
Smart Images

Figure CN118967704B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image semantic segmentation, and in particular to a method for semantic segmentation of sentinel lymph node ultrasound images based on deep learning. Background Art
[0002] In the treatment of breast cancer, it is crucial to know whether the cancer has spread to the lymphatic system. The sentinel lymph node (SLN), as the "first stop" of cancer cell metastasis, is the key to evaluating whether breast cancer has spread. The dual-modal imaging technology of two-dimensional ultrasound and contrast-enhanced ultrasound plays a crucial role in the detection of SLN canceration. Although the dual-modal imaging technology of two-dimensional ultrasound and contrast-enhanced ultrasound has played a key role in SLN canceration detection, its inherent limitations and traditional diagnostic methods relying on manual review urgently need to be improved by introducing new technical means.
[0003] The literature "Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]. Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, 3431 - 440" proposed the Fully Convolutional Networks (FCN). The core contribution of FCN is to transform the traditional Convolutional Neural Network (CNN) into a fully convolutional form, enabling it to accept input images of any size and output dense prediction maps of corresponding sizes. However, the segmentation results of this method are not fine enough and the cost is relatively high. At the same time, the literature "Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation[C]. Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5 - 9, 2015, Proceedings, Part III18, 2015, 234 - 241" proposed the U-Net network. Its most prominent feature is the unique "U" - shaped structure. This structure enables the network to achieve high - precision segmentation by effectively using the context information and accurate localization information of the image with less data. Although the segmentation accuracy of U-Net is improved compared with FCN, it also has the disadvantages of large redundancy and insufficient global information extraction ability. Summary of the Invention
[0004] The purpose of the present invention is to provide a semantic segmentation method for sentinel lymph node ultrasound images based on deep learning, aiming at the deficiencies of the existing technology, so as to realize the semantic segmentation of dual - modality sentinel lymph node ultrasound images, and solve the problems that a large amount of manual intervention is required to identify the sentinel lymph node part in ultrasound images at the current stage, manual segmentation is time - consuming and laborious, and the segmentation results are not good.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A semantic segmentation method for sentinel lymph node ultrasound images based on deep learning, the method comprising the following steps:
[0007] Step 1: Obtain the ultrasound image dataset of sentinel lymph nodes, and preprocess the dataset to obtain the training set;
[0008] Step 2: Construct the RA-U-Net++ network for image semantic segmentation, and tune the network parameters of the RA-U-Net++ network based on the training set to obtain a medical image semantic segmentation model for target lesions;
[0009] Among them, the RA-U-Net++ network adopts a U-shaped structure, which includes L (L≥2) levels, and an encoding node is set at each level. The encoding node adopts a residual atrous pyramid module, and the output feature map of the encoding node at the current level is input into the encoding node of the next level after downsampling; from the first layer to the L-1 layer of the RA-U-Net++ network, decoding nodes are set. The first layer includes L-1 decoding nodes, and one decoding node is reduced layer by layer starting from the first layer. In each level, the decoding node is located after the encoding node. The input of each decoding node includes: the feature map after skip connection of the output feature maps of each node before the current decoding node in the same level, and the feature map after upsampling of the output feature map of the previous adjacent node in the next level; the final segmentation result is obtained based on the output feature map of the last decoding node in the first layer. Preferably, the output feature map of the last decoding node in the first layer can be sent into a 1×1 convolutional layer with an output channel of 2 to obtain the final segmentation result.
[0010] Further, in Step 2, when tuning the network parameters of the RA-U-Net++ network, deep supervision training is carried out in combination with the output feature maps of all decoding nodes in the first layer. The output feature maps of each decoding node are respectively sent into a 1×1 convolutional layer with an output channel of 2, and the segmentation losses of the output results of each convolutional branch are calculated and fused to obtain a comprehensive segmentation loss.
[0011] Further, the loss function used to calculate the segmentation loss is:
[0012]
[0013] Among them, N is the number of samples, y i represents the true label of the i-th sample, and p i represents the probability that the network predicts the i-th sample as the positive class, and σ() represents the Sigmoid function.
[0014] Further, the decoding node sequentially includes: a splicing layer and two convolutional modules, where the convolutional module sequentially includes a convolutional layer with a convolution kernel of 3×3, an activation function, and a batch normalization layer.
[0015] Further, the preprocessing includes image screening, image classification, image cropping, image enhancement processing, and image annotation.
[0016] Furthermore, the residual atrous pyramid module includes a convolutional module with residual connection and an atrous spatial pyramid pooling structure;
[0017] Among them, the convolutional module with residual connection includes a main branch and a shortcut connection branch. The main branch includes two 3×3 convolutional layers, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function; the shortcut connection branch sequentially includes a 1×1 convolutional layer and a batch normalization layer;
[0018] The input of the atrous spatial pyramid pooling structure is the output feature map of the convolutional module with residual connection, and it includes a global average pooling branch, an atrous convolutional branch, and a convolutional branch;
[0019] Among them, the global average pooling branch sequentially includes a 1×1 global average pooling layer, a 1×1 convolutional layer, a batch normalization layer, a ReLu activation function, and upsampling;
[0020] The atrous convolutional branch includes three atrous convolutional modules with different atrous rates in parallel. Each atrous convolutional module sequentially includes an atrous convolutional layer, a batch normalization layer, and a ReLu activation function;
[0021] The convolutional branch sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a ReLu activation function;
[0022] Concatenate the output feature maps of the global average pooling branch, the atrous convolutional branch, and the convolutional branch of the atrous spatial pyramid pooling structure, and then obtain the output feature map of the residual atrous pyramid module through a convolutional module. Among them, this convolutional module sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a ReLu activation function.
[0023] The technical solution provided by the present invention at least brings the following beneficial effects:
[0024] The present invention realizes the automatic segmentation of sentinel lymph node ultrasound images, getting rid of the cumbersome, time-consuming and laborious manual segmentation; and the segmentation effect of the present invention is good, and a relatively high intersection over union IoU can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0026] Figure 1 It is a schematic diagram of the implementation process of the sentinel lymph node ultrasound image semantic segmentation method based on deep learning provided by the embodiment of the present invention;
[0027] Figure 2 Schematic diagram of the network model structure adopted in the embodiment of the present invention;
[0028] Figure 3 Schematic diagram of the standard directory format of the picture annotation tool labelme adopted in the embodiment of the present invention;
[0029] Figure 4 Schematic diagram of the structure of the residual dilated pyramid module adopted in the embodiment of the present invention. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described in detail and completely below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Generally, the components of the embodiments of the present invention described and shown in the accompanying drawings can be arranged and designed in different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the present application that is claimed, but merely represents the selected embodiments of the present invention.
[0031] The embodiment of the present invention provides a method for semantic segmentation of sentinel lymph node ultrasound images based on deep learning. First, the sentinel lymph node data is preprocessed to establish a data set, and a labeled data set is obtained. Then, based on this data set, the constructed semantic segmentation network (RA-U-Net++ network) is trained to obtain a network model (medical image semantic segmentation model) for sentinel lymph node segmentation, so as to realize the semantic segmentation of the target ultrasound image. That is, the preprocessed image data is used as the input of the RA-U-Net++ network, and the network parameters are adjusted through backpropagation to obtain the optimal network structure. Furthermore, the manually set gold standard can be evaluated with the segmentation result of the model to quantify the segmentation effect of the medical image semantic segmentation model of the present invention. When applying the method for semantic segmentation of sentinel lymph node ultrasound images based on deep learning provided in the embodiment of the present invention, as Figure 1 shown, its main process includes:
[0032] (1) Obtain the original data set and perform preprocessing on it (such as image screening, image classification, preprocessing, image annotation, etc.) to obtain data sets SLN2DUS and SLNCEUS that can be used for image semantic segmentation. In the embodiment of the present invention, the multi-modal ultrasound images of sentinel lymph nodes adopted include conventional two-dimensional ultrasound and contrast-enhanced ultrasound images.
[0033] In this embodiment, in order to implement a Python script for format conversion, the annotated JSON files are converted into the semantic segmentation label format of PASCAL VOC. The JSON files in the SLN2D_Annotated and SLNCE_Annotated folders are processed by a Python script and converted into two folders, SLN2DUS and SLNCEUS. The file directory of each folder includes JPEGImages, SegmentationClass, class_name.txt, SegmentationClassBinary, and SegmentationClassVisualization. Among them, the SLN contrast-enhanced ultrasound images are stored in JPEGImages; the annotation information corresponding to the original images is stored in SegmentationClass, usually a color image, where different colors represent different categories; the binary annotation images are stored in SegmentationClassBinary, where each pixel has only two possible values (0 or 255); the visualization version of the annotation images is stored in SegmentationClassVisualization; and class_name.txt records the list of names of all target categories participating in the segmentation, with each line representing a category.
[0034] (2) Data augmentation is performed on the datasets SLN2DUS and SLNCEUS to enrich the sample data. For example, the Albumentations library is used for data augmentation, and the data augmentation methods include but are not limited to: RandomRotate90 (randomly rotate 90 degrees), Flip (randomly horizontally or vertically flip the image), OneOf (randomly adjust the image hue, brightness, contrast, saturation), etc.
[0035] The processed dataset is divided into a training set Train and a test set Test. Among them, the training set Train is used for parameter tuning of the network model and includes M original medical images and M gold standards; the test set Test is used to evaluate the segmentation performance of the trained model and includes N original medical images Test M , and N gold standards Test G .
[0036] (3) Based on the processed dataset, the RA-U-Net++ network constructed in the embodiment of the present invention is trained through multiple data iterations to obtain an optimal model, that is, a medical image semantic segmentation model;
[0037] (4) Performance testing is carried out on the medical image semantic segmentation model based on the test set Test, and the evaluation results are obtained based on the specified evaluation metrics. That is, one by one, the sentinel lymph node ultrasound images are used as the input of the medical image semantic segmentation model, the model gives the segmentation results of each voxel, and then the segmentation results of the model are evaluated with the gold standard marked by doctors. The evaluation criteria include but are not limited to: IoU (Intersection over Union), mIoU (Mean Intersection over Union), fwIoU (Frequency Weighted Intersection over Union), Acc (Accuracy), mAcc (mean pixel-wise accuracy).
[0038] In the embodiment of the present invention, the constructed RA-U-Net++ network adopts a symmetric U-shaped structure, as Figure 2 shown, it includes several layers (node layers), for example, including 5 layers. The overall architecture of the RA-U-Net++ network can be divided into three parts: an encoder - nested skip connections - decoder, and has a deep supervision mechanism; in addition, by integrating the ASPP structure, the RA-U-Net++ expands the receptive field of the network, enhances the semantic expression and the processing ability of multi-scale features. At the same time, the addition of residual connections effectively alleviates the problem of gradient disappearance and further improves the accuracy of segmentation. Figure 2 In, the green arrow represents downsampling, which is used to gradually reduce the spatial resolution of the image while increasing the depth of the feature channels, thereby allowing the network to capture more global features; the blue arrow represents upsampling, which is used to gradually restore the resolution of the feature map, reconstruct the details and structure of the image, and enable the network to perform accurate pixel-level predictions; the red dashed arrow is a skip connection (i.e., a shortcut connection), which specifically represents a concatenation operation, for the feature map of the current layer node and the upsampled feature map of the next level, and can reduce the semantic gap between the encoder and decoder feature maps.
[0039] The encoder of the RA-U-Net++ network is the backbone network part, which is the node X 0,0 、X 1,0 、X 2,0 、X 3,0 、X 4,0 and the downsampling path composed of downsampling; at the same time, a composite module that combines residual connections and dilated spatial pyramid pooling structure (in the embodiment of the present invention, it is named ResidualASPPBlock) is adopted in the encoder, and in Figure 2 it is represented by red nodes to enhance the feature transfer ability of the RA-U-Net++ network and improve its ability in feature fusion.
[0040] The decoder of the RA-U-Net++ network is an upsampling path composed of upsampling, skip connections, and convolutional modules. The feature maps from the same level and the upsampled feature maps from the previous level are concatenated and input into a convolutional network. Each convolutional network contains two identical convolutional modules, and each convolutional module consists of a convolutional layer with a 3×3 convolutional kernel, an activation function (usually ReLU), and a batch normalization layer (BN). Subsequently, the feature maps output by the convolutional network are upsampled, and then concatenated with other feature maps and input into the next decoder node until X 4,0 The decoded feature maps are then fed into a 1×1 convolutional layer with an output channel of 2 to obtain a segmentation image of size h×w×2, where h and w are the same as the original image size.
[0041] During network training, the RA-U-Net++ network implements deep supervision by adding a 1×1 convolutional layer with an output channel of 2 after the outputs of X 1,0 、X 2,0 、X 3,0 、X 4,0 After convolutional calculation, feature maps of the same size can be obtained for each branch, and since the size of the feature maps is the same as that of the original image, they can be used to directly predict the final segmentation map. In the embodiment of the present invention, the loss function is as shown in the following formula:
[0042]
[0043] where N is the number of samples, y i represents the true label of the i-th sample, and p i represents the probability that the model predicts the i-th sample as the positive class. σ() represents the Sigmoid function, which is used to map the original output to the interval (0,1).
[0044] Embodiment
[0045] Environment setup and language selection: The operating system is Ubuntu20.04, the CPU model is AMD EPYC7542 32-Core, the GPU is RTX 3090 (24GB), the memory size is 64GB. This embodiment is written in Python, the Python version is Python 3.8, the Pytorch version is PyTorch 1.11.0, and the CUDA version is Cuda 11.3. In addition, dependent libraries such as opencv-python 4.9.0.80 and albumentations 1.3.1 are also installed.
[0046] 1. Dataset preparation:
[0047] In this embodiment, the ultrasound image data of sentinel lymph nodes is from relevant departments of Sichuan Provincial People's Hospital. The medical expert team collected multi-modal ultrasound images of sentinel lymph nodes in the affected axilla of 763 breast cancer patients who visited the hospital from June 2017 to May 2022, including conventional two-dimensional ultrasound and contrast-enhanced ultrasound images.
[0048] 1-1) Picture screening. Since there is some unusable data in the pictures collected from the hospital, in this embodiment, they are deleted based on manual selection.
[0049] 1-2) Image classification. The images exported from the ultrasound instrument include two types: SLN contrast-enhanced ultrasound and SLN two-dimensional ultrasound. Due to the significant differences in color and features between these two types of images, in this embodiment, they are distinguished so that they can be accurately labeled separately and appropriate model parameters can be applied for training.
[0050] 1-3) Preprocessing. Preprocessing includes two parts: ① converting the images from JPEG format to PNG format, ② ReSize the images and cut off the useless parts. The picture format exported from the ultrasound instrument is JPEG. However, in the subsequent binarization process, due to the compression characteristics of the JPEG format, some pixel points cannot be correctly binarized. Therefore, to avoid this situation, this embodiment uses the lossless format PNG. In addition, the images exported from the ultrasound instrument are original images, and many regions in the pictures are not required for training (such as machine names, hospital Logos). These additional regions will increase the picture size and reduce the training speed. Therefore, this embodiment cuts off the useless information and only retains the central part of the image. For SLN contrast-enhanced ultrasound images, the size after ReSize is 448×480, and for SLN two-dimensional ultrasound images, the size after ReSize is 736×496.
[0051] 1-4) Construct the labelme standardized directory format. To ensure the efficiency and orderliness of the annotation process, this embodiment constructs the standardized directory structure as shown in Figure 3 The SLN2D_img and SLNCE_img directories store SLN two-dimensional contrast images and SLN contrast-enhanced ultrasound images respectively. Correspondingly, the SLN2D_Annotated and SLNCE_Annotated directories are used to save the completed annotated json label files.
[0052] 1-5) Put the preprocessed SLN contrast-enhanced ultrasound images and SLN two-dimensional ultrasound images into the SLNCE_img and SLN2D_img folders respectively.
[0053] 1-6) Label the images using Labelme. First, click the CreatePolygons button to start drawing polygons, and then use the mouse to draw points to mark the boundaries of the targets (when the mouse is placed on the first point and clicked, the boundary will be automatically closed). After labeling, a selection box for choosing the category will pop up, and you can select the corresponding category.
[0054] 1-7) The labeling results are json label files. SLN2D_Annotated and SLNCE_Annotated store the labeling results of SLN two-dimensional ultrasound and SLN contrast-enhanced ultrasound respectively.
[0055] 1-8) Submit the labeling results to professional doctors for review on Labelme to ensure that the sentinel lymph nodes in each image are correctly segmented.
[0056] 2. Construction of the RA-U-Net++ network:
[0057] Implement the code writing and running according to the network structure using the Python language and the PyTorch framework.
[0058] 2-1) Implementation of the encoder module. The RA-U-Net++ network constructed in this embodiment is as Figure 2 shown. Its network layer consists of 5 layers, and the number of channels in each layer is set to 32, 64, 128, 256, and 512 respectively. The encoder uses a residual atrous pyramid module, which is represented by red nodes in Figure 2 . This module integrates residual connections and ASPP (Atrous Spatial Pyramid Pooling) on the basis of the standard convolutional layer, uses atrous convolutions with different atrous rates to expand the receptive field of the model, enabling it to capture more extensive context information. At the same time, introducing residual connections can help the gradient propagate more effectively through the network, alleviate the problem of gradient disappearance, and ensure that deep features are effectively learned. In the selection of downsampling, the RA-U-Net++ network uses the classic Max Pooling.
[0059] Implementation of the convolutional network structure (convolutional module with residual connection). See Figure 4, this structure contains two 3×3 convolutional layers with a stride of 1 and an output channel number of O (customizable) on the main line, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function. On the shortcut connection branch, there is a 1×1 convolutional layer with a stride of 1 and an output channel number of O as well, followed by a batch normalization layer. The 1×1 convolution is mainly used to adjust the channel number of the input data xi so that its size is equal to the output on the main line. Subsequently, the feature map on the main line is added to the feature map on the shortcut connection branch, and then the final feature map is output through the ReLu activation function.
[0060] Implementation of the Atrous Spatial Pyramid Pooling structure. As Figure 4 shown, its structure can be divided into three parts: ① Global average pooling branch; ② Atrous convolution branches with three different dilation rates; ③ 1x1 convolution branch; among them, the global average pooling branch includes a 1×1 global average pooling layer, a 1×1 convolutional layer, a batch normalization layer, a ReLu activation function, and upsampling. The global average pooling branch is used to capture the global context information of the image. This branch first performs global average pooling on the input feature map to obtain a single feature vector, then adjusts the channel number through a 1x1 convolution, then performs batch normalization and introduces non-linearity through the ReLu function, and finally restores to the same size as the original feature map through upsampling. This can ensure that the model can utilize the global information of the entire image; the atrous convolution branches with different dilation rates consist of three parallel atrous convolutions with different dilation rates, and each module consists of an atrous convolutional layer, a batch normalization layer, and a ReLu activation function, where the dilation rates are set to 6, 12, and 18 respectively. The introduction of the atrous convolution branches with three different dilation rates expands the receptive field of the convolutional layer, enabling the model to observe a wider area in the input image; at the same time, it can capture multi-scale context information in the image, improving the accuracy and robustness of segmentation or recognition; in addition, ASPP also includes a 1×1 convolution branch, which is used to capture the feature information at the original scale.
[0061] Implementation of the Residual Atrous Pyramid Module. As Figure 4 shown, the Residual Atrous Pyramid Module consists of a convolutional module with a residual connection and an Atrous Spatial Pyramid Pooling structure. The upper part is a convolutional module with a residual connection, and the lower part is an Atrous Spatial Pyramid Pooling structure, that is, the output feature maps of the three parts of the Atrous Spatial Pyramid Pooling structure are concatenated, and then passed through a 1×1 convolutional module (sequentially including a 1×1 convolutional layer, a batch normalization layer, and a ReLu activation function) to obtain the output feature map of the Residual Atrous Pyramid Module.
[0062] 2-2) Implementation of the decoder module.
[0063] See Figure 2, in this embodiment, there are a total of 5 node layers, and its backbone network is node X in the figure n,0 (n = 0, …, 4), that is, encoding nodes; each layer corresponds to a number of decoding nodes X i,j , i is used to identify the node layer identifier, j is used to identify the node identifier of each node layer. The first layer includes 4 decoding nodes, and the number of decoding nodes decreases layer by layer. Only one encoding node is set in the last layer. That is, each decoding node sequentially includes a splicing layer and a convolutional network including two identical convolutional modules. An upsampling layer is provided between the decoding nodes of adjacent node layers to realize the splicing of the feature map from the same level and the upsampled feature map from the previous level by the current decoding node. The output feature map of the last decoding node in the first layer is sent into a 1×1 convolutional layer with an output channel of 2 to obtain the final segmentation result.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0065] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the creative concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for semantic segmentation of sentinel lymph node ultrasound images based on deep learning, characterized in that, It includes the following steps: Step 1: Obtain the ultrasound image dataset of sentinel lymph nodes, and preprocess the dataset to obtain a training set; Step 2: Construct a RA-U-Net++ network for image semantic segmentation, and optimize the network parameters of the RA-U-Net++ network based on the training set to obtain a medical image semantic segmentation model for the target lesion; Among them, the RA-U-Net++ network adopts a U-shaped structure, which includes L levels, where L≥2; an encoding node is set at each level, and the encoding node adopts a residual atrous pyramid module. The output feature map of the encoding node at the current level is input into the encoding node of the next level after downsampling; From the first layer to the L-1 layer of the RA-U-Net++ network, decoding nodes are set. The first layer includes L-1 decoding nodes, and one decoding node is reduced layer by layer starting from the first layer. In each level, the decoding node is located after the encoding node. The input of each decoding node includes: the feature map after skip connection of the output feature maps of each node before the current decoding node in the same level, and the feature map after upsampling of the output feature map of the previous adjacent node in the next level; Obtain the final segmentation result based on the output feature map of the last decoding node in the first layer; The residual atrous pyramid module includes a convolutional module with residual connection and an atrous spatial pyramid pooling structure; Among them, the convolutional module with residual connection includes a main branch and a shortcut connection branch. The main branch includes two 3×3 convolutional layers, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function; the shortcut connection branch sequentially includes a 1×1 convolutional layer and a batch normalization layer; The input of the atrous spatial pyramid pooling structure is the output feature map of the convolutional module with residual connection, and it includes a global average pooling branch, an atrous convolution branch, and a convolution branch; Among them, the global average pooling branch sequentially includes a 1×1 global average pooling layer, a 1×1 convolutional layer, a batch normalization layer, a ReLu activation function, and upsampling; The atrous convolution branch includes three atrous convolution modules with different atrous rates in parallel. Each atrous convolution module sequentially includes an atrous convolutional layer, a batch normalization layer, and a ReLu activation function; The convolution branch sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a ReLu activation function; Concatenate the output feature maps of the global average pooling branch, the atrous convolution branch, and the convolution branch of the atrous spatial pyramid pooling structure, and then obtain the output feature map of the residual atrous pyramid module through a convolutional module. Among them, this convolutional module sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a ReLu activation function.
2. The method according to claim 1, wherein Send the output feature map of the last decoding node in the first layer into a 1×1 convolutional layer with an output channel of 2 to obtain the final segmentation result.
3. The method according to claim 1, wherein In Step 2, when optimizing the network parameters of the RA-U-Net++ network, perform deep supervision training by combining the output feature maps of all decoding nodes in the first layer. Send the output feature maps of each decoding node into a 1×1 convolutional layer with an output channel of 2, and separately calculate the segmentation loss of the output result of each convolution branch and fuse them to obtain a comprehensive segmentation loss.
4. The method according to claim 3, wherein The loss function used to calculate the segmentation loss is as follows: Among them, N is the number of samples, and y i represents the true label of the i-th sample, and p i represents the probability that the network predicts the i-th sample as the positive class, and σ() represents the Sigmoid function.
5. The method according to claim 1, wherein The decoding nodes sequentially include: a splicing layer and two convolutional modules. Among them, the convolutional module sequentially includes a convolutional layer with a convolution kernel of 3×3, an activation function, and a batch normalization layer.
6. The method according to claim 1, wherein The preprocessing includes image screening, image classification, image cropping, image enhancement processing, and image annotation.
7. The method according to claim 1, wherein The three dilation rates of the dilated convolution branch are respectively set to: 6, 12, 18.
8. The method according to claim 1, characterized in that The value of hierarchical level L is set to 5, and the number of channels of each hierarchical level is respectively set to 32, 64, 128, 256, 512.
Citation Information
Patent Citations
Image segmentation method, system and storage medium
CN114066903A