Medical Image Segmentation Model Based on Pseudo-Mask Guided Feature Aggregation

By introducing pseudo-mask-guided feature enhancement and multi-scale multi-stage feature aggregation modules in the medical image segmentation model, the problems of reduced boundary accuracy and incompatibility in the existing technology are solved, and higher-quality nucleus and gland segmentation results are achieved.

CN115909326BActive Publication Date: 2025-06-20NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211342921.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-06-20
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have problems such as reduced boundary accuracy, incompatibility of features and neglect of spatial relationships in cell nucleus and gland segmentation, resulting in unsatisfactory segmentation results.

Method used

A medical image segmentation model (PG-FANet) based on pseudo-mask-guided feature aggregation is proposed. Through convolution blocks, pseudo-mask-guided feature enhancement modules and multi-scale multi-stage feature aggregation modules, multi-scale and multi-stage feature aggregation modules, multi-scale and multi-stage feature aggregation modules, multi-scale and multi-stage features are extracted and aggregated to improve the segmentation effect.

Benefits of technology

This model effectively extracts the contextual characteristics of the cell nucleus and glands, improves the accuracy of segmentation boundaries and the quality of segmentation results, which is better than existing advanced methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115909326B_ABST
    Figure CN115909326B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical image segmentation model based on pseudo-mask-guided feature aggregation. The medical image segmentation model includes: a convolutional block, a second-order network model structure, a pseudo-mask-guided feature enhancement module, a multi-scale multi-stage feature aggregation module, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a loss function. The second-order network model structure includes: a first-order sub-network and a second-order sub-network. The medical image segmentation model of the present invention is tested on the nucleus segmentation MoNuSeg dataset and the gland segmentation CRAG dataset using multi-scale and multi-stage features, not only extracting detailed nucleus and gland segmentation results, but also achieving an advanced quality assessment effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and particularly relates to a medical image segmentation model based on pseudo-mask-guided feature aggregation. Background Art

[0002] Biomedical image analysis is usually the first step in the diagnosis, stratification, and clinical management of diseases such as cancer. The size, shape, and some other morphological appearances of histological image structures such as cell nuclei and glands are highly correlated with the presence and severity of diseases [1]. Trained pathologists usually manually examine these images, which is laborious and subjective. Therefore, automated computer methods [2] have been developed for the quantitative and objective analysis of histopathological images. Image segmentation for extracting cells, cell nuclei, or glands from histopathological images precedes all analyses and is a key step in the entire automated diagnosis and analysis process.

[0003] The current state-of-the-art cell nucleus / gland segmentation methods [1, 3, 4, 5, 6] adopt the fully convolutional network (FCN) [7] and its variant U-Net [8], using the encoder to extract features and recover from low-resolution feature maps to generate high-resolution segmentation predictions. Compared with traditional methods, they have achieved good segmentation performance. The encoder part of the FCN is inspired by the structure originally designed for image classification, that is, the receptive field can be greatly increased through pooling operations to extract more abstract features. However, this pooling downsampling may reduce the accuracy of the boundaries in the segmentation task. In addition, the skip connections of U-Net may introduce feature incompatibility [9] and bring differences throughout the propagation process. Another problem in these methods is the cross-entropy loss used for training, which only cares about whether the classification of a single pixel is correct and does not care about the spatial relationship between pixels. Although deep neural networks can learn some high-level features from real labels, when objects have extremely similar colors, textures, and shapes, especially in medical images, the segmentation results are still not satisfactory. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a medical image segmentation model based on pseudo-mask-guided feature aggregation, which effectively extracts the context features of cell nuclei (MoNuSeg) / glands (CRAG), segments the corresponding cell nucleus / gland instances, and is used for downstream task analysis.

[0005] The purpose of the present invention is achieved by the following technical solutions.

[0006] A medical image segmentation model (PG-FANet) based on pseudo-mask-guided feature aggregation, comprising: a convolutional block, a second-order network model structure, a pseudo-mask-guided feature enhancement module (MGFE), a multi-scale multi-stage feature aggregation module (MMFA), a first convolutional layer, a second convolutional layer, a third convolutional layer, and a loss function L seg , the second-order network model structure includes: a first-order sub-network and a second-order sub-network;

[0007] The convolutional block is used to input initial data thereto and respectively flow the rough features output from the convolutional block to the first-order sub-network and the pseudo-mask-guided feature enhancement module;

[0008] The architectures of the second-order sub-network and the first-order sub-network are the same, each including: I + 1 residual blocks (RB i_s ), and an atrous spatial pyramid pooling (ASPP) module. The I + 1 residual blocks (RB i_s ) of the first-order sub-network are used to finely adjust the rough features and then convey the first-order refined features to the atrous spatial pyramid pooling (ASPP) module of the first-order sub-network; the atrous spatial pyramid pooling (ASPP) module of the first-order sub-network is used to extract high-order latent features from the first-order refined features;

[0009] The first convolutional layer is used to generate a pseudo-mask for the high-order latent features obtained by the first-order sub-network, and the pseudo-mask is respectively conducted to the pseudo-mask-guided feature enhancement module and the loss function L seg ;

[0010] The pseudo-mask-guided feature enhancement module is used to enhance the expression ability of the rough features by using the pseudo-mask to obtain pseudo-mask-guided fused features;

[0011] The I + 1 residual blocks (RB i_s ) of the second-order sub-network are used to input the fused features and output second-order refined features, and the atrous spatial pyramid pooling (ASPP) module of the second-order sub-network is used to receive the second-order refined features output from the (I + 1)-th residual block of the second-order sub-network and output high-order latent features;

[0012] The multi-scale multi-stage feature aggregation module (MMFA) includes: a multi-scale feature aggregation module and a multi-stage feature aggregation module. The multi-scale feature aggregation module is used to perform multi-scale feature aggregation on the low-level features output from the i-th residual block of the first-order sub-network and the low-level features output from the i-th residual block of the second-order sub-network to obtain multi-scale aggregated features, where i = 1,..., I;

[0013] The second convolutional layer is used to fuse the multi-scale aggregated features to output high-order features;

[0014] The multi-stage feature aggregation module is used to perform multi-stage feature aggregation on the feature output of the (I + 1)-th residual block of the first-order sub-network, the feature output of the (I + 1)-th residual block of the second-order sub-network, and the high-order features, and then output multi-scale multi-stage aggregated features;

[0015] The third convolutional layer is used to perform feature concatenation and then fusion on the multi-scale multi-stage aggregated features and the high-order latent features obtained from the second-order sub-network to obtain a prediction result;

[0016] The loss function L seg is used to calculate based on the pseudo-mask and the prediction result to obtain the loss function value.

[0017] In the above technical solution, the first convolutional layer includes: an upsampling layer and a convolutional layer, and the calculation process of the first convolutional layer is as follows:

[0018] Y s = Conv(Up(X c ))

[0019] where X c is the high-order latent feature obtained from the first-order sub-network, Up is the upsampling layer in the first convolutional layer, Con is the convolutional layer, and Y s is the pseudo-mask.

[0020] In the above technical solution, the calculation formula of the multi-scale feature aggregation module is:

[0021]

[0022] where X m is the multi-scale aggregated feature, is the i-th residual block (RB i_s ) of the s-th order sub-network, is the low-level feature output by the (i - 1)-th residual block of the s-th order sub-network, s = 1, 2, i = 1,..., I, where is the rough feature output by the convolutional block, is the fusion feature guided by the pseudo-mask, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, and Conv is the convolutional layer.

[0023] In the above technical solution, the operation process of the second convolutional layer is as follows:

[0024] X′ m = Conv(X m )

[0025] where X′ m is the high-order feature, Conv is the convolutional layer, and X mis multi-scale aggregated features;

[0026] In the above technical solution, the calculation formula of the multi-stage feature aggregation module is as follows:

[0027]

[0028] where X' m is the high-order feature, X h is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, Conv is the convolutional layer, is the (I + 1)-th residual block of the s-th order sub-network, s = 1, 2, is the feature output of the I-th residual block of the s-th order sub-network.

[0029] In the above technical solution, the third convolutional layer includes: upsampling, feature concatenation and convolutional layer, and the calculation formula of the third convolutional layer is as follows:

[0030] Y s = Conv(concat(X h , Up(X f )))

[0031] where Y s is the prediction result, Conv is the convolutional layer, concat is the feature concatenation operation, X h is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, X f is the high-order latent feature obtained from the second-order sub-network.

[0032] In the above technical solution, the loss function where λ Dice is the weight value for adjusting L Dice and λ vcc is the weight value for adjusting L vcc ;

[0033]

[0034] where B is the minimum batch in training, D is the number of instances in B, B d is all the pixels in the minimum batch (MiniBatch) that belong to instance d, |B d | is the number of pixels in B d , μ d is the average value of the probabilities of the correct classes of all pixels in B d , p h is the probability of the correct class of pixel h in the prediction result obtained by the third convolutional layer, h = 1,..., |B d |.

[0035] The medical image segmentation model of the present invention is tested on the nucleus segmentation MoNuSeg dataset and the gland segmentation CRAG dataset using multi-scale and multi-stage features. It not only extracts detailed nucleus and gland segmentation results, but also achieves advanced quality assessment effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic structural diagram of the medical image segmentation model of the present invention;

[0037] Figure 2 It is a schematic structural diagram of the pseudo-mask guided feature enhancement module;

[0038] Figure 3 It is a schematic structural diagram of the multi-scale multi-stage feature aggregation module;

[0039] Figure 4 It is the nucleus segmentation effect on the MoNuSeg dataset and the gland segmentation effect diagram on the CRAG dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The technical solution of the present invention will be further described below with specific embodiments.

[0041] Embodiment 1

[0042] A medical image segmentation model based on pseudo-mask guided feature aggregation (PG-FANet), comprising: a convolutional block, a second-order network model structure, a pseudo-mask guided feature enhancement module (MGFE), a multi-scale multi-stage feature aggregation module (MMFA), a first convolutional layer, a second convolutional layer, a third convolutional layer, and a loss function L seg , and the second-order network model structure includes: a first-order sub-network and a second-order sub-network;

[0043] The convolutional block is used to input initial data thereto and respectively flow the rough features output from the convolutional block to the first-order sub-network and the pseudo-mask guided feature enhancement module;

[0044] The second-order sub-network and the first-order sub-network have the same architecture, each including: I + 1 residual blocks (RB i_s ) and an atrous spatial pyramid pooling (ASPP) module. The I + 1 residual blocks (RB i_s ) of the first-order sub-network are used to refine the rough features and then send the first-order refined features to the atrous spatial pyramid pooling (ASPP) module of the first-order sub-network; the atrous spatial pyramid pooling (ASPP) module of the first-order sub-network is used to extract high-order latent features from the first-order refined features; in the embodiment of the present invention, I = 3;

[0045] The first convolutional layer is used to generate a pseudo-mask for the high-order latent features obtained by the first-order sub-network, and the pseudo-mask is respectively transmitted to the pseudo-mask guided feature enhancement module and the loss function L seg ;

[0046] The first-order sub-network is used for rough pseudo-mask generation, and the pseudo-mask guided feature enhancement module is used to enhance the expression ability of rough features by using the pseudo-mask to obtain pseudo-mask guided fusion features; the pseudo-mask guided feature enhancement module moves the attention of the second-order sub-network to the region of interest under the guidance of the pseudo-mask. As Figure 2 shown, the pseudo-mask guided feature enhancement module stitches together the rough features generated by the convolutional block and the pseudo-mask features of the first-stage sub-network and fuses them through a 1×1 convolutional layer as the input of the second-stage sub-network.

[0047] The I + 1 residual blocks (RB i_s ) of the second-order sub-network are used to input the fusion features and output second-order refined features, and the atrous spatial pyramid pooling (ASPP) module of the second-order sub-network is used to receive the second-order refined features output by the I + 1 residual block of the second-order sub-network and output high-order latent features for refining the prediction results;

[0048] In view of the fact that the second-order sub-network and the first-order sub-network extract features for different scales, shapes and sizes, the present invention uses a multi-scale multi-stage feature aggregation module (MMFA) to aggregate multi-scale and multi-stage features, improve the feature expression ability of the model, and avoid the feature incompatibility problem in the U-shaped skip connection. The multi-scale multi-stage feature aggregation module (MMFA) includes: a multi-scale feature aggregation module and a multi-stage feature aggregation module. The multi-scale feature aggregation module is used to perform multi-scale feature aggregation on the low-level features output by the i-th residual block of the first-order sub-network and the low-level features output by the i-th residual block of the second-order sub-network to obtain multi-scale aggregated features, where i = 1,..., I;

[0049] The second convolutional layer is used to fuse the multi-scale aggregated features to output high-order features;

[0050] The multi-stage feature aggregation module is used to perform multi-stage feature aggregation on the feature output of the I + 1 residual block of the first-order sub-network, the feature output of the I + 1 residual block of the second-order sub-network and the high-order features, and then output multi-scale multi-stage aggregated features;

[0051] The third convolutional layer is used to perform feature stitching and then fusion on the multi-scale multi-stage aggregated features and the high-order latent features obtained by the second-order sub-network to obtain the prediction result;

[0052] The loss function L seg is used to calculate according to the pseudo-mask and the prediction result to obtain the loss function value.

[0053] In the above technical solution, the first convolutional layer includes: an upsampling layer and a convolutional layer. The calculation process of the first convolutional layer is as follows:

[0054] Y s = Conv(Up(X c ))

[0055] where X c is the high-order latent feature obtained by the first-order sub-network, Up is the upsampling layer in the first convolutional layer, Con is the convolutional layer, and Y s is the pseudo-mask.

[0056] In the above technical solution, the calculation formula of the multi-scale feature aggregation module is:

[0057]

[0058] where X m is the multi-scale aggregation feature, is the i-th residual block (RB i_s ) of the s-th order sub-network, is the low-level feature output by the (i - 1)-th residual block of the s-th order sub-network, s = 1, 2, i = 1,..., I, where, is the rough feature output by the convolutional block, is the fusion feature guided by the pseudo-mask, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, and Conv is the convolutional layer. The multi-scale feature aggregation module re-uses the information guided by the pseudo-mask and obtains a better feature representation for further propagation. As Figure 3 shown, the convolutional layer in the multi-scale feature aggregation module is used to receive the low-level features.

[0059] In the above technical solution, the second convolutional layer is used to further improve the feature expression and provide more expressive high-level features for the multi-stage feature aggregation module. The operation process of the second convolutional layer is as follows:

[0060] X' m = Conv(X m )

[0061] where X' m is the high-order feature, Conv is the convolutional layer, and X m is the multi-scale aggregation feature;

[0062] In the above technical solution, as Figure 3As shown, the convolutional layer in the multi-stage feature aggregation module is used to receive the feature output of the (I + 1)-th residual block of the first-order sub-network and the feature output of the (I + 1)-th residual block of the second-order sub-network. As the network depth increases, the spatial information of low-level features (such as region boundaries) may be lost. The present invention uses a multi-stage feature aggregation module to fuse high-order features, the feature output of the (I + 1)-th residual block of the first-order sub-network, and the feature output of the (I + 1)-th residual block of the second-order sub-network, avoiding the introduction of U-shaped skip connections with incompatible features.

[0063] The calculation formula of the multi-stage feature aggregation module is as follows:

[0064]

[0065] Among them, X′ m is the high-order feature, X h is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, Conv is the convolutional layer, is the (I + 1)-th residual block of the s-th order sub-network, s = 1, 2, is the feature output of the I-th residual block of the s-th order sub-network.

[0066] In the above technical solution, the third convolutional layer includes: upsampling, feature concatenation, and a convolutional layer. The calculation formula of the third convolutional layer is as follows:

[0067] Y s = Conv(concat(X h , Up(X f )))

[0068] Among them, Y s is the prediction result, Conv is the convolutional layer, concat is the feature concatenation operation, X h is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, X f is the high-order latent feature obtained from the second-order sub-network.

[0069] In the above technical solution, the loss function L seg adopts the cross-entropy loss function (Cross Entropy Loss) (L ce ), the dice loss function (Dice Loss) (L Dice) and the variance-constrained cross-loss function are used together as the loss function. Among them, the variance-constrained cross-loss function performs local constraints on the pixels belonging to the same segmentation instance, aiming to solve the problem that the network cannot completely segment the entire segmentation instance when the segmentation instances in the image have uneven colors or textures. Since the medical image segmentation model of the present invention has a second-order network model structure, the present invention optimizes the first-order sub-network and the second-order sub-network through the cross-entropy loss function (L ce ) and the dice loss function (L Dice ). Therefore, the main loss of supervised learning is:

[0070]

[0071] Among them, Y is the gold standard of the initial data segmentation result, and Y s is the prediction result.

[0072] Although the cross-entropy loss function, the dice loss function, or their combination are usually used in many medical image segmentation methods and have achieved remarkable success. However, histological image segmentation not only requires segmenting the cell nuclei / glands from the background but also separating each object from other instances. The spatial relationship between objects needs to be considered. Therefore, the present invention uses the variance-constrained cross-loss function (L vcc ) as an auxiliary loss constraint for training.

[0073] L vcc is defined as [4]:

[0074]

[0075] Among them, B is the minimum batch in training, D is the number of instances in B, B d is all the pixels belonging to instance d in the mini-batch, |B d | is the number of pixels in B d , μ d is the average value of the probabilities of the correctly predicted classes of all pixels in B d , p h is the probability of the correctly predicted class of pixel h in the prediction result obtained from the third convolutional layer, where h = 1,..., |B d |.

[0076] The loss function L seg is:

[0077] The loss function

[0078] Among them, λ Dice is the weight value for adjusting L Dice , and λvcc To adjust the weight value of L vcc .

[0079] Example 2

[0080] Production of the cell nucleus dataset: The multi-organ nucleus segmentation dataset (MoNuSeg)

[23] consists of 44 H&E-stained histopathological images. The resolution of the images is 1000×1000 pixels. The present invention screens the electron microscopy images of cell nuclei to be trained, and selects the histopathological images with manually annotated cell nuclei to form a training dataset (30 images) and a test dataset (14 images). Subsequently, the images in the training dataset are randomly divided into a training set A (27 images) and a validation set B (3 images), and the test dataset serves as the test set C. Image patches of size 128×128 are cropped from each image using a sliding window. A total of 1728 image patches are obtained from MoNuSeg. Online data augmentation is performed on the image patches, including random scaling, flipping, rotation, and affine operations. All image patches are normalized using the mean and standard deviation of the images in ImageNet

[24] .

[0081] Training a neural network for cell nucleus segmentation: The image patches of the training set A are fed into a medical image segmentation model for training, and the optimal weights of the medical image segmentation model are selected using the image patches of the validation set B. Data augmentation such as flipping and rotation is performed on the test set C, and the image patches of the test set C are fed into the medical image segmentation model for prediction, and the prediction results are re-stitched in the original order of the image patches to obtain accurate cell nucleus segmentation results.

[0082] Set the number of image patches in one training batch to 16, and the number of training iterations to 500.

[0083] The cell nucleus segmentation quality score metrics include: F1-score (F1), intersection over union (IoU), average Dice coefficient (Dice), aggregated Jaccard index (AJI), and 95% Hausdorff distance (95HD).

[0084] Example 3

[0085] Gland dataset creation: The colorectal adenocarcinoma gland (CRAG) dataset contains a total of 38 whole slide images (WSIs), from which 213 H&E CRA images with different cancer grades were obtained

[25] . In the present invention, gland electron microscopy images to be trained were screened, and all gland electron microscopy images were selected to form a training dataset of 173 images and a test dataset of 40 images. Subsequently, the training dataset was randomly divided into 153 images as the training set D and 20 images as the validation set E, and the test dataset was used as the test set F. The resolution of these images is 1512×1516. For CRAG, the present invention extracted 5508 image patches of 480×480 pixels from 153 images in the training set D and performed online data augmentation, including random scaling, flipping, rotation, and affine operations. All these training images were normalized by using the mean and standard deviation of the images in ImageNet

[24] .

[0086] Training a neural network for gland segmentation: The image patches of the training set D were fed into a medical image segmentation model for training, and the optimal weights of the medical image segmentation model were selected using the image patches of the validation set E. The image patches of the test set F were augmented with data such as flipping and rotation, and were fed into the medical image segmentation model in the order of the image patches for prediction, and the prediction results were re-stitched in the original order of the image patches to obtain accurate gland segmentation results.

[0087] The number of image patches in one training batch was set to 8, and the number of training iterations was 300.

[0088] Gland segmentation quality score metrics include: F1-score (F1), object-level Dice coefficient (Dice obj ), object-level Hausdorff distance (Haus obj ), and 95% object-level Hausdorff distance (95HD obj ).

[0089] Evaluate Example 2 and Example 3.

[0090] The present invention conducts fully supervised medical image segmentation model training and testing experiments on MoNuSeg and CRAG, and re-implements some algorithms under the same experimental environment for fair comparison, such as Micro-Net

[10] , U2-Net

[11] , R2U-Net

[12] , LinkNet

[13] , FullNet

[14] , BiO-Net

[15] , U-Net

[16] . On the other hand, for complex algorithms without source code, such as M-Net

[17] , BiX-NAS

[18] , DCAN

[19] , MILD-Net[3], DSE

[20] , PRS 2

[21] , the present invention directly copies the quality score evaluation values from the original work of the above-mentioned comparative documents. All re-implemented algorithms use the same data augmentation and training strategies to ensure the fairness of the comparison.

[0091] Nucleus segmentation: As Figure 4 shown, the medical image segmentation model of the present invention successfully extracts nucleus instances. As shown in Table 1, the medical image segmentation model is superior to all other models in all evaluation metrics under the same and different experimental environment settings. Under the same experimental settings, compared with the outstanding algorithm FullNet

[14] , the medical image segmentation model of the present invention improves by 2% in the AJI metric, which is a key metric for nucleus segmentation. Moreover, compared with the models under different experimental settings, the performance of the medical image segmentation model of the present invention is also very excellent in the Dice and IoU metrics.

[0092] Table 1 Quality evaluation indices for nucleus segmentation. * indicates directly citing the original data

[0093]

[0094] Gland segmentation: As Figure 4 shown, the medical image segmentation model of the present invention successfully extracts gland instances. The present invention uses other methods to evaluate the gland segmentation performance of the medical image segmentation model. As shown in Table 2, among all models, the medical image segmentation model always achieves the best performance in terms of F1, Dice obj and Haus obj .

[0095] Table 2 Quality evaluation indices for gland segmentation. * indicates directly citing the original data

[0096]

[0097]

[0098] In summary, the experimental results show that the medical image segmentation model of the present invention is superior to the current state-of-the-art methods and improves the fully supervised segmentation results.

[0099] Example 4

[0100] The present invention further conducts ablation studies on the proposed MGFE and MMFA on the nuclear segmentation dataset MoNuSeg to evaluate the contributions of different modules in the framework of the present invention. First, all modules are removed, and the second-order network model structure is degraded to a single-stage network (i.e., DeepLabV2

[22] with additional convolutional layers), and then the proposed modules (i.e., MGFE, MMFA) are gradually added to the model. As shown in Table 3, when the present invention adds the first-order backbone to DeepLabV2, the AJI score, an overall performance evaluation metric on MoNuSeg, increases by 0.8%, indicating that the additional backbone cannot improve the model's ability. However, when the MGFE module is introduced, the AJI metric increases by 1.4%. The MGFE module uses the coarse segmentation result as an enhancement, ultimately improving the model's learning ability. With the application of the MMFA module, a continuous improvement in performance can be observed.

[0101] Table 3 Quality evaluation indices of ablation experiments for each module of PG-FANet.

[0102]

[0103] References

[0104] [1] Gurcan M N, Boucheron L E, Can A, et al. Histopathological image analysis: A review [J]. IEEE reviews in biomedical engineering, 2009, 2: 147-171.

[0105] [2] Janowczyk A, Madabhushi A. Deep learning for digital pathology image analysis: A comprehensive tutorial with selected use cases [J]. Journal of pathology informatics, 2016, 7(1): 29.

[0106] [3]Graham S,Chen H,Gamper J,et al.MILD-Net:Minimal information lossdilated network for gland instance segmentation in colon histology images[J].Medical image analysis,2019,52:199-211.

[0107] [4]Qu H,Riedlinger G,Wu P,et al.Joint segmentation and fine-grainedclassification of nuclei in histopathology images[C] / / 2019 IEEE 16thinternational symposium on biomedical imaging(ISBI 2019).IEEE,2019:900-904.

[0108] [5]Qu H,Wu P,Huang Q,et al.Weakly supervised deep nuclei segmentationusing points annotation in histopathology images[C] / / International Conferenceon Medical Imaging with Deep Learning.PMLR,2019:390-400.

[0109] [6]Xu Y,Li Y,Liu M,et al.Gland instance segmentation by deepmultichannel side supervision[C] / / International Conference on Medical ImageComputing and Computer-Assisted Intervention.Springer,Cham,2016:496-504.

[0110] [7]Long J,Shelhamer E,Darrell T.Fully convolutional networks for semantic segmentation[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition.2015:3431-3440.

[0111] [8]Ronneberger O,Fischer P,Brox T.U-net:Convolutional networks for biomedical image segmentation[C] / / International Conference on Medical image computing and computer-assisted intervention.Springer,Cham,2015:234-241.

[0112] [9]Ibtehaz N,Rahman M S.MultiResUNet:Rethinking the U-Net architecture for multimodal biomedical image segmentation[J].Neural networks,2020,121:74-87.

[0113]

[10] S.E.A.Raza,L.Cheung,M.Shaban,S.Graham,D.Epstein,S.Pelengaris,M.Khan,and N.M.Rajpoot,“Micro-Net:A unified model for segmentation of various objects in microscopy images,”Medical Image Analysis,vol.52,pp.160–173,2019.

[0114]

[11] X. Qin, Z. Zhang, C. Huang, M. Dehghan, O. R. Zaiane, and M. Jagersand, “U2-Net: Going deeper with nested U-structure for salient object detection,” Pattern Recognition, vol. 106, p. 107404, 2020.

[0115]

[12] M. Z. Alom, C. Yakopcic, T. M. Taha, and V. K. Asari, “Nuclei segmentation with recurrent residual convolutional neural networks based U-Net (R2U-Net),” in NAECON 2018 - IEEE National Aerospace and Electronics Conference. IEEE, 2018, pp. 228–233.

[0116]

[13] A. Chaurasia and E. Culurciello, “LinkNet: Exploiting encoder representations for efficient semantic segmentation,” in 2017 IEEE Visual Communications and Image Processing (VCIP). IEEE, 2017, pp. 1–4.

[0117]

[14] H. Qu, Z. Yan, G. M. Riedlinger, S. De, and D. N. Metaxas, “Improving nuclei / gland instance segmentation in histopathology images by full resolution neural network and spatial constrained loss,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2019, pp. 378–386.

[0118]

[15] T. Xiang, C. Zhang, D. Liu, Y. Song, H. Huang, and W. Cai, “BiO-Net: Learning Recurrent Bi-directional Connections for Encoder-Decoder Architecture,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2020, pp. 74–84.

[0119]

[16] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2015, pp. 234–241.

[0120]

[17] R. Mehta and J. Sivaswamy, “M-Net: A convolutional neural network for deep brain structure segmentation,” in 2017 IEEE 14th International Symposium on Biomedical Imaging (ISBI 2017). IEEE, 2017, pp. 437–440.

[0121]

[18] T. Xiang, C. Zhang, X. Wang, Y. Song, D. Liu, H. Huang, and W. Cai, “Towards bi-directional skip connections in encoder-decoder architectures and beyond,” Medical Image Analysis, vol. 78, p. 102420, 2022.

[0122]

[19] H. Chen, X. Qi, L. Yu, and P.-A. Heng, “DCAN: deep contour-aware networks for accurate gland segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2487–2496.

[0123]

[20] Y. Xie, H. Lu, J. Zhang, C. Shen, and Y. Xia, “Deep Segmentation Emendation Model for Gland Instance Segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2019, pp. 469–477.

[0124]

[21] Y. Xie, J. Zhang, Z. Liao, J. Verjans, C. Shen, and Y. Xia, “Pairwise Relation Learning for Semi-supervised Gland Segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2020, pp. 417–427.

[0125]

[22] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 4, pp. 834–848, 2017.5

[0126]

[23] N. Kumar, R. Verma, D. Anand, Y. Zhou, O. F. Onder, E. Tsougenis, H. Chen, P.-A. Heng, J. Li, Z. Hu et al., “A multi-organ nucleus segmentation challenge,” IEEE Transactions on Medical Imaging, vol. 39, no. 5, pp. 1380–1391, 2019.

[0127]

[24] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2009, pp. 248–255.

[0128]

[25] R. Awan, K. Sirinukunwattana, D. Epstein, S. Jefferyes, U. Qidwai, Z. Aftab, I. Mujeeb, D. Snead, and N. Rajpoot, “Glandular morphometrics for objective grading of colorectal adenocarcinoma histology images,” Scientific Reports, vol. 7, no. 1, pp. 1–12, 2017.

[0129] The above has given an exemplary description of the present invention. It should be noted that without departing from the core of the present invention, any simple deformation, modification, or equivalent replacement that can be made by those skilled in the art without creative labor falls within the protection scope of the present invention.

Claims

1. A medical image segmentation model based on pseudo-mask-guided feature aggregation, characterized in that, Including: Convolution block, second-order network model structure, pseudo-mask guided feature enhancement module, multi-scale multi-stage feature aggregation module, first convolutional layer, second convolutional layer, third convolutional layer, and loss function L seg , the second-order network model structure includes: a first-order sub-network and a second-order sub-network; The convolutional block is used to input initial data thereto and respectively flow the rough features output from the convolutional block to the first-order sub-network and the pseudo-mask guided feature enhancement module; The second-order sub-network and the first-order sub-network have the same architecture, each including: I + 1 residual blocks and a dilated spatial pyramid pooling module. The I + 1 residual blocks of the first-order sub-network are used to finely adjust the rough features and then send the first-order refined features to the dilated spatial pyramid pooling module of the first-order sub-network; the dilated spatial pyramid pooling module of the first-order sub-network is used to extract high-order latent features from the first-order refined features; The first convolutional layer is used to generate a pseudo-mask for the high-order latent features obtained by the first-order sub-network, and the pseudo-masks are respectively transmitted to the pseudo-mask guided feature enhancement module and the loss function L seg ; The pseudo-mask guided feature enhancement module is used to enhance the expression ability of the rough features by using the pseudo-mask to obtain pseudo-mask guided fused features; The I + 1 residual blocks of the second-order sub-network are used to input the fused features and output second-order refined features, and the dilated spatial pyramid pooling module of the second-order sub-network is used to receive the second-order refined features output from the (I + 1)-th residual block of the second-order sub-network and output high-order latent features; The multi-scale multi-stage feature aggregation module includes: a multi-scale feature aggregation module and a multi-stage feature aggregation module. The multi-scale feature aggregation module is used to perform multi-scale feature aggregation on the low-level features output from the i-th residual block of the first-order sub-network and the low-level features output from the i-th residual block of the second-order sub-network to obtain multi-scale aggregated features, where i = 1, ……, I; The second convolutional layer is used to fuse the multi-scale aggregated features to output high-order features; The multi-stage feature aggregation module is used to perform multi-stage feature aggregation on the feature outputs of the (I + 1)-th residual block of the first-order sub-network, the feature outputs of the (I + 1)-th residual block of the second-order sub-network, and the high-order features, and then output multi-scale multi-stage aggregated features; The third convolutional layer is used to perform feature concatenation and then fusion on the multi-scale multi-stage aggregated features and the high-order latent features obtained from the second-order sub-network to obtain a prediction result; The loss function L seg is used to calculate based on the pseudo-mask and the prediction result to obtain the loss function value.

2. The medical image segmentation model according to claim 1, characterized in that, The first convolutional layer includes: an upsampling layer and a convolutional layer.

3. The medical image segmentation model according to claim 2, characterized in that, The calculation process of the first convolutional layer is as follows: Y s = Conv(Up(X c )) Among them, X c is the high-order latent feature obtained from the first-order sub-network, Up is the upsampling layer in the first convolutional layer, Conv is the convolutional layer, and Y s is the pseudo-mask.

4. The medical image segmentation model according to claim 1, characterized in that, The calculation formula of the multi-scale feature aggregation module is: Among them, X m is the multi-scale aggregation feature, is the i-th residual block of the s-th order subnet, is the low-level feature output by the (i-1)-th residual block of the s-th order subnet, s = 1, 2, i = 1, ……, I, where is the rough feature output by the convolutional block, is the pseudo-mask-guided fusion feature, Up is the upsampling layer, δ is the parametric rectified linear unit, is the batch normalization process, Conv is the convolutional layer.

5. The medical image segmentation model according to claim 1, characterized in that, The operation process of the second convolutional layer is as follows: X′ m = Conv(X m ) Among them, X' m is a high-order feature, Conv is a convolutional layer, and X m is a multi-scale aggregated feature.

6. The medical image segmentation model according to claim 1, characterized in that, The calculation formula of the multi-stage feature aggregation module is as follows: Among them, X′ m is a high-order feature, X h is a multi-scale multi-stage aggregation feature, Up is the upsampling layer, δ is the parametric rectified linear unit, is batch normalization processing, Conv is the convolutional layer, is the (I + 1)-th residual block of the s-th order sub-network, s = 1, 2, is the feature output of the I-th residual block of the s-th order sub-network.

7. The medical image segmentation model according to claim 1, characterized in that, The third convolutional layer includes: upsampling, feature concatenation, and a convolutional layer.

8. The medical image segmentation model according to claim 7, characterized in that, The calculation formula of the third convolutional layer is as follows: Y s = Conv(concat(X h , Up(X f ))) where Y s is the prediction result, Conv is the convolutional layer, concat is the feature concatenation operation, X h is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, and X f is the high-order latent feature obtained by the second-order sub-network.

9. The medical image segmentation model according to claim 1, characterized in that, Loss function Among them, λ Dice is the weight value for adjusting L Dice and λ vcc is the weight value for adjusting L vcc as well.

10. The medical image segmentation model according to claim 9, characterized in that, Among them, B is the smallest batch in training, D is the number of instances in B, and B d is all the pixels belonging to instance d in the smallest batch, |B d | is the number of pixels in B d , and μ d is the average value of the probabilities of the correctly predicted classes of all pixels in B d . p h is the probability of the correct class of pixel h in the prediction result obtained by the third convolutional layer, where h = 1, ……, |B d |.

Citation Information

Patent Citations

  • Pancreatic cell image segmentation method based on improved U-Net network

    CN109191471A

  • Histology image segmentation model based on semi-supervised learning

    CN115131565A