MR image segmentation algorithm based on selective feature interactive fusion network
Through the MR image segmentation algorithm based on the selective feature interaction fusion network, combined with the dual encoder, the selective feature interaction module and the multi-scale guided feature reconstruction module, the problem of difficulty in segmenting the pancreatic region in the MR image is solved, and higher segmentation accuracy and model performance are achieved.
Patent Information
- Application Number
- CN202111582063.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The prior art is difficult to accurately and automatically segment the pancreatic region in the MR image, mainly because the pancreas accounts for a small proportion in the image and is similar to the image intensity of the surrounding tissue, so the boundaries are difficult to distinguish.
The MR image segmentation algorithm based on the selective feature interaction fusion network is adopted, and the WP and OOP sequence features in the multi-sequence MR images are extracted through a dual encoder, and combined with the selective feature interaction module and the multi-scale guided feature reconstruction module, feature fusion and reconstruction are performed, which enhances the extraction ability of the segmentation network, and improves the performance of the segmentation model through semi-supervised strategies and uncertainty estimation and morphological algorithm correction.
The accurate and automatic segmentation of the pancreatic region in the MR image is achieved, the segmentation ability and accuracy of the segmentation model are improved, and the problem of pancreatic segmentation difficulties in the prior art is solved.
Smart Images

Figure CN114299083B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an MR image segmentation algorithm based on a selective feature interactive fusion network, and belongs to the technical field of image processing. Background Art
[0002] Pancreatic cancer is one of the most common cancers in the world. It is a devastating malignant disease with a median survival of 3-6 months and a 5-year survival rate of less than 5%. Currently, the most effective treatment for pancreatic cancer is surgical resection. With the development of medical imaging technology, computed tomography (CT) and magnetic resonance (MR) images have become important imaging tools for doctors to diagnose pancreatic cancer. MR images with good soft tissue contrast are more suitable for observing the pancreas and tumors than CT images. Therefore, accurately and automatically segmenting the pancreas from MR images can assist doctors in diagnosing pancreatic diseases and increase patients' survival rates.
[0003] However, since the pancreas occupies a small proportion in the MR image, the image intensity of the pancreas and the surrounding tissues are similar and the boundary is difficult to distinguish, and the pancreatic structure of different patients is different, it is difficult to segment the pancreatic area in the MR image in the existing scheme. Summary of the invention
[0004] The purpose of the present invention is to provide an MR image segmentation algorithm based on a selective feature interactive fusion network to solve the problems existing in the prior art.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] In a first aspect, an embodiment of the present invention provides an MR image segmentation algorithm based on a selective feature interactive fusion network, the algorithm comprising:
[0007] Inputting the multi-sequence MR labeled images into a dual-branch segmentation network for training to obtain a segmentation model that can automatically segment the multi-sequence MR images;
[0008] The dual-branch segmentation network includes a dual encoder, a decoder, a selective feature interaction module and a multi-scale guided feature reconstruction module. The dual encoder is used to extract WP (water phase) sequence features and OOP (out of phase) sequence features in the multi-sequence MR images respectively. The decoder is used to fuse and decode the image features extracted by the dual encoder and perform segmentation.
[0009] A semi-supervised strategy is adopted to input unlabeled multi-sequence MR images into the segmentation model to obtain pseudo-labels predicted by the network. The pseudo-labels are then estimated for uncertainty and corrected using morphological algorithms. After correction, they are sent into the dual-branch segmentation network together with the images as pseudo-labels for fine-tuning training to obtain a segmentation model with stronger segmentation capabilities.
[0010] Optionally, a selective feature interaction module is provided between the dual encoders, and the selective feature interaction module is used to selectively exchange image information in the WP sequence and the OOP sequence.
[0011] Optionally, the selective feature interaction module includes:
[0012] For the features extracted from each layer in each branch encoder, the extracted features are corrected so that the strength of the corrected features is consistent with the strength of the corrected features from another branch encoder;
[0013] Calculate the weight corresponding to each feature based on the corrected features;
[0014] According to the calculated weights, the extracted features are interactively fused with the features extracted by another branch encoder, and the fused features are returned to the corresponding branch encoder.
[0015] Optionally, calculating the weight corresponding to each feature according to the corrected feature includes:
[0016] Perform pyramid pooling operation on the corrected features to obtain each feature after pooling;
[0017] Flatten each feature after pooling and concatenate them to get the feature vector;
[0018] Perform matrix multiplication on the feature vector corresponding to the WP sequence image and the feature vector corresponding to the OOP sequence image to obtain a similarity relationship matrix;
[0019] According to the similarity relationship matrix, the similarity between the WP sequence features of each channel and the overall OOP features, and the similarity between the OOP sequence features of each channel and the overall WP features are obtained;
[0020] The weight corresponding to each feature is determined based on the obtained similarities.
[0021] Optionally, a feature reconstruction module is provided between the skip connection and the decoding layer in the decoder, and the feature reconstruction module is used to reduce the redundancy of low-level semantic features.
[0022] Optional, including:
[0023] Normalize the high-level semantic features obtained from the previous layer and the low-level semantic features obtained from the skip connection;
[0024] The normalized high-level semantic features and the normalized low-level semantic features are fused to obtain fused semantic features;
[0025] Extract the fused semantic features through different spatial extraction methods to obtain a variety of spatial attention features;
[0026] Synthesizing the spatial attention features of the pancreatic region according to the multiple spatial attention features obtained;
[0027] Acquire a semantic feature after noise suppression according to the spatial attention feature and the low-level semantic feature, and pass the semantic feature after noise suppression through a channel attention module;
[0028] The channel attention module generates a channel weight according to the pancreatic region intensity of the semantic feature after noise suppression;
[0029] Weighting the low-level semantic features according to the channel weights to obtain channel attention features;
[0030] According to the spatial attention features and the channel attention features, the reconstructed low-level semantic features are determined.
[0031] Optionally, the uncertainty estimation is used to judge and correct the credibility of pseudo labels generated using unlabeled data, and the morphological algorithm is used to correct the pseudo labels to make them more credible.
[0032] Optional, including:
[0033] According to the characteristics of MR images, the window width and window position are changed to obtain multiple data with different contrasts of the same patient; the data are sent to a double-branch network to obtain different pseudo labels;
[0034] Calculating uncertainty estimates of original pseudo labels according to the differences between the different pseudo labels and sorting the original pseudo labels;
[0035] For the data with high uncertainty estimation, information entropy is calculated as the weight, and the pseudo labels under the various comparisons are weighted to obtain labels with higher credibility.
[0036] Optional, including:
[0037] Obtaining a mask based on context information of the MR image;
[0038] removing the portion with low mask value, and performing an erosion and dilation operation on the processed mask;
[0039] Correcting each pseudo label according to the mask after corrosion and expansion;
[0040] A final credible pseudo label is determined according to the adjacent pseudo labels of the corrected pseudo label.
[0041] The network model capable of automatically segmenting the multi-sequence MR images is obtained by inputting the labeled multi-sequence MR images into the segmentation network for training; wherein the segmentation network comprises a dual encoder, a decoder, a selective feature interaction module and a multi-scale guided feature reconstruction module, the dual encoder is used to extract the image features of the WP (waterphase) sequence images and the image features of the OOP (out of phase) sequence images in the multi-sequence MR images, the selective feature interaction module is used to selectively exchange information useful to each other, and enhance the ability of the dual branches to extract pancreatic information, the multi-scale guided feature reconstruction module is used to reconstruct low-level semantic features that are conducive to the segmentation of small target pancreatic regions, and the decoder is used to fuse and decode the image feature information extracted by the dual encoder, and segment the pancreatic region; the unlabeled multi-sequence MR images are input into the trained network model using semi-supervised technology to obtain the segmented pancreatic region, the more reliable pancreatic region obtained after correction using the uncertainty estimation and morphological algorithm is then sent to the segmentation network for training, and the segmentation performance of the network is further improved. The problem of segmenting the diseased pancreas in the prior art is solved, and the image features of the WP sequence images and the OOP sequence images are integrated, thereby achieving feature complementarity and accurate segmentation.
[0042] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A flowchart of an image segmentation algorithm provided by one embodiment of the present invention;
[0044] Figure 2 A schematic diagram of a multi-sequence MR image provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of the network architecture of a dual-branch pancreatic segmentation network provided by an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of the structure of a selective feature interaction module provided by an embodiment of the present invention;
[0047] Figure 5 The low-level semantic features before and after reconstruction by the multi-scale guided feature reconstruction module provided by one embodiment of the present invention;
[0048] Figure 6 A schematic diagram of the structure of a multi-scale guided feature reconstruction module provided by an embodiment of the present invention;
[0049] Figure 7 A schematic diagram of a segmented network update provided by an embodiment of the present invention;
[0050] Figure 8 This is a diagram of the segmentation results of the pancreatic region obtained by using different segmentation algorithms provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0051] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0053] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0054] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0055] Please refer to Figure 1 , which shows a flow chart of an MR image segmentation algorithm based on a selective feature interactive fusion network provided by an embodiment of the present application, such as Figure 1 As shown, the algorithm includes:
[0056] 1. Input the multi-sequence MR labeled images into a dual-branch segmentation network for training to obtain a segmentation model that can automatically segment the images;
[0057] Please refer to Figure 2 , which shows a possible multi-sequence MR image acquired by the present application. Figure 2 In (a)-(d) four sequences of T1 phase: WP (Water Phase), FP (Fat Phase), IP (In Phase), OOP (Out of Phase). (e) The arrow indicates a severe lesion that has changed the shape of the pancreas. (f) The contrast between the pancreas and the surrounding tissue is low.
[0058] Among them, the dual-branch segmentation network includes a dual encoder, a decoder, a selective feature interaction module and a multi-scale guided feature reconstruction module. The dual encoder is used to extract the WP sequence features and OOP sequence features in the multi-sequence MR images respectively, and the decoder is used to fuse and decode the image features extracted by the dual encoder and segment the pancreatic area.
[0059] 2. Input the unlabeled multi-sequence MR images into the segmentation model to obtain the pseudo labels predicted by the network;
[0060] 3. The uncertainty of the pseudo-labels is estimated and the morphological algorithm is corrected. After correction, the pseudo-labels are sent together with the images into the dual-branch segmentation network for fine-tuning training to obtain a pancreatic segmentation model with stronger segmentation ability.
[0061] In actual implementation, the segmentation network needs to be trained before this step. For details, please refer to Figure 3 , which shows the principle diagram of training the segmentation network in this application, such as Figure 3 As shown in the figure, a dual-branch network SFI-Net is constructed based on the ResUNet architecture, and a ResNet branch is added to the encoder; ResNet uses the ResNet-34 version, downsampled 5 times in total, and the number of convolution blocks used in each layer is [3,4,6,3]. The bottom layer splices the outputs of the two layers together. This application uses a network architecture constructed by training sample multi-sequence MR images, and the trained network is the segmentation network. Figure 3 As shown, the present application uses the WP sequence images and the OOP sequence images in the multi-sequence MR images for training, and there is no limitation on this.
[0062] In one possible embodiment, the loss function L of the segmentation network is total It is the sum of the binary cross entropy loss function (BCE) and the DICE loss function:
[0063] Ltotal =L BCE +L DICE
[0064] Where L BCE With L DICE The expression is (s=108):
[0065]
[0066] Where y is the gold standard (manually labeled), p is the segmentation result, j is the jth pixel value of the image, H is the height of the image, and W is the width of the image.
[0067] After the segmentation network is trained, the segmented pancreatic region can be obtained by inputting the multi-sequence MR images into the trained segmentation network.
[0068] In summary, by inputting labeled multi-sequence MR images into the segmentation network for training, a network model that can automatically segment the image is obtained; wherein, the segmentation network includes a dual encoder, a decoder, a selective feature interaction module and a multi-scale guided feature reconstruction module, the dual encoder is used to extract the image features of the WP sequence image and the image features of the OOP sequence image in the multi-sequence MR image, the selective feature interaction module is used to selectively exchange information useful to each other, and enhance the ability of the dual branch to extract pancreatic information, the multi-scale guided feature reconstruction module is used to reconstruct low-level semantic features that are conducive to the segmentation of small target pancreatic regions, and the decoder is used to fuse and decode the image feature information extracted by the dual encoder, and segment the pancreatic region; the unlabeled multi-sequence MR images are input into the trained network model using semi-supervised technology to obtain the segmented pancreatic region, and the more reliable pancreatic region obtained after correction using the uncertainty estimation and morphological algorithm is sent to the segmentation network for training, so as to further improve the segmentation performance of the network. The segmentation problem of the diseased pancreas in the prior art is solved, and the image features of the WP sequence image and the OOP sequence image are integrated, thereby achieving feature complementarity and accurate segmentation effect.
[0069] In a possible implementation of this embodiment, a selective feature interaction module is provided between the first branch encoder and the second branch encoder, and the selective feature interaction module is used to exchange image features in the WP sequence image and the OOP sequence image. Specifically, the selective feature interaction module includes:
[0070] First, for the features extracted from each layer in each branch encoder, the extracted features are corrected so that the strength of the corrected features is consistent with the strength of the features corrected by another branch encoder;
[0071] Get the features f of the two sequences WP and OOP passed by the previous layer of network in,W ,f in,O Since the feature intensities of the two sequences are inconsistent, the selective feature interaction module modifies the intensity range of the original feature by convolution to obtain a new feature f W ,f O .
[0072] f W =Conv 1×1 (BNR(Conv 3×3 (f in,W )))
[0073] Among them, Conv 1×1 and Conv 3×3 They are 1×1 and 3×3 convolution operations respectively. and C, W and H represent the number of channels, width and height of the feature respectively. BNR is batch normalization and ReLU. OOP sequence feature f O The calculation of f W resemblance.
[0074] Second, calculate the weight corresponding to each feature based on the corrected features;
[0075] (1) Perform pyramid pooling on the corrected features to obtain each feature after pooling;
[0076] The corrected features Use 4-layer pyramid pooling operation. After pooling, each feature {f Wa ,f Wb ,f Wc ,f Wd}, where a, b, c, and d are the sizes of the pooled images. OOP sequence feature f O The pooling method is similar to f W Similar, no further elaboration.
[0077] (2) Flatten each pooled feature and concatenate them to obtain a feature vector;
[0078] The pooled features are flattened to obtain representation vectors {f″ at different scales Wa ,f″ Wb ,f″ Wc ,f″ Wd}, All vectors are connected to get the representation vector representing each channel Similarly, the OOP sequence feature can obtain the feature vector f′ O .
[0079] (3) performing matrix multiplication on the feature vector corresponding to the WP sequence image and the feature vector corresponding to the OOP sequence image to obtain a similarity relationship matrix;
[0080] Similarity relationship matrix X:
[0081] X=f′ W ×(f′ O ) T
[0082] Where T represents the matrix transpose, the relation matrix The element X in ij Represents the similarity between the feature map of the i-th channel in the WP sequence and the feature map of the j-th channel in the OOP sequence. The larger the element value, the more similar the two features are.
[0083] (4) obtaining the similarity of the WP sequence features of each channel to the overall OOP features, and the similarity of the OOP sequence features of each channel to the overall WP features according to the similarity relationship matrix;
[0084] The relationship matrix is averaged in two dimensions to obtain two similarity vectors S W ={S Wi |i∈[1,C]},S O ={S Oj |j∈[1,C]}, where:
[0085]
[0086] Where S Wi is the similarity of the WP sequence feature of the i-th channel to the overall OOP feature; S Oj is the similarity of the OOP sequence features of the jth channel to the overall WP features.
[0087] (5) Determine the weight corresponding to each feature based on the obtained similarities.
[0088] Two vectors S W ,S O After the convolution layer, batch normalization and Sigmoid activation layer, the feature weights can be obtained.
[0089] Third, the extracted features are interactively fused with the features extracted by another branch encoder according to the calculated weights, and the fused features are returned to the corresponding branch encoder.
[0090] The feature weights calculated using the above steps are used to calculate the original information f in,W ,f in,O Weighting can extract complementary information features f W,sup ,fO,sup The extracted features are interactively fused with the features of the other sequence and finally sent back to the encoding network.
[0091]
[0092]
[0093] Where F() is the cascade of 1×1 convolutional layer, Relu layer, 1×1 convolutional layer and sigmoid layer. * is element-wise multiplication. represents a 3×3 convolution operation that changes the channel, cat represents the splicing operation by channel. Final output features Characteristics of OOP sequences f o The construction of is similar.
[0094] Please refer to Figure 4 , which shows a schematic diagram of the principle of the selective feature interaction module.
[0095] By setting up a selective feature interaction module to exchange and fuse the features of two complementary sequences, the encoder is helped to extract richer pancreatic information and further improve the accuracy of the extracted pancreatic region.
[0096] In another possible implementation of this embodiment, a feature reconstruction module is provided between the jump connection and the decoding layer in the decoder, and the feature reconstruction module is used to reduce the redundancy of low-level semantic features. Specifically, the feature reconstruction module is used to:
[0097] First, the high-level semantic features obtained from the previous layer and the low-level semantic features obtained from the skip connection are normalized;
[0098] High-level semantic features obtained from the previous layer and low-level semantic features obtained by skip connections All are normalized by 1×1 convolution.
[0099] Second, the normalized high-level semantic features and the normalized low-level semantic features are fused to obtain fused semantic features;
[0100] After normalization, f H Upsample to f L The scale of f L Perform additive fusion to obtain fusion features
[0101] f Fus =Up(Conv1×1 (f H ))+Conv 1×1 (f L )
[0102] Where Up is the upsampling operation.
[0103] Third, extract the fused semantic features through different extraction methods to obtain multiple attention maps;
[0104] For example, three attention maps are obtained by extracting using three different extraction methods.
[0105] Fourth, obtaining the attention value of the pancreatic region according to the obtained multiple attention maps;
[0106] The 3 channels are fused into a 1-channel image through a 7×7 convolutional layer. The image is normalized to [0,1] using the Sigmoid function to obtain the final attention map Att; The value of will be close to 1 in the pancreas and close to 0 elsewhere. Att will effectively weaken f L The noise information present in .
[0107] Fifth, obtaining a semantic feature after noise suppression according to the attention value and the low-level semantic feature, and the semantic feature after noise suppression passes through a channel attention module;
[0108] The semantic feature after noise suppression can be the product of the attention value and the ground semantic feature. That is: f Sam =f L ×Att.
[0109] Sixth, the channel attention module generates a weight according to the attention value and the pancreatic region strength;
[0110] Because the effect of the channel attention module is related to the intensity of the feature pixels, the noise intensity of the pancreatic features containing redundant information may be greater than the intensity of the pancreatic region itself, making the channel attention module invalid. Att allows the channel attention module to generate channel weights based on the intensity of the pancreatic region. This removes redundant information in the channel dimension.
[0111] Seventh, obtaining the features of the semantic features after noise suppression after passing through the channel attention module according to the weights;
[0112] The feature hypothesis after the channel attention module is f Cam =f Sam ×α.
[0113] Eighth, low-level semantic features are determined based on the semantic features after noise suppression and the features after passing through the channel attention module.
[0114] Reconstructed low-level semantic features The feature f Sam With f Cam The sum of: f Rec =f Sam +
[0115] f Cam .
[0116] Att is generated by the MGFR module (feature reconstruction module) at 4 scales. The features before and after the MGFR module reconstructs the low-level semantic features are as follows: Figure 5 As shown. The image is 32 2 The pancreas is relatively small in size, so the module focuses more on the reconstruction of background information, while the 64 2 and 128 2 The module can reconstruct more pancreatic region information. 2 At this size, the module can reconstruct the edge information features of the pancreas from low-level semantics.
[0117] Please refer to Figure 6 , which shows a schematic diagram of the principle of the feature reconstruction module.
[0118] By setting up a feature reconstruction module to eliminate the redundancy of low-level semantic features and narrow the gap between low-level semantic features and high-level semantic features, the segmentation accuracy of the pancreatic region is further improved.
[0119] In another possible embodiment of the present application, the semi-supervised algorithm of the above embodiment further includes a correction of the pseudo-label uncertainty estimation and the morphological algorithm, specifically including:
[0120] First, for the dataset D = L ∪ U, use the labeled dataset L{x i ,y i}Train SIF-Net to get a pre-trained model;
[0121] Second, in the semi-supervised stage, this initially trained model is used to train the unlabeled data U{x i} perform prediction segmentation and obtain a set of i The pseudo label y' i According to the characteristics of MR images, the window width and window position are adjusted to change the contrast of the original data, making the tissue structure that was originally difficult to distinguish with the naked eye clear. The transformed image is sent to the pre-trained model for prediction to obtain different pseudo labels.
[0122] Third, the predicted pancreatic region is corrected;
[0123] (1) Preliminarily correcting the predicted pancreatic region using an uncertainty estimation algorithm to obtain a first correction result;
[0124] The differences of these pseudo labels are used to calculate the uncertainty estimate of the original pseudo labels, and the pseudo labels are sorted; the calculation formula of uncertainty E is as follows:
[0125]
[0126] Where C is the number of images with different contrasts, y' k is the pseudo label predicted by the network under the kth contrast, y' j Represents the jth pixel of the image. The closer E is to 0, the more realistic the pseudo-label is. On the contrary, pseudo-labels with high E need some correction.
[0127] For data with high uncertainty, the information entropy is calculated as the weight to be added, and the pseudo labels generated under various comparisons are used to complement each other to obtain labels with higher credibility. The specific correction process is defined as follows:
[0128]
[0129] Where j represents the jth pixel in the prediction image. When the network is not sure whether the target area is the pancreas area or the background area, its p value is close to 0.5, e k The value tends to a peak value; on the contrary, when it is determined, the p value tends to 0 or 1, e k Approaches 0. For all contrast images, e={e k |k∈[1,C]} becomes the weight value w={w k |k∈[1,C]} performs weighted summation on the pseudo labels to obtain the final pseudo label y′.
[0130] (2) The first correction result is corrected according to the morphological algorithm and the context information of the MR image to obtain a final corrected pancreatic region.
[0131] Using the context information of the MR image, a mask is obtained by summing up all the pseudo-labels of the same patient, and the part with low mask value (over-segmented area or the outer contour of the pancreas) is removed first; secondly, the mask is corroded to eliminate the noise area around the pancreas; finally, the mask is expanded and the expanded mask is applied to each pseudo-label to remove part of the over-segmented area. The expansion rate needs to be set greater than the corrosion rate to restore part of the outer contour area removed by the mask.
[0132]
[0133] y″i =Dil(Ero(mask))×y′ i
[0134] Where t is a preset threshold used to remove over-segmented areas. Ero is an erosion operation and Dil is a dilation operation.
[0135] For the pseudo label y″ i For the under-segmented areas in , we use an algorithm to supplement the context information. The specific steps are as follows:
[0136] y″ i =y″ i +(y″ i-2 +y″ i-1 )∩(y″ i+1 +y″ i+2 )
[0137] where y″ i-2 ,y″ i-1 ,y″ i+1 ,y″ i+2 is with y″ i Adjacent slices. This method is used to supplement the unpredicted areas and obtain more reliable pseudo labels y″ i .
[0138] Fourth, the segmentation network is updated according to the unlabeled dataset and the corrected pancreatic region.
[0139] Please refer to Figure 7 , which shows the revised process of generating pseudo labels for unlabeled data.
[0140] By updating the segmentation network, a segmentation network with higher segmentation accuracy can be obtained, thereby improving the segmentation accuracy of the pancreatic region.
[0141] In one possible embodiment, the experiment selected 395 3D images from abdominal MR of 395 pancreatic cancer patients (351 ductal adenocarcinomas, 22 neuroendocrine carcinomas, 20 adenosquamous carcinomas, 1 intraductal papillary mucinous tumor, and 1 mucinous cystic tumor with invasive carcinoma) from Changhai Hospital of Naval Medical University. The size of the 3D MR image is W×H×D, where W, H, and D correspond to the number of slices in the sagittal, coronal, and transverse planes, respectively. In order to accurately segment the abnormal pancreas, slices are used in the transverse view, and the size of each slice is W×H. There are two different image sizes, 512×512 and 320×260. The 320×260 image is upsampled to 512×512 by bilinear interpolation and zero padding in the H dimension. The labeled data is divided into 4 subsets L1{x 1i ,y 1i}, L2{x2i ,y 2i}, L3{x 3i ,y 3i}, L4{x 4i ,y 4i The number of cases is 26, 25, 25, and 25, and the number of slice pairs is 732, 702, 687, and 707, respectively. The unlabeled data has a total of 294 patients and 5300 slice pairs. The unlabeled data is only used during network fine-tuning.
[0142] The experimental results of the present invention are as follows:
[0143] In order to quantitatively evaluate the performance of the algorithm proposed in this paper, the segmentation results are compared with the gold standard according to the following four indicators: Jac coefficient (Jaccard similarity coeffificient), DSC coefficient (Dice similarity coefficient), precision (Precision), recall (Recall). DSC calculates the overlap between the segmentation result and the gold standard and is defined as:
[0144]
[0145] Where TP is the number of true positive segmented pixels, FP is the number of false positive segmented pixels, and FN is the number of false negative segmented pixels. The Jac coefficient, precision, and recall rate are calculated as follows:
[0146]
[0147] The ablation experiment of the present invention is shown in Table 1. The comparative experiment of the present invention and others' algorithms is shown in Table 2.
[0148] As shown in Table 1, this is the ablation experiment of the algorithm adopted by the present invention.
[0149] Methods Jac(%) DSC(%) Prec(%) Rec(%) Baseline(W) 67.47±13.11 79.71±10.56 80.59±14.29 80.59±9.12 Baseline(O) 70.54±11.82 82.07±9.10 82.70±12.60 82.99±7.74 Baseline 71.49±12.29 82.66±9.56 83.09±13.38 83.80±7.47 Baseline+SFI 72.51±11.21 83.48±8.39 83.63±11.31 84.72±7.66 Baseline+MGFR 72.85±9.84 83.88±7.11 84.14±9.73 84.14±9.73 SIF-Net 73.95±9.78 84.61±7.11 86.51±9.14 83.63±7.86 Baseline+Semi 73.86±10.80 84.45±8.10 86.08±11.54 85.36±6.69 SIF-Net+Nocorrect 74.62±9.83 85.05±7.13 86.32±8.94 84.54±7.71 SIF-Net+Semi 76.21±8.59 86.19±6.09 88.18±7.35 84.73±7.48
[0150] Table 1 is shown in Table 2, which is a comparison of the segmentation quantitative results of the algorithm of the present invention and the existing algorithm.
[0151] Methods Jac(%) p-value DSC (%) Prec(%) Rec(%) FuseNet 65.44±11.27 <0.001 78.45±9.20 78.04±13.14 80.90±7.58 nnUnet(O) 66.32±16.67 <0.001 78.31±13.75 77.79±15.59 80.43±13.72 CLFF 66.97±13.91 <0.001 79.23±11.44 79.13±15.66 81.63±8.81 nnUnet(O+W) 67.76±15.11 <0.001 79.68±12.16 79.11±14.53 81.90±12.02 HPN 71.55±11.77 <0.05 82.62±8.77 82.49±12.50 84.53±7.71 SIF-Net 73.95±9.78 <0.05 84.61±7.11 86.51±9.14 83.63±7.86 CoraNet 67.56±12.93 <0.001 79.84±10.23 79.57±13.46 81.63±9.41 70.72±11.83 <0.001 82.22±8.99 82.51±12.36 83.24±8.49 UDA 71.22±12.12 <0.01 82.52±9.20 81.89±12.90 84.67±7.60 72.01±11.80 <0.01 83.08±9.21 80.71±12.65 87.24±6.71 SIF-Net+Semi 76.21±8.59 86.19±6.09 88.45±7.35 84.73±7.48
[0152] Table 2
[0153] CoraNet and UDA have two training processes, the first row is the pre-training process with labeled data, and the second row is the result of the final semi-supervised algorithm.
[0154] like Figure 8As shown in the figure, the segmentation results of different algorithms are as follows: (a) original image (OOPWP), (b) ground truth, (ck) segmentation results of nnUnet(O), FuseNet, CLFF, nnUnet(W+O), HPN, SIF-Net, CoraNet, UDA and SIFNet+semi.
[0155] Combination Figure 8 It can be seen that the pancreas segmentation algorithm provided in the present application can accurately segment the pancreas region in the MR image.
[0156] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0157] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A MR image segmentation algorithm based on a selective feature interactive fusion network, characterized in that: The algorithm includes: Inputting multiple sequence MR labeled images into a dual-branch segmentation network for training to obtain a segmentation model that can automatically segment the images; The dual-branch segmentation network includes a dual encoder, a decoder, a selective feature interaction module and a multi-scale guided feature reconstruction module. The dual encoders are used to extract WP sequence features and OOP sequence features from the multi-sequence MR images respectively. The decoder is used to fuse and decode the image features extracted by the dual encoders and segment the pancreatic region. A semi-supervised strategy is adopted to input unlabeled multi-sequence MR images into the segmentation model to obtain pseudo labels predicted by the network. The pseudo labels are then estimated for uncertainty and corrected using morphological algorithms. After correction, they are sent together with the images as pseudo labels into the dual-branch segmentation network for fine-tuning training, thus obtaining a segmentation model with stronger segmentation capabilities. A selective feature interaction module is provided between the dual encoders, and the selective feature interaction module is used to selectively exchange image information in the WP sequence and the OOP sequence; The selective feature interaction module comprises: For the features extracted from each layer in each branch encoder, the extracted features are corrected so that the strength of the corrected features is consistent with the strength of the corrected features from another branch encoder; Calculate the weight corresponding to each feature based on the corrected features; According to the calculated weights, the extracted features are interactively fused with the features extracted by another branch encoder, and the fused features are returned to the corresponding branch encoder; In the decoder, a multi-scale guided feature reconstruction module is arranged between the skip connection and the decoding layer, and the reconstruction module is used to reduce the redundancy of low-level semantic features; Normalize the high-level semantic features obtained from the previous layer and the low-level semantic features obtained from the skip connection; The normalized high-level semantic features and the normalized low-level semantic features are fused to obtain fused semantic features; Extract the fused semantic features through different spatial extraction methods to obtain a variety of spatial attention features; Synthesizing the spatial attention features of the pancreatic region according to the multiple spatial attention features obtained; Acquire a semantic feature after noise suppression according to the spatial attention feature and the low-level semantic feature, and pass the semantic feature after noise suppression through a channel attention module; The channel attention module generates a channel weight according to the pancreatic region intensity of the semantic feature after noise suppression; Weighting the low-level semantic features according to the channel weights to obtain channel attention features; According to the spatial attention features and the channel attention features, the reconstructed low-level semantic features are determined.
2. The algorithm according to claim 1, characterized in that The step of calculating the weight corresponding to each feature according to the corrected features includes: Perform pyramid pooling operation on the corrected features to obtain each feature after pooling; Flatten each feature after pooling and concatenate them to get the feature vector; Perform matrix multiplication on the feature vector corresponding to the WP sequence image and the feature vector corresponding to the OOP sequence image to obtain a similarity relationship matrix; According to the similarity relationship matrix, the similarity between the WP sequence features of each channel and the overall OOP features, and the similarity between the OOP sequence features of each channel and the overall WP features are obtained; The weight corresponding to each feature is determined based on the obtained similarities.
3. The algorithm according to claim 1, characterized in that The uncertainty estimation is used to judge and correct the credibility of the pseudo-labels generated using the unlabeled data, and the morphological algorithm is used to correct the pseudo-labels to make them more credible.
4. The algorithm according to claim 3, characterized in that include: According to the characteristics of MR images, the window width and window position are changed to obtain multiple data with different contrasts for the same patient; The data is fed into a dual-branch network to obtain different pseudo labels; Calculating uncertainty estimates of original pseudo labels according to the differences between the different pseudo labels and sorting the original pseudo labels; For the data with high uncertainty estimation, information entropy is calculated as the weight, and the pseudo labels under the various comparisons are weighted to obtain labels with higher credibility.
5. The algorithm according to claim 3, characterized in that include: Obtaining a mask based on context information of the MR image; removing the portion with low mask value, and performing an erosion and dilation operation on the processed mask; Correcting each pseudo label according to the mask after corrosion and expansion; A final credible pseudo label is determined according to the adjacent pseudo labels of the corrected pseudo label.
Citation Information
Patent Citations
Medical image automatic segmentation method based on multi-path attention fusion
CN111681252A
Semi-supervised remote sensing image semantic segmentation method and device, and computer equipment
CN113298815A