Liver Tumor Segmentation Method Based on Multi-Temporal Fusion and Dual Attention Mechanism
Through the liver tumor segmentation method with multi-temporal feature fusion and dual attention mechanism, the problem of insufficient segmentation accuracy in the prior art is solved, and higher segmentation accuracy is achieved.
Patent Information
- Application Number
- CN202210881264.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-26
AI Technical Summary
In the prior art, the segmentation network alone uses the portal venous phase image training to ignore other phase image information, while the multi-phase image training method does not fully consider the binding of each phase feature at different depths of the network, resulting in insufficient liver tumor segmentation accuracy.
The multi-phase feature fusion mechanism and the dual attention mechanism are adopted to integrate the superficial features of the arterial and portal venous images, and adjust the attention in the early fusion of the encoded part, and build the MFDA-Net network to improve the accuracy of liver tumor segmentation.
It improves the accuracy of segmentation of liver tumors in multi-phase CT images, reduces attention to irrelevant information, and improves the ability of liver tumor segmentation.
Smart Images

Figure CN115272357B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of graphics processing technology, and in particular, to a liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism, which can be used for liver tumor segmentation of multi-temporal CT images. Background Art
[0002] As one of the most common tumor diseases, liver cancer is characterized by insidious onset, long latency, strong metastasis, and rapid disease progression. With the development of computer technology, medical imaging technology is an essential tool for doctors in the diagnosis and treatment process of patients. Among them, computed tomography (CT) technology is the most commonly used means in the diagnosis, treatment, and follow-up of liver cancer patients. Manually identifying liver tumors in CT images is not only time-consuming and laborious but also has the disadvantage that the subjectivity in the identification process affects the accuracy. Therefore, using a computer to implement an accurate and reliable automatic segmentation algorithm can effectively improve the work efficiency of doctors in clinical scenarios and reduce the workload of doctors.
[0003] In the existing methods for segmenting liver tumors using deep convolutional neural networks, there are mainly two types of methods: training a segmentation network solely using portal venous phase images and training a segmentation network using multi-temporal images. The first type of method usually adds a network feature fusion module and an attention mechanism for extracting multi-scale information of the image in a standard encoder-decoder structure segmentation network to enhance the network's ability to segment liver tumors. The second type of method mostly splices the three images obtained from multi-temporal CT into a three-channel image and directly inputs it into the segmentation network for training, or inputs two or three images into a dual-branch or triple-branch network for training respectively. Among them, each branch shares the network structure but does not share the network weights and performs feature fusion at a certain stage in the network decoding layer. Both of the above two types of methods have their own deficiencies. The first type of method only extracts tumor segmentation information from portal venous phase images and ignores the supplementary information contained in the other two images. In the second type of method, the feature fusion of each phase image obtained from multi-temporal CT is too simple and crude, without considering the combinability of the features of each temporal image at different depths of the network, resulting in the network learning a lot of redundant and irrelevant information and affecting the segmentation performance of liver tumors. Summary of the Invention
[0004] An embodiment of the present disclosure provides a liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism, which can segment liver tumors in multi-temporal CT images, solve the problem of poor accuracy of liver tumor segmentation in the prior art, and improve the segmentation accuracy of liver tumors in multi-temporal CT images. The technical solution is as follows:
[0005] According to the first aspect of the embodiments of the present disclosure, there is provided a liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism, the method comprising:
[0006] Step 1: Obtain L case data. Take the arterial phase image and the portal venous phase image of the same layer obtained from each case's multi-phase CT scans each time as a group of input image pairs, and obtain the liver tumor labels of the portal venous phase images. Randomly take P case data from the L case data as the training set, and the remaining Q case data as the test set;
[0007] Step 2: Construct a multi-phase feature fusion mechanism MFF to fuse the two features obtained by shallow convolution of the arterial phase and portal venous phase images to obtain a fused feature map;
[0008] Step 3: Construct a dual attention mechanism DAM including a position attention module PAM and a channel attention module CAM. Input the feature map obtained after convolution operation on the fused feature map into the dual attention mechanism DAM for processing to obtain a dual attention feature map;
[0009] Step 4: Embed the multi-phase feature fusion mechanism MFF and the dual attention mechanism DAM into the U-Net structure to form a liver tumor segmentation network MFDA-Net based on multi-phase fusion and dual attention mechanism;
[0010] Step 5: Train the MFDA-Net with the input image pairs composed of arterial phase images and portal venous phase images in the training set and the corresponding portal venous phase liver tumor labels to obtain the trained MFDA-Net;
[0011] Step 6: Input the image pairs composed of arterial phase images and portal venous phase images in the test set into the trained MFDA-Net, use the binarization method to process the output results of the network to obtain the segmentation results, and represent the segmentation results in the form of contours on the portal venous phase images.
[0012] In one embodiment, constructing a multi-phase feature fusion mechanism MFF to fuse the two features obtained by shallow convolution of the arterial phase and portal venous phase images to obtain a fused feature map includes:
[0013] The feature map A extracted from the arterial phase image and the feature map V extracted from the portal venous phase image are respectively passed through a convolutional layer with a kernel size of 1×1, a sliding step of 1, and a number of 1 / 4 of the input feature map channels to obtain the feature map A1 and the feature map V1;
[0014] After adding the feature maps A1 and V1, they are successively passed through an activation layer with a rectified linear unit ReLU function as the activation function, a convolutional layer with a kernel size of 1×1, a sliding step of 1, and a number of 1, and an activation layer with a Sigmoid activation function to obtain an attention weight map;
[0015] Multiply the feature map A by the attention weight map to obtain the adjusted attention feature map A2;
[0016] After performing a fusion operation on the feature map V and the feature map A2 in the channel direction of the feature map, it is successively passed through a convolutional layer with a convolutional kernel size of 3×3, a stride of 1, and the number of kernels equal to the number of channels of the feature map V, and an activation layer with a rectified linear unit ReLU function as the activation function to obtain the fused feature map V2.
[0017] In one embodiment, the feature map obtained after performing a convolution operation on the fused feature map is input into the dual attention mechanism DAM for processing, and the dual attention feature map obtained includes:
[0018] The feature map obtained after performing a convolution operation on the fused feature map is respectively input into the position attention module PAM and the channel attention module CAM for processing to obtain the position attention feature map and the channel attention feature map;
[0019] The position attention feature map and the channel attention feature map are added together and then successively passed through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1, and an activation layer with a rectified linear unit ReLU function as the activation function to obtain the dual attention feature map.
[0020] In one embodiment, the feature map obtained after performing a convolution operation on the fused feature map is input into the position attention module PAM, and the position attention feature map obtained includes:
[0021] The feature map F obtained after performing a convolution operation on the fused feature map is respectively passed through two convolutional layers with a convolutional kernel size of 1×1, the number of kernels equal to C / 4, and a stride of 1 to obtain the feature map FX and the feature map FY. The size of the feature map F is H×W×C, where H, W, and C are the height, width, and number of channels of the feature map F, respectively;
[0022] The feature map F is passed through a convolutional layer with a convolutional kernel size of 1×1, the number of kernels equal to C, and a stride of 1 to obtain the feature map FZ;
[0023] The feature map FY and the feature map FZ are respectively reshaped to obtain a feature map FY1 with a size of and a feature map FZ1 with a size of (H×W)×C;
[0024] The feature map FX is reshaped to obtain a feature map FX1 with a size of and then transposed to obtain a feature map FX2 with a size of ;
[0025] The feature map FY1 and the feature map FX2 are multiplied and then passed through the Softmax function to obtain an inter-channel correlation map X1 with a size of (H×W)×(H×W);
[0026] Multiply the inter-channel correlation map X1 with the feature map FZ1, and then obtain the feature map FZ2 with the size of H×W×C through Reshape;
[0027] Multiply the feature map FZ2 by the scale correlation coefficient α, and then add it to the feature map F to obtain the position attention feature map FP.
[0028] In one embodiment, the feature map obtained after performing a convolution operation on the fused feature map is input into the channel attention module CAM for processing, and the obtained channel attention feature map includes:
[0029] Perform Reshape on the feature map F to obtain the feature map F1 with the size of (H×W)×C;
[0030] Perform Transpose on the feature map F1 to obtain the feature map F2 with the size of C×(H×W);
[0031] Multiply the feature map F1 and the feature map F2, and then obtain the inter-channel correlation map X2 with the size of C×C through the Softmax function;
[0032] Multiply the inter-channel correlation map X2 with the feature map F1, and then obtain the feature map F3 with the size of H×W×C through Reshape;
[0033] Multiply the feature map F3 by the scale correlation coefficient β, and then add it to the feature map F to obtain the final channel attention feature map FC.
[0034] In one embodiment, MFDA-Net includes an encoding part, and the encoding part includes 6 encoding blocks. Each encoding block consists of two consecutive convolutional layers, a ReLU activation layer, and a pooling layer using the max pooling method. A ReLU activation layer is connected after the convolutional layer, and the convolutional kernel size of the convolutional layer is 3×3 with a sliding step of 1. The encoding part of MFDA-Net includes:
[0035] In the first stage, input the arterial phase image and the portal venous phase image in the input image pair of the training set into two first-layer encoding blocks respectively;
[0036] In the second stage, input the arterial phase feature map and the portal venous phase feature map obtained in the first stage into the multi-temporal feature fusion mechanism MFF to obtain the portal venous phase feature map that fuses the arterial phase image features for the first time;
[0037] In the third stage, input the arterial phase feature map obtained in the first stage and the portal venous phase feature map obtained in the second stage into two second-layer encoding blocks respectively;
[0038] In the fourth stage, input the arterial phase feature map and the portal venous phase feature map obtained in the third stage into the multi-temporal feature fusion mechanism MFF to obtain the portal venous phase feature map that fuses the arterial phase image features for the second time;
[0039] In the fifth stage, the output feature map obtained in the fourth stage is input into the third encoding block;
[0040] In the sixth stage, the output feature map obtained in the fifth stage is input into the fourth encoding block;
[0041] In the seventh stage, the output feature map obtained in the sixth stage is input into the Dual Attention Mechanism (DAM) to obtain a feature map with both channel correlation and position correlation adjusted;
[0042] In the eighth stage, the output feature map obtained in the seventh stage is input into the fifth encoding block;
[0043] In the ninth stage, the output feature map obtained in the eighth stage is input into the sixth encoding block;
[0044] The number of convolution kernels in the six-layer encoding block is 32, 64, 128, 256, 512, and 1024 in sequence.
[0045] In one embodiment, MFDA-Net includes a decoding part. The decoding part includes 5 decoding blocks. Each decoding block consists of an upsampling layer, a concatenate feature fusion layer, two consecutive convolutional layers, and a ReLU activation layer. A ReLU activation layer is connected after the convolutional layer. The size of the convolutional kernel is 3×3, and the sliding step is 1. The decoding part of MFDA-Net includes:
[0046] In the first stage, the output feature map of the ninth stage and the output feature map of the eighth stage are input into the first decoding block;
[0047] In the second stage, the output feature map of the first stage and the output feature map of the seventh stage are input into the second decoding block;
[0048] In the third stage, the output feature map of the second stage and the output feature map of the fifth stage are input into the third decoding block;
[0049] In the fourth stage, the output feature map of the third stage and the output feature map of the fourth stage are input into the fourth decoding block;
[0050] In the fifth stage, the output feature map of the fourth stage and the output feature map of the second stage are input into the fifth decoding block;
[0051] The number of convolution kernels in the five-layer decoding block is 512, 256, 128, 64, and 32 in sequence;
[0052] In the sixth stage, after the decoding part of MFDA-Net, there is a convolutional layer and an activation layer connected. The convolutional layer is used to reduce the number of channels of the feature map. The size of the convolutional kernel is 3×3, the number is 1, and the sliding step is 1. The activation layer uses the Sigmoid activation function to normalize the output result of the convolutional layer.
[0053] In one embodiment, the input image pairs composed of arterial phase images and portal venous phase images in the training set and the corresponding portal venous phase liver tumor labels are used to train MFDA-Net, and the trained MFDA-Net obtained includes:
[0054] The input image pairs composed of arterial phase images and portal venous phase images in the training set and the corresponding portal venous phase liver tumor labels are input into MFDA-Net for training. The parameters in the network are updated by backpropagation of the loss value calculated by the loss function using the network output generated each time training and the corresponding tumor labels, and the trained MFDA-Net is obtained.
[0055] In one embodiment, the loss function is the Combo Loss function, and its formula is as follows:
[0056]
[0057] where α1 and β1 are both weighting coefficients, α1 = 0.3, β1 = 0.8; n represents the total number of pixel points of the input image; g i is the value of the i-th pixel point of the tumor label corresponding to the input image; p i is the predicted value of the i-th pixel point of the segmentation result obtained after the input image is input into the network; ε is a constant, ε = 1.
[0058] In one embodiment, the maximum number of training epochs is set to 100, the learning rate is 0.0001, and the batch size is 4; it is set that if there is no improvement within 5 generations of training, the learning rate is multiplied by 0.1, and if there is no improvement within 10 generations of training, the training is stopped in advance.
[0059] The present invention has the following advantages compared with the prior art:
[0060] The embodiments of the present disclosure utilize a multi-temporal feature fusion mechanism to fuse the shallow features extracted by convolution of arterial phase images and portal venous phase images in multi-temporal CT. The positions and deformations of each organ in these two phases of images are relatively small, and there are obvious differences in the contrast between liver tumors and normal liver tissues; the feature fusion mechanism introduces more shape features and texture features of the liver and liver tumors into the segmentation of liver tumors, and the addition of arterial phase images improves the accuracy of liver tumor segmentation.
[0061] Meanwhile, by adding a multi-temporal feature fusion mechanism in the early stage of the encoding part of the segmentation network to fuse the arterial phase image features and the portal venous phase image features, and then adding a dual attention mechanism that takes into account both position correlation and channel correlation in the later stage of the encoding part, the attention of the segmentation network to liver tumors is further adjusted, enabling the network to pay more attention to information related to liver tumor segmentation during the training process, reducing the attention of the segmentation network to irrelevant information, further improving the ability of the network to segment liver tumors, and enhancing the accuracy of liver tumor segmentation.
[0062] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0064] Figure 1 is a flowchart of a liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism provided by an embodiment of the present disclosure;
[0065] Figure 2 is a schematic structural diagram of a multi-temporal feature fusion mechanism MFF provided by an embodiment of the present disclosure;
[0066] Figure 3 is a schematic structural diagram of a dual attention mechanism DAM provided by an embodiment of the present disclosure;
[0067] Figure 4 is a schematic structural diagram of an MFDA-Net based on multi-temporal fusion and dual attention mechanism provided by an embodiment of the present disclosure;
[0068] Figure 5 is a pair of arterial phase images and portal venous phase images in a test set provided by an embodiment of the present disclosure;
[0069] Figure 6 is obtained by using the present invention to Figure 5 the portal venous phase image in and performing liver tumor segmentation;
[0070] Figure 7 is to Figure 6 mark the segmentation result of on the portal venous phase image;
[0071] Figure 8 is a comparison result graph obtained by using a U-Net network with five downsamplings and the network of the present invention to Figure 5 perform liver tumor segmentation on the portal venous phase image in; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0073] Referring to Figure 1 , the liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism provided by the embodiments of the present disclosure includes:
[0074] Step 1: Make a training set and a test set, specifically including:
[0075] Obtain L case data, take the same layer of the arterial phase image and the portal venous phase image obtained by each case's multi-temporal CT scan each time as a group of input image pairs, and obtain the liver tumor label of the portal venous phase image. The liver tumor label of the portal venous phase image is used as the label for training the MFDA-Net. Randomly take P case data from the L case data as the training set, and the remaining Q case data as the test set. Exemplarily, 80% of all case data can be used as the training set, and the remaining 20% of the case data as the test set.
[0076] In the process of diagnosis and treatment of liver cancer, etc., the multi-temporal images obtained by the multi-temporal CT technology are superior in reflecting the pathological morphology of the liver, bringing more information for the detection of liver tumors. After each multi-temporal CT scan of a patient, arterial phase, portal venous phase, and delayed phase images can be obtained. Among them, the arterial phase image is obtained about 30 seconds after the patient is intravenously injected with a contrast agent, and the portal venous phase image is obtained about 60 - 90 seconds after the intravenous injection of the contrast agent. Since the time interval between the arterial phase image and the portal venous phase image is short, the positions and deformations of each organ between the two phases of images are small, and there are obvious differences in the enhancement degrees between the liver tumor and the normal liver tissue; the delayed phase image is taken about 3 - 5 minutes after the patient is intravenously injected with the contrast agent, and the time interval from the previous two phases is long, and the position and deformation of the organ are larger compared with the previous two-phase images. Therefore, effectively fusing the information of tumors, organs, etc. in the arterial phase image and the portal venous phase image obtained by the same patient's multi-temporal CT scan is beneficial to segment the liver tumor in its CT image.
[0077] Step 2: Construct a multi-temporal feature fusion mechanism MFF to fuse the two features obtained by shallow convolution of the arterial phase and portal venous phase images to obtain a fused feature map.
[0078] Among them, the fused feature map fuses the arterial phase image features and the portal vein image features, and can provide more favorable shape features and texture features for MFDA-Net.
[0079] Referring to Figure 2 As shown, step 2 specifically includes: the feature map A extracted from the arterial phase image and the feature map V extracted from the portal vein phase image are respectively passed through a convolutional layer with a kernel size of 1×1, a sliding stride of 1, and a number of 1 / 4 of the number of input feature map channels to obtain the feature map A1 and the feature map V1;
[0080] After adding the feature maps A1 and V1, they are passed through an activation layer with a rectified linear unit (ReLU) function as the activation function, then through a convolutional layer with a kernel size of 1×1, a sliding stride of 1, and a number of 1, and then through an activation layer using a sigmoid activation function to obtain the attention weight map;
[0081] Multiply the feature map A by the attention weight map to obtain the feature map A2 with adjusted attention;
[0082] Perform a concat (fusion) operation on the feature map V and the feature map A2 in the feature map channel direction, then pass through a convolutional layer with a kernel size of 3×3, a sliding stride of 1, and a number of channels equal to the number of channels of the feature map V, and then through an activation layer with a rectified linear unit (ReLU) function as the activation function, and finally obtain the fused feature map V2.
[0083] Step 3: Construct a dual attention mechanism (DAM) that includes a position attention module (PAM) and a channel attention module (CAM), and process the feature map obtained by performing a convolution operation on the fused feature map through the dual attention mechanism DAM to obtain a dual attention feature map.
[0084] Among them, the dual attention mechanism DAM focuses on the correlation between global positions of features and the correlation between channels, and through this mechanism, a dual attention feature map with higher attention to the features of the liver tumor region can be obtained.
[0085] Referring to Figure 3 , the dual attention mechanism (DAM) consists of a position attention module (PAM) and a channel attention module (CAM). Denote the size of the feature map F output after performing a convolution operation on the fused feature map output by the upper layer as H×W×C, where H, W, and C are the height, width, and number of channels of the feature map F respectively. Step 3 is specifically implemented as follows:
[0086] 3.1) In the position attention module PAM:
[0087] The feature map F obtained after performing a convolution operation on the fused feature map is respectively passed through two convolutional layers with a convolutional kernel size of 1×1, a number of C / 4, and a sliding stride of 1 to obtain the feature map FX and the feature map FY. The size of the feature map F is H×W×C, where H, W, and C are the height, width, and number of channels of the feature map F, respectively.
[0088] The feature map F is passed through a convolutional layer with a convolutional kernel size of 1×1, a number of C, and a sliding stride of 1 to obtain the feature map FZ.
[0089] The feature map FY and the feature map FZ are respectively reshaped (size adjusted) to obtain the feature map FY1 with a size of and the feature map FZ1 with a size of (H×W)×C.
[0090] The feature map FX is reshaped to obtain the feature map FX1 with a size of and then transposed to obtain the feature map FX2 with a size of The feature map FY1 and the feature map FX2 are multiplied and then passed through the Softmax function to obtain the inter-channel correlation map X1 with a size of (H×W)×(H×W).
[0091] The inter-channel correlation map X1 and the feature map FZ1 are multiplied and then reshaped to obtain the feature map FZ2 with a size of H×W×C.
[0092] The feature map FZ2 is multiplied by the scale correlation coefficient α and then added to the feature map F to obtain the position attention feature map FP.
[0093] The feature map FZ2 is multiplied by the scale correlation coefficient α and then added to the feature map F to obtain the position attention feature map FP.
[0094] 3.2) In the channel attention module CAM:
[0095] The feature map F is reshaped to obtain the feature map F1 with a size of (H×W)×C.
[0096] The feature map F1 is transposed to obtain the feature map F2 with a size of C×(H×W).
[0097] The feature map F1 and the feature map F2 are multiplied and then passed through the Softmax function to obtain the inter-channel correlation map X2 with a size of C×C.
[0098] The inter-channel correlation map X2 and the feature map F1 are multiplied and then reshaped to obtain the feature map F3 with a size of H×W×C.
[0099] The feature map F3 is multiplied by the scale correlation coefficient β and then added to the feature map F to obtain the final channel attention feature map FC.
[0100] 3.3) Add the channel attention feature map FC and the position attention feature map FP, then pass them through a convolutional layer with a convolutional kernel size of 3×3 and a sliding step of 1, and then through an activation layer with the rectified linear unit ReLU function as the activation function to finally obtain the dual attention feature map.
[0101] Step 4: Construct an MFDA-Net based on multi-temporal fusion and dual attention mechanism.
[0102] Construct an MFDA-Net based on multi-temporal fusion and dual attention mechanism. The basic structure of this segmentation network is a U-Net structure with five downsamplings. This segmentation network includes an encoding part and a decoding part. After the first two encoding blocks in the encoding part, add the MFF in step 2 to fuse the features extracted from the arterial phase image and the features extracted from the portal venous phase image. After the fourth encoding block in the encoding part, add the DAM in 3) to further improve the network's attention to liver tumors and reduce the attention to irrelevant information.
[0103] 4.1) Refer to Figure 4 For the backbone network of the MFDA-Net based on multi-temporal fusion and dual attention mechanism, a U-Net structure with five downsamplings is adopted. This segmentation network includes an encoding part and a decoding part. The encoding part contains 6 encoding blocks, and the decoding part contains 5 decoding blocks. The specific encoding blocks and decoding blocks are as follows:
[0104] The encoding block consists of two consecutive convolutional layers, a pooling layer using the max pooling method, and a ReLU activation layer. The convolutional kernel size is 3×3, the sliding step is 1, and a ReLU activation layer is connected after each convolutional layer.
[0105] The decoding block consists of an upsampling layer, a concatenate feature fusion layer, two consecutive convolutional layers, and a ReLU activation layer. The convolutional kernel size is 3×3, the sliding step is 1, and a ReLU activation layer is connected after each convolutional layer.
[0106] 4.2) Refer to Figure 4 The encoding part of the MFDA-Net includes the following 9 stages:
[0107] In the first stage, the arterial phase image and the portal venous phase image in the input image pair in the training set are respectively input into two first-layer encoding blocks.
[0108] In the second stage, input the arterial phase feature map and the portal venous phase feature map obtained in the first stage into the multi-temporal feature fusion mechanism MFF in 2) to obtain the portal venous phase feature map that fuses the arterial phase image features for the first time.
[0109] In the third stage, input the arterial phase feature map obtained in the first stage and the portal venous phase feature map obtained in the second stage into two second-layer encoding blocks respectively.
[0110] In the fourth stage, the arterial phase feature map and the portal venous phase feature map obtained in the third stage are input into the multi-temporal feature fusion mechanism MFF in 2) to obtain the portal venous phase feature map that fuses the arterial phase image features for the second time;
[0111] In the fifth stage, the output feature map obtained in the fourth stage is input into the third layer of the encoding block;
[0112] In the sixth stage, the output feature map obtained in the fifth stage is input into the fourth layer of the encoding block;
[0113] In the seventh stage, the output feature map obtained in the sixth stage is input into the dual attention mechanism DAM in 3) to obtain the feature map with both channel correlation and position correlation adjusted;
[0114] In the eighth stage, the output feature map obtained in the seventh stage is input into the fifth layer of the encoding block;
[0115] In the ninth stage, the output feature map obtained in the eighth stage is input into the sixth layer of the encoding block;
[0116] The number of convolution kernels of the above six-layer encoding block is 32, 64, 128, 256, 512, and 1024 in sequence.
[0117] In the embodiment of the present disclosure, the multi-temporal feature fusion mechanism MFF is added after the first two encoding blocks in the encoding part to fuse the features extracted from the arterial phase image and the features extracted from the portal venous phase image, and the dual attention mechanism DAM is added after the fourth encoding block in the encoding part to further improve the network's attention to liver tumors and reduce the attention to irrelevant information.
[0118] 4.3) Refer to Figure 4 , the decoding part of MFDA-Net includes the following six stages:
[0119] In the first stage, the output feature map of the ninth stage and the output feature map of the eighth stage are input into the first layer of the decoding block;
[0120] In the second stage, the output feature map of the first stage and the output feature map of the seventh stage are input into the second layer of the decoding block;
[0121] In the third stage, the output feature map of the second stage and the output feature map of the fifth stage are input into the third layer of the decoding block;
[0122] In the fourth stage, the output feature map of the third stage and the output feature map of the fourth stage are input into the fourth layer of the decoding block;
[0123] In the fifth stage, the output feature map of the fourth stage and the output feature map of the second stage are input into the fifth layer of the decoding block;
[0124] The number of convolution kernels of the above five-layer decoding block is 512, 256, 128, 64, and 32 in sequence.
[0125] In the sixth stage, a convolutional layer and an activation layer are connected after the decoding part of MFDA-Net. Among them, the convolutional layer is used to reduce the number of channels of the feature map. The size of the convolutional kernel is 3×3, the number is 1, and the sliding step is 1; the activation layer uses the Sigmoid activation function to normalize the output result of the convolutional layer.
[0126] Step 5: Train the MFDA-Net constructed in Step 4.
[0127] Specifically, take the input image pairs composed of arterial phase images and portal venous phase images and the corresponding liver tumor labels from the training set and input them into the constructed MFDA-Net for training. The process is as follows:
[0128] 5.1) Input the input image pairs composed of arterial phase images and portal venous phase images in the training set and the corresponding portal venous phase liver tumor labels into the MFDA-Net constructed in 4) for training;
[0129] 5.2) Set the maximum number of training epochs to 100, the learning rate to 0.0001, and the batch size to 4. Update the parameters in the network by backpropagation of the loss value calculated by the Combo Loss function using the network output generated by each training and the corresponding tumor label. Finally, obtain the trained MFDA-Net. If there is no improvement within 5 generations during training, the learning rate is multiplied by 0.1, and if there is no improvement within 10 generations, an early stopping mechanism is used to prevent overfitting of network training.
[0130] Among them, the loss function used when training the segmentation network is the Combo Loss function, and its formula is as follows:
[0131]
[0132] Among them, both α1 and β1 are weighting coefficients, α1 = 0.3, β1 = 0.8; n represents the total number of pixel points of the input image; g i is the value of the i-th pixel point of the tumor label corresponding to the input image; p i is the predicted value of the i-th pixel point of the segmentation result obtained after the input image is input into the network; ε is a constant, ε = 1.
[0133] Step 6: Segment the liver tumors in the portal venous phase images of the test set image pairs.
[0134] Specifically, it includes: inputting the image pairs composed of arterial phase images and portal venous phase images in the test set into the trained MFDA-Net obtained in Step 5, and performing binarization processing on the network output result using a threshold of 0.5 to finally obtain the segmentation result of the liver tumors in the test set.
[0135] In the embodiments of the present disclosure, a multi-temporal feature fusion mechanism is utilized to fuse the shallow features extracted by convolution from the arterial phase image and the portal venous phase image in multi-temporal CT. The positions and deformations of each organ in these two phases of images are relatively small, and there are obvious differences in the contrast between liver tumors and normal liver tissues. The feature fusion mechanism introduces more shape features and texture features of the liver and liver tumors into the segmentation of liver tumors, and the addition of the arterial phase image improves the accuracy of liver tumor segmentation.
[0136] Meanwhile, by adding a multi-temporal feature fusion mechanism at the early stage of the encoding part of the segmentation network to fuse the arterial phase image features and the portal venous phase image features, and then adding a dual attention mechanism that takes into account both position correlation and channel correlation at the later stage of the encoding part, the attention of the segmentation network to liver tumors is further adjusted, enabling the network to pay more attention to the information related to liver tumor segmentation during the training process, reducing the attention of the segmentation network to irrelevant information, further improving the ability of the network to segment liver tumors, and enhancing the accuracy of liver tumor segmentation.
[0137] The effects of the present invention are realized through the following simulations.
[0138] Simulation 1: Input a pair of arterial phase image and portal venous phase image shown in Figure 5 into the trained dual-temporal input MFDA-Net in step 6. The result of liver tumor segmentation shown on the portal venous phase image is as shown in Figure 6 shown.
[0139] Simulation 2: Mark the segmentation result of Figure 6 onto the portal venous phase image of Figure 5 . The result is as shown in Figure 7 . It can be seen from Figure 7 that the present invention can effectively segment liver tumors in the portal venous phase image.
[0140] Simulation 3: Input a pair of arterial phase image and portal venous phase image shown in Figure 5 into the U-Net network with five downsamplings. The result of liver tumor segmentation shown on the portal venous phase image is compared with Figure 6 and shown as in Figure 8 shown. It can be seen that the present invention can effectively improve the segmentation effect of liver tumors in the portal venous phase.
[0141] The embodiments of the present disclosure also provide a liver tumor segmentation device based on multi-temporal fusion and dual attention mechanism. The liver tumor segmentation device based on multi-temporal fusion and dual attention mechanism includes a receiver, a transmitter, a memory, and a processor. The transmitter and the memory are respectively connected to the processor. The memory stores at least one computer instruction, and the processor is configured to load and execute at least one computer instruction to implement the aboveFigure 1 The liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism described in the corresponding embodiment.
[0142] Based on the above Figure 1 Based on the liver tumor segmentation method described in the corresponding embodiment, the embodiments of the present disclosure further provide a computer-readable storage medium. For example, a non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. Computer instructions are stored on the storage medium for executing the above Figure 1 The liver tumor segmentation method described in the corresponding embodiment, which will not be elaborated here.
[0143] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0144] After considering the specification and practicing the disclosure herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
Claims
1. A liver tumor segmentation method based on multi-temporal fusion and dual attention mechanism, characterized in that The method includes: Step 1: Obtain L case data. Take the arterial-phase image and the portal-venous-phase image of the same layer obtained from each case during multi-phase CT scans as a group of input image pairs, and obtain the liver tumor labels of the portal-venous-phase images. Randomly take P case data from the L case data as the training set, and the remaining Q case data as the test set; Step 2: Construct a multi-phase feature fusion mechanism MFF to fuse the two features obtained by shallow convolution of the arterial-phase and portal-venous-phase images to obtain a fused feature map; Step 3: Construct a dual attention mechanism DAM including a position attention module PAM and a channel attention module CAM. Input the feature map obtained after convolution operation on the fused feature map into the dual attention mechanism DAM for processing to obtain a dual attention feature map; Step 4: Embed the multi-phase feature fusion mechanism MFF and the dual attention mechanism DAM into the U-Net structure to form a liver tumor segmentation network MFD A-Net based on multi-phase fusion and dual attention mechanism; Step 5: Train the MFD A-Net with the input image pairs composed of arterial-phase images and portal-venous-phase images in the training set and the corresponding portal-venous-phase liver tumor labels to obtain a trained MFD A-Net; Step 6: Input the image pairs composed of arterial-phase images and portal-venous-phase images in the test set into the trained MFD A-Net, use a binarization method to process the output result of the network to obtain a segmentation result, and represent the segmentation result in the form of a contour on the portal-venous-phase image.
2. The method according to claim 1, wherein The construction of the multi-phase feature fusion mechanism MFF to fuse the two features obtained by shallow convolution of the arterial-phase and portal-venous-phase images to obtain a fused feature map includes: The feature map A extracted from the arterial-phase image and the feature map V extracted from the portal-venous-phase image are respectively passed through a convolutional layer with a kernel size of 1×1, a sliding step of 1, and a number of 1 / 4 of the number of input feature map channels to obtain the feature map A1 and the feature map V1; After adding the feature maps A1 and V1, successively pass through an activation layer with a rectified linear unit ReLU function as the activation function, a convolutional layer with a kernel size of 1×1, a sliding step of 1, and a number of 1, and an activation layer with a Sigmoid activation function to obtain an attention weight map; Multiply the feature map A by the attention weight map to obtain an adjusted attention feature map A2; Perform a feature map channel-wise fusion operation on the feature map V and the feature map A2, and then successively pass through a convolutional layer with a kernel size of 3×3, a sliding step of 1, and a number of channels equal to the number of channels of the feature map V and an activation layer with a rectified linear unit ReLU function as the activation function to obtain a fused feature map V2.
3. The method according to claim 1, wherein Inputting the feature map obtained after convolution operation on the fused feature map into the dual attention mechanism DAM for processing to obtain a dual attention feature map includes: The feature maps obtained after performing convolution operations on the fused feature maps are respectively input into the position attention module PAM and the channel attention module CAM for processing to obtain a position attention feature map and a channel attention feature map; The position attention feature map and the channel attention feature map are added together and then successively passed through a convolutional layer with a convolutional kernel size of 3×3 and a sliding step of 1 and an activation layer with a rectified linear unit ReLU function as the activation function to obtain a dual attention feature map.
4. The method according to claim 3, characterized in that, The step of inputting the feature maps obtained after performing convolution operations on the fused feature maps into the position attention module PAM to obtain a position attention feature map includes: The feature map F obtained after performing convolution operations on the fused feature map is respectively passed through two convolutional layers with a convolutional kernel size of 1×1, a number of C / 4, and a sliding step of 1 to obtain a feature map FX and a feature map FY. The size of the feature map F is H×W×C, where H, W, and C are the height, width, and number of channels of the feature map F respectively; The feature map F is passed through a convolutional layer with a convolutional kernel size of 1×1, a number of C, and a sliding step of 1 to obtain a feature map FZ; Reshape the feature map FY and the feature map FZ respectively to obtain a feature map FY1 with a size of and a feature map FZ1 with a size of (H×W)×C; Reshape the feature map FX to obtain a feature map FX1 with a size of , and then perform Transpose to obtain a feature map FX2 with a size of ; The feature map FY1 and the feature map FX2 are multiplied and then passed through the Softmax function to obtain an inter-channel correlation map X1 of size (H×W)×(H×W); The inter-channel correlation map X1 is multiplied by the feature map FZ1 and then passed through Reshape to obtain a feature map FZ2 of size H×W×C; The feature map FZ2 is multiplied by the scale correlation coefficient α and then added to the feature map F to obtain a position attention feature map FP.
5. The method according to claim 3, wherein The step of inputting the feature maps obtained after performing convolution operations on the fused feature maps into the channel attention module CAM for processing to obtain a channel attention feature map includes: The feature map F is reshaped to obtain a feature map F1 of size (H×W)×C; The feature map F1 is transposed to obtain a feature map F2 of size C×(H×W); The feature map F1 and the feature map F2 are multiplied and then passed through the Softmax function to obtain an inter-channel correlation map X2 of size C×C; The inter-channel correlation map X2 is multiplied by the feature map F1 and then passed through Reshape to obtain a feature map F3 of size H×W×C; The feature map F3 is multiplied by the scale correlation coefficient β and then added to the feature map F to obtain a channel attention feature map FC.
6. The method according to claim 1, wherein The MFDA-Net includes an encoding part, and the encoding part includes 6 encoding blocks. Each encoding block consists of two consecutive convolutional layers, a ReLU activation layer, and a pooling layer using the max pooling method. A ReLU activation layer is connected after the convolutional layer. The convolutional kernel size of the convolutional layer is 3×3, and the sliding step is 1; The encoding part of the MFDA-Net includes: In the first stage, the arterial phase image and the portal venous phase image in the input image pair in the training set are respectively input into two first-layer encoding blocks; In the second stage, the arterial phase feature map and the portal venous phase feature map obtained in the first stage are input into the multi-temporal feature fusion mechanism MFF to obtain a portal venous phase feature map that fuses the arterial phase image features for the first time; In the third stage, the arterial phase feature map obtained in the first stage and the portal venous phase feature map obtained in the second stage are respectively input into two second-layer encoding blocks; In the fourth stage, the arterial phase feature map and the portal venous phase feature map obtained in the third stage are input into the multi-temporal feature fusion mechanism MFF to obtain the portal venous phase feature map that fuses the arterial phase image features for the second time; In the fifth stage, the output feature map obtained in the fourth stage is input into the third-layer encoding block; In the sixth stage, the output feature map obtained in the fifth stage is input into the fourth-layer encoding block; In the seventh stage, the output feature map obtained in the sixth stage is input into the dual attention mechanism DAM to obtain the feature map with both channel correlation and position correlation adjusted; In the eighth stage, the output feature map obtained in the seventh stage is input into the fifth-layer encoding block; In the ninth stage, the output feature map obtained in the eighth stage is input into the sixth-layer encoding block; The number of convolution kernels of the six-layer encoding block is 32, 64, 128, 256, 512, and 1024 in sequence.
7. The method according to claim 6, characterized in that, The MFDA-Net includes a decoding part, and the decoding part includes 5 layers of decoding blocks. Each layer of decoding block consists of an upsampling layer, a concatenate feature fusion layer, two consecutive convolutional layers, and a ReLU activation layer. A ReLU activation layer is connected after the convolutional layer. The size of the convolutional kernel is 3×3, and the sliding step is 1; The decoding part of the MFDA-Net includes: In the first stage, the output feature map of the ninth stage and the output feature map of the eighth stage are input into the first-layer decoding block; In the second stage, the output feature map of the first stage and the output feature map of the seventh stage are input into the second-layer decoding block; In the third stage, the output feature map of the second stage and the output feature map of the fifth stage are input into the third-layer decoding block; In the fourth stage, the output feature map of the third stage and the output feature map of the fourth stage are input into the fourth-layer decoding block; In the fifth stage, the output feature map of the fourth stage and the output feature map of the second stage are input into the fifth-layer decoding block; The number of convolution kernels of the five-layer decoding block is 512, 256, 128, 64, and 32 in sequence; In the sixth stage, a convolutional layer and an activation layer are connected after the decoding part of the MFDA-Net. The convolutional layer is used to reduce the number of channels of the feature map. The size of the convolutional kernel is 3×3, the number is 1, and the sliding step is 1; The activation layer uses the Sigmoid activation function to normalize the output result of the convolutional layer.
8. The method according to claim 1, characterized in that, Training the MFDA-Net with the input image pairs composed of arterial phase images and portal venous phase images in the training set and the corresponding portal venous phase liver tumor labels includes: Inputting the input image pairs composed of arterial phase images and portal venous phase images in the training set and the corresponding portal venous phase liver tumor labels into the MFDA-Net for training. The parameters in the network are updated by backpropagation of the loss value calculated by the loss function using the network output generated each time training and the corresponding tumor labels, and the trained MFDA-Net is obtained.
9. The method according to claim 8, characterized in that The loss function is the Combo Loss function, and its formula is as follows: Among them, both α1 and β1 are weighting coefficients, α1 = 0.3, β1 = 0.8; n represents the total number of pixel points of the input image; g i is the value of the i-th pixel point of the tumor label corresponding to the input image; p i is the predicted value of the i-th pixel point of the segmentation result obtained after the input image is input into the network; ε is a constant, ε = 1.
10. The method according to claim 8, wherein Set the maximum number of training epochs to 100, the learning rate to 0.0001, and the batch size to 4; set the learning rate to be multiplied by 0.1 if there is no improvement within 5 training epochs, and stop training early if there is no improvement within 10 training epochs.
Citation Information
Patent Citations
Multi-temporal CT image liver tumor segmentation method based on semantic migration
CN112785605A
Multi-temporal remote sensing image cloud region reconstruction method based on collaborative attention mechanism
CN114140357A
Cited By
A deep learning liver lesion image recognition method
CN122820572A