Multi-Phase Liver Tumor Segmentation Method Based on Multi-Head Cross-Attention Transformation Network
By constructing a multi-phase liver tumor segmentation model of multi-head cross-attention conversion network, the problem of insufficient segmentation accuracy in the prior art is solved, and higher segmentation accuracy and consistency are achieved.
Patent Information
- Application Number
- CN202210973043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-15
AI Technical Summary
When existing computer-assisted diagnostic methods use multi-phase CT images for liver tumor segmentation, there is a convolutional neural network that cannot effectively capture long-distance dependencies, resulting in information loss of underlying features such as edges, texture shapes, etc., affecting segmentation accuracy.
A multi-phase liver tumor segmentation model based on a multi-head cross attention conversion network was constructed. The characteristics of the arterial and portal vein images were analyzed and selectively fused through multi-head cross attention blocks, and model training was performed on combining segmentation loss, consistency loss and level set loss to improve segmentation accuracy.
It reduces the loss of underlying characteristic information of liver tumors, improves the accuracy and consistency of liver tumor segmentation, and improves segmentation accuracy.
Smart Images

Figure CN115330816B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a liver tumor segmentation method, which can be used for automatically detecting and segmenting human liver tumor lesions in CT images. Background Art
[0002] Liver cancer refers to a malignant tumor disease that occurs in the liver. Research shows that there were approximately more than 900,000 new cases of liver cancer and approximately more than 800,000 death cases globally in 2020. In China, the incidence and mortality rates of liver cancer have been at a high level all year round. Therefore, the screening of liver tumors, the diagnosis and treatment of liver cancer patients are important links to improve the survival period of liver cancer patients. Currently, liver cancer patients usually determine the situation of their own liver tumors through multi-phase CT scans during the treatment process.
[0003] However, during the treatment of liver cancer, a single CT scan can produce dozens to hundreds of images. Physicians need to layer by layer label the liver tumors in the images. This process not only takes time and effort, but also physicians are prone to missing the labeling of some lesions due to factors such as fatigue, resulting in too large differences in tumor labeling between adjacent layers, and the labeling results vary depending on the practical experience and subjective judgment of the physicians. Therefore, there is an urgent need to establish a reliable computer-aided diagnosis method for automatically segmenting liver tumors to provide a basis for radiologists to objectively evaluate the situation of liver tumors. With the development of computer-aided diagnosis in clinical medicine, a large number of researchers have devoted themselves to the computer analysis of liver tumor segmentation using CT images.
[0004] In the existing computer-aided diagnosis liver tumor segmentation methods, many researchers have made a lot of efforts in how to use multi-phase CT images to improve the accuracy of liver tumor segmentation. Tao et al. proposed in "Liver segmentation from registered multiphase CT data sets with EM clustering and GVF level set" published in Medical Imaging that first register other phase images with arterial phase images or portal venous phase images, and then use the rich information of multi-phase images combined with Gaussian mixture model and level set method to improve the accuracy and robustness of complex liver segmentation. Hasegawa et al. proposed in "Automatic detection and segmentation of liver tumors in multi-phase ct images by phase attention mask r-cnn" published in IEEE a method of fusing multi-phase image features in enhanced CT in the feature extraction stage under the MaskR-CNN framework. This method makes full use of the information of multi-phase CT images to improve the segmentation effect of liver tumors. Xu et al. proposed in "Study on the segmentation method of multi-phase CT liver tumor based on dual-channel U-Nets" published in Journal of Physics a multi-level multi-phase CT image liver tumor segmentation algorithm with dual-channel cascading. First, use U-net to segment the liver region, multiply the arterial phase and portal venous phase images by the obtained liver region, and then use two new U-nets to segment the liver tumors in the liver regions of the two-phase images respectively. Finally, the two segmentation results are fused through a fusion layer to obtain the final liver tumor segmentation result.
[0005] All of the above methods use multi-phase CT images for liver tumor segmentation. Although they can provide more information about liver tumors for the segmentation task, due to the use of convolutional neural networks, these methods have limitations in capturing long-distance dependence relationships, cannot perform good global information correlation analysis, and will cause losses of underlying features such as edges, texture shapes, etc. in the network feature extraction stage, thus affecting the segmentation accuracy of the network. Summary of the Invention
[0006] The purpose of the present invention is to propose a multi-phase liver tumor segmentation method based on a multi-head cross-attention conversion network in view of the deficiencies of the prior art, so as to reduce the loss of underlying features of liver tumors and improve the segmentation accuracy.
[0007] The technical solution for achieving the object of the present invention is as follows: By utilizing the multi-head cross-attention, the correlation analysis of the features of the arterial phase and portal venous phase images can be performed and selectively fused, reducing the loss of the underlying feature information of liver tumors. By introducing the dual-task consistency loss to constrain the segmentation results of the network, the accuracy of the liver tumor segmentation results is improved. The implementation steps are as follows:
[0008] (1) Obtain the multi-phase CT image dataset and perform preprocessing:
[0009] (1a) Adopt the window method to narrow the range of CT values, increase the contrast between the liver and surrounding tissues, set the window level to 50 HU and the window width to 300 HU, and limit the CT values of the CT images within the range of [-100 HU, 200 HU];
[0010] (1b) Perform normalization processing on the CT images by using the minimum-maximum normalization method, and map the CT values into the range of [0, 1];
[0011] (1c) Downsample the normalized arterial phase images, portal venous phase images and the corresponding portal venous phase liver tumor labels to 256×256;
[0012] (2) Apply the randomly selected method to divide the preprocessed multi-phase CT image dataset into a training set and a test set according to the ratio of 3:1. Among them, there are 7774 arterial phase and portal venous phase images in the training set, and 1323 arterial phase and portal venous phase images in the test set. And the arterial phase and portal venous phase images share the annotated portal venous phase liver tumor labels;
[0013] (3) Construct a multi-phase liver tumor segmentation model based on the multi-head cross-attention conversion network:
[0014] (3a) Construct a multi-head cross-attention block composed of multiple linearly connected layers in parallel, multiple cross-attention layers in parallel, a concatenation layer and a linear layer cascaded in sequence. Among them, the cross-attention layer is composed of a MatMul matrix multiplication layer, a scale normalization layer, a Softmax layer and a MatMul matrix multiplication layer cascaded in sequence;
[0015] (3b) Construct a feature fusion module MCAF composed of two position encoding layers in parallel, two normalization layers in parallel, a multi-head cross-attention block, a concatenation layer and a linear layer cascaded in sequence;
[0016] (3c) Select the existing TransUNet network, connect the feature fusion module MCAF between the encoder and the Transformer Layer of the network, and construct a multi-phase liver tumor segmentation model based on the multi-head cross-attention conversion network;
[0017] (4) Input the multi - temporal CT training set into the constructed liver tumor segmentation model, and set the segmentation loss Loss of the model seg , consistency loss Loss sdf and level - set loss Loss lsf , and use the backpropagation method to perform iterative training on it to obtain a trained liver tumor segmentation model;
[0018] (5) Input the multi - temporal CT test set into the trained liver tumor segmentation model to obtain the segmentation results on the test set.
[0019] The present invention has the following advantages compared with the prior art:
[0020] First, since the present invention constructs a multi - temporal liver tumor segmentation model based on a multi - head cross - attention conversion network, it can perform correlation analysis on the features of the arterial - phase and portal - venous - phase images through the multi - head cross - attention block therein and selectively fuse them, reducing the loss of underlying feature information of liver tumors.
[0021] Second, by setting the segmentation loss, consistency loss, and level - set loss of the model, the present invention calculates the model loss value using the obtained pixel - level segmentation results, level - set results, and liver tumor labels, and updates the model parameters through the backpropagation method, thereby performing consistency constraints on the output results of the model and improving the segmentation accuracy of liver tumors. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the implementation flow block diagram of the present invention;
[0023] Figure 2 is the framework diagram of the liver tumor segmentation model in the present invention;
[0024] Figure 3 is Figure 2 the schematic diagram of the segmentation network structure in the liver tumor segmentation model;
[0025] Figure 4 is Figure 3 the schematic diagram of the feature fusion module MCAF in the segmentation network;
[0026] Figure 5 is Figure 4 the schematic diagram of the multi - head cross - attention block in the feature fusion module MCAF;
[0027] Figure 6 is the attention comparison diagram of liver tumor feature extraction before and after using the MCAF module in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] The following further describes the examples and effects of the present invention in detail with reference to the drawings.
[0029] See Figure 1 , the specific implementation steps of this example are as follows:
[0030] Step 1, obtain a multi-temporal CT image dataset and perform preprocessing.
[0031] 1.1) Collect a multi-temporal CT image dataset with labels from liver cancer patients in the hospital. Each patient has 1 to 10 multi-temporal CT images, and each multi-temporal CT contains complete arterial phase and portal venous phase images;
[0032] 1.2) Use the window method to narrow the range of CT values, increase the contrast between the liver and surrounding tissues, set the window level to 50 HU, the window width to 300 HU, and limit the CT values of the CT images within the interval [-100 HU, 200 HU];
[0033] 1.3) Perform normalization processing on the CT images using min-max normalization, so that the CT values are mapped to the interval [0, 1], and obtain the pixel values Y of the normalized CT images:
[0034]
[0035] where X represents the CT value of the CT image, X min represents the minimum CT value, with a value of -100 HU, X max represents the maximum CT value, with a value of 200 HU;
[0036] 1.4) Downsample the normalized arterial phase images, portal venous phase images, and the corresponding portal venous phase liver tumor labels to 256×256.
[0037] Step 2, divide the multi-temporal CT image dataset.
[0038] Apply the random selection method to divide the preprocessed multi-temporal CT image dataset into a training set and a test set according to a ratio of 3:1. Among them, the training set contains 7774 arterial phase and portal venous phase images each, and the test set contains 1323 arterial phase and portal venous phase images each, and the arterial phase and portal venous phase images share the labeled portal venous phase liver tumor labels.
[0039] Step 3, construct a multi-temporal liver tumor segmentation model based on the multi-head cross-attention conversion network.
[0040] Refer to Figure 2 , the specific implementation of this step is as follows:
[0041] 3.1) Construct the feature fusion module MCAF:
[0042] 3.1.1) Design the multi-head cross-attention block
[0043] It is composed of a cascade of multiple linearly connected layers in parallel, multiple cross-attention layers in parallel, a concatenation layer, and a linear layer in sequence, forming a multi-head cross-attention block. Among them, the cross-attention layer is composed of a cascade of a MatMul matrix multiplication layer, a scale normalization layer, a Softmax layer, and a MatMul matrix multiplication layer, as Figure 5 shown;
[0044] 3.1.2) A feature fusion module MCAF is composed of a cascade of two position encoding layers in parallel, two normalization layers in parallel, a multi-head cross-attention block, a concatenation layer, and a linear layer in sequence, as Figure 4 shown;
[0045] 3.2) Select the TransUNet network:
[0046] Select the TransUNet network composed of a cascade of an encoder, a Transformer Layer, a Reshape layer, a convolutional layer, a decoder, and a convolutional block, as Figure 3 shown, where:
[0047] The encoder mentioned above is composed of two CNN convolutional blocks with the same structure and shared weights. Each convolutional block is composed of the first 4 residual blocks conv1, conv2_x, conv3_x, conv4_x of ResNet-50:
[0048] The decoder of the network is composed of a cascade of four decoding blocks. Each decoding block includes an upsampling layer, two convolutional layers, and two ReLU activation layers;
[0049] The convolutional block is composed of two branches connected in parallel. The first branch includes two convolutional layers with a kernel size of 3×3 and a stride of 1. The second branch includes a convolutional layer with a kernel size of 3×3 and a stride of 1 and a Tanh activation layer.
[0050] 3.3) Connect the feature fusion module MCAF constructed in 3.1) between the encoder of the TransUNet network selected in 3.2) and the Transformer Layer of the TransUNet network to form a multi-temporal liver tumor segmentation model based on the multi-head cross-attention conversion network.
[0051] 3.4) Construct the loss function of the liver tumor segmentation model:
[0052] 3.4.1) Set the segmentation loss Loss seg , which is expressed as follows:
[0053]
[0054] Among them, G represents the liver tumor label image, S represents the predicted pixel-level segmentation map, n represents the total number of image pixel points, the subscript i represents the index of the pixel point in the image, α represents a constant with a value of 0.7, and ε represents a constant;
[0055] 3.4.2) Set the consistency loss Loss sdf , which is expressed as follows:
[0056]
[0057] Among them, represents the converted pixel-level segmentation map, S represents the predicted pixel-level segmentation map, R represents the predicted level set map, n represents the total number of image pixel points, and the subscript i represents the index of the pixel point in the image;
[0058] 3.4.3) Set the level set loss Loss lsf , which is expressed as follows:
[0059]
[0060]
[0061] Among them, G represents the liver tumor label image, R represents the predicted level set map, G’ represents the true level set map, n represents the total number of image pixel points, and the subscripts i and j respectively represent the indices of different pixel points in the image, represents the boundary contour of the target liver tumor, S in , S out respectively represent the internal area and the external area of the target liver tumor;
[0062] 3.4.4) Add the segmentation loss Loss seg in 3.4.1), the consistency loss Loss sdf in 3.4.2) and the level set loss Loss lsf in 3.4.3) to form the loss function of the liver tumor segmentation model.
[0063] Step 4, use the backpropagation method to iteratively train the liver tumor segmentation model.
[0064] 4.1) Set the total number of iteration rounds to 200, the batch size to 8, the initial learning rate lr0 to 0.0001, and the learning rate changes according to with the number of training iterations, where t represents the current training iteration number, and t max represents the total number of training iterations, and the optimizer is the Adam optimizer;
[0065] 4.2) Input the multi - temporal CT training set into the liver tumor segmentation model in batches according to the batch size set in 4.1), and obtain the pixel - level segmentation result and the level - set result;
[0066] 4.3) According to the pixel - level segmentation result, the level - set result obtained in 4.2), and the liver tumor labels of the training set, calculate the loss value of the model through the loss function constructed in 4.1);
[0067] 4.4) Perform backpropagation on the loss value obtained in 4.3), update the model parameters, and obtain the preliminarily trained model W;
[0068] 4.5) Loop through the processes of 4.2) - 4.4) for the preliminarily trained model W until the number of iteration rounds reaches 200, then stop training to obtain the trained liver tumor segmentation model W';
[0069] Step 5, perform liver tumor segmentation on the multi - temporal CT test set.
[0070] Input the multi - temporal CT test set into the trained liver tumor segmentation model W' to obtain the segmentation result of the test set.
[0071] The technical effects of the present invention are further described below in combination with the simulation experiment results:
[0072] I. Simulation conditions:
[0073] The simulation experiment platform of the present invention is the Win10 operating system, configured with a 3.6GHz×8 CoreTM i7 - 9700K CPU and Nvidia RTX2080Ti GPU, using the PyTorch deep - learning framework, and the development language is Python.
[0074] The simulation experiment data of the present invention comes from a multi - temporal CT image data set with labels collected from liver cancer patients in the hospital. Each patient has 1 - 10 multi - temporal CT images. Each multi - temporal CT contains complete arterial - phase and portal - vein - phase images, and the arterial - phase and portal - vein - phase images share the labeled portal - vein - phase liver tumor labels. The image size in the data set is 512×512, the pixel pitch is 0.6426mm - 0.75mm, and the slice thickness is 2.5mm.
[0075] II. Simulation content and result analysis:
[0076] Simulation 1, respectively use the existing medical image segmentation method and the method of the present invention to perform liver tumor segmentation on the above - mentioned multi - temporal CT image data set, and the results are shown in Table 1:
[0077] Table 1 Comparison of different quantitative indicators of the liver tumor segmentation effects of different methods
[0078] Method DSC (%) HD (pixel) JC (%) UR (%) SEN (%) Existing method 82.05 119.20 72.96 14.96 84.14 Method of the present invention 85.56 72.64 76.83 11.48 87.55
[0079] In Table 1, DSC represents the Dice coefficient, HD represents the Hausdorff distance, JC represents the Jaccard similarity coefficient, UR represents the under-segmentation rate, and SEN represents the sensitivity. Among them, the higher the DSC, JC, and SEN indicators, the better the segmentation effect; the lower the HD and UR indicators, the better the segmentation effect.
[0080] As can be seen from Table 1, the DSC index, HD index, JC index, UR index, and SEN index of the liver tumor segmentation of the method of the present invention for liver cancer patients are all improved compared with the existing methods, which proves that the multi-phase liver tumor segmentation method based on the multi-head cross-attention conversion network of the present invention can improve the performance of liver tumor segmentation.
[0081] Simulation 2: To better observe the fusion effect of MCAF, 4 images were selected to visually display the feature maps before and after fusion superimposed on the input portal venous phase image, and the results are as Figure 6 shown, where:
[0082] 6(a) represents the portal venous phase image, and the solid line frame represents the contour of the liver tumor label.
[0083] 6(b) represents the arterial phase feature map before fusion.
[0084] 6(c) represents the portal venous phase feature map before fusion.
[0085] 6(d) represents the feature map after fusing the two-phase images of 6(b) and 6(c).
[0086] Comparing 6(b), 6(c), and 6(d), the attention to the liver tumor in the fused feature map is significantly improved, and the liver tumor area noticed is more complete. It shows that the present invention can effectively fuse the features of the liver tumor in the arterial phase and portal venous phase images, and selectively retain the useful features of the liver tumor in the two-phase images during the fusion process, so that the feature information amount of the liver tumor after fusion increases but is not redundant.
[0087] In summary, aiming at the problem that the existing methods will cause information loss of underlying features such as edges, texture shapes, etc. when extracting features of liver tumors, the feature fusion module MCAF constructed by the present invention can perform correlation analysis on the features of the arterial phase and portal venous phase images and selectively fuse them, which can strengthen the information related to liver tumors and shield irrelevant information, and improve the liver tumor segmentation accuracy of the segmentation network.
Claims
1. A multi-temporal liver tumor segmentation method based on a multi-head cross-attention conversion network, characterized in that Including: (1) Obtain a multi-temporal CT image dataset and perform preprocessing: (1a) Use the window method to narrow the range of CT values, increase the contrast between the liver and surrounding tissues, set the window level to 50 HU, the window width to 300 HU, and limit the CT values of the CT image within the interval [-100 HU, 200 HU]; (1b) Perform normalization processing on the CT image using min-max normalization, and map the CT values to the interval [0, 1]; (1c) Downsample the normalized arterial phase image, portal venous phase image, and the corresponding portal venous phase liver tumor label to 256×256; (2) Apply the random selection method to divide the preprocessed multi-temporal CT image dataset into a training set and a test set according to a ratio of 3:
1. Among them, the training set contains 7774 arterial phase and portal venous phase images each, and the test set contains 1323 arterial phase and portal venous phase images each. And the arterial phase and portal venous phase images share the annotated portal venous phase liver tumor label; (3) Construct a multi-temporal liver tumor segmentation model based on the multi-head cross-attention transformation network: (3a) Construct a multi-head cross-attention block composed of multiple linearly connected layers in parallel, multiple cross-attention layers in parallel, a splicing layer, and a linear layer cascaded in sequence. Among them, the cross-attention layer is composed of a MatMul matrix multiplication layer, a scale normalization layer, a Softmax layer, and a MatMul matrix multiplication layer cascaded in sequence; (3b) Construct a feature fusion module MCAF composed of two position encoding layers in parallel, two normalization layers in parallel, a multi-head cross-attention block, a splicing layer, and a linear layer cascaded in sequence; (3c) Select the existing TransUNet network, connect the feature fusion module MCAF between the encoder and the TransformerLayer layer of this network to form a multi-temporal liver tumor segmentation model based on the multi-head cross-attention transformation network; (4) Input the multi-temporal CT training set into the constructed liver tumor segmentation model, and set the segmentation loss Loss seg and the consistency loss Loss sdf and the level set loss Loss lsf , and use the backpropagation method to perform iterative training on it to obtain a trained liver tumor segmentation model; (5) Input the multi-temporal CT test set into the trained liver tumor segmentation model to obtain the segmentation result on the test set.
2. The method according to claim 1, wherein, In the above (1b), min-max normalization is used to perform normalization processing on the CT image, and the formula is as follows; Among them, X represents the CT value of the CT image, and X min represents the minimum CT value, with a value of -100 HU, and X max represents the maximum CT value, with a value of 200 HU, and Y represents the pixel value of the normalized CT image.
3. The method according to claim 1, wherein The TransUNet network in the above (3c) is composed of an encoder, a Transformer Layer layer, a Reshape layer, a convolutional layer, a decoder, and a convolutional block cascaded in sequence; The above encoder is composed of two CNN convolutional blocks with the same structure and shared weights. Each convolutional block is composed of the first 4 residual blocks conv1, conv2_x, conv3_x, conv4_x of ResNet-50: The above decoder is composed of four decoding blocks cascaded in sequence. Each decoding block includes an upsampling layer, two convolutional layers, and two ReLU activation layers; The above convolutional block is composed of two branches connected in parallel. Among them, the first branch contains two convolutional layers with a convolution kernel size of 3×3 and a stride of 1, and the second branch contains a convolutional layer with a convolution kernel size of 3×3 and a stride of 1 and a Tanh activation layer.
4. The method according to claim 1, wherein The segmentation loss Loss in (4) seg , is expressed as follows: Among them, G represents the liver tumor label image, S represents the predicted pixel-level segmentation map, n represents the total number of image pixels, the subscript i represents the index of the pixel in the image, α represents a constant with a value of 0.7, and ε represents a constant between 0 and 1.
5. The method according to claim 1, characterized in that The consistency loss Loss in (4) sdf is expressed as follows: Among them, represents the converted pixel-level segmentation map, S represents the predicted pixel-level segmentation map, R represents the predicted level set map, n represents the total number of image pixel points, and the subscript i represents the index of the pixel point in the image.
6. The method according to claim 1, wherein The level set loss Loss in (4) lsf is expressed as follows: Among them, G represents the liver tumor label image, R represents the predicted level set map, G' represents the true level set map, n represents the total number of image pixels, and the subscripts i and j respectively represent the indices of different pixels in the image. represents the boundary contour of the target liver tumor, S in and S out represent the internal region and the external region of the target liver tumor respectively.
7. The method according to claim 1, characterized in that, In the above (4), the liver tumor segmentation model is iteratively trained using the backpropagation method, and its implementation is as follows; (4a) Set the total number of iteration rounds to 200, the batch size to 8, the initial learning rate lr0 to 0.0001, and the learning rate changes according to the number of training iterations. Among them, t represents the current training iteration number, and t max represents the total number of training iterations. The optimizer is the Adam optimizer; (4b) Set the combined loss function as UnionLoss = Loss seg + Loss lsf + λ d Loss sdf , Among them, t represents the current training iteration number, t max represents the total number of training iterations; (4c) Input the multi-phase CT training set into the liver tumor segmentation model to obtain the pixel-level segmentation result and the level set result; (4d) According to the pixel-level segmentation result, the level set result obtained in (4c), and the liver tumor label of the training set, calculate the loss value of the model through the combined loss function; (4e) Backpropagate the loss value obtained in (4d) to update the model parameters to obtain the preliminarily trained model W; (4f) Loop through the processes of (4c)-(4e) for the preliminarily trained model W until the number of iterations reaches 200, then stop training to obtain the trained liver tumor segmentation model W'.
Citation Information
Patent Citations
Ultrasonic breast tumor automatic segmentation method based on attention enhancement improved U-shaped network
CN112785598A
Multi-temporal CT image liver tumor segmentation method based on semantic migration
CN112785605A