Liver tumor automatic segmentation method based on UNet model fused with SENet attention mechanism

By integrating the SENet attention mechanism in the UNet model, the feature extraction ability of liver tumor segmentation is enhanced, and the UNet model's efficiency and accuracy in liver tumor segmentation is solved, achieving more efficient diagnosis and better diagnosis and treatment experience.

CN120298682APending Publication Date: 2025-07-11XI AN JUNENG MEDICAL ENGINEERING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510296385.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing UNet model cannot meet the high-efficiency rapid diagnosis needs in liver tumor segmentation, cannot adapt to individual cases of different types of cancer, and lacks strong generalization ability, resulting in a decrease in diagnostic accuracy.

Method used

The automatic segmentation method of liver tumors fusion with SENet attention mechanism is adopted based on the UNet model. By introducing the SENet attention mechanism into the encoder, the feature extraction ability is enhanced, and combined with the upsampling and jump connection of the decoder, high-precision tumor segmentation results are generated.

Benefits of technology

It improves the targeted identification and prediction capabilities of tumor areas, improves the accuracy and efficiency of doctors' diagnosis, and provides a better diagnosis and treatment experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298682A_ABST
    Figure CN120298682A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic liver tumor segmentation method based on a UNet model fused with an SENet attention mechanism. The method comprises the following steps that an encoder of the UNet model receives input image data; the method comprises the following steps: gradually extracting the features of an input image through an encoder of a UNet model, reducing the spatial resolution, enabling the encoder to specifically extract the features of the input image with higher resolution capability by using an SENet attention mechanism, and screening out effective representative image feature information; the bottom convolution block performs convolution processing on the image feature information output by the encoder and then inputs the image feature information into the decoder; and performing up-sampling operation through a decoder of the UNet model, and generating a segmentation result through a 1 * 1 convolution kernel. According to the method, the subject Unet model and the channel attention mechanism SENet are included, and the targeted recognition and prediction capability of the model on the tumor existence area can be effectively improved through the attention mechanism, so that the diagnosis accuracy and efficiency of a doctor are improved, and better diagnosis and treatment experience is brought to a patient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to an automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism. Background Art

[0002] At present, the detection methods for liver cancer mainly rely on medical imaging examinations such as MRI, CT, and ultrasound images, and the treatment methods for liver cancer mainly include surgical resection of the cancerous part and chemotherapy and radiotherapy. As the largest internal organ in the human body, the liver is particularly important for human metabolism, which further increases the importance of the surgical resection accuracy for liver cancer treatment. To achieve precise resection, in addition to the judgment and operation of professional doctors, scientific and technological means are also required. However, each set of CT images will generate hundreds to thousands of pictures. Relying on doctors to operate one by one will waste a large amount of human and material resources, and the diagnostic accuracy will also decrease. Therefore, an image segmentation technology with high accuracy and high efficiency is needed to assist doctors in performing high-efficiency diagnostic work.

[0003] Image segmentation is an important basic technology in the field of computer vision, an important part of image understanding, and also an important technology applied to the field of medical image processing. It is a process of subdividing digital images into multiple image sub-regions, which can make the images easier to understand by simplifying or changing the representation form of the images. In other words, image segmentation is to attach labels to each pixel in the digital image so that pixels with the same label have certain common visual characteristics. With the continuous innovation of computer technology and artificial intelligence technology, machine learning and deep learning algorithms have gradually been applied to the field of medical image processing and achieved good results. This enables the full-automatic segmentation of medical images, which not only improves the segmentation efficiency but also has higher accuracy than the semi-automatic segmentation that combines manual operation and computer processing. The UNet model, as a widely used medical image segmentation model at present, can accurately and effectively screen and segment specific regions, so as to achieve the effect of delineating the target area and segmenting the tumor area.

[0004] However, although the UNet model has quite high accuracy, it still cannot meet the requirements of doctors in many clinical practice situations and cannot perform high-efficiency and rapid diagnosis, which indicates that the UNet model has considerable room for optimization and improvement. Different individual cases of different types of cancers require a model with stronger generalization ability to capture different types of features, so as to achieve more targeted prediction and segmentation. SENet (Squeeze-and-Excitation Networks) enhances the feature extraction ability of the UNet model and thus its prediction ability by introducing an attention mechanism in the channel dimension. Summary of the Invention

[0005] The object of the present invention is to provide an automatic liver tumor segmentation method based on the UNet model fused with the SENet attention mechanism, including the main UNet model and the channel attention mechanism SENet. The attention mechanism can effectively improve the targeted recognition and prediction ability of the model for the tumor presence area (i.e., the target area), thereby improving the accuracy and efficiency of doctor diagnosis and bringing a better diagnosis and treatment experience for patients.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] An automatic liver tumor segmentation method based on the UNet model fused with the SENet attention mechanism, comprising the following steps:

[0008] S1, the encoder of the UNet model receives the input image data;

[0009] S2, the encoder of the UNet model gradually extracts the features of the input image and reduces the spatial resolution. The SENet attention mechanism is used to make the encoder more discriminative in extracting the features of the input image, and the effective and representative image feature information is screened out;

[0010] S3, after the bottom convolution block performs convolution processing on the image feature information output by the encoder, it is input into the decoder;

[0011] S4, through the decoder of the UNet model for upsampling operation, and then passing through a 1*1 convolution kernel, the segmentation result is generated.

[0012] Preferably, in S2, the encoder structure contains four convolution blocks, and each of the four convolution blocks contains two convolution layers, an attention mechanism module and a pooling part in sequence.

[0013] Preferably, in S2, when passing through the four convolution blocks, in each convolution block, the data enters the SENet attention mechanism module after every two convolution layers, and the output data after passing through the attention mechanism module enters the pooling part.

[0014] Preferably, in S2, the two convolution layers conv1 and conv2 in the first convolution block each have 64 3*3 convolution kernels, the two convolution layers conv3 and conv4 in the second convolution block each have 128 3*3 convolution kernels, the two convolution layers conv5 and conv6 in the third convolution block each have 256 3*3 convolution kernels, and the two convolution layers conv7 and conv8 in the fourth convolution block each have 512 3*3 convolution kernels.

[0015] Preferably, in S2, the SENet attention mechanism is divided into a compression part and an excitation part. The compression part uses "GlobalAveragePooling2D()" for global average pooling. The excitation part uses the non-linear activation functions ReLU and Sigmoid in two fully connected layers respectively, and both select "he_normal" as the convolution kernel weight initialization method.

[0016] Preferably, in S2, all pooling is 2*2 max pooling.

[0017] Preferably, in S3, the bottom convolution block has only two convolutional layers. The two convolutional layers conv9 and conv10 contained in the bottom convolution block each have 1024 3*3 convolutional kernels.

[0018] Preferably, in S4, the decoder structure contains four upsampling modules. Each upsampling module sequentially contains: a 2*2 transposed convolution, and two convolutional layers.

[0019] Preferably, in S4, in the first upsampling part, a transposed convolution operation is first performed on conv10 to obtain up1. up1 is skip-connected to conv8 in the encoder part. conv11 and conv12 each have 512 3*3 convolutional kernels; in the second upsampling part, a transposed convolution operation is first performed on conv12 to obtain up2. up2 is skip-connected to conv6 in the encoder part. conv13 and conv14 each have 256 3*3 convolutional kernels; in the third upsampling part, a transposed convolution operation is first performed on conv14 to obtain up3. up3 is skip-connected to conv4 in the encoder part. conv15 and conv16 each have 128 3*3 convolutional kernels; in the fourth upsampling part, a transposed convolution operation is first performed on conv16 to obtain up4. up4 is skip-connected to conv2 in the encoder part. conv17 and conv18 each have 64 3*3 convolutional kernels.

[0020] Preferably, in S4, 1 1*1 convolutional kernel makes the output image size return to the input image size.

[0021] The beneficial effects of the present invention are as follows: The targeted recognition and prediction ability of the model for the tumor presence area (i.e., the target area) can be effectively improved through the attention mechanism, thereby improving the accuracy and efficiency of doctors' diagnosis and bringing a better medical experience for patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the overall flowchart of the present invention.

[0023] Figure 2This is the detailed flowchart of the present invention. Specific embodiments

[0024] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention.

[0025] As Figure 1 - Figure 2 shown, an automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism includes the following steps:

[0026] S1, the encoder of the UNet model receives the input image data;

[0027] The input data is the preprocessed patient CT image and the supporting mask image, which is a DICOME file. The size of the input image data is (512, 512, 1), representing a single-channel grayscale image of 512*512.

[0028] S2, the encoder (Encoder) of the UNet model gradually extracts the features of the input image and reduces the spatial resolution. The SENet attention mechanism is used to make the encoder more discriminative in extracting the features of the input image, and filter out effective and representative image feature information;

[0029] The UNet model, as the main part of the algorithm, is mainly composed of an encoder and a decoder, and is connected by skip connections. The encoder part is similar to a traditional convolutional neural network and is mainly used to extract the features of the image. It gradually reduces the size of the image through multiple convolutional operations and retains deeper feature information. Each layer will perform convolution and activation to extract important features.

[0030] The loss function of the UNet model is selected as "binary_crossentropy", the optimization function is selected as "adam", and the evaluation metric is selected as "accuracy" correct rate.

[0031] The encoder structure contains four convolutional blocks. Each of the four convolutional blocks contains two convolutional layers, an attention mechanism module and a pooling part in sequence. When passing through the four convolutional blocks, in each convolutional block, the data enters the SENet attention mechanism module after every two convolutional layers, and the output data after passing through the attention mechanism module enters the pooling part. All poolings are 2*2 max pooling to output a feature map with half the size.

[0032] The first convolutional block contains two convolutional layers, conv1 and conv2, each with 64 3x3 convolutional kernels. The second convolutional block contains two convolutional layers, conv3 and conv4, each with 128 3x3 convolutional kernels. The third convolutional block contains two convolutional layers, conv5 and conv6, each with 256 3x3 convolutional kernels. The fourth convolutional block contains two convolutional layers, conv7 and conv8, each with 512 3x3 convolutional kernels.

[0033] The SENet attention mechanism, as an auxiliary part added to the UNet algorithm, is added to the encoder in the UNet model in this algorithm. It consists of a squeeze part and an excitation part. The squeeze part uses "GlobalAveragePooling2D()" for global average pooling to obtain the global information of each channel of the input feature map (i.e., the average value of each channel). This process compresses the feature map of each channel into a scalar, thereby extracting the global information of that channel. The excitation part uses the nonlinear activation functions ReLU and Sigmoid in two fully connected layers respectively, and both choose "he_normal" as the convolutional kernel weight initialization method. The excitation operation learns the channel weights through two fully connected (Dense) layers. The first Dense layer is used for dimensionality reduction and nonlinear transformation, and the second Dense layer is used to restore the channel dimension and generate weights.

[0034] After going through four convolutional blocks, the image size becomes (512, 512, 64), (256, 256, 128), (128, 128, 256), and (64, 64, 512) in sequence. This shows that through convolution and pooling, features of the original image are gradually extracted, and the size of the feature map is also gradually decreasing. And after each convolutional block, the size of the feature image is reduced by half.

[0035] S3. After the bottom convolutional block performs convolutional processing on the image feature information output by the encoder, it is input into the decoder.

[0036] The bottom convolutional block has no attention mechanism module and pooling part, only two convolutional layers. The two convolutional layers, conv9 and conv10, contained in the bottom convolutional block each have 1024 3x3 convolutional kernels.

[0037] The activation function used in all convolutional layers involved in the four convolutional blocks and the bottom convolutional block in the encoder structure is ReLU, and padding is selected as "same" to ensure that the size of the output feature map is the same as the input size. At the same time, the convolutional kernel weight initialization method is "he_normal".

[0038] After passing through the bottom convolutional block, the image size becomes (32, 32, 1024).

[0039] S4. Perform upsampling operations through the decoder of the UNet model, and then go through a 1×1 convolutional kernel to generate the segmentation result.

[0040] The role of the decoder is to gradually restore the features extracted by the encoder into a segmentation result of the same size as the input image. This process uses upsampling technology, and the features of the corresponding layers in the encoder are spliced into the decoder through skip connections, retaining more details. The role of skip connections is to ensure that the decoder can utilize the features in the encoder during upsampling, prevent the loss of details, and further improve the accuracy of segmentation.

[0041] The decoder structure contains four upsampling modules. Each upsampling module contains, in sequence: a 2×2 transposed convolution, two convolutional layers. The purpose of the transposed convolution is to double the size of the feature map. Among them, in the first upsampling part, first perform a transposed convolution operation on conv10 to obtain up1, and up1 makes a skip connection with conv8 in the encoder part. conv11 and conv12 each have 512 3×3 convolutional kernels; in the second upsampling part, first perform a transposed convolution operation on conv12 to obtain up2, and up2 makes a skip connection with conv6 in the encoder part. conv13 and conv14 each have 256 3×3 convolutional kernels; in the third upsampling part, first perform a transposed convolution operation on conv14 to obtain up3, and up3 makes a skip connection with conv4 in the encoder part. conv15 and conv16 each have 128 3×3 convolutional kernels; in the fourth upsampling part, first perform a transposed convolution operation on conv16 to obtain up4, and up4 makes a skip connection with conv2 in the encoder part. conv17 and conv18 each have 64 3×3 convolutional kernels. The activation function used in each convolutional layer in the four upsampling modules is ReLU, the padding is selected as "same", and the convolutional kernel weight initialization method is "he_normal".

[0042] In the decoder part of the model, the size of the feature map is gradually restored through transposed convolution and convolution, and multi-scale features are fused through skip connections to make the output image contain as rich information as possible. The image size changes from (32, 32, 1024) to (64, 64, 512), (128, 128, 256), (256, 256, 128), (512, 512, 64) in sequence. After passing through each upsampling layer, the size of the feature image increases to twice the original size.

[0043] After passing through four upsampling layers, finally go through a 1*1 convolutional kernel and use the Sigmoid activation function to make the output image size return to (512, 512, 1), which is the same as the input image size.

Claims

1. An automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism, characterized in that It includes the following steps: S1, the encoder of the UNet model receives the input image data; S2, the encoder of the UNet model gradually extracts the features of the input image and reduces the spatial resolution. The SENet attention mechanism is used to make the encoder extract the features of the input image more discriminatively, and filter out the effective and representative image feature information; S3, after the bottom convolution block performs convolution processing on the image feature information output by the encoder, it is input into the decoder; S4, through the decoder of the UNet model, an upsampling operation is performed, and then it goes through a 1*1 convolution kernel to generate the segmentation result.

2. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 1, characterized in that, In S2, the encoder structure contains four convolution blocks, and each of the four convolution blocks sequentially contains two convolution layers, an attention mechanism module, and a pooling part.

3. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 2, characterized in that In S2, when going through the four convolution blocks, in each convolution block, the data enters the SENet attention mechanism module after every two convolution layers, and the output data after passing through the attention mechanism module enters the pooling part.

4. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 3, characterized in that, In S2, the two convolution layers conv1 and conv2 in the first convolution block each have 64 3*3 convolution kernels, the two convolution layers conv3 and conv4 in the second convolution block each have 128 3*3 convolution kernels, the two convolution layers conv5 and conv6 in the third convolution block each have 256 3*3 convolution kernels, and the two convolution layers conv7 and conv8 in the fourth convolution block each have 512 3*3 convolution kernels.

5. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 3, wherein In S2, the SENet attention mechanism is divided into a compression part and an excitation part. The compression part uses "GlobalAveragePooling2D()" for global average pooling. The excitation part uses the non-linear activation functions ReLU and Sigmoid in two fully connected layers respectively, and both choose "he_normal" as the convolution kernel weight initialization method.

6. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 3, wherein In S2, all pooling is 2*2 max pooling.

7. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 4, characterized in that, In S3, the bottom convolution block has only two convolution layers. The two convolution layers conv9 and conv10 in the bottom convolution block each have 1024 3*3 convolution kernels.

8. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 7, wherein, In S4, the decoder structure contains four upsampling modules, and each upsampling module sequentially contains: a 2*2 transposed convolution, two convolution layers.

9. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 8, characterized in that, In S4, in the first upsampling part, first perform a transposed convolution operation on conv10 to obtain up1. up1 is skip-connected to conv8 in the encoder part. conv11 and conv12 each have 512 3×3 convolutional kernels. In the second upsampling part, first perform a transposed convolution operation on conv12 to obtain up2. up2 is skip-connected to conv6 in the encoder part. conv13 and conv14 each have 256 3×3 convolutional kernels. In the third upsampling part, first perform a transposed convolution operation on conv14 to obtain up3. up3 is skip-connected to conv4 in the encoder part. conv15 and conv16 each have 128 3×3 convolutional kernels. In the fourth upsampling part, first perform a transposed convolution operation on conv16 to obtain up4. up4 is skip-connected to conv2 in the encoder part. conv17 and conv18 each have 64 3×3 convolutional kernels.

10. The automatic liver tumor segmentation method based on the UNet model integrating the SENet attention mechanism according to claim 1, characterized in that, In S4, 1 1×1 convolutional kernel makes the output image size return to the input image size.