UNet liver tumor automatic segmentation method for transfer learning based on VGG16 pre-training model
By combining the transfer learning of VGG16 pre-trained model and UNet model, the problem of insufficient generalization ability of UNet model in liver tumor segmentation is solved, and efficient and accurate automatic segmentation of liver tumors is achieved, which improves diagnostic efficiency and accuracy.
Patent Information
- Application Number
- CN202510297521.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-11
AI Technical Summary
The existing UNet model cannot meet the diagnostic needs of high efficiency and high accuracy in liver tumor segmentation, especially in individual cases of different types of cancer.
The UNet liver tumor automatic segmentation method based on the VGG16 pre-trained model for transfer learning is adopted. The encoder of the VGG16 model and the decoder of the UNet model are used, combined with jump connections, and the prediction and generalization capabilities of the model are improved through transfer learning.
It improves the accuracy and efficiency of liver tumor segmentation, reduces the demand for sample size and labeling, reduces the cost of computing power, and provides patients with a more efficient diagnosis and treatment experience.
Smart Images

Figure CN120298684A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for automatically segmenting liver tumors by UNet based on transfer learning of a VGG16 pre-trained model. Background Art
[0002] At present, the detection methods for liver cancer mainly rely on medical imaging examinations such as MRI, CT, and ultrasound images, and the treatment methods for liver cancer mainly include surgical resection of cancerous parts and chemotherapy and radiotherapy. As the largest internal organ in the human body, the liver is particularly important for human metabolism, which further increases the importance of the accuracy of surgical resection for liver cancer treatment. To achieve precise resection, in addition to the judgment and operation of professional doctors, scientific and technological means are also needed. However, each set of CT images will generate hundreds to thousands of pictures. Relying on doctors to operate one by one will waste a large amount of human and material resources, and the diagnostic accuracy will also decline. Therefore, an image segmentation technology with high accuracy and high efficiency is needed to assist doctors in performing high-efficiency diagnostic work.
[0003] Image segmentation is an important basic technology in the field of computer vision, an important part of image understanding, and also an important technology applied to the field of medical image processing. It is a process of dividing a digital image into multiple image sub-regions. By simplifying or changing the representation form of the image, the image can be more easily understood. In other words, image segmentation is to attach labels to each pixel in a digital image so that pixels with the same label have certain common visual characteristics. With the continuous innovation of computer technology and artificial intelligence technology, machine learning and deep learning algorithms have gradually been applied to the field of medical image processing and achieved good results. This enables the full-automatic segmentation of medical images, not only improving the segmentation efficiency but also being more accurate than semi-automatic segmentation that combines manual operation and computer processing. The UNet model, as a currently widely used medical image segmentation model, can accurately and effectively screen and segment specific regions, so as to achieve the effect of delineating the target area and segmenting the tumor area.
[0004] However, although the UNet model has quite high accuracy, it still cannot meet the requirements of doctors in many clinical practice situations and cannot perform high-efficiency and rapid diagnosis, which indicates that the UNet model has considerable room for optimization and improvement. Different individual cases of different types of cancer require a model with stronger generalization ability to capture different types of features, so as to achieve more targeted prediction and segmentation.
[0005] VGG16, short for Visual Geometry Group 16 - layer network, was proposed by the Visual Geometry Group at the University of Oxford. It achieved excellent results in the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2014, demonstrating the powerful capabilities of deep convolutional neural networks in image classification tasks. The VGG16 model has the ability to extract rich hierarchical features through multiple layers of convolution and pooling operations, and performs particularly well in image classification tasks. By using some VGG16 modules to replace the encoder part of the UNet model, the general features learned on large datasets (such as ImageNet) can be inherited, so that good feature representations can also be obtained on small datasets. The decoder part of UNet is retained to restore spatial information for high - precision image segmentation. The general features extracted by VGG16 can be directly used for the segmentation task of UNet to improve the overall performance of the model, and the skip - connection mechanism of UNet can fuse features at different levels extracted by VGG16 to enhance the model's ability to capture details. In addition, using a pre - trained VGG16 model can greatly reduce the training time and effectively avoid the overfitting problem on small datasets. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for automatic segmentation of liver tumors in UNet based on transfer learning with a VGG16 pre - trained model, including the pre - trained model VGG16 and the model UNet. Grafting the VGG16 model, which has been trained and matured with a large number of samples in other fields, onto the UNet model for transfer learning can effectively improve the prediction ability of the UNet model, enhance the accuracy and efficiency of doctors' diagnosis, and bring a better medical treatment experience for patients.
[0007] To achieve the above - mentioned purpose, the technical solutions adopted by the present invention are as follows:
[0008] A method for automatic segmentation of liver tumors in UNet based on transfer learning with a VGG16 pre - trained model. The encoder of the automatic segmentation method uses the first four convolutional blocks in the VGG16 pre - trained model, the bottom convolutional block uses the bottom convolutional block of the UNet model, and the decoder uses the decoder of the UNet model, including the following steps:
[0009] S1, the encoder receives the input image data;
[0010] S2, the encoder gradually extracts the features of the input image and reduces the spatial resolution, screening out effective and representative image feature information;
[0011] S3, after the bottom convolutional block performs convolutional processing on the image feature information output by the encoder, it is input into the decoder;
[0012] At S4, the decoder performs an upsampling operation and then goes through a 1×1 convolutional kernel to generate the segmentation result.
[0013] Preferably, in the four convolutional blocks of the encoder, the first convolutional block contains two convolutional layers, block1conv1 and block1conv2, each having 64 3×3 convolutional kernels; the second convolutional block contains two convolutional layers, block2conv1 and block2conv2, each having 128 3×3 convolutional kernels; the third convolutional block contains three convolutional layers, block3conv1, block3conv2, and block3conv3, each having 256 3×3 convolutional kernels; and the fourth convolutional block contains three convolutional layers, block4conv1, block4conv2, and block4conv3, each having 512 3×3 convolutional kernels.
[0014] Preferably, in S3, the bottom convolutional block contains two convolutional layers, conv6 and conv7, each having 1024 3×3 convolutional kernels.
[0015] Preferably, in S4, the decoder structure includes four upsampling modules, and each upsampling module sequentially contains: a 2×2 transposed convolution and two convolutional layers.
[0016] Preferably, in S4, in the first upsampling part, a transposed convolution operation is first performed on conv7 to obtain up1, up1 is skip-connected to block4conv3 in the encoder part, and the convolutional layers conv8 and conv9 each have 512 3×3 convolutional kernels; in the second upsampling part, a transposed convolution operation is first performed on conv9 to obtain up2, up2 is skip-connected to block3conv3 in the encoder part, and the convolutional layers conv10 and conv11 each have 256 3×3 convolutional kernels; in the third upsampling part, a transposed convolution operation is first performed on conv11 to obtain up3, up3 is skip-connected to block2conv2 in the encoder part, and the convolutional layers conv12 and conv13 each have 128 3×3 convolutional kernels; in the fourth upsampling part, a transposed convolution operation is first performed on conv13 to obtain up4, up4 is skip-connected to block1conv2 in the encoder part, and the convolutional layers conv14 and conv15 each have 64 3×3 convolutional kernels.
[0017] Preferably, in S4, a 1×1 convolutional kernel makes the output image size return to the input image size.
[0018] The beneficial effects of the present invention are as follows: It includes the pre-trained model VGG16 and the model UNet. By grafting the VGG16 model, which has been trained and matured with a large number of samples in other fields, onto the UNet model for transfer learning, the prediction ability of the UNet model can be effectively improved, its generalization ability can be enhanced, and the requirements for the number of samples and annotation needs can be reduced. This can effectively reduce the computing power cost, improve the accuracy and efficiency of doctor diagnosis, and bring a better diagnosis and treatment experience for patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is the overall flowchart of the present invention.
[0020] Figure 2 It is the detailed flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention.
[0022] As Figure 1 - Figure 2 shown, a method for automatically segmenting liver tumors of UNet based on transfer learning with a VGG16 pre-trained model. The encoder of the automatic segmentation method uses the first four convolutional blocks in the VGG16 pre-trained model, the bottom convolutional block uses the bottom convolutional block of the UNet model, and the decoder uses the decoder of the UNet model, including the following steps:
[0023] S1, the encoder receives the input image data.
[0024] The input data is the preprocessed CT images of patients and the supporting mask images, which are DICOM files, and the size of the input data is (512, 512, 1).
[0025] S2, the encoder gradually extracts the features of the input image and reduces the spatial resolution, and filters out the effective and representative image feature information.
[0026] The encoder structure includes four convolutional blocks, and the four convolutional blocks belong to the first four convolutional blocks of the VGG16 pre-trained model. The VGG16 pre-trained model, as one of the main parts of the algorithm, mainly serves as the encoder part. This part is mainly used to extract the features of the image. It gradually reduces the size of the image through convolution, activation, and max pooling operations in multiple convolutional blocks, and retains deeper feature information, so as to extract important features.
[0027] The model loss function selects "binary_crossentropy", the optimization function selects "adam", and the evaluation metric selects "binary_accuracy" binary accuracy rate.
[0028] Among the four convolutional blocks, the first convolutional block contains two convolutional layers, block1conv1 and block1conv2, each having 64 3*3 convolutional kernels. The second convolutional block contains two convolutional layers, block2conv1 and block2conv2, each having 128 3*3 convolutional kernels. The third convolutional block contains three convolutional layers, block3conv1, block3conv2, and block3conv3, each having 256 3*3 convolutional kernels. The fourth convolutional block contains three convolutional layers, block4conv1, block4conv2, and block4conv3, each having 512 3*3 convolutional kernels.
[0029] Enter the encoder part, which is the part of the VGG16 pre-trained model. Pass through the four convolutional blocks in sequence, and the image size becomes (512,512,64), (256,256,128), (128,128,256), (64,64,512) in turn. The steps that the encoder image goes through show that the original image is gradually extracting features through convolution and pooling, and the size of the feature map is also gradually decreasing. And for each convolutional block passed through, the size of the feature image is reduced by half.
[0030] S3, after the bottom convolutional block performs convolution processing on the image feature information output by the encoder, it is input into the decoder.
[0031] The two convolutional layers, conv6 and conv7, contained in the bottom convolutional block each have 1024 3*3 convolutional kernels.
[0032] The activation function used by all convolutional layers involved in the four convolutional blocks of the encoder and the bottom convolutional block is ReLU, the padding is selected as "same", and the convolutional kernel weight initialization method is "he_normal".
[0033] After passing through the bottom convolutional block, the image size becomes (32,32,1024).
[0034] S4, the decoder performs an upsampling operation and then passes through 1 1*1 convolutional kernel to generate the segmentation result.
[0035] The decoder part of the UNet model, as another main part of the algorithm, is connected to the encoder part through skip connections. The role of the decoder is to gradually restore the features extracted by the encoder into a segmentation result of the same size as the input image. This process uses the upsampling technique, and the features of the corresponding layers in the encoder are spliced into the decoder through skip connections, retaining more details. The role of the skip connection is to ensure that the decoder can utilize the features in the encoder during upsampling, prevent detail loss, and further improve the segmentation accuracy. The combination of the VGG16 model and UNet can make full use of the advantages of both, reduce the training time and data requirements, improve the model performance, and is suitable for medical image processing scenarios with limited data volume but high-precision segmentation requirements.
[0036] The decoder structure includes four upsampling modules. Each upsampling module sequentially contains: a 2*2 transposed convolution, two convolutional layers. The purpose of the transposed convolution is to double the size of the feature map. In the first upsampling part, first perform a transposed convolution operation on conv7 to obtain up1. up1 is connected to block4conv3 in the encoder part through a skip connection. The convolutional layers conv8 and conv9 each have 512 3*3 convolutional kernels; in the second upsampling part, first perform a transposed convolution operation on conv9 to obtain up2. up2 is connected to block3conv3 in the encoder part through a skip connection. The convolutional layers conv10 and conv11 each have 256 3*3 convolutional kernels; in the third upsampling part, first perform a transposed convolution operation on conv11 to obtain up3. up3 is connected to block2conv2 in the encoder part through a skip connection. The convolutional layers conv12 and conv13 each have 128 3*3 convolutional kernels; in the fourth upsampling part, first perform a transposed convolution operation on conv13 to obtain up4. up4 is connected to block1conv2 in the encoder part through a skip connection. The convolutional layers conv14 and conv15 each have 64 3*3 convolutional kernels. The activation function used in each convolutional layer in the four upsampling modules is ReLU, the padding is selected as "same", and the convolutional kernel weight initialization method is "he_normal".
[0037] The decoder part (the decoder part of the UNet model) gradually restores the size of the feature map through transposed convolution and convolution, and combines skip connections to fuse multi-scale features so that the output image contains as much rich information as possible. The image size changes from (32, 32, 1024) to (64, 64, 512), (128, 128, 256), (256, 256, 128), (512, 512, 64) in turn. Each time it passes through an upsampling layer, the size of the feature image increases to twice the original.
[0038] After passing through four upsampling modules, the image will then go through a 1*1 convolutional kernel using the Sigmoid activation function, so that the size of the output image returns to (512, 512, 1), which is the same as the input image.
Claims
1. An automatic segmentation method for liver tumors of UNet based on transfer learning of VGG16 pre-trained model, characterized in that The encoder of the automatic segmentation method uses the first four convolutional blocks in the VGG16 pre-trained model. The bottom convolutional block uses the bottom convolutional block of the UNet model, and the decoder uses the decoder of the UNet model, including the following steps: S1, the encoder receives the input image data; S2, the encoder gradually extracts the features of the input image and reduces the spatial resolution, screening out effective and representative image feature information; S3, after the bottom convolutional block performs convolutional processing on the image feature information output by the encoder, it is input into the decoder; S4, the decoder performs an upsampling operation, and then goes through a 1*1 convolutional kernel to generate the segmentation result.
2. The automatic segmentation method of UNet liver tumors based on transfer learning using the VGG16 pre-trained model according to claim 1, characterized in that, In the four convolutional blocks of the encoder, the first convolutional block contains two convolutional layers, block1conv1 and block1conv2, each having 64 3*3 convolutional kernels. The second convolutional block contains two convolutional layers, block2conv1 and block2conv2, each having 128 3*3 convolutional kernels. The third convolutional block contains three convolutional layers, block3conv1, block3conv2, and block3conv3, each having 256 3*3 convolutional kernels. The fourth convolutional block contains three convolutional layers, block4conv1, block4conv2, and block4conv3, each having 512 3*3 convolutional kernels.
3. The UNet liver tumor automatic segmentation method based on transfer learning using the VGG16 pre-trained model according to claim 2. In S3, the bottom convolutional block contains two convolutional layers, conv6 and conv7, each having 1024 3*3 convolutional kernels.
4. The automatic segmentation method of UNet liver tumors based on transfer learning with a pre-trained VGG16 model according to claim 3, characterized in that, In S4, the decoder structure includes four upsampling modules, and each upsampling module sequentially contains: a 2*2 transposed convolution and two convolutional layers.
5. The automatic segmentation method of UNet liver tumors based on transfer learning using the VGG16 pre-trained model according to claim 4, characterized in that In S4, in the first upsampling part, first perform a transposed convolution operation on conv7 to obtain up1. up1 is skip-connected to block4conv3 in the encoder part. The convolutional layers conv8 and conv9 each have 512 3*3 convolutional kernels. In the second upsampling part, first perform a transposed convolution operation on conv9 to obtain up2. up2 is skip-connected to block3conv3 in the encoder part. The convolutional layers conv10 and conv11 each have 256 3*3 convolutional kernels. In the third upsampling part, first perform a transposed convolution operation on conv11 to obtain up3. up3 is skip-connected to block2conv2 in the encoder part. The convolutional layers conv12 and conv13 each have 128 3*3 convolutional kernels. In the fourth upsampling part, first perform a transposed convolution operation on conv13 to obtain up4. up4 is skip-connected to block1conv2 in the encoder part. The convolutional layers conv14 and conv15 each have 64 3*3 convolutional kernels.
6. The automatic segmentation method of UNet liver tumors based on transfer learning with the VGG16 pre-trained model according to claim 1, characterized in that, In S4, a 1*1 convolutional kernel makes the output image size return to the input image size.