Pancreas automatic segmentation method and system based on abdomen CT image

By adopting an improved U-Net network in the pancreatic segmentation task, combining deformable convolution and attention gate modules, using the focus generalized dice loss function, the problems of low pancreatic segmentation accuracy and high computing resource consumption are solved, and more efficient pancreatic segmentation is achieved.

CN120047387APending Publication Date: 2025-05-27YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017566.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has problems in pancreatic segmentation tasks that are not high in accuracy, consume large computing resources, and are difficult to adapt to the characteristics of highly deformed targets.

Method used

Using a segmentation network based on U-Net, combined with a deformable convolution module and an attention gate module, the pancreas is achieved through the focus generalized dice loss function.

Benefits of technology

It improves the accuracy of pancreatic segmentation, reduces the consumption of computing resources, and can better adapt to the highly deformed characteristics of pancreatic organs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047387A_ABST
    Figure CN120047387A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an automatic pancreas segmentation method and system based on an abdominal CT image, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a pancreas CT image data set; acquiring a coarse segmentation bounding box according to the pancreas CT image data set; acquiring pancreas fine segmentation result data according to the coarse segmentation bounding box; and completing automatic pancreas segmentation according to the pancreas fine segmentation result data. According to the invention, computing resource consumption can be reduced and pancreas segmentation precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an automatic pancreas segmentation method and system based on abdominal CT images. Background Art

[0002] In recent years, with the iteration of computer hardware devices such as GPUs and the development of deep learning technology, medical image analysis has gradually entered the era of automation and precision. The U-Net network has significant advantages in pixel-level segmentation tasks of medical images. However, due to the inherent anatomical characteristics of the pancreas tissue, such as: the volume ratio of the pancreas in abdominal CT images is less than 0.5%; the contrast between the pancreas and surrounding organs is weak, resulting in blurred boundaries of the pancreas tissue; there is a high degree of heterogeneity in size, shape, and position among different patients, and the segmentation of the inflamed and tumorous pancreas is more challenging than that of the normal pancreas, etc., the accuracy of pancreas segmentation is still not ideal, and its segmentation method still has a large exploration space.

[0003] In addition, in current advanced pancreas segmentation algorithms, although fusing multi-scale features, combining context information, or adding attention mechanisms can achieve richer, more targeted, and stronger feature expressions, the fixed geometric structures of convolutional and pooling units in conventional convolutional neural networks limit the network's adaptability to the features of highly deformed targets, enabling it to sample only in fixed regions, and the commonly used Dice loss in segmentation tasks is insensitive to the accuracy of the target segmentation boundary region, which all lead to rough predicted boundaries of the pancreas tissue and impaired segmentation accuracy. At the same time, most 2D models cannot capture spatial information along the three-dimensional space and have the disadvantages of small scale and low accuracy, while 3D models require more computing resources and occupy more GPU memory, and are restricted by resource conditions in practical applications. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to provide an automatic pancreas segmentation method and system based on abdominal CT images.

[0005] The technical solution adopted by the present invention is as follows:

[0006] On the one hand, the embodiments of the present invention provide an automatic pancreas segmentation method based on abdominal CT images, and the method includes the following steps:

[0007] Obtain a pancreas CT image dataset;

[0008] Obtain a rough segmentation bounding box according to the pancreas CT image dataset;

[0009] Obtain pancreas fine segmentation result data according to the rough segmentation bounding box;

[0010] Complete automatic pancreas segmentation according to the pancreas fine segmentation result data.

[0011] Further, the obtaining of the pancreatic CT image dataset includes the following steps:

[0012] Obtain a medical image dataset containing pancreatic CT images, mark the area of the pancreas, and obtain pancreatic region label data;

[0013] According to the pancreatic region label data, perform slicing processing on each three-dimensional pancreatic CT image, and remove the slices without the pancreatic region to obtain pancreatic region sliced images;

[0014] Perform intensity value truncation processing and normalization processing on the pancreatic region sliced images to obtain a pancreatic CT image dataset.

[0015] Further, the obtaining of the rough segmentation bounding box according to the pancreatic CT image dataset includes the following steps:

[0016] Construct a segmentation network improved based on U-Net;

[0017] Set the focal generalized dice loss function;

[0018] According to the segmentation network and the focal generalized dice loss function, use multiple views for training to obtain a rough segmentation model; the views include the sagittal plane, the coronal plane, and the axial plane;

[0019] According to the rough segmentation model and the pancreatic CT image dataset, obtain a rough segmentation result;

[0020] Use the majority voting method to process the rough segmentation result to obtain fused data;

[0021] Perform bounding box cropping on the fused data to obtain a rough segmentation bounding box.

[0022] Further, the constructing of the segmentation network improved based on U-Net includes the following steps:

[0023] Set the deformable convolution module;

[0024] Set the encoder branch; the encoder branch includes a downsampling module, an encoder three-dimensional convolution block, and an encoder two-dimensional convolution block;

[0025] Set the decoder branch; the decoder branch includes an upsampling module, a decoder three-dimensional convolution block, and a decoder two-dimensional convolution block;

[0026] Set that both the encoder two-dimensional convolution block and the decoder two-dimensional convolution block include the deformable convolution module;

[0027] Set that each layer of the encoder branch is connected to the corresponding layer of the decoder branch through a skip connection;

[0028] Set an attention gate module; set the attention gate module on the skip connection between the encoder branch and the decoder branch;

[0029] Complete the construction of a segmentation network improved based on U-Net according to the encoder branch, the decoder branch and the attention gate module.

[0030] Further, the formulas used for setting the deformable convolution module include:

[0031]

[0032] where R is the convolution kernel; p n is the enumeration of the sampling area of the convolution kernel; W(p n ) is the weight at the corresponding position of the convolution kernel; p 0 represents the position on the output feature map after sampling on the input feature map; Δp n represents the added offset in the convolution kernel; X(p 0 +p n +Δp n ) represents the element value at the position p 0 +p n +Δp n on the input feature map; Y(p 0 ) represents the element value at the position p 0 on the output feature map.

[0033] Further, the formulas used for setting the focal generalized dice loss function include:

[0034]

[0035] where FGDL represents the focal generalized dice loss function; represents the weight for balancing the region size; N represents the total number of voxels in the image; i represents the voxel ordinal number in the image; ∈ represents the numerical factor for stable training; p li ∈[0,1] represents the probability value of the voxel in the network prediction; g li ∈[0,1] represents the probability value of the voxel in the gold standard; γ varies in the range of [1,3].

[0036] Further, the obtaining of the pancreatic fine segmentation result data according to the rough segmentation bounding box includes the following steps:

[0037] Obtain a fine segmentation model according to the segmentation network;

[0038] Cut the rough segmentation bounding box along the sagittal plane, the coronal plane, and the axial plane to obtain cutting data;

[0039] Input the cutting data into the fine segmentation model to obtain a view segmentation result;

[0040] According to the view segmentation result, use the weighted voting method to fuse the results of different views to obtain three-dimensional volume block data;

[0041] Optimize the three-dimensional volume block data through morphological closing operation and maximum connected component processing to obtain the pancreatic fine segmentation result data.

[0042] On the other hand, an embodiment of the present invention also provides a pancreatic automatic segmentation system based on abdominal CT images, and the system includes:

[0043] A first module for obtaining a pancreatic CT image data set;

[0044] A second module for obtaining a rough segmentation bounding box according to the pancreatic CT image data set;

[0045] A third module for obtaining pancreatic fine segmentation result data according to the rough segmentation bounding box;

[0046] A fourth module for completing pancreatic automatic segmentation according to the pancreatic fine segmentation result data.

[0047] Further, the pancreatic automatic segmentation system based on abdominal CT images further includes:

[0048] A data processing module for inputting the preprocessed CT images into a network training model and a test model;

[0049] A data optimization processing module for performing fine segmentation and re-optimization on the segmentation result of the three-dimensional volume block data through morphological closing operation and maximum connected component processing to obtain a pancreatic segmentation result.

[0050] On the other hand, an embodiment of the present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the pancreatic automatic segmentation method based on abdominal CT images as described above.

[0051] The embodiments of the present application at least include the following beneficial effects: The present application provides a pancreatic automatic segmentation method and system based on abdominal CT images. The present invention obtains a pancreatic CT image data set; obtains a rough segmentation bounding box according to the pancreatic CT image data set; obtains pancreatic fine segmentation result data according to the rough segmentation bounding box; and completes pancreatic automatic segmentation according to the pancreatic fine segmentation result data. The present invention can reduce the consumption of computing resources and improve the accuracy of pancreatic segmentation. Description of the Drawings

[0052] Figure 1 It is a flowchart of the automatic pancreas segmentation method based on abdominal CT images provided by an embodiment of the present invention;

[0053] Figure 2 It is a flowchart of the automatic pancreas CT segmentation method based on the improved U-Net provided by an embodiment of the present invention;

[0054] Figure 3 It is a network structure diagram of an automatic pancreas CT segmentation method based on the improved U-Net provided by an embodiment of the present invention;

[0055] Figure 4 It is a schematic diagram of the deformable convolution module structure provided by an embodiment of the present invention;

[0056] Figure 5 It is a schematic diagram of the attention gate module structure provided by an embodiment of the present invention. Detailed implementation manners

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.

[0058] It can be understood that the terms "first", "second", etc. used in the present application can be used in this document to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".

[0059] The terms "at least one", "a plurality", "each", "any one", etc. used in the present application, at least one includes one, two or more, a plurality includes two or more, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0061] Before elaborating on the embodiments of this application in detail, some nouns and terms involved in the embodiments of this application are first explained, and the nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0062] 1) CT image, computed tomography image, a medical imaging technique;

[0063] 2) Focal Generalized Dice Loss (FGDL), Focal Generalized Dice Loss, which is a loss function that combines Focal Loss and Generalized Dice Loss. This loss focuses on the pixels on the mask boundary, helps alleviate the class imbalance problem in pancreas segmentation to improve the segmentation accuracy of the boundary;

[0064] 3) U-Net, a deep learning network architecture;

[0065] 4) 2D, two-dimensional;

[0066] 5) 3D, three-dimensional;

[0067] 6) Deformable conv, Deformable convolution module (Deformable conv2D), both composed of a 2D convolutional offset layer, a convolutional layer, BatchNorm, and a ReLU layer; the role of the convolutional offset layer is to tell the network how to deform and how to sample the feature map. Through deformable convolution, the network can refine the extracted pancreas region and better extract the geometric perception features of the pancreas;

[0068] 7) Attention-Gated (AG), Attention-Gated mechanism, a deep learning model that can be used for medical image analysis. It can adaptively focus on target structures, even if these structures have different shapes and sizes, without introducing a large number of parameters and computational amounts, while improving the segmentation performance. Through this mechanism, the model can automatically suppress irrelevant regions in the input image and highlight the significant features useful for specific tasks. This framework is applicable to image classification and segmentation tasks and is particularly suitable for improving the detection accuracy in the field of medical imaging;

[0069] 8) BatchNorm, batch normalization, a commonly used technique in deep learning;

[0070] 9) ReLU, Rectified Linear Unit, a rectified linear unit, an activation function;

[0071] 10) ROI, Region of Interest, region of interest;

[0072] 11) Adam optimizer, Adaptive Moment Estimation, an optimization algorithm;

[0073] 12) Epoch, training cycle, which refers to the process of passing the entire training dataset through the neural network once during training;

[0074] 13) batch size, batch size, the number of samples input into the neural network in one training;

[0075] 14) NIH, National Institutes of Health, the National Institutes of Health of the United States;

[0076] 15) Conv, convolution, an operation in deep learning;

[0077] 16) Sigmoid, an S-shaped curve activation function;

[0078] 17) Gold standard, in the field of medical imaging, etc., refers to the reference standard considered to be the most accurate and reliable;

[0079] 18) Network prediction, refers to the inference process of the deep learning model based on training data and input information;

[0080] 19) 2.5D U-Net, the network consists of two parts: 3D U-Net and 2D U-Net. Based on the traditional U-Net network architecture, they are located in the upper and lower parts of the network architecture respectively, still retaining skip connections, but the encoder and decoder branches are composed of 2D and 3D convolution blocks;

[0081] 20) HU, an abbreviation for Hounsfield Units, is the standard unit for CT (Computed Tomography) images.

[0082] The following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings.

[0083] On the one hand, the embodiments of the present invention provide a method for automatic pancreas segmentation based on abdominal CT images, referring to Figure 1 , the method includes the following steps:

[0084] S100. Obtain a pancreatic CT image dataset;

[0085] S200. Obtain a rough segmentation bounding box according to the pancreatic CT image dataset;

[0086] S300. Obtain the pancreatic fine segmentation result data according to the rough segmentation bounding box;

[0087] S400. Complete the automatic segmentation of the pancreas according to the pancreatic fine segmentation result data.

[0088] The embodiment of the present invention discloses that step S100 obtains the pancreatic CT image dataset, including the following steps:

[0089] S110. Obtain a medical image dataset containing pancreatic CT images, mark the area of the pancreas, and obtain the pancreatic area label data;

[0090] S120. According to the pancreatic area label data, perform slicing processing on each three-dimensional pancreatic CT image, and eliminate the slices without the pancreatic area to obtain the pancreatic area sliced images;

[0091] S130. Perform intensity value truncation processing and normalization processing on the pancreatic area sliced images to obtain the pancreatic CT image dataset.

[0092] The embodiment of the present invention discloses that step S200 obtains a rough segmentation bounding box according to the pancreatic CT image dataset, including the following steps:

[0093] S210. Construct a segmentation network improved based on U-Net;

[0094] S220. Set the focal generalized dice loss function;

[0095] S230. According to the segmentation network and the focal generalized dice loss function, use multiple views for training to obtain a rough segmentation model; the views include the sagittal plane, the coronal plane, and the axial plane;

[0096] S240. According to the rough segmentation model and the pancreatic CT image dataset, obtain the rough segmentation result;

[0097] S250. Use the majority voting method to process the rough segmentation result to obtain the fusion data;

[0098] S260. Perform bounding box cropping on the fusion data to obtain the rough segmentation bounding box.

[0099] The embodiment of the present invention discloses that step S210 constructs a segmentation network improved based on U-Net, including the following steps:

[0100] S211. Set the deformable convolution module;

[0101] S212. Set the encoder branch; the encoder branch includes a downsampling module, an encoder three-dimensional convolution block, and an encoder two-dimensional convolution block;

[0102] S213. Set the decoder branch; the decoder branch includes an upsampling module, a decoder three-dimensional convolutional block, and a decoder two-dimensional convolutional block;

[0103] S214. Set that both the encoder two-dimensional convolutional block and the decoder two-dimensional convolutional block include a deformable convolution module;

[0104] S215. Set that each layer of the encoder branch is connected to the corresponding layer of the decoder branch through a skip connection;

[0105] S216. Set an attention gate module; set the attention gate module on the skip connection between the encoder branch and the decoder branch;

[0106] S217. Complete the construction of a segmentation network improved based on U-Net according to the encoder branch, the decoder branch, and the attention gate module.

[0107] The embodiment of the present invention discloses that the formula used for setting the deformable convolution module in step S211 includes:

[0108]

[0109] Among them, R is the convolution kernel; p n is the enumeration of the sampling area of the convolution kernel; W(p n ) is the weight at the corresponding position of the convolution kernel; p 0 represents the position on the output feature map after sampling on the input feature map; Δp n represents the added offset in the convolution kernel; X(p 0 +p n +Δp n ) represents the element value at the position p 0 +p n +Δp n on the input feature map; Y(p 0 ) represents the element value at the position p 0 on the output feature map.

[0110] The embodiment of the present invention discloses that the formula used for setting the focal generalized dice loss function in step S220 includes:

[0111]

[0112] Among them, FGDL represents the focal generalized dice loss function; represents the weight for balancing the region size; N represents the total number of voxels in the image; i represents the voxel ordinal number in the image; ∈ represents the numerical factor for stable training; p li ∈[0,1] represents the probability value of the voxel in the network prediction; gli ∈[0,1] represents the probability value of the voxel in the gold standard; γ varies within the range of [1, 3].

[0113] The embodiment of the present invention discloses that step S300 obtains the pancreatic fine segmentation result data according to the rough segmentation bounding box, including the following steps:

[0114] S310. Obtain a fine segmentation model according to the segmentation network;

[0115] S320. Cut the rough segmentation bounding box along the sagittal plane, coronal plane, and axial plane to obtain cutting data;

[0116] S330. Input the cutting data into the fine segmentation model to obtain a view segmentation result;

[0117] S340. According to the view segmentation result, use the weighted voting method to fuse the results of different views to obtain three-dimensional volume block data;

[0118] S350. Optimize the three-dimensional volume block data through morphological closing operation and maximum connected component processing to obtain the pancreatic fine segmentation result data.

[0119] On the other hand, the embodiment of the present invention also provides a pancreatic automatic segmentation system based on abdominal CT images, and the system includes:

[0120] The first module is used to obtain a pancreatic CT image dataset;

[0121] The second module is used to obtain a rough segmentation bounding box according to the pancreatic CT image dataset;

[0122] The third module is used to obtain the pancreatic fine segmentation result data according to the rough segmentation bounding box;

[0123] The fourth module is used to complete the pancreatic automatic segmentation according to the pancreatic fine segmentation result data.

[0124] The pancreatic automatic segmentation system based on abdominal CT images disclosed in the embodiment of the present invention further includes:

[0125] The data processing module is used to input the preprocessed CT images into the network training model and the test model;

[0126] The data optimization processing module is used to perform morphological closing operation and maximum connected component processing to perform fine segmentation and re-optimization on the segmentation result of the three-dimensional volume block data to obtain the pancreatic segmentation result.

[0127] As an optional implementation manner, there is also an embodiment of the present invention:

[0128] This embodiment aims to provide a pancreatic CT automatic segmentation method based on improved U-Net to solve the technical pain points such as the fixed sampling area that is not conducive to improving boundary fidelity, the poor sensitivity of Dice loss to the imbalance problem between pancreatic and background regions, the difficulty of most 2D models in fully utilizing slice spatial information, and the 3D model occupying more computing resources and GPU memory.

[0129] In order to overcome the defects of the prior art, the present invention proposes a pancreatic CT automatic segmentation method and system based on improved U-Net.

[0130] In the first aspect, the present invention constructs a coarse-to-fine two-stage pancreas segmentation framework, including:

[0131] (1) Obtain a pancreatic CT image dataset;

[0132] (2) The pancreatic CT image dataset is preprocessed and used as the input of the above-mentioned coarse segmentation model for training to obtain an optimized coarse segmentation bounding box, which is cropped as the input of the fine segmentation model. Fine segmentation training and fine segmentation optimization are then performed to obtain the pancreatic fine segmentation result, and the finally trained coarse-to-fine model is used for pancreatic segmentation.

[0133] refer to Figure 2 The specific steps of the pancreatic CT automatic segmentation method based on improved U-Net are as follows:

[0134] Step 1: Obtain a pancreatic CT image dataset, annotate the pancreatic organ, use the annotated pancreatic organ as a label, and perform corresponding preprocessing on it. To reduce the data differences caused by the medical image acquisition process and display more pancreatic details, the operation is as follows:

[0135] (1) Slice the 3D pancreatic CT and remove slices that do not contain the pancreatic area;

[0136] (2) The original intensity value of each CT image was truncated to [-100, 240] HU, and the CT values ​​below -100 were set to -100 and those above 240 were set to 240;

[0137] (3) Normalize the original CT data to obtain the preprocessed image.

[0138] Step 2: Model construction. The pancreas segmentation network follows a coarse-to-fine framework. The backbone network is a 2.5D U-Net variant of the traditional U-Net. The deformable convolution module (Deformable conv2D) and the attention gating mechanism (Attention-Gated, AG) are integrated into the network, which are used for both coarse and fine segmentation stages.

[0139] The encoder branch contains four downsampling modules. The first two layers are 3D convolutions, and the last three layers are 2D convolutions. The first two layers are 3D convolution blocks composed of two 3D convolutions (3×3×3), BatchNorm 3D, and ReLU activation functions. The last three layers are 2D convolution blocks composed of two deformable convolution modules with a size of 3×3. A max-pooling layer is set after each convolution block for downsampling operations. The max-pooling in the upper and lower parts of the network is processed on a 3D pooling layer with a size of 2×2×2 and a 2D pooling layer with a size of 2×2 respectively. After the fourth downsampling operation, the fifth layer is replaced with a deformable convolution module;

[0140] The decoder branch includes four upsampling modules, and the positions of the deformable convolution modules correspond to those in the encoder branch. That is, the 2D U-Net part at the lower part of the network architecture is a 2D convolution block composed of two deformable convolution modules with a size of 3×3, BatchNorm 2D, and ReLU activation functions. After each 2D convolution block, an upsampling operation is performed using an upsampling layer with a size of 2×2, BatchNorm 2D, a conv2D with a size of 2×2, and BatchNorm 2D. The 3D U-Net part at the upper part of the network architecture is a 3D convolution block composed of two 3D convolution layers with a size of 3×3×3, BatchNorm 3D, and ReLU activation functions. After each 3D convolution block, an upsampling operation is performed using an upsampling layer with a size of 2×2×2, BatchNorm 3D, a conv3D with a size of 2×2×2, and BatchNorm 3D. After the last 3D convolution block of the network, a 2D convolution of 1×1 is connected to obtain the predicted segmentation map. Each layer of the encoder branch is connected to the corresponding layer of the decoder branch through skip connections to introduce feature fusion at different levels;

[0141] To balance efficiency and accuracy, the deformable convolution modules replace the conventional convolution blocks in the third, fourth, and fifth layers of the encoder and the corresponding layers of the decoder in the 2.5D U-Net. The variable receptive field is used to effectively learn pancreatic features of various shapes and scales, and a deformable 2.5D U-Net for constructing an adaptive highly deformable target is built.

[0142] The attention gate module is set on the skip connection between the encoder branch and the decoder branch to eliminate irrelevant and noisy responses generated by the skip connection, strengthen the extraction of the region of interest (ROI), that is, highlight the significant features of the pancreas and automatically focus on the feature learning of the pancreas;

[0143] Step 3: Use the preprocessed dataset in Step 1 as the input. The loss function adopts the Focal Generalized Dice Loss (FGDL), and train the model constructed in Step 2 until convergence to obtain a trained coarse-to-fine model, and use this model for pancreatic segmentation.

[0144] Use multiple groups of slices obtained by cutting a three-dimensional volume block along the sagittal plane, coronal plane, and axial plane to train two groups of models (each group of models contains three models, namely M S , M C and M A for the three views) respectively for the coarse segmentation and fine segmentation stages. To prevent the model from being severely affected by the background, in the coarse segmentation stage, this invention uses slices where the pancreas occupies at least 100 pixels in the entire abdominal CT image, and combines 3 adjacent slices into a 3-channel input to fully utilize the context information between slices. Through the above-constructed coarse segmentation model obtain the segmentation results of the three views, and fuse the prediction masks of these different views into a three-dimensional volume block through majority voting. In the mode where FGDL focuses on the learning of boundary pixels, realize the re-optimization of the coarse segmentation to obtain the bounding box of the coarse segmentation. Crop the bounding box as the input of a smaller area of the fine segmentation model, and cut it again along the sagittal plane, coronal plane, and axial plane, and input it into the fine segmentation model to obtain the segmentation results of the three views. Through the weighted voting fusion strategy, give a larger weight to the plane with the short-axis slice to obtain a better two-dimensional effect and then synthesize a three-dimensional volume block. To make the pancreatic contour obtained in the fine segmentation stage more delicate, post-process the pancreatic boundary. Through morphological closing operation and the largest connected component processing, re-optimize the fine segmentation of the segmentation results of the three-dimensional volume block data.

[0145] In the model training stage, use operations such as rotation, translation, and scaling to perform data augmentation on the training data to prevent overfitting. Adopt four-fold cross-validation. In each round, take three parts as the training set, and the remaining one part as the test set, for a total of four rounds. Set the learning rate to 0.0001, adopt the Adam optimizer, train for 20 epochs in total, set the batch size to 1 for gradient update, and the loss function FGDL used is as follows:

[0146]

[0147] where, is the weight for balancing the region size, N and ∈ respectively represent the total number of voxels in the image and the numerical factor for stable training (set to 1). i represents the voxel ordinal number in the image. g li ∈[0,1], p li∈[0,1] corresponds to the probability values of the voxels in the gold standard G and the network prediction P respectively. γ varies within the range of [1, 3], and it is set to 1.5. Finally, the trained coarse-to-fine two-stage pancreatic segmentation model is used to test the CT images that have undergone the same preprocessing operations to evaluate the model segmentation performance.

[0148] In a second aspect, the present invention provides a pancreatic CT automatic segmentation system based on an improved U-Net. Specifically, it includes:

[0149] (1) A data acquisition module for acquiring a pancreatic CT image dataset;

[0150] (2) A data processing module for inputting the preprocessed CT images into a network training model and a test model;

[0151] (3) A data optimization processing module for performing morphological closing operation and maximum connected component processing, and then performing fine segmentation and optimization on the segmentation result of the three-dimensional volume block data to obtain the final pancreatic segmentation result.

[0152] The specific manners of the data acquisition module, the data processing module, and the data optimization processing module respectively correspond to the specific steps of the pancreatic CT automatic segmentation based on the improved U-Net described above.

[0153] Correspondingly, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned abdominal CT image analysis method are implemented. At the same time, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the steps of the above-mentioned abdominal CT image analysis method are implemented.

[0154] The beneficial effects of the present invention:

[0155] (1) Among 2D and 3D models, a 2.5D U-Net variant that combines 2D and 3D convolutional layers is applied to alleviate the problems that the 2D model cannot fully utilize the context information between slices and is easily confused by background regions, and the 3D model has high requirements for GPU memory, thus restricting the depth of the network and the number of feature maps, and affecting model training.

[0156] (2) In the 2.5D U-Net network, by replacing the conventional convolutional layers in the encoder branch with deformable convolutional layers, the receptive field range can be expanded, enabling the convolutional kernel to generate offsets according to the input feature map, concentrating on the target regions of interest, and being able to adapt to the characteristics of the highly deformed pancreatic organ.

[0157] (3) In the 2.5D U-Net network, an attention gate module is integrated on the skip connection to suppress the feature response in irrelevant background areas, reduce the learning of redundant features, improve the segmentation performance to a certain extent, and reduce the consumption of computing resources.

[0158] (4) In the pancreatic segmentation task of abdominal CT, there is a class imbalance problem between the foreground object and the background of the pancreas. The commonly used Dice loss is relatively insensitive to the class imbalance problem. The pancreatic boundary determines the shape of the pancreas. The use of FGDL can alleviate the class imbalance problem in pancreatic segmentation, so that the network will focus on learning boundary pixels to improve boundary accuracy and achieve re-optimization of the segmentation results of coarse segmentation.

[0159] (5) The morphological closing operation and the maximum connected domain processing are used to perform fine segmentation and optimization operations on the segmentation results of the three-dimensional volume block data. The former can eliminate incoherent segmentation results, and the latter can remove small false positive areas and smooth the pancreatic boundary information.

[0160] As an optional embodiment, the present invention also has embodiments:

[0161] The dataset used in this embodiment is the NIH pancreatic CT dataset, which contains 82 contrast-enhanced three-dimensional CT data. The voxel size of each data is 512×512×L, L∈([181,466]), the background label is 0, and the pancreas label is 1.

[0162] A pancreatic CT automatic segmentation method based on improved U-Net, comprising:

[0163] (1) Obtain a pancreatic CT image dataset;

[0164] (2) The pancreatic CT image dataset is preprocessed and used as the input of the above-mentioned coarse segmentation model for training to obtain an optimized coarse segmentation bounding box, which is cropped as the input of the fine segmentation model. Fine segmentation training and fine segmentation optimization are then performed to obtain the pancreatic fine segmentation result, and the finally trained coarse-to-fine model is used for pancreatic segmentation.

[0165] The specific steps of data processing, model construction and training of the pancreatic CT automatic segmentation method based on improved U-Net are as follows:

[0166] Step 1: Obtain and preprocess the medical CT image dataset. The specific steps are as follows:

[0167] (1) Slice the 3D pancreatic CT and remove the slices that do not contain the pancreatic area to obtain a two-dimensional slice image;

[0168] (2) Intercept the original intensity values of each CT image to [-100, 240] HU, set the CT values lower than -100 to -100, and set the CT values higher than 240 to 240;

[0169] (3) Normalize the original CT data to [0, 1] to obtain the preprocessed image.

[0170] Step 2, Figure 3 The network structure diagram of an automatic pancreatic CT segmentation method based on an improved U-Net in the embodiment of the present invention is given. The pancreatic segmentation network includes: an encoder branch, a decoder branch, Deformable conv2D, and AG. Specifically, it includes:

[0171] First, the backbone network 2.5D U-Net model is composed of an encoder, a decoder, and skip connections. The encoder contains four downsampling modules. The first two layers are 3D convolutions, and the last three layers are 2D convolutions. The first two layers are 3D convolution blocks composed of two 3D convolutions (3×3×3), BatchNorm 3D, and ReLU activation functions. The last three layers are 2D convolution blocks composed of two deformable 2D convolutions with a size of 3×3. A max pooling layer is set after each convolution block for downsampling operations. The max pooling of the upper and lower parts of the network is processed on a 3D pooling layer with a size of 2×2×2 and a 2D pooling layer with a size of 2×2 respectively. After the fourth downsampling operation, the fifth layer is replaced with a deformable convolution module. The decoder includes four upsampling modules, and the positions of the deformable convolution modules correspond to the settings of the encoder branch. That is, the 2D U-Net part at the lower part of the network architecture is a 2D convolution block composed of two deformable convolution modules with a size of 3×3, BatchNorm 2D, and ReLU activation functions. After each 2D convolution block, an upsampling operation is performed using an upsampling layer with a size of 2×2, BatchNorm 2D, a conv2D with a size of 2×2, and BatchNorm 2D. The 3D U-Net part at the upper part of the network architecture is a 3D convolution block composed of two 3D convolution layers with a size of 3×3×3, BatchNorm 3D, and ReLU activation functions. After each 3D convolution block, an upsampling operation is performed using an upsampling layer with a size of 2×2×2, BatchNorm 3D, a conv3D with a size of 2×2×2, and BatchNorm 3D. After the last 3D convolution block of the network, a 2D convolution of 1×1 is connected to obtain the predicted segmentation map. During upsampling, the feature maps output by the corresponding downsampling modules passed through the skip connections are fused, introducing feature fusion at different levels.

[0172] Replace the 3x3 convolutions in the third, fourth, and fifth layers of the 2.5D U-Net encoder branch and the convolutions on the decoder branch corresponding to the corresponding layers of the encoder branch with 3x3 deformable convolutions. Different from conventional convolutions, deformable convolutions have additional learnable offsets, and the convolution kernel no longer samples in a fixed square region, but can focus on the target regions of interest and adaptively locate objects of different shapes.

[0173] The specific construction process of the deformable convolution module is as follows: Conventional convolution can be regarded as a weighted sum of a 2D sampling grid with weights W. After regular sampling on the input feature map A, for any position p on the output feature map B 0 we can get:

[0174] In the above formula, R is the grid, i.e., the convolution kernel, and p n is the enumeration of the sampling area of the convolution kernel, and W(p n ) represents the weight at the corresponding position of the convolution kernel, X(p 0 +p n ) represents the element value at the position p 0 +p n on the input feature map, and Y(p 0 ) represents the element value at the position p 0 on the output feature map, which is obtained by convolving the convolution kernel with the input feature map.

[0175] For deformable convolution, we get:

[0176]

[0177] In the above formula, Δp n is the additional offset in the grid R. After the offset, the sampling area of the convolution kernel becomes an irregular area. Figure 4 The schematic diagram of deformable convolution is shown as follows. By replacing the 3x3 conventional convolutions in the third, fourth, and fifth layers of the above-mentioned 2.5D U-Net encoder part with deformable convolutions with learnable offsets, a deformable 2.5D U-Net for adaptively highly deformed targets is constructed.

[0178] The attention gate module scales the input feature x p through the attention coefficient α∈(0,1]. Its schematic diagram is as shown in Figure 5 . The specific construction process is as follows: Please refer to the network structure schematic diagram in Figure 3 , and take the attention gate module in the fourth-layer skip connection as an example for detailed explanation. The output result F 4 ×H 4 ×W 4 ×D 4 of the fourth layer of the encoder branch is denoted as xp Input, the output result F of the fifth layer 5 ×H 5 ×W 5 ×D 5 Denoted as g_input, as the gate signal, the two are respectively convolved by 1×1 and then added. The obtained result is successively passed through the ReLU activation function, 1×1 convolution, Sigmoid activation function (the purpose is to assign a value to each part of the feature map, that is, the attention weight), and resampled to obtain a 1D weight matrix α, and α is multiplied by x p input to obtain the new feature map x p’ , obtaining the new feature map with attention weights assigned. The introduction of this module is beneficial to highlighting the significant features of the pancreas and enabling the network to automatically focus on the feature learning of the pancreas.

[0179] Step 3: Use multiple groups of slices cut from the three-dimensional volume block along the sagittal plane, coronal plane, and axial plane to train two groups of models (each group of models contains three models, namely M S , M C and M A three views) for the coarse segmentation and fine segmentation stages respectively. In the coarse segmentation stage of the present invention, slices in which the pancreas occupies at least 100 pixels in the entire abdominal CT image are used, and 3 adjacent slices are combined into a 3-channel input to fully utilize the context information between slices. The segmentation results of the three views are obtained through the coarse segmentation model . These prediction masks of different views are fused into a three-dimensional volume block through majority voting. In the mode where FGDL focuses on the learning of boundary pixels, the coarse segmentation is re-optimized to obtain the bounding box of the coarse segmentation. The input of the smaller area of the fine segmentation model is obtained by cropping the bounding box of the coarse segmentation. The input is cut again along the sagittal plane, coronal plane, and axial plane and input into the fine segmentation model to obtain the segmentation results of the three views. Through the weighted voting fusion strategy, a larger weight is given to the plane with the short-axis slice and then synthesized into a three-dimensional volume block. Through morphological closing operation and maximum connected component processing, the segmentation result of the three-dimensional volume block data is re-optimized for fine segmentation.

[0180] For the fine segmentation re-optimization operation, the post-processing of the present invention uses morphological closing operation and maximum connected component processing to optimize the segmentation result of the three-dimensional volume block data, which plays a role in smoothing the boundary information of the pancreas. Based on the details described above, the segmentation model constructed in step 2 is trained, and the finally trained coarse-to-fine model is used for pancreas segmentation.

[0181] The code of the present invention is implemented based on the Keras framework, and the model is trained on an NVIDIA GeForce RTX 3090. During the training phase, to mitigate overfitting in the training process of the model, data augmentation operations such as random flipping, translation, and scaling are performed on the training data to improve the generalization ability of the model. Four-fold cross-validation is adopted. In each round, three parts are taken as the training set, and the remaining one part is taken as the test set, for a total of four rounds. The learning rate is set to 0.0001, the Adam optimizer is used, 20 epochs are trained, and the batch size is 1 for gradient update. The loss function is optimized using FGDL to make the model pay more attention to the segmentation effect of the target region boundary:

[0182]

[0183] Among them, is the weight for balancing the region size. N and ∈ respectively represent the total number of voxels in the image and the numerical factor for stable training (set to 1). i represents the voxel ordinal number in the image. g li ∈[0,1], p li ∈[0,1] respectively correspond to the probability values of the voxels in the gold standard G and the network prediction P. γ varies within the range of [1,3] and is set to 1.5. Different from the commonly used Dice loss, the FGDL loss focuses on the pixels on the mask boundary, which helps to alleviate the class imbalance problem in pancreatic segmentation to improve the segmentation accuracy of the boundary.

[0184] An automatic pancreatic CT segmentation system based on an improved U-Net. Specifically, it includes:

[0185] (1) A data acquisition module for acquiring a pancreatic CT image dataset;

[0186] (2) A data processing module for inputting the preprocessed CT images into a network training model and a test model;

[0187] (3) A data optimization processing module for performing morphological closing operations and maximum connected component processing on the segmentation result of the three-dimensional volume block data, and then optimizing the fine segmentation to obtain the final pancreatic segmentation result.

[0188] The specific methods of the data acquisition module, the data processing module, and the data optimization processing module respectively correspond to the specific steps of the above-mentioned automatic pancreatic CT segmentation based on the improved U-Net.

[0189] Accordingly, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned abdominal CT image analysis method are implemented. At the same time, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the steps of the above-mentioned abdominal CT image analysis method are implemented.

[0190] On the other hand, an embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute the pancreatic automatic segmentation method based on abdominal CT images as described above.

[0191] Those of ordinary skill in the art can understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0192] On the other hand, an embodiment of the present invention also provides a pancreatic automatic segmentation device based on abdominal CT images, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the pancreatic automatic segmentation method based on abdominal CT images as described above is implemented.

[0193] The processor and the memory can be connected via a bus or other means. The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory that is remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0194] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for automatic pancreas segmentation based on abdominal CT images, characterized in that: The method comprises the following steps: Obtain a pancreatic CT image dataset; According to the pancreatic CT image dataset, obtaining a coarse segmentation bounding box; Acquiring pancreas fine segmentation result data according to the coarse segmentation boundary box; According to the pancreas fine segmentation result data, automatic pancreas segmentation is completed.

2. The method according to claim 1, characterized in that: The step of acquiring a pancreatic CT image dataset comprises the following steps: Obtain a medical image dataset containing pancreatic CT images, mark the pancreatic region, and obtain pancreatic region label data; According to the pancreatic region label data, each three-dimensional pancreatic CT image is sliced, and slices not containing the pancreatic region are removed to obtain pancreatic region slice images; The pancreatic region slice image is subjected to intensity value interception processing and normalization processing to obtain a pancreatic CT image data set.

3. The method according to claim 1, characterized in that The step of obtaining a rough segmentation boundary box according to the pancreatic CT image dataset comprises the following steps: Build a segmentation network based on the improved U-Net; Set the focal generalized dice loss function; According to the segmentation network and the focal generalized dice loss function, a plurality of views are used for training to obtain a coarse segmentation model; the views include sagittal plane, coronal plane, and axial plane; Obtaining a coarse segmentation result according to the coarse segmentation model and the pancreatic CT image dataset; The rough segmentation results are processed using a majority voting method to obtain fused data; The fused data is subjected to bounding box cropping to obtain a coarse segmentation bounding box.

4. The method according to claim 3, characterized in that The construction of the segmentation network based on the improved U-Net includes the following steps: Set up the deformable convolution module; An encoder branch is set; the encoder branch includes a downsampling module, an encoder three-dimensional convolution block and an encoder two-dimensional convolution block; Setting a decoder branch; the decoder branch includes an upsampling module, a decoder three-dimensional convolution block and a decoder two-dimensional convolution block; Setting the encoder two-dimensional convolution block and the decoder two-dimensional convolution block to include the deformable convolution module; Setting each layer of the encoder branch to be connected to a corresponding layer of the decoder branch via a skip connection; Setting an attention gate module; setting the attention gate module on a jump connection between the encoder branch and the decoder branch; According to the encoder branch, the decoder branch and the attention gate module, a segmentation network based on the improvement of U-Net is constructed.

5. The method according to claim 4, characterized in that The deformable convolution module is set up, and the formula used includes: Among them, R is the convolution kernel; p n is the enumeration of the convolution kernel sampling area; W(p n ) is the weight of the corresponding position of the convolution kernel; p0 represents the position on the output feature map after sampling on the input feature map; Δp n represents the offset added to the convolution kernel; X(p0+p n +Δp n ) represents p0+p on the input feature map n +Δp n The element value at position; Y(p0) represents the element value at position p0 on the output feature map.

6. The method according to claim 3, characterized in that The formula used in setting the focal generalized dice loss function includes: Among them, FGDL stands for focal generalized dice loss function; represents the weight of the balanced region size; N represents the total number of voxels in the image; i represents the voxel number in the image; ∈ represents the numerical factor for stable training; p li ∈[0,1] represents the probability value of the voxel in the network prediction; g li ∈[0,1] represents the probability value of the voxel in the gold standard; γ varies in the range of [1,3].

7. The method according to claim 3, characterized in that The step of obtaining pancreas fine segmentation result data according to the coarse segmentation boundary box comprises the following steps: According to the segmentation network, a fine segmentation model is obtained; Cutting the rough segmentation boundary box along the sagittal plane, the coronal plane, and the axial plane to obtain cutting data; Inputting the cutting data into the fine segmentation model to obtain a view segmentation result; According to the view segmentation results, a weighted voting method is used to fuse the results of different views to obtain three-dimensional volume block data; The three-dimensional volume block data is optimized through morphological closing operation and maximum connected domain processing to obtain pancreas fine segmentation result data.

8. The automatic pancreas segmentation system based on abdominal CT images is characterized by: The system comprises: The first module is used to obtain the pancreatic CT image dataset; A second module is used to obtain a coarse segmentation boundary box according to the pancreatic CT image dataset; A third module is used to obtain pancreas fine segmentation result data according to the coarse segmentation boundary box; The fourth module is used to complete automatic pancreas segmentation according to the pancreas fine segmentation result data.

9. The system according to claim 8, characterized in that The system further comprises: A data processing module is used to input the pre-processed CT images into the network training model and the test model; The data optimization processing module is used to adopt morphological closing operation and maximum connected domain processing to perform fine segmentation and then optimize the segmentation result of the three-dimensional volume block data to obtain the pancreas segmentation result.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the automatic pancreas segmentation method based on abdominal CT images as described in any one of claims 1 to 7.