A full-heart segmentation method, device, equipment and medium based on computer tomography 3D images

Through the improved 3D Faster R-CNN and 3D U-Net network combined with CIoU loss function, the problem of waste of computing resources and inaccurate bounding box loss function in full heart segmentation is solved, and efficient and accurate full heart automatic segmentation is achieved.

CN115760894BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211532394.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-07-22
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

In the prior art, there is a problem that computational resource waste and segmentation quality depends on the extraction quality of the region of interest in full-cardiac segmentation, and the bounding box loss function is inaccurately affecting the segmentation accuracy.

Method used

The improved 3D Faster R-CNN network is used for detection and positioning, and combined with P3D ResNet and FPN networks to enhance feature learning capabilities, optimize bounding boxes using CIoU loss function, and segmentation combined with the improved 3D U-Net network to improve segmentation performance through residual links and deep supervision paths.

Benefits of technology

It realizes high-precision and fast automatic segmentation of the whole heart, improves segmentation efficiency and accuracy, and reaches the most advanced level in the field, and takes only about 6 seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760894B_ABST
    Figure CN115760894B_ABST
Patent Text Reader

Abstract

The present invention discloses a full heart segmentation method, device, equipment and medium based on computer tomography 3D images, relating to the field of medical image segmentation. The present invention proposes a two-stage segmentation strategy for heart segmentation. In the first stage, the Faster R-CNN network is used to detect the bounding box of the heart. In the second stage, the original cardiac CT image aligned with the bounding box is input into 3D U-Net for heart substructure segmentation. In addition, the present invention also redefines the bounding box loss function and adopts the CIoU loss function. The experimental results show that this scheme has achieved the most advanced segmentation effect on the 2017 Multi-Modal Whole Heart Segmentation Challenge (MM-WHS) dataset, reaching an average Dice score of 91.1%. At the same time, the segmentation time of a single cardiac CT image has been significantly improved from several minutes to less than 6 seconds, realizing faster and more accurate automatic full heart segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image segmentation, and in particular to a full-heart segmentation method, device, equipment and medium based on 3D images of computed tomography. Background Art

[0002] Full-heart segmentation refers to extracting the shape and volume of the heart substructures, including seven parts, namely: left ventricle (LV), left atrium (LA), left ventricular myocardium (Myo), right ventricle (RV), right atrium (RA), ascending aorta (AA), and pulmonary artery (PA). In clinical diagnosis, it is often necessary to obtain the functional parameter indicators of the heart. Improving the segmentation accuracy of the heart substructures helps to improve the reliability of calculating the cardiac structure function parameters.

[0003] With the rapid development of convolutional neural networks in intelligent medical image computing, the method based on the 3D U-Net network has shown powerful segmentation capabilities. In recent years, directly using 3D images as the input of the 3D U-Net network has gradually become a research trend in this field. However, it is usually used together with the tiling strategy, that is, the entire image is divided into small slices for separate processing, and finally the segmentation results are merged. The foreground information in the heart image only accounts for a very small proportion, which will lead to a huge waste of computing resources. To solve the above problems, a method of two cascaded 3D U-Net networks has been proposed to segment the entire heart structure by further segmenting the region of interest (RoI) dynamically extracted in the first stage. Although the original resolution is retained, the final segmentation quality depends on the extraction quality of the RoI extracted in the first stage. In addition, scholars applied 3D U-Net as the network model and combined principal component analysis as a data augmentation technique to change the input and output of the network. The disadvantage of this method is that the manual selection of the principal components will affect the effect of PCA data augmentation, thus affecting the segmentation quality.

[0004] In recent years, the detection-based segmentation method has shown its superiority in medical image segmentation. An improved Faster R-CNN shows good localization accuracy and time efficiency. It obtains multiple proposal boxes by inputting the image into the region proposal network (RPN), and then inputs the multiple proposal boxes into the classification regression network to obtain the classification results and the target bounding boxes. On this basis, Mask R-CNN proposed by scholars adds a branch for predicting the target mask after the RPN structure of Faster R-CNN, and has been proven to have strong generality in medical image segmentation. In terms of the application in the field of medical images, Mask R-CNN is combined with the ray casting volume rendering algorithm to realize 3D diagnosis of lung nodule detection and segmentation. Compared with segmenting lung nodules, the entire heart segmentation task is more complex and challenging for the segmentation head of Mask R-CNN.

[0005] The previous method consisted of a detection and localization module based on Faster R-CNN and a segmentation module based on 3D U-Net, achieving high-precision automatic whole-heart segmentation. This method retained the smooth L1 loss function of the bounding box loss function in Faster R-CNN. However, in practice, the intersection over union (IoU) loss function was used to predict the bounding box. Although multiple bounding boxes had the same loss value, their IoU could vary significantly.

[0006] Therefore, the present invention proposes a fully automatic and efficient segmentation method, device, equipment, and medium for cardiac CT images. Summary of the Invention

[0007] Aiming at the above problems, the present invention aims to provide a fully automatic and efficient segmentation method, device, equipment, and medium for cardiac CT images, considering the distance between the center points of the ground truth bounding box and the predicted bounding box, as well as the aspect ratio of the widths and heights of the two bounding boxes. In addition, considering the difference between the image size and the frame size of cardiac CT images, the feature extraction network ResNet was replaced with P3D ResNet. And an FPN network was added after the P3D ResNet network, which could fuse features of different resolutions, enhance the feature learning ability, adapt to cardiac images of different scales, and reach the state-of-the-art level in the task of fully automatic and rapid segmentation of cardiac CT images.

[0008] The main idea of the technical solution adopted by the present invention: The present invention proposes a fully automatic and efficient segmentation method, device, equipment, and medium for cardiac CT images, consisting of a detection and localization method based on the Faster R-CNN network and a segmentation method based on the 3D U-Net network, achieving high-precision automatic whole-heart segmentation.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A method for whole-heart segmentation based on computer tomography 3D images, characterized in that: the method comprises the following steps:

[0011] S1. Detection and localization

[0012] Detect and localize the cardiac region in the CT image through an improved 3D Faster R-CNN network;

[0013] S2. Segmentation

[0014] Segment the cardiac region in the CT image detected and localized in step S1 through an improved 3D U-Net network;

[0015] S3. Set a loss function according to the stage results obtained from the detection and localization in step S1 and the stage results obtained from the segmentation in step S2;

[0016] S4. The loss function in step S3 guides the segmentation of the CT image in the correct direction, making the predicted value of the CT image segmentation approach the true value.

[0017] Furthermore, the improved 3D Faster R-CNN network in step S1 includes the following improvement steps:

[0018] S101. Expand the existing Faster R-CNN to 3D Faster R-CNN and replace all 2D convolutional kernels with 3D convolutional kernels;

[0019] S102. Replace the ResNet part of the existing Faster R-CNN with a P3D ResNet structure. The P3D network divides the three-dimensional space into two orthogonal bases and simulates a 3×3×3 convolutional kernel with 1×3×3 and 3×1×1 convolutional kernels;

[0020] S103. Add an FPN structure after P3D ResNet to combine feature maps of different resolutions.

[0021] Furthermore, the improved 3D Faster R-CNN network in step S1 detects and locates the heart region in the CT image, including the following steps:

[0022] S1011. The input cardiac CT image is processed through a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer to obtain a feature map with only one-fourth of the original image size, denoted as C1;

[0023] S1021. After being processed by two P3D structures, a feature map C2 with one-eighth of the original image size is obtained; on this basis, after being processed by three P3D structures, a feature map C3 with one-sixteenth of the original image size is obtained;

[0024] S1031. Convolve and upsample C3 and sum it with C2 after convolution to obtain P2; perform two convolution operations on C3 to obtain P3, that is, complete the FPN part to adjust the input image size to different sizes and obtain feature maps of different resolutions;

[0025] S1041. Connect P2 and P3 to obtain a feature map as the input of the RPN part to extract different candidate boxes, and fine-tune the obtained proposed ROI by aligning it with the input feature map;

[0026] S1051. Finally, a predicted bounding box containing the heart, which is very close to the true bounding box, and the probability of containing the heart are obtained, i.e., the detection and localization tasks of the heart region are completed.

[0027] Further, the improved 3D U-Net network in step S2 includes the following improvement steps:

[0028] S201. Residual links are added on the encoding path;

[0029] S202. An additional deconvolution layer is added for upsampling;

[0030] S203. A deep supervision path is added on the decoding path to combine feature information of different scales.

[0031] Further, the improved 3D U-Net network in step S2 segments the heart region, including the following steps:

[0032] S2012. The segmentation consists of an encoding path, a decoding path, and a deep supervision path;

[0033] S2022. The original CT image aligned with the bounding box is used as the input of the segmentation network. Through layers of networks on the encoding path, a feature map containing high-resolution information is obtained. At the same time, residual connections are added so that the segmentation of the CT image can also learn simple features in the deep network to enhance the feature learning ability;

[0034] S2032. Skip connections are introduced in the decoding path to connect the feature information of corresponding scales in the encoding and decoding paths to improve the segmentation performance;

[0035] S2042. On the deep supervision path, the outputs of different resolutions on the decoding path are superimposed on the final output as the final prediction output. At the same time, a deconvolution layer is added at the end to expand the output feature size of the segmentation network to twice the size of the input feature map to compensate for the accuracy loss caused by downsampling;

[0036] S2052. The final output is a one-hot encoded feature map with 8 channels and a size of (128, 128, 128).

[0037] Further, the setting of the loss function in step S3 includes the following steps:

[0038] S301. In detection and localization, there are two points in the phased results:

[0039] The loss function for measuring the error between the candidate bounding boxes obtained through the RPN part and the true bounding boxes: the classification loss function in the RPN network stage , the bounding box loss function in the RPN stage ;

[0040] A loss function for measuring the error between the candidate bounding box obtained through the bounding box fine-tuning part and the ground truth bounding box: the classification loss function for bounding box optimization , the bounding loss function for bounding box optimization ;

[0041] S302. In the segmentation, there are two points in the intermediate result:

[0042] A loss function for measuring the error between the predicted value of the heart segmentation mask obtained through the 3D U-Net segmentation network and the ground truth value of the heart segmentation mask: the segmentation mask loss function ;

[0043] A loss function for measuring the error between the predicted heart edge map obtained through edge extraction and the ground truth heart edge map: the segmentation edge loss function ;

[0044] S303. Set the loss function according to the intermediate results obtained in steps S301 and S302;

[0045] S304. The loss function is defined as:

[0046] (2)

[0047] (3)

[0048] (4)

[0049] (5)

[0050] where B represents the predicted bounding box , represents the actual bounding box , B, represents B, the center point of, is the diagonal length of the smallest rectangle covering the two bounding boxes, represents the Euclidean distance;

[0051] S305. The loss function can guide the segmentation of the CT image in the correct direction, making the predicted value of the CT image segmentation approach the ground truth value.

[0052] A full heart segmentation device based on computer tomography 3D images, characterized by including the following modules:

[0053] The detection and localization module, through the improved 3D Faster R-CNN network, detects and locates the heart region in the CT image;

[0054] A segmentation module that segments the heart region in the detected and located CT image through an improved 3D U-Net network;

[0055] A prediction module that makes the segmentation of the CT image proceed in the correct direction and makes the predicted value of the segmentation of the CT image approach the true value.

[0056] Furthermore, a whole heart segmentation device based on a computer tomography 3D image, characterized in that

[0057] The detection and location module includes:

[0058] A processing unit 1 that processes the cardiac CT image through processing module 1 to obtain feature maps of different sizes, denoted as C1, C2, and C3;

[0059] A processing unit 2 that sums the convolution and upsampling processing of C3 and C2 after convolution processing to obtain P2; performs two convolution processes on C3 to obtain P3, and obtains feature maps of different resolutions;

[0060] An extraction unit that uses the feature maps of different resolutions obtained by connecting P2 and P3 as the input of the RPN part to extract different candidate boxes, fine-tunes the obtained proposed ROI to align with the input feature map, and obtains a predicted bounding box containing the heart and its probability of containing the heart that is very close to the true bounding box;

[0061] The segmentation module includes:

[0062] An encoding path unit that obtains a feature map containing high-resolution information through layers of networks on the encoding path, and at the same time adds residual connections so that the segmentation of the CT image can also learn simple features in the deep network to strengthen the feature learning ability;

[0063] A decoding path unit that introduces skip connections in the decoding path to connect the feature information of corresponding scales of the encoding and decoding paths to improve the segmentation performance;

[0064] A deep supervision path unit that superimposes the outputs of different resolutions on the decoding path on the final output as the final prediction output, and at the same time adds a transposed convolution layer at the end to expand the output feature size of the segmentation network to twice the size of the input feature map to make up for the accuracy loss caused by downsampling; the final output is a one-hot encoded feature map.

[0065] An electronic device includes a processor and a memory, characterized in that the memory stores computer instructions, and the processor is used to run the computer instructions stored on the memory to implement the steps of the above-mentioned whole heart segmentation method based on a computer tomography 3D image.

[0066] A computer-readable storage medium stores computer instructions thereon, wherein the computer instructions are used to cause a computer to execute the above-mentioned full-heart segmentation method based on a computed tomography 3D image.

[0067] The beneficial effects of the present invention are as follows: Compared with the prior art, the improvements of the present invention are as follows.

[0068] 1. The present invention proposes a fully automatic and efficient segmentation method, device, equipment and medium for cardiac CT images, replacing the feature extraction network ResNet with P3D ResNet. And an FPN network is added after the P3D ResNet network. This structure can fuse features of different resolutions, enhance the feature learning ability, can adapt to cardiac images of different scales, and reaches the state-of-the-art level in the task of fully automatic and rapid segmentation of cardiac CT images.

[0069] 2. The present invention replaces all bounding box loss functions with the CIoU loss function. The CIoU loss function mainly considers the overlapping area between the predicted box and the ground truth box, and adds the distance information of the center points of the bounding boxes and the ratio information of the aspect ratio as penalty terms. Its application in this application effectively improves the convergence speed and regression accuracy, and demonstrates strong segmentation performance in the cardiac CT image segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a flowchart of the full-heart efficient segmentation method of the present invention.

[0071] Figure 2 It is a schematic diagram of overlapping target boxes of the present invention.

[0072] Figure 3 It is the original images, their labels and corresponding segmentation results of three CT images 1008, 1009 and 1016 of the present invention.

[0073] Figure 4 It is the architecture of the detection and positioning module of the present invention.

[0074] Figure 5 It is the 3D U-Net network architecture in the segmentation module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] In order to enable those of ordinary skill in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0076] Please refer to Figures 1-5 , a full-heart segmentation method based on a computed tomography 3D image, characterized in that: the method includes the following steps:

[0077] S1. Detection and localization

[0078] Detect and localize the heart region in the CT image through an improved 3D Faster R-CNN network. Compared with the traditional U-Net network that uses a tiling strategy to process the image into small pieces for input into the network, the method in this paper takes the entire three-dimensional CT image as the network input, retaining more comprehensive heart image information. Moreover, the present invention uses a bounding box centered on the heart for localization, enabling the subsequent segmentation module to focus more on the most critical part of the heart CT image and saving a large amount of computing resources.

[0079] S2. Segmentation

[0080] Segment the heart region in the CT image detected and localized in step S1 through an improved 3D U-Net network. After obtaining the target region output in the previous stage, segment it to obtain the region of interest.

[0081] S3. Set a loss function according to the stage results obtained from the detection and localization in step S1 and the stage results obtained from the segmentation in step S2. The loss function is used to evaluate the degree of difference between the predicted value and the true value of the model (it should be noted that the model in this application is the method proposed in this application, that is, in this application, model = method). The smaller the loss function value, the better the model performance, that is, the predicted value is closer to the true value. During the model training process, use the loss function to guide the model in the correct direction, making the predicted value of the model approach the true value, thereby achieving an improvement in the model performance.

[0082] S4. The loss function in step S3 guides the model in the correct direction, making the predicted value of the model approach the true value.

[0083] It should be noted that the improved 3D Faster R-CNN network in step S1 includes the following improvement steps:

[0084] S101. Expand the existing Faster R-CNN to 3D Faster R-CNN and replace all 2D convolutional kernels with 3D convolutional kernels.

[0085] S102. Replace the ResNet part of the existing Faster R-CNN with a P3D ResNet structure. The P3D network divides the three-dimensional space into two orthogonal bases and uses 1×3×3 convolutional kernels and 3×1×1 convolutional kernels to simulate 3×3×3 convolutional kernels.

[0086] S103. Add an FPN structure after the P3D ResNet to combine feature maps of different resolutions.

[0087] Based on the improved 3D Faster R-CNN network in step S1, detect and locate the heart region in the CT image as Figure 4 shown, including the following steps:

[0088] S1011. The input cardiac CT image is processed through a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer to obtain a feature map with only one-fourth of the original image size, denoted as C1;

[0089] S1021. After being processed by two P3D structures, a feature map C2 with one-eighth of the original image size is obtained; on this basis, after being processed by three P3D structures, a feature map C3 with one-sixteenth of the original image size is obtained;

[0090] S1031. Convolve and upsample C3 and sum it with C2 after convolution processing to obtain P2; perform two convolution processes on C3 to obtain P3, that is, complete the FPN part to adjust the input image size to different sizes and obtain feature maps with different resolutions;

[0091] S1041. Connect P2 and P3 to obtain a feature map as the input of the RPN part, which is used to extract different candidate boxes, and fine-tune the obtained proposed ROIs to align with the input feature map;

[0092] S1051. Finally, a predicted bounding box containing the heart and its probability of containing the heart that is very close to the true bounding box are obtained, that is, the detection and location tasks of the heart region are completed.

[0093] It should be noted that the improved 3D U-Net network in step S2 includes the following improvement steps:

[0094] S201. Add residual connections on the encoding path to make it better learn the feature map;

[0095] S202. Add an additional deconvolution layer for upsampling to make the output size twice the input size to compensate for the accuracy loss caused by downsampling;

[0096] S203. Add a deep supervision path on the decoding path to combine feature information at different scales to achieve better segmentation performance.

[0097] Based on the improved 3D U-Net network in step S2, segment the heart region, as Figure 5 shown. This figure shows that in the ablation experiment, the size of the ROI is set to (64, 64, 64), and C represents the number of channels. It includes the following steps:

[0098] S2012. The segmentation consists of an encoding path, a decoding path, and a deep supervision path;

[0099] The original CT image aligned with the bounding box is used as the input of the segmentation network. Through layers of networks in the encoding path, a feature map containing high-resolution information is obtained. At the same time, residual connections are added so that the model can also learn simple features in the deep network to enhance the feature learning ability.

[0100] S2032. Introduce skip connections in the decoding path to connect the feature information of corresponding scales in the encoding and decoding paths to improve the segmentation performance.

[0101] S2042. On the deep supervision path, the outputs of different resolutions on the decoding path are stacked onto the final output as the final prediction output. At the same time, a transposed convolution layer is added at the end to expand the output feature size of the segmentation network to twice the size of the input feature map to compensate for the accuracy loss caused by downsampling.

[0102] S2052. The final output is a one-hot encoded feature map with 8 channels and a size of (128, 128, 128).

[0103] The setting of the loss function in step S3 includes the following steps:

[0104] First, the loss function is used to evaluate the degree of difference between the predicted value and the true value of the model. The smaller the loss function value, the better the model performance, that is, the predicted value is close to the true value. During the model training process, the loss function is used to guide the model in the correct direction, making the predicted value of the model approach the true value, thereby achieving the improvement of the model performance.

[0105] S301. In detection and localization, there are two points in the phased results:

[0106] The loss function for measuring the error between the candidate boxes obtained through the RPN part and the true bounding boxes: the classification loss function in the RPN network stage , the bounding box loss function in the RPN stage ;

[0107] The loss function for measuring the error between the candidate boxes obtained through the bounding box fine-tuning part and the true bounding boxes: the classification loss function for bounding box optimization , the bounding loss function for bounding box optimization ;

[0108] S302. In segmentation, there are two points in the phased results:

[0109] The loss function for measuring the error between the predicted value of the heart segmentation mask obtained through the 3D U-Net segmentation network and the true value of the heart segmentation mask: the segmentation mask loss function ;

[0110] The loss function for measuring the error between the predicted heart edge map obtained by edge extraction and the true heart edge map: the segmentation edge loss function ;

[0111] S303. Set the loss function based on the intermediate results obtained in steps S301 and S302; the weight coefficients of the loss function are set as: w1:w2:w3:w4:w5:w6 = 100:50:1:20:1:1, (1).

[0112] Replace all bounding box loss functions with the CIoU loss function. The bounding box loss functions include and . The CIoU loss function mainly considers the overlapping area between the predicted box and the true box, and adds the distance information between the center points of the bounding boxes and the ratio information of the aspect ratio as penalty terms. Its application in this application effectively improves the convergence speed and regression accuracy, and demonstrates strong segmentation performance in the cardiac CT image segmentation task.

[0113] S304. The loss function is defined as:

[0114] (2)

[0115] (3)

[0116] (4)

[0117] (5)

[0118] where B represents the predicted bounding box , represents the actual bounding box , B, represents B, the center point of is the diagonal length of the smallest rectangle covering the two bounding boxes, represents the Euclidean distance;

[0119] S305. The loss function can guide the model in the correct direction, making the predicted value of the model approach the true value, thereby improving the performance of the model.

[0120] Segment the cardiac CT image according to the above steps. Embodiment

[0121] This application conducts experiments on the dataset of the 2017 Multi-modal Whole Heart Segmentation Challenge (MM-WHS). This dataset provides 60 sets of real heart CT image data, among which 20 sets of manually labeled data are used as training data, and 40 sets of unlabeled data are used as test data. Since there are only 20 sets of publicly annotated data, it is necessary to re-partition the dataset to avoid overfitting. For the original 40 sets of unlabeled data, a deep supervised 3D U-Net network with a weighted loss function is used to automatically segment and generate pseudo-labels as the training set. The other 20 sets of manually labeled data are randomly divided into 15 sets as the test set, and the remaining 5 sets are used as the validation set.

[0122] The average size of each original heart CT image is (512, 512, 265), and the frame size ranges from [177, 363]. This application adjusts the images to a unified size of (320, 320, 192), and then inputs them into the network. For the labeled data, since the segmentation categories are not continuous, the voxel values of the seven substructures and the background are mapped to an 8-state one-hot encoding. The final size of the input label data is (320, 320, 192, 8).

[0123] Table 1 shows that the average Dice score of the method proposed in this invention in the whole heart segmentation test set is 91.1%. Among the average Dice scores for segmenting the seven structures of the heart, it beats the method of the competition champion, leading by 0.3%. In terms of segmentation speed, the fully automatic segmentation method for the heart CT images of this invention only takes one-twentieth of the time of the SEG-CNN competition champion algorithm, which is only less than 6 seconds, greatly improving the efficiency of segmenting the seven substructures of a single heart. It should be noted that the time-consuming statistics of this invention do not consider the time for data loading and post-processing. If considered, the average time-consuming is less than 15 seconds.

[0124] Table 1. Comparison of test results between this application and excellent models on the CT dataset (T: average running time)

[0125]

[0126] As can be seen from this application Figure 3 (a)-(c) show the original images, their labels, and the corresponding segmentation results of three CT images, namely 1008, 1009, and 1016, in the test set from top to bottom. In addition, (d)-(e) respectively describe the three-dimensional visualizations of the ground truth labels and the predicted segmentation results. From the visual effect, the whole heart segmentation result of this application is close to the ground truth labels manually annotated by doctors.

[0127] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A full heart segmentation method based on computer tomography 3D images, characterized in that: The method includes the following steps: S1. Detection and localization Detect and localize the heart region in the CT image through an improved 3D Faster R-CNN network; The improved 3D Faster R-CNN network in step S1 includes the following improvement steps: S101. Expand the existing Faster R-CNN into 3D Faster R-CNN and replace all 2D convolutional kernels with 3D convolutional kernels; S102. Replace the ResNet part of the existing Faster R-CNN with a P3D ResNet structure. The P3D network divides the three-dimensional space into two orthogonal bases and simulates a 3×3×3 convolutional kernel with 1×3×3 and 3×1×1 convolutional kernels; S103. Add an FPN structure after the P3D ResNet to combine feature maps of different resolutions; S2. Segmentation Segment the heart region in the CT image detected and localized in step S1 through an improved 3D U-Net network; The improved 3D U-Net network in step S2 includes the following improvement steps: S201. Add residual links on the encoding path; S202. Add an additional deconvolution layer for upsampling; S203. Add a deep supervision path on the decoding path to combine feature information of different scales; S3. Set a loss function according to the stage results obtained from detection and localization in step S1 and the stage results obtained from segmentation in step S2; S4. The loss function in step S3 guides the segmentation of the CT image in the correct direction, making the predicted value of the CT image segmentation approach the true value.

2. The full heart segmentation method based on computer tomography 3D images according to claim 1, characterized in that: The improved 3D Faster R-CNN network in step S1 detecting and localizing the heart region in the CT image includes the following steps: S1011. The input heart CT image is processed through a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer to obtain a feature map with only one-fourth of the original image size, denoted as C1; S1021. After being processed by two P3D structures, a feature map C2 with one-eighth of the original image size is obtained; on this basis, after being processed by three P3D structures, a feature map C3 with one-sixteenth of the original image size is obtained; S1031. Perform convolution and upsampling on C3 and sum it with C2 after convolution processing to obtain P2; perform two convolution processes on C3 to obtain P3, that is, complete the adjustment of the input image size to different sizes in the FPN part to obtain feature maps of different resolutions; S1041. Connect P2 and P3 to obtain a feature map as the input of the RPN part to extract different candidate boxes, and fine-tune the obtained proposed ROI by aligning it with the input feature map; S1051. Finally, obtain a predicted bounding box containing the heart that is very close to the true bounding box and the probability of containing the heart, that is, complete the detection and localization tasks of the heart region.

3. The whole heart segmentation method based on computer tomography 3D images according to claim 2, characterized in that: The improved 3D U-Net network in step S2 segmenting the heart region includes the following steps: S2012. The segmentation consists of an encoding path, a decoding path, and a deep supervision path; The original CT image aligned with the bounding box is used as the input of the segmentation network. Through layers of networks in the encoding path, a feature map containing high-resolution information is obtained. At the same time, residual connections are added so that the segmentation of the CT image can also learn simple features in the deep network to enhance the feature learning ability. S2032. In the decoding path, skip connections are introduced to connect the feature information of corresponding scales in the encoding and decoding paths to improve the segmentation performance. S2042. On the deep supervision path, the outputs of different resolutions on the decoding path are superimposed on the final output as the final prediction output. At the same time, a deconvolution layer is added at the end to expand the output feature size of the segmentation network to twice the size of the input feature map to compensate for the accuracy loss caused by downsampling. S2052. The final output is a one-hot encoded feature map with 8 channels and a size of (128, 128, 128).

4. A full-heart segmentation method based on computer tomography 3D images according to claim 3, characterized in that: The setting of the loss function in step S3 includes the following steps: S301. In detection and localization, there are two points in the phased results: Loss function for measuring the error between the candidate bounding boxes obtained through the RPN part and the ground truth bounding boxes: Classification loss function in the RPN network stage , Bounding box loss function in the RPN stage ; Loss functions for measuring the error between the candidate bounding boxes obtained from the bounding box fine-tuning part and the ground truth bounding boxes: classification loss function for bounding box optimization , bounding box loss function for bounding box optimization ; S302. In segmentation, there are two points in the phased results: Loss function for measuring the error between the predicted value of the heart segmentation mask obtained by the 3D U-Net segmentation network and the true value of the heart segmentation mask: segmentation mask loss function ; Loss function for measuring the error between the predicted heart edge map obtained through edge extraction and the true heart edge map: segmentation edge loss function ; S303. Set the loss function according to the phased results obtained through step S301 and step S302. S304. The loss function is defined as: (2) (3) (4) (5) Among them, B represents the predicted bounding box , represents the actual bounding box , B, represents the center point of B, and is the diagonal length of the smallest rectangle covering the two bounding boxes, representing the Euclidean distance; S305. The loss function can guide the segmentation of the CT image in the correct direction, making the predicted value of the segmentation of the CT image approach the true value.

5. A full heart segmentation device based on computer tomography 3D images, characterized in that, Implemented by the method for full heart segmentation based on computed tomography 3D images according to claim 1, including the following modules: The detection and localization module, through the improved 3D Faster R-CNN network, detects and locates the heart region in the CT image. The segmentation module, segments the heart region in the CT image detected and located through the improved 3D U-Net network. The prediction module, makes the segmentation of the CT image in the correct direction, making the predicted value of the segmentation of the CT image approach the true value.

6. The full heart segmentation device based on computed tomography 3D images according to claim 5, characterized in that The detection and localization module includes: The processing unit 1 processes the cardiac CT image through the processing module 1 to obtain feature maps of different sizes, denoted as C1, C2, and C3. The processing unit 2 sums the convolution and upsampling processing of C3 and the convolution-processed C2 to obtain P2; performs two convolution processes on C3 to obtain P3, obtaining feature maps of different resolutions. The extraction unit uses the feature maps of different resolutions obtained by connecting P2 and P3 as the input of the RPN part to extract different candidate boxes, and fine-tunes the obtained proposed ROIs to align with the input feature map to obtain a predicted bounding box containing the heart that is very close to the true bounding box and the probability of containing the heart. The segmentation module includes: The encoding path unit obtains a feature map containing high-resolution information through layers of networks in the encoding path. At the same time, residual connections are added so that the segmentation of the CT image can also learn simple features in the deep network to enhance the feature learning ability. The decoding path unit introduces skip connections in the decoding path to connect the feature information of corresponding scales in the encoding and decoding paths, so as to improve the segmentation performance; The deep supervision path unit superimposes the outputs of different resolutions on the decoding path onto the final output as the final prediction output. At the same time, a transposed convolutional layer is added at the end to expand the output feature size of the segmentation network to twice the size of the input feature map to compensate for the accuracy loss caused by downsampling; the final output is a one-hot encoded feature map.

7. An electronic device, comprising a processor and a memory, characterized in that, Computer instructions are stored on the memory, and the processor is used to run the computer instructions stored on the memory to implement the steps of the whole heart segmentation method based on computer tomography 3D images according to any one of claims 1-4.

8. A computer-readable storage medium having computer instructions stored thereon, wherein, The computer instructions are used to cause the computer to execute the whole heart segmentation method based on computer tomography 3D images according to any one of claims 1-4.

Citation Information

Patent Citations

  • Liver tumor segmentation method and system based on convolutional neural network

    CN111627019A

  • Residual 3D U-Net medical image segmentation method based on multi-scale depth supervision

    CN115249250A