A whole heart segmentation method for CT images based on deep learning
Through the CT image full-cardiac segmentation method based on deep learning, and using technical means such as deep neural network and attention mechanism, the problems of inaccurate whole-cardiac segmentation results in the existing technology and the long test time are solved, and efficient and accurate whole-cardiac segmentation is achieved to meet clinical needs.
Patent Information
- Application Number
- CN202210253312.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-03-15
AI Technical Summary
The existing technology has insufficient medical image data in total cardiac segmentation, resulting in inaccurate segmentation results, and the use of 3D methods requires cutting data to waste resources, and the test time is too long to meet clinical needs.
Using the CT image full-heart segmentation method based on deep learning, the whole heart is automatically divided into 7 substructures using a deep neural network. By introducing residual modules and multi-scale fusion modules based on attention mechanisms, the extraction ability and segmentation accuracy of the network are improved, and a mixed loss function is proposed to solve the problem of category imbalance.
It significantly improves the accuracy and efficiency of full-heart segmentation, greatly reduces the time and learning costs invested by doctors, and achieves fast and accurate full-heart segmentation to meet the needs of clinical diagnosis.
Smart Images

Figure CN114596317B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical images, and in particular to a method for whole-heart segmentation of CT images based on deep learning. Background Art
[0002] The heart is one of the important organs of the human body and undertakes important functions of the human body. Whole heart segmentation refers to dividing the heart into 7 substructures, mainly the left ventricle (LV), left ventricular myocardium (Myo), left atrium (LA), right atrium (RA), right ventricle (RV), pulmonary artery (PA), and ascending artery (AA). The literature (China Cardiovascular Health and Disease Report Writing Group. Summary of China Cardiovascular Health and Disease Report 2020 [J]. Chinese Journal of Circulation, 2021, 36(6): 521-545.) points out that cardiovascular disease ranks first among urban and rural residents in my country, accounting for 46.6% in rural areas and 43.81% in cities. It is estimated that about 330 million people in my country suffer from cardiovascular disease. Whole heart segmentation can extract clinical indicators such as ejection fraction, cardiac chamber volume, and myocardial thickness. It is the basis for doctors to conduct functional analysis, diagnosis and treatment, and is of great significance in clinical research. The difficulty of whole heart segmentation lies mainly in the similar grayscale values between the substructures of the heart, which makes it difficult to classify the boundary pixels. Secondly, the number of background pixels is greater than the number of foreground pixels, resulting in serious class imbalance. Before the rise of deep learning, the main methods for whole heart segmentation were deformable models and atlas methods. However, it is very difficult to build models on limited data sets. At the same time, this processing is very time-consuming and cannot meet the requirements of doctors for reading films. With the rapid development of deep learning in the field of computer vision and medical images, many researchers have applied deep learning to whole heart segmentation. The literature [Yang X, Bian C, Yu L, et al. Hybrid lossguided convolutional networks for whole heart parsing [C] / / International workshop on statistical atlases and computational models of the heart. Springer, Cham, 2017: 215-223.] uses a 3D Unet network model and fuses dice loss and cross entropy loss to form a new loss function to achieve good segmentation results.
[0003] However, using the 3D method requires cutting the whole heart 3D volume data into small pieces in sequence, inputting them into the network to obtain outputs, and then splicing the outputs together to obtain the final segmentation results. Most of the heart data belongs to the background, and this strategy wastes resources. At the same time, it takes more than several minutes to test a data, which wastes manpower and time costs and cannot meet clinical needs. Summary of the invention
[0004] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and to provide a method for whole-heart segmentation in CT images based on deep learning. A deep neural network is used to automatically divide the whole heart into 7 substructures, including the left atrium, left ventricle, left ventricular myocardium, right atrium, right ventricle, pulmonary artery, and ascending artery. This overcomes the shortcomings of insufficient medical image data that lead to inaccurate segmentation results, greatly reduces the time and learning costs invested by doctors, and improves diagnostic efficiency.
[0005] The technical solution adopted by the present invention is:
[0006] The present invention provides a method for whole heart segmentation of CT images based on deep learning, comprising the following steps:
[0007] Step 1: Input the 3D medical image to be segmented and define it as a 3D feature array of size C×H×W according to the size of the image, expressed as: image X(C×H×W);
[0008] Step 2: Preprocess the image to be segmented in step 1 as a training sample;
[0009] Step 3: Establish a whole heart segmentation network, use the whole heart segmentation network encoder to generate 5 feature maps of different depths Out0, Out1, Out2, Out3, Out4, and use the multi-scale fusion module based on the attention mechanism and the deep supervision module to restore the features in the decoding stage to obtain the feature maps, Y0, Y1, Y2, Y3, Y4;
[0010] Step 4: Use the training samples to train the whole heart segmentation network established in S3 to obtain a trained deep convolutional neural network;
[0011] Step 5: Input the whole heart CT image preprocessed in step 2 into the deep convolutional neural network trained in step 4 for image segmentation, and output the segmented whole heart CT image, including 7 substructures: left atrium, left ventricle, left ventricular myocardium, right atrium, right ventricle, pulmonary artery, and ascending artery.
[0012] The specific process of step 1 is as follows: data input, input the image to be segmented, and define it as a 3D feature array of size C×H×W, expressed as: image X (C×H×W); the input CT volume data is obtained using conventional angiography, each volume data covers all substructures of the entire heart, and each volume data consists of M 2D slices.
[0013] The specific process of step 2 is as follows: preprocessing the input CT data set; recoding the original label value of the label data, and one-hot encoding the newly encoded label value; training the neural network using a 2D convolution method, performing 2D slicing on each 3D volume data, normalizing the data, and obtaining the maximum grayscale value A of each volume data. max and the minimum gray value A min , according to the normalization formula to update the gray value A of each point new , as shown in formula (1):
[0014]
[0015] Among them, A old is the original grayscale value, and the image is rotated clockwise or counterclockwise to perform data enhancement and data scaling operations.
[0016] The specific process of step 3 includes the following steps:
[0017] Step 3.1: For the encoding stage in the network, retain the first four feature modules of the residual network, remove the last fully connected layer and average pooling layer, retain 4 layers and the first 7×7 convolution layer and pooling layer, and each layer contains two residual modules; the residual module uses a 3×3 convolution kernel for convolution operation, and each convolution layer is followed by a BN layer and ReLU operation to normalize and activate the extracted features;
[0018] According to the image X obtained in step 1, the image X is subjected to feature extraction operation to change the number of feature map channels, and the maximum pooling operation is performed to change the size of the feature map to obtain the feature map , and then through the layer layer, we get the feature map Out i+1 , where i is the feature map index, i = 1, 2, 3, 4, the Out i+1 They are
[0019]
[0020]
[0021]
[0022]
[0023] Step 3.2: After upsampling in the decoding stage, a multi-scale fusion module based on the attention mechanism is added. The multi-scale fusion module consists of two branches. One branch uses three 3×3 convolutions for convolution operations. The outputs of the three convolution operations are feature-fused to extract spatial features of different scales. The other branch introduces a 1×1 convolution layer to further capture additional spatial information.
[0024] Combined with the feature maps Out0~Out4 obtained in the encoding stage, upsampling operation is performed to change the size of the feature map, and splicing operation is performed with the feature maps in the encoding stage in turn to strengthen the semantic information between the image contexts. Then, through the multi-scale fusion module based on the attention mechanism, the number of feature map channels is changed, and the deep supervision mechanism is injected to amplify the hidden layer features. After injecting the deep supervision mechanism, the output feature maps are Y0~Y4, which are:
[0025]
[0026] The specific process of step 4 is: calculate the predicted value through the network, compare the predicted value with the true value, the true value is the label data annotated with all target related information, calculate the loss value through the loss function, and then perform back propagation to update the network, and save the updated training weights to the specified location.
[0027] The up-sampling operation adopts a bilinear interpolation algorithm.
[0028] The 2D slice size is 512×512.
[0029] The specific process of adding the attention mechanism includes the following steps:
[0030] Step 3.2.1: Perform a global average pooling operation on the multi-scale semantic information fusion module to compress the spatial information into a vector with the same dimension as the number of channels, and obtain the attention weight of each feature channel;
[0031] Step 3.2.2: In order to capture the nonlinear relationship between feature channels, two fully connected layers, relu and sigmoid activation functions are introduced, and weight normalization is achieved through sigmoid;
[0032] Step 3.2.3: Use the scale operation to multiply the weights by the input feature map channel by channel.
[0033] Beneficial technical effects
[0034] 1. The present invention introduces a residual module in the Unet network encoding stage, and uses the idea of residual connection to increase the number of network layers, accelerate network convergence faster, and enhance the network's ability to extract the substructure features of the whole heart. In the decoding stage, a multi-scale fusion module based on the attention mechanism is introduced. The module fuses multi-scale features after deconvolution and reuses features, better integrating low-level features and high-level features. At the same time, a deep supervision mechanism is injected into the decoder to improve the accuracy of the Unet network image segmentation method.
[0035] 2. The present invention proposes a new hybrid loss function, which combines the weighted cross entropy loss function and the weighted DICE hybrid loss function to solve the category imbalance problem and also plays a good driving role in the segmentation details between subclasses.
[0036] 3. The whole heart automatic segmentation method described in the present invention was tested on a cardiac CT data set and compared with the manual segmentation results of experts. The quantitative analysis results showed that the segmentation results of the present invention were consistent with the segmentation results manually calibrated by experts, and the error evaluation was also within the error range of manual calibration.
[0037] 4. The implementation method of the present invention is simple, and it only takes a few seconds to test a data. The result is highly accurate, and the processing process does not require human interaction, thus meeting the application requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A diagram of a method for whole heart segmentation in CT images based on deep learning provided by an embodiment of the present invention;
[0039] Figure 2 A schematic diagram of a network structure provided by an embodiment of the present invention;
[0040] Figure 3 A schematic diagram of the structure of a multi-scale fusion module based on an attention mechanism provided in an embodiment of the present invention;
[0041] Figure 4 A schematic diagram of the structure of the SE-Net channel attention mechanism provided by an embodiment of the present invention;
[0042] Figure 5 A schematic diagram of whole heart segmentation results provided by an embodiment of the present invention;
[0043] in, Figure 5 (a) is an original cardiac CT image selected from the test set; Figure 5 (b) Figure 5 (a) Corresponding label image; Figure 5 (c) Figure 5 (a) The corresponding segmentation result image. DETAILED DESCRIPTION
[0044] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments;
[0045] A deep learning-based whole heart segmentation method for CT images, which uses a deep neural network to automatically divide the whole heart into seven substructures. This overcomes the shortcomings of insufficient medical image data, which leads to inaccurate segmentation results, greatly reduces the time and learning costs invested by doctors, and improves diagnostic efficiency.
[0046] In this embodiment, a method for whole heart segmentation of CT images based on deep learning is provided. Figure 1 As shown, the following steps are included:
[0047] Step 1: Read the CT image to be segmented in the training set and define it as a 3D feature array of C×H×W according to the size of the image, expressed as: X(C×H×W);
[0048] In this embodiment, according to the size of the read CT image to be segmented, the read CT image is defined as a 3D feature array of size 1×512×512, expressed as: X(1×512×512), where the first dimension is 1, representing the number of channels of X, and the second and third dimensions are both 512, representing the feature map size of X; the elements in the array X are expressed as X i,j,k , represents the pixel value at the position with coordinate (i, j, k) in the 3D array X;
[0049] Step 2: Image preprocessing: Convert 3D volume data into 2D slices. According to each mask image, obtain the sequence image with the foreground part, discard the part with only the background image, perform data denoising and data enhancement operations on the data to enhance the generalization of the network model. In order to adapt the CT image size to the input size of the network model, the slices are cropped to 224×224. The cropped image is represented as X(1×224×224), covering the entire heart structure.
[0050] Step 3: Build a whole heart segmentation network model, which is built based on the UNet network of the codec. The network model is shown in the figure Figure 2 As shown; in order to better capture the detailed features of the heart substructure, the present invention adopts a residual network structure as the encoder feature extraction network. In order to improve the detection accuracy, a feature fusion method based on the attention mechanism is used to fuse the feature map. The multi-scale fusion module based on the attention mechanism is as follows Figure 3 As shown, the attention mechanism SE-Net details are as follows Figure 4 As shown;
[0051] The encoder is based on the idea of residual network and uses the ResNet18 model pre-trained by the ImageNet dataset. The first four feature blocks of ResNet18 are retained and the last fully connected layer and pooling layer are removed. Compared with the original block, the residual network adds a shortcut mechanism. This residual structure can avoid gradient disappearance, accelerate the rapid convergence of the network, and enhance the network's ability to capture the substructure features of the whole heart. The feature maps of the encoder obtained by ResNet18 are Out0~Out4, Out0(64×112×112), Out1(64×56×56), Out2(128×28×28), Out3(256×14×14), Out4(512×7×7);
[0052] In the decoder stage, the feature map output by the encoder is upsampled, and a multi-scale fusion module based on the attention mechanism and a deep supervision output module are added to each layer. The output feature maps are Y0~Y4, Y0(8×224×224), Y1(16×112×112), Y2(32×56×56, Y3(64×28×28), Y4(128×14×14). At the same time, the label is downsampled 4 times, with coefficients of 1 / 16, 1 / 8, 1 / 4, and 1 / 2, respectively, so that the number of channels in each layer is equal to 8, that is, the number of segmented categories;
[0053] Step 4: Train the network. Before training, load the pre-trained model into the encoder for initialization. The pre-trained model uses the ResNet18 structure and trains the classification model on the ImageNet dataset.
[0054] Set the loss function, which is expressed as L hybrid , use two hyperparameters α and β to control the weight ratio of the two loss functions, as shown in formula (2):
[0055]
[0056] In this embodiment, when α=0.5 and β=0.5, the model effect is the best;
[0057] Select a suitable optimization learning method, use the Adam optimizer for optimization, and continue training until the loss function converges; set the relevant hyperparameters, set the maximum number of iterations epochs to 20, the batchsize (the number of training samples in each batch) to 8, and the initial learning rate lr to 10 -4 ; The process of updating the weights of each convolution kernel in the training operation network is to iteratively update the weights until the accuracy meets the requirements of the present invention. In this embodiment, the core process of training is completed by calling the training function, and the training results are saved in a specified directory;
[0058] Step 5: Data testing: input the preprocessed whole heart CT image to be segmented into the trained neural network to obtain the segmentation result as follows Figure 5 As shown, Figure 5 (a) is an original cardiac CT image selected from the test set; Figure 5 (b) Figure 5 (a) Corresponding label image; Figure 5 (c) Figure 5 (a) The corresponding segmentation result image.
[0059] In this embodiment, the data set used in the experiment of the method of the present invention is obtained by conventional cardiac angiography in Shanghai Shuguang Hospital, China, and the data set includes 20 labeled volume data; each volume data is composed of about 200 to 300 2D slices, and the size of each 2D slice is 512×512; each 2D slice is manually annotated by an experienced doctor; 10 individual data are randomly selected as training sets, and another 10 individual data are used as test sets; after statistics, there are 4330 groups of images, of which the training set has 2039 groups of images and the test set has 2291 groups of images, and one group of images represents a CT image and a corresponding label image; first, the label data is re-encoded, and the grayscale value list of the re-encoded label image is shown in Table 1:
[0060] Table 1 List of grayscale values of re-encoded label images
[0061]
[0062] Then one-hot encode the newly encoded label value, normalize the pixel value of each 2D slice image to 0-225, rotate each image 180 degrees clockwise for data augmentation, and finally scale the CT image and the corresponding label image to 224×224;
[0063] Input the CT images in the training set into the neural network, set the configuration parameters, set the number of images input per training to 8, and set the learning rate to 1e -4 , set the optimization algorithm to Adam, set the parameter β1 to 0.9, β2 to 0.999, set the number of iterations to 20, and set the loss function to the improved loss function of the present invention;
[0064] Where L hybrid The definition of is as follows, and two hyperparameters α and β are used to control the weight ratio of the two loss functions, as shown in formula (3):
[0065]
[0066] In this embodiment, when α=0.5, β=0.5, the model effect is the best; among them, Lcwce is the weighted cross entropy loss function, as shown in formula (4):
[0067]
[0068] L CWDICE is the weighted DICE loss function, as shown in formula (5):
[0069]
[0070] C represents the number of categories (C = 8), including the 7 substructures of the whole heart and the background category, M represents the batch size, in this experiment M = 8, w c Represents the weight coefficient of each category, y jc represents the ground truth image annotated by the doctor, The predicted probability value of each pixel in the model prediction result;
[0071] Because the experiment conducted in the present invention is an image segmentation experiment, in order to quantitatively analyze the accuracy of the experimental results, DICE is used to measure the experimental results to evaluate the performance of the network; the DICE indicator is used to calculate the difference between the predicted value (X) and the true value (Y), and the value range is [0,1], where the numerator represents the intersection between the predicted value and the true value; the larger the DICE value, the better the segmentation performance, and the formula definition is shown in formula (6):
[0072]
[0073] The whole heart segmentation network and the Unet network were tested on the CT images of the test set to obtain the segmentation results. The segmentation results and the labeled image data in the test set were used as the input of the DICE index for calculation. The results are shown in Table 2:
[0074] Table 2 DICE values of different methods
[0075] Method LV) Myo RV LA RA AA PA Mean Unet 0.878 0.818 0.778 0.845 0.815 0.941 0.826 0.843 Our 0.916 0.876 0.874 0.905 0.877 0.928 0.826 0.886
[0076] According to the quantitative analysis of the data in Table 2, it can be analyzed that the whole heart segmentation network proposed in the present invention can reach 0.886 in the DICE index for measuring the similarity between images, which greatly exceeds the Unet network in the background technology. Except for the pulmonary artery (PA) DICE value, the DICE values of other substructures are also significantly higher than the Unet method. The use of the deep learning-based CT image whole heart segmentation method has achieved excellent segmentation effect.
[0077] The running time is recorded, and the results are shown in Table 3:
[0078] Table 3 Average running time and detailed information of computer systems
[0079]
[0080] According to the running time recorded in Table 3, it can be analyzed that the method proposed in the present invention only takes 10 seconds, while the Unet method takes more than 1 minute, which is too long and cannot meet the clinical reading needs of doctors. The deep learning-based CT image whole heart segmentation method can also achieve good results in terms of time.
[0081] according to Figure 5 As shown, a qualitative analysis shows that the whole heart segmentation image obtained by using the deep learning-based CT image whole heart segmentation method can accurately segment the various substructures of the heart, and the boundaries of each substructure are also processed very smoothly.
[0082] To sum up, it can be shown that compared with the Unet network in the background material, this method has achieved good segmentation results in terms of time and performance.
Claims
1. A method for whole heart segmentation in CT images based on deep learning, characterized by: The following steps are involved: Step 1: Input the 3D medical image to be segmented and define it as a 3D feature array of size C×H×W according to the size of the image, expressed as: image X(C×H×W); Step 2: Preprocess the image to be segmented in step 1 as a training sample; Step 3: Establish a whole heart segmentation network, use the whole heart segmentation network encoder to generate 5 feature maps of different depths Out0, Out1, Out2, Out3, Out4, and use the multi-scale fusion module based on the attention mechanism and the deep supervision module to restore the features in the decoding stage to obtain the feature maps, Y0, Y1, Y2, Y3, Y4; The following steps are involved: Step 3.1: For the encoding stage in the network, retain the first four feature modules of the residual network, remove the last fully connected layer and average pooling layer, retain 4 layers and the first 7×7 convolution layer and pooling layer, and each layer contains two residual modules; the residual module uses a 3×3 convolution kernel for convolution operation, and each convolution layer is followed by a BN layer and ReLU operation to normalize and activate the extracted features; According to the image X obtained in step 1, the image X is subjected to feature extraction operation to change the number of feature map channels, and the maximum pooling operation is performed to change the size of the feature map to obtain the feature map After the layer layer, we get the feature map Out i+1 , where i is the feature map index, i = 1, 2, 3, 4, the Out i+1 They are Step 3.2: After upsampling in the decoding stage, a multi-scale fusion module based on the attention mechanism is added. The multi-scale fusion module consists of two branches. One branch uses three 3×3 convolutions for convolution operations. The outputs of the three convolution operations are feature-fused to extract spatial features of different scales. The other branch introduces a 1×1 convolution layer to further capture additional spatial information. Combined with the feature maps Out0~Out4 obtained in the encoding stage, upsampling operation is performed to change the size of the feature map, and splicing operation is performed with the feature maps in the encoding stage in turn to strengthen the semantic information between the image contexts. Then, through the multi-scale fusion module based on the attention mechanism, the number of feature map channels is changed, and the deep supervision mechanism is injected to amplify the hidden layer features. After injecting the deep supervision mechanism, the output feature maps are Y0~Y4, which are: Step 4: Use the training samples to train the whole heart segmentation network established in S3 to obtain a trained deep convolutional neural network; Step 5: Input the whole heart CT image preprocessed in step 2 into the deep convolutional neural network trained in step 4 for image segmentation, and output the segmented whole heart CT image, including 7 substructures: left atrium, left ventricle, left ventricular myocardium, right atrium, right ventricle, pulmonary artery, and ascending artery.
2. The method for whole heart segmentation in CT images based on deep learning as claimed in claim 1, characterized in that: The specific process of step 1 is as follows: data input, input the image to be segmented, and define it as a 3D feature array of size C×H×W, expressed as: image X (C×H×W); the input CT volume data is obtained using conventional angiography, each volume data covers all substructures of the entire heart, and each volume data consists of M 2D slices.
3. The method for whole heart segmentation in CT images based on deep learning as claimed in claim 1, characterized in that: The specific process of step 2 is as follows: preprocessing the input CT data set; recoding the original label value of the label data, and one-hot encoding the newly encoded label value; training the neural network using a 2D convolution method, performing 2D slicing on each 3D volume data, normalizing the data, and obtaining the maximum grayscale value A of each volume data. max and the minimum gray value A min , according to the normalization formula to update the gray value A of each point new , as shown in formula (1): Among them, A old is the original grayscale value, and the image is rotated clockwise or counterclockwise to perform data enhancement and data scaling operations.
4. The method for whole heart segmentation in CT images based on deep learning as claimed in claim 1, characterized in that: The specific process of step 4 is: calculate the predicted value through the network, compare the predicted value with the true value, the true value is the label data annotated with all target related information, calculate the loss value through the loss function, and then perform back propagation to update the network, and save the updated training weights to the specified location.
5. The method for whole heart segmentation in CT images based on deep learning as claimed in claim 1, characterized in that: The up-sampling operation adopts a bilinear interpolation algorithm.
6. The method for whole heart segmentation in CT images based on deep learning as claimed in claim 1, characterized in that: The 2D slice size is 512×512.
7. The method for whole heart segmentation in CT images based on deep learning as claimed in claim 1, characterized in that: The specific process of adding the attention mechanism includes the following steps: Step 3.2.1: Perform a global average pooling operation on the multi-scale semantic information fusion module to compress the spatial information into a vector with the same dimension as the number of channels, and obtain the attention weight of each feature channel; Step 3.2.2: In order to capture the nonlinear relationship between feature channels, two fully connected layers, relu and sigmoid activation functions are introduced, and weight normalization is achieved through sigmoid; Step 3.2.3: Use the scale operation to multiply the weights by the input feature map channel by channel.
Citation Information
Patent Citations
Multi-task multi-classification chest organ segmentation model establishment and segmentation method and system
CN112241966A
Three-dimensional liver image semantic segmentation method based on context attention strategy
CN112927255A