A Lower Limb Bone CT Image Segmentation Algorithm Based on Improved U-Net
By improving the U-Net algorithm, the densely connected hollow convolution module and the fusion module of attention mechanism are solved, the problem of low segmentation accuracy in bone CT image segmentation is achieved, the precise segmentation of bone CT images is achieved, and the accuracy of medical image segmentation is improved.
Patent Information
- Application Number
- CN202210535354.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-05-17
AI Technical Summary
The existing segmentation algorithm has the problem of low segmentation accuracy in bone CT image segmentation, especially the threshold segmentation algorithm depends on data quality, and neural network-based algorithms cannot fully utilize the features of bone CT image.
Improve the U-Net algorithm, use densely connected hollow convolution modules to extract bone features through data augmentation and cropping, and use a fusion module combining attention mechanisms in the network decoding stage to make full use of spatial information and semantic information, and optimize network training parameters to improve segmentation accuracy.
The precise segmentation of bone CT images is realized, the problem of bone information loss is improved, and the accuracy of medical image segmentation is improved.
Smart Images

Figure CN114913327B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image semantic segmentation, and particularly relates to a lower limb bone CT image segmentation algorithm based on an improved U-Net. Background Art
[0002] Existing segmentation algorithms are roughly divided into two categories: algorithms based on threshold segmentation; algorithms based on neural networks. The core of the algorithms based on threshold segmentation lies in finding a relatively appropriate threshold to segment bones from the rest of the tissues, but such algorithms rely too much on the quality of the data. Even though many studies start from the image itself to improve the accuracy of image clustering and segmentation, there are still problems with relatively low segmentation accuracy. Compared with traditional segmentation algorithms, neural network algorithms for segmentation have a large number of learnable neuron parameters and non-linear expressions, and can well remove the noise in CT data and extract bone features. However, the existing network models cannot utilize the characteristics of bone CT images and cannot achieve accurate segmentation. Summary of the Invention
[0003] Object of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides a lower limb bone CT image segmentation algorithm based on an improved U-Net. This method is an improvement based on the U-Net algorithm, making up for some deficiencies of the U-Net algorithm and making full use of the characteristics of bone CT images, achieving relatively accurate segmentation ability.
[0004] Technical Solution: In a first aspect, the present invention provides a lower limb bone CT image segmentation algorithm based on an improved U-Net, including:
[0005] Label the collected CT images of the lower limb bones of patients to obtain a CT data set;
[0006] Divide the CT data set according to a ratio, construct a training set and a test set, and perform data augmentation and random cropping on the training set to obtain the training set after data augmentation and cropping;
[0007] Import the training set after data augmentation and cropping into the improved U-Net network to extract bone features of multiple different-dimensional channels, and fuse the bone features of multiple different-dimensional channels to obtain a prediction map;
[0008] Compare the prediction map with the corresponding gold standard comparison map to obtain a network loss function;
[0009] Import the network loss function into the backpropagation training model for calculation to obtain network training parameters;
[0010] Improve the U-Net network using training parameter optimization, import the test set data into the optimized U-Net network, obtain the segmented bone images for testing, and score the performance of the optimized U-Net network through the segmented bone images for testing;
[0011] Select the optimized U-Net network with a performance score higher than others and import it into the CT dataset for CT image segmentation.
[0012] In a further embodiment, divide the CT dataset proportionally, construct it into a training set and a test set, and perform data augmentation and random cropping on the training set to obtain the training set after data augmentation and cropping, including:
[0013] Slice the CT dataset along the Axial direction to generate a total of 8000 two-dimensional CT images in dcm format, and divide the CT dataset into a training set and a test set at a ratio of 8:2; where the window width of the HU value of the CT dataset is set to [1000 - 1500];
[0014] Crop the images in the training set that contain more bones to obtain small-sized training samples of images containing more bones;
[0015] Perform data augmentation on the small-sized training samples of images containing more bones to obtain the training dataset after data augmentation and cropping;
[0016] Among them, the constraint formula for the cropping area is:
[0017]
[0018] In the formula, N represents the total number of bone pixels in the current area, i represents the current random cropping times, and i takes values in [1, 100];
[0019] The data augmentation method performs three operations each time it is input into the network for training, including: random rotation, random horizontal flipping, and photometric distortion.
[0020] In a further embodiment, import the training set after data augmentation and cropping into the improved U-Net network to extract bone features of multiple different dimensions of channels, and fuse the bone features of multiple different dimensions of channels. The method for obtaining the prediction map includes:
[0021] Multiple small-sized training samples in the training set are input into the downsampling module of the improved U-Net network for multi-layer convolution calculations to extract the bone feature maps encoded by the network. Among them, batch normalization operations and Relu activation function operations are performed on the feature maps output by each layer of convolution to obtain the bone feature maps;
[0022] Input the network-encoded skeletal feature map into the densely connected dilated convolution module for feature extraction to extract fine skeletal features;
[0023] Among them, the convolutional layer of the downsampling module is 3, the convolutional kernel size is 3×3, the stride is 1, the number of channels of the first downsampling module is 64, the number of channels of the second downsampling module is 128; the number of channels of the third downsampling module is 256;
[0024] Input the network-encoded skeletal feature map into the upsampling module for multi-layer convolutional calculation to output the network-decoded skeletal feature map. Among them, the feature information output by each layer of convolution is subjected to batch normalization operation and Relu activation function operation to obtain the skeletal feature map; the number of convolutional layers of the upsampling module is 3, the convolutional kernel size is 1×1, the stride is 1, the number of channels of the first upsampling module is 256; the number of channels of the second upsampling module is 128; the number of channels of the third upsampling module is 64;
[0025] Input the skeletal feature maps output by the upsampling module and the downsampling module into the fusion module combined with the attention mechanism for fusion to generate a prediction map.
[0026] In a further embodiment, the method for inputting the network-encoded skeletal feature map into the densely connected dilated convolution module for feature extraction to extract fine skeletal features includes:
[0027] The feature map X is input into the first dilated convolutional layer to output and generate the feature map X1; among them, the dilation rate of the first dilated convolutional layer is 3, the convolutional kernel size is 3, the stride is 1, and the input and output channels are the same.
[0028] The feature maps X and X1 are fused by channel dimension and input into the second dilated convolutional layer to output and generate the feature map X2; among them, the dilation rate of the second dilated convolutional layer is 5, the convolutional kernel size is 3, the stride is 1, the input channel is 2n, and the output channel is n;
[0029] The feature maps X, X1, and X2 are fused by channel dimension and input into the third dilated convolutional layer to output and generate the feature map X3; among them, the dilation rate of the third dilated convolutional layer is 7, the convolutional kernel size is 3, the stride is 1, the input channel is 3n, and the output channel is n;
[0030] The feature maps X, X1, X2, and X3 are fused by channel dimension and input into the fourth dilated convolutional layer to output fine skeletal features; among them, the convolutional kernel size of the fourth dilated convolutional layer is 1, the stride is 1, the input channel number is 4n, and the output is a convolution with a channel of n;
[0031] Among them, the model operation formula is:
[0032]
[0033] Y = Conv 3x3 ([X3, X2, X1, X])
[0034] Wherein, X represents the input, Xi represents the output of the intermediate operation, Y represents the final output, di represents the dilation rate, Conv represents the dilated convolution operation, [X i-1 , X i-2 ,..., X1] or [X3, X2, X1, X] represents the channel dimension connection.
[0035] The input of each layer is the output channel connection of all previous intermediate operations, and finally the dimension is reduced through the convolution operation as the output; the selection of the dilation rate also determines the quality of information extraction, and a poor combination of dilation rates will bring about a grid effect.
[0036] In a further embodiment, the method for fusing the bone feature maps output by the upsampling module and the downsampling module by inputting them into a fusion module incorporating an attention mechanism to generate a prediction map includes:
[0037] Extract high-dimensional features H and low-dimensional features L from the bone feature maps output by the upsampling module and the downsampling module respectively;
[0038] Randomly divide the high-dimensional feature H obtained from the upsampling module equally by channel dimension to generate categories of feature maps H1 and H2, and implement cross-domain operation on the low-dimensional feature through a convolutional layer with a convolutional kernel size of 1×1, a stride of 1, and the same number of input and output channels to generate L;
[0039] The feature map H1 and L generate channel attention feature Fc through the channel attention branch, and the feature map H2 and L generate spatial attention feature Fs through the spatial attention branch;
[0040] Obtain channel weights and spatial weights through the channel attention feature Fc and the spatial attention feature Fs;
[0041] Fuse the channel weights and spatial weights respectively by channel dimension and then perform channel dimension reduction through convolution operation to obtain the fused channel weights and spatial weights;
[0042] Output and generate a prediction map according to the fused channel weights and spatial weights.
[0043] In a further embodiment, the method for the channel attention feature Fc to obtain channel weights includes:
[0044] Perform global average pooling on each channel of the feature map H1 to obtain the global feature of the current channel;
[0045] Perform a convolution operation with a kernel size of 1×1 on the global features of the current channel, with n input channels and n / r output channels, to generate the feature map of the first convolutional layer of channel attention;
[0046] Perform a convolution operation with a kernel size of 1×1 on the feature map of the first convolutional layer of channel attention, with n / r input channels and n output channels; generate the feature map of the second convolutional layer of channel attention;
[0047] The feature map of the second convolutional layer of channel attention is transformed through the Relu activation function and the Sigmoid function to generate the channel output weight;
[0048] The channel branch formula is:
[0049] ω C =Conv 1x1 (Conv 1x1 (GAP([H1,L])))
[0050] F C =Sigmoid(ω C )*[H1,L]
[0051] In the formula, ω C is the channel weight obtained from the feature map, H1 is the randomly assigned high-dimensional feature, L is the low-dimensional feature, Conv represents the convolution operation with a kernel of 1x1; GAP represents global average pooling; Sigmoid represents the S function, and [H1,L] represents the concatenation of the channel dimensions.
[0052] In a further embodiment, the method for obtaining the spatial weight of the spatial attention feature Fs includes:
[0053] Extract the high-dimensional feature H and the low-dimensional feature L from the skeleton feature maps output by the upsampling module and the downsampling module respectively; and add the high-dimensional feature H and the low-dimensional feature L;
[0054] Perform two operations of maximum pooling and average pooling on the channels of the added features respectively to obtain two feature map channels;
[0055] Connect the two obtained feature maps in channels to generate a feature map with 1 channel, where the channel connection is through a convolutional layer operation with a kernel size of 7×7;
[0056] Use the Sigmoid function to transform the feature map with 1 channel as the spatial weight after the fusion of the high-dimensional and low-dimensional features;
[0057] The spatial branch formula is:
[0058] ω S= Conv([(AvgPool(H2 + L), MaxPool(H2 + L)])
[0059] F S = Sigmoid(ω S ) * (H2 + L);
[0060] Where ω S is the spatial weight obtained from the feature map, H2 is the randomly assigned high-dimensional feature, MaxPool() is the max pooling operation, and AvgPool() is the average pooling operation.
[0061] In a further embodiment, the backpropagation training model uses the SGD optimizer to optimize the model parameters and uses the poly learning strategy to adjust the learning rate. The expression is:
[0062]
[0063] Where lr is the initial learning rate, power is the degree of change in the learning rate decay, total_epoch is the maximum number of training times, epoch is the current number of training times, and batch_size is the number of samples in each training during the training process.
[0064] In a further embodiment, when importing the prediction map that meets the gold standard into the backpropagation training model to obtain the network training parameters, the backpropagation training model uses the loss function to calculate the imported prediction map. The expression is:
[0065] L loss = L Dice + αL Focal
[0066] Where L loss is the total loss of the model, L Dice is the dice coefficient loss, L Focal is the focal loss, and α is the weight coefficient for neutralizing the two losses.
[0067] In a further embodiment, the improved U-Net network is optimized using the training parameters, and the test set data is imported into the optimized improved U-Net network to obtain the segmented bone image for testing. The expression for scoring the performance of the optimized improved U-Net network using the segmented bone image for testing is:
[0068] Dice = 2TP / (2TP + FP + FN)
[0069] IoU = TP / (TP + FP + FN)
[0070] Recall = TP / (TP + FN)
[0071] Precision = TP / (TP + FP)
[0072] In the formula, TP (True Positives), TN (True Negatives), FP (False Positives), and FN (False Negatives) represent the number of bone pixel points where both the prediction and the label are (true positives), the number of background pixel points where both the prediction and the label are (true negatives), the number of pixel points where the prediction is bone and the label is background (false positives), and the number of pixel points where the prediction is background and the label is bone (false negatives) in sequence; Dice is the dice evaluation coefficient, IoU is the intersection over union of bone pixels, Recall is the recall rate of bone pixels, and Precision is the accuracy rate of bone pixels.
[0073] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0074] Based on the improvement of the U-Net algorithm, in the network encoding stage, a densely connected atrous convolution module is used to strengthen the extraction of bone features; in the network decoding stage, a fusion module combined with an attention mechanism is used to make full use of spatial information and semantic information, improve the problem of bone information loss, make up for some deficiencies of the U-Net algorithm, and make full use of the features of bone CT images, realizing accurate medical CI image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a comparison diagram of CT data images and the gold standard;
[0076] Figure 2 It is an example diagram of CT data images processed by the U-Net network;
[0077] Figure 3 It is a schematic structural diagram of a densely connected multi-scale convolution module;
[0078] Figure 4 It is a schematic structural diagram of a fusion module combined with an attention mechanism;
[0079] Figure 5 It is a comparison diagram of 2D prediction results of multiple samples;
[0080] Figure 6 It is a comparison diagram of 3D prediction results of multiple samples. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] To more fully understand the technical content of the present invention, the technical solutions of the present invention will be further introduced and described below in conjunction with specific embodiments, but not limited thereto.
[0082] Such asFigures 1 to 6 The following further describes a lower limb bone CT image segmentation algorithm based on an improved U-Net in this embodiment, including the following steps:
[0083] Label the collected CT images of the patient's lower limb bones to obtain a CT data set;
[0084] Divide the CT data set proportionally to construct a training set and a test set, and perform data augmentation and random cropping on the training set to obtain the augmented and cropped training set;
[0085] Import the augmented and cropped training set into the improved U-Net network to extract bone features of multiple different dimensional channels, and fuse the bone features of multiple different dimensional channels to obtain a prediction map;
[0086] Compare the prediction map with the corresponding gold standard comparison map to obtain a network loss function;
[0087] Import the network loss function into the backpropagation training model for calculation to obtain network training parameters;
[0088] Use the training parameters to optimize the improved U-Net network, import the test set data into the optimized improved U-Net network to obtain the segmented bone images for testing, and score the performance of the optimized improved U-Net network through the segmented bone images for testing;
[0089] Select the optimized improved U-Net network with a performance score higher than others and import it into the CT data set for CT image segmentation.
[0090] Dividing the CT data set proportionally to construct a training set and a test set, and performing data augmentation and random cropping on the training set to obtain the augmented and cropped training set includes:
[0091] Slice the CT data set along the Axial direction to generate a total of 8000 two-dimensional CT images in dcm format, and divide the CT data set into a training set and a test set in a ratio of 8:2; where the window width of the HU value of the CT data set is set to [1000 - 1500];
[0092] Crop the images in the filtered training set that contain more bones to obtain small-sized training samples of images containing more bones;
[0093] Perform data augmentation on the small-sized training samples of images containing more bones to obtain the augmented and cropped training data set;
[0094] Among them, the constraint formula for the cropping area is:
[0095]
[0096] In the formula, N represents the total number of bone pixels in the current area, and i represents the current number of random croppings, where i takes values in [1, 100]; refer to Figure 1 Observing the original dataset, it can be seen that the bone data accounts for a small proportion and is relatively concentrated in the original image, resulting in most of the image being useless information. Therefore, during the training process, planned random cropping of the dataset can save training time and play a role in data augmentation; the purpose of the planned random cropping in this embodiment is to crop the original image into small-sized training samples of 128×128 that contain more bone images.
[0097] Each time the data augmentation method is input into the network for training, three operations are performed, including: random rotation, random horizontal flipping, and photometric distortion. The augmentation method uses the random online augmentation method.
[0098] The method of importing the data-augmented and cropped training set into the improved U-Net network to extract bone features of multiple different-dimensional channels and fusing the bone features of multiple different-dimensional channels to obtain the prediction map includes:
[0099] Multiple small-sized training samples in the training set are input into the downsampling module of the improved U-Net network for multi-layer convolution calculations to extract the bone feature map encoded by the network. Among them, batch normalization operations are performed on the feature maps output by each layer of convolution, and Relu activation function operations are used to obtain the bone feature map;
[0100] The bone feature map encoded by the network is input into the densely connected dilated convolution module for feature extraction to extract fine bone features;
[0101] Among them, the number of convolutional layers in the downsampling module is 3, the convolutional kernel size is 3×3, the stride is 1, the number of channels in the first downsampling module is 64, the number of channels in the second downsampling module is 128; the number of channels in the third downsampling module is 256;
[0102] The bone feature map encoded by the network is input into the upsampling module for multi-layer convolution calculations to output the bone feature map decoded by the network. Among them, batch normalization operations are performed on the feature information output by each layer of convolution, and Relu activation function operations are used to obtain the bone feature map; the number of convolutional layers in the upsampling module is 3, the convolutional kernel size is 1×1, the stride is 1, the number of channels in the first upsampling module is 256; the number of channels in the second upsampling module is 128; the number of channels in the third upsampling module is 64;
[0103] The bone feature maps output by the upsampling module and the downsampling module are input into a fusion module that combines an attention mechanism for fusion to generate a prediction map; in this embodiment, during the network encoding stage, the data sequentially passes through the downsampling module and the dilated convolutional module with dense connections to extract bone features; during the network decoding stage, the data sequentially passes through the upsampling module and the fusion module that combines the attention mechanism to restore the resolution of the bone features to the original size.
[0104] Input the bone feature maps encoded by the network into the dilated convolutional module with dense connections for feature extraction. The method for extracting fine bone features includes:
[0105] The feature map X is input into the first dilated convolutional layer to generate the feature map X1; where the dilation rate of the first dilated convolutional layer is 3, the convolutional kernel size is 3, the stride is 1, and the number of input and output channels is the same.
[0106] The feature maps X and X1 are fused along the channel dimension and input into the second dilated convolutional layer to generate the feature map X2; where the dilation rate of the second dilated convolutional layer is 5, the convolutional kernel size is 3, the stride is 1, the input channel is 2n, and the output channel is n.
[0107] The feature maps X, X1, and X2 are fused along the channel dimension and input into the third dilated convolutional layer to generate the feature map X3; where the dilation rate of the third dilated convolutional layer is 7, the convolutional kernel size is 3, the stride is 1, the input channel is 3n, and the output channel is n.
[0108] The feature maps X, X1, X2, and X3 are fused along the channel dimension into the fourth dilated convolutional layer to output fine bone features; where the convolutional kernel size of the fourth dilated convolutional layer is 1, the stride is 1, the number of input channels is 4n, and the output is a convolution with n channels.
[0109] Among them, the model operation formula is:
[0110]
[0111] Y = Conv 3x3 ([X3, X2, X1, X])
[0112] In the formula, X represents the input, Xi represents the output of the intermediate operation, Y represents the final output, di represents the dilation rate, Conv represents the dilated convolution operation, and [X i-1 , X i-2 ,..., X1] or [X3, X2, X1, X] represents the connection along the channel dimension.
[0113] The input of each layer is the connection of the output channels of all previous intermediate operations, and finally the dimension is reduced through a convolution operation as the output; the selection of the dilation rate also determines the quality of information extraction, and a poor combination of dilation rates will bring about a grid effect. According to the theory of hybrid dilated convolution and the comparison results of this embodiment, a combination of dilation rates of 3, 5, and 7 is selected.
[0114] The method of fusing the skeleton feature maps output by the upsampling module and the downsampling module into the fusion module combined with the attention mechanism to generate a prediction map includes:
[0115] Extract high-dimensional features H and low-dimensional features L from the skeleton feature maps output by the upsampling module and the downsampling module respectively;
[0116] Randomly divide the high-dimensional feature H obtained from the upsampling module into two categories of feature maps H1 and H2 according to the channel dimension, and perform cross-domain operations on the low-dimensional features through a convolutional layer with a kernel size of 1×1, a stride of 1, and the same number of input and output channels to generate L;
[0117] The feature map H1 and L generate channel attention feature Fc through the channel attention branch, and the feature map H2 and L generate spatial attention feature Fs through the spatial attention branch;
[0118] Obtain channel weights and spatial weights through the channel attention feature Fc and the spatial attention feature Fs;
[0119] Fuse the channel weights and spatial weights according to the channel dimension respectively, and then perform channel dimension reduction through a convolution operation to obtain the fused channel weights and spatial weights;
[0120] According to the fused channel weights and spatial weights, output and generate a prediction map. In this embodiment, the channel attention branch in the fusion module is a classic squeeze-and-excitation module, which aims to obtain the relationship between feature channels. Even after the low-dimensional features perform cross-domain operations through a 1×1 convolution, there is still a lot of useless information in them. If directly fused with high-dimensional features, there is a possibility of destroying semantic information. Therefore, the channel attention mechanism is used to suppress useless information channels; the spatial attention branch in the fusion module aims to obtain the relationship in the space of the feature map, and learn the spatial weight parameters shared by channels from itself to highlight the skeleton features. The spatial attention branch realizes the fusion of semantic information and spatial information by adding feature maps of different dimensions, and then through the attention mechanism, focuses the attention of the network on the skeleton information.
[0121] The method for the channel attention feature Fc to obtain the channel weight includes:
[0122] Perform global average pooling on each channel of the feature map H1 to obtain the global feature of the current channel;
[0123] Perform a convolution operation with a kernel size of 1×1, an input channel of n, and an output channel of n / r on the global features of the current channel to generate the feature map of the first convolutional layer of channel attention;
[0124] Perform a convolution operation with a kernel size of 1×1, an input channel of n / r, and an output channel of n on the feature map of the first convolutional layer of channel attention; generate the feature map of the second convolutional layer of channel attention;
[0125] The feature map of the second convolutional layer of channel attention is transformed through the Relu activation function and the Sigmoid function to generate the channel output weight;
[0126] The channel branch formula is:
[0127] ω C = Conv 1x1 (Conv 1x1 (GAP([H1,L])))
[0128] F C = Sigmoid(ω C )*[H1,L]
[0129] In the formula, ω C is the channel weight obtained from the feature map, H1 is the randomly assigned high-dimensional feature, L is the low-dimensional feature, Conv represents the convolution operation with a kernel of 1x1; GAP represents global average pooling; Sigmoid represents the S function, and [H1,L] represents the concatenation of the channel dimensions.
[0130] The method for obtaining the spatial weight of the spatial attention feature Fs includes:
[0131] Extract the high-dimensional feature H and the low-dimensional feature L from the skeleton feature maps output by the upsampling module and the downsampling module respectively; and add the high-dimensional feature H and the low-dimensional feature L;
[0132] Perform two operations of max pooling and average pooling on the channels of the added features respectively to obtain two feature map channels;
[0133] Connect the two obtained feature maps in channels to generate a feature map with 1 channel, where the channel connection is through a convolutional layer operation with a kernel size of 7×7;
[0134] Use the Sigmoid function to transform the feature map with 1 channel as the spatial weight after the fusion of high-dimensional and low-dimensional;
[0135] The spatial branch formula is:
[0136] ω S=Conv([(AvgPool(H2+L),MaxPool(H2+L)])
[0137] F S =Sigmoid(ω S )*(H2+L);
[0138] In the formula, ω S is the spatial weight obtained according to the feature map, H2 is the randomly assigned high-dimensional feature, MaxPool() is the maximum pooling operation, and AvgPool() is the average pooling operation.
[0139] In a further embodiment, the back propagation training model uses the SGD optimizer to train and optimize the model parameters, and uses the poly learning strategy to adjust the learning rate, which is expressed as:
[0140]
[0141] Wherein, lr is the initial learning rate, power is the degree of change of learning rate attenuation, total_epoch is the maximum number of training times, epoch is the current number of training times, batch_size is the number of samples for each training during the training process. In this embodiment, the initial learning rate lr is set to 0.001, power is 0.9, the maximum number of training times total_epoch is 100, and batch_size is set to 64.
[0142] The Dice coefficient loss function can be used to measure the similarity between the predicted value and the true value. Due to the particularity that bone data accounts for a relatively small proportion in CT data, the Focal loss function is introduced to mine difficult samples and adjust the scenarios where the ratio of positive and negative samples is unbalanced.
[0143] The prediction graph that meets the gold standard is imported into the back propagation training model for calculation. In the process of obtaining the network training parameters, the back propagation training model uses the loss function to calculate the imported prediction graph. The expression is:
[0144] L loss =L Dice +αL Focal
[0145] Where, L loss is the total loss of the model, L Dice is the dice coefficient loss, L Focal is the focal loss, α is the weight coefficient for neutralizing the two losses, and in this embodiment, α is set to 0.5;
[0146] The U-Net network improved by optimizing training parameters is used, and the test set data is imported into the optimized U-Net network to obtain the segmented bone images for testing. The expression for scoring the performance of the optimized U-Net network using the segmented bone images obtained from testing is as follows:
[0147] Dice = 2TP / (2TP + FP + FN)
[0148] IoU = TP / (TP + FP + FN)
[0149] Recall = TP / (TP + FN)
[0150] Precision = TP / (TP + FP)
[0151] In the formula, TP (True Positives), TN (True Negatives), FP (False Positives), and FN (False Negatives) represent, in sequence, the number of bone pixel points where both the prediction and the label are the same (true positives), the number of background pixel points where both the prediction and the label are the same (true negatives), the number of pixel points where the prediction is a bone and the label is a background (false positives), and the number of pixel points where the prediction is a background and the label is a bone (false negatives); Dice is the dice evaluation coefficient, IoU is the intersection over union of bone pixels, Recall is the recall rate of bone pixels, and Precision is the accuracy rate of bone pixels; the ranges of the above indicators are all between 0 and 1, and the closer to 1, the stronger the prediction ability of the model. According to the above indicators, representative network models in the field of semantic segmentation: U-Net, Attention U-Net, and BiSeNet were selected as comparison network models. The average value was calculated through multiple experiments on the CT dataset as the final experimental result, and the experimental results shown in Table 1 prove the superiority of the present invention.
[0152] Table 1 Experimental results on the CT dataset
[0153]
[0154] In summary, the present invention is based on the improvement of the U-Net algorithm. In the network encoding stage, a dilated convolutional module with dense connections is used to strengthen the extraction of bone features; in the network decoding stage, a fusion module combined with an attention mechanism is used to fully utilize spatial information and semantic information, improve the problem of bone information loss, make up for some deficiencies of the U-Net algorithm, and make full use of the features of bone CT images, achieving accurate medical CI image segmentation.
[0155] Embodiments of the present application may be provided as a method, a system, or a computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0156] Embodiments of the present application may be provided as a method, a system, or a computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0157] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0160] The above are only the preferred embodiments of the present invention. Without departing from the technical principle of the present invention, several improvements and modifications can also be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An algorithm for segmenting CT images of lower limb bones based on improved U-Net, characterized in that, Including: Label the collected CT images of the patient's lower limb bones to obtain a CT data set; Divide the CT data set proportionally to construct a training set and a test set, and perform data augmentation and random cropping on the training set to obtain the training set after data augmentation and cropping; Import the training set after data augmentation and cropping into an improved U-Net network to extract bone features of multiple different-dimensional channels, and fuse the bone features of multiple different-dimensional channels to obtain a prediction map; Compare the prediction map with the corresponding gold standard comparison map to obtain a network loss function; Import the network loss function into the backpropagation training model for calculation to obtain network training parameters; Use the training parameters to optimize the improved U-Net network, import the test set data into the optimized improved U-Net network to obtain a segmented bone image for testing, and score the performance of the optimized improved U-Net network through the segmented bone image for testing; Select the optimized improved U-Net network with a performance score higher than others and import it into the CT data set for CT image segmentation; Among them, the improved U-Net network includes an upsampling module and a downsampling module; the bone feature maps output by the upsampling module and the downsampling module are input into a fusion module combined with an attention mechanism for fusion to generate a prediction map, including: Respectively extract high-dimensional features H and low-dimensional features L from the bone feature maps output by the upsampling module and the downsampling module; Randomly divide the high-dimensional feature H obtained from the upsampling module equally by channel dimension to generate feature map H1 and feature map H2 categories, and perform cross-domain operation on the low-dimensional feature through a convolutional layer with a convolutional kernel size of 1×1, a stride of 1, and the same number of input and output channels to generate L; Feature map H1 and L generate channel attention feature Fc through the channel attention branch, and feature map H2 and L generate spatial attention feature Fs through the spatial attention branch; Obtain channel weights and spatial weights through channel attention feature Fc and spatial attention feature Fs; Fuse the channel weights and spatial weights respectively according to the channel dimension and perform channel dimensionality reduction through convolution operations to obtain the fused channel weights and spatial weights; According to the fused channel weights and spatial weights, output and generate a prediction map; The method for obtaining channel weights by channel attention feature Fc includes: Perform global average pooling on each channel of feature map H1 to obtain the global feature of the current channel; Perform a convolution operation with a convolutional kernel size of 1×1, an input channel of n, and an output channel of n / r on the global feature of the current channel to generate the feature map of the first convolutional layer of channel attention; Perform a convolution operation with a convolutional kernel size of 1×1, an input channel of n / r, and an output channel of n on the feature map of the first convolutional layer of channel attention; generate the feature map of the second convolutional layer of channel attention; The feature map of the second convolutional layer of channel attention is transformed through the Relu activation function and the Sigmoid function to generate channel output weights; The channel branch formula is: ω C = Conv 1x1 (Conv 1x1 (GAP([H1,L]))) F C = Sigmoid(ω C ) * [H1, L] where ω C is the channel weight obtained from the feature map, H1 is the randomly assigned high-dimensional feature, L is the low-dimensional feature, Conv represents the convolution operation with a 1x1 convolution kernel; GAP represents global average pooling; Sigmoid represents the S function, and [H1, L] represents the concatenation in the channel dimension; The method for obtaining spatial weights by spatial attention feature Fs includes: Extract high-dimensional features H and low-dimensional features L from the bone feature maps output by the upsampling module and the downsampling module respectively; and add the high-dimensional features H and the low-dimensional features L; Perform two operations of max pooling and average pooling on the channels of the added features respectively to obtain two feature map channels; Connect the two obtained feature maps in channels to generate a feature map with 1 channel number, where the channel connection is an operation through a convolutional layer with a convolutional kernel size of 7×7; Use the Sigmoid function to transform the feature map with 1 channel number as the spatial weight after the fusion of high dimensions and low dimensions; The formula for the spatial branch is: ω S = Conv([(AvgPool(H2 + L), MaxPool(H2 + L)]) F S = Sigmoid(ω S ) * (H2 + L); where ω S is the spatial weight obtained from the feature map, H2 is the randomly assigned high-dimensional feature, MaxPool() is the max pooling operation, and AvgPool() is the average pooling operation.
2. The lower limb bone CT image segmentation algorithm based on the improved U-Net according to claim 1, wherein, Divide the CT dataset according to a certain proportion to construct a training set and a test set, and perform data augmentation and random cropping on the training set to obtain the training set after data augmentation and cropping, including: Slice the CT dataset along the Axial direction to generate a total of 8000 two-dimensional CT images in dcm format, and divide the CT dataset into a training set and a test set at a ratio of 8:2; where the window size of the HU value of the CT dataset is set to [1000 - 1500]; Crop the images in the training set that contain more bones to obtain small-sized training samples of images containing more bones; Perform data augmentation on the small-sized training samples of images with more bones to obtain the training dataset after data augmentation and cropping; Among them, the constraint formula for the cropping area is: In the formula, N represents the total number of bone pixels in the current area, i represents the current random cropping times, and i takes values in [1, 100]; The data augmentation method performs three operations each time it is input into the network for training, including: random rotation, random horizontal flipping, and photometric distortion.
3. The lower limb bone CT image segmentation algorithm based on the improved U-Net according to claim 1, characterized in that, The method of importing the training set after data augmentation and cropping into the improved U-Net network to extract bone features of multiple different dimension channels and fusing the bone features of multiple different dimension channels to obtain a prediction map includes: Multiple small-sized training samples in the training set are input into the downsampling module of the improved U-Net network for multi-layer convolution calculation to extract the bone feature map encoded by the network. Among them, batch normalization operation and Relu activation function operation are performed on the feature map output by each layer of convolution to obtain the bone feature map; Input the bone feature map encoded by the network into the densely connected dilated convolution module for feature extraction to extract fine bone features; Among them, the number of convolutional layers in the downsampling module is 3, the convolutional kernel size is 3×3, the stride is 1, the number of channels in the first layer of the downsampling module is 64, the number of channels in the second layer of the downsampling module is 128; the number of channels in the third layer of the downsampling module is 256; Input the bone feature map encoded by the network into the upsampling module for multi-layer convolution calculation to output the bone feature map decoded by the network. Among them, batch normalization operation and Relu activation function operation are performed on the feature information output by each layer of convolution to obtain the bone feature map; the number of convolutional layers in the upsampling module is 3, the convolutional kernel size is 1×1, the stride is 1, the number of channels in the first layer of the upsampling module is 256; the number of channels in the second layer of the upsampling module is 128; the number of channels in the third layer of the upsampling module is 64; The bone feature maps output by the upsampling module and the downsampling module are input into a fusion module incorporating an attention mechanism for fusion to generate a prediction map.
4. The lower limb bone CT image segmentation algorithm based on the improved U-Net according to claim 3, wherein The bone feature maps encoded by the network are input into a dilated convolutional module with dense connections for feature extraction. The methods for extracting fine bone features include: The feature map X is input into the first dilated convolutional layer, and the generated feature map X1 is output; wherein the dilation rate of the first dilated convolutional layer is 3, the convolutional kernel size is 3, the stride is 1, and the number of input and output channels is the same. The feature map X and X1 are fused along the channel dimension and input into the second dilated convolutional layer, and the generated feature map X2 is output; wherein the dilation rate of the second dilated convolutional layer is 5, the convolutional kernel size is 3, the stride is 1, the input channel is 2n, and the output channel is n. The feature maps X, X1, and X2 are fused along the channel dimension and input into the third dilated convolutional layer, and the generated feature map X3 is output; wherein the dilation rate of the third dilated convolutional layer is 7, the convolutional kernel size is 3, the stride is 1, the input channel is 3n, and the output channel is n. The feature maps X, X1, X2, and X3 are fused along the channel dimension into the fourth dilated convolutional layer, and fine bone features are output; wherein the convolutional kernel size of the fourth dilated convolutional layer is 1, the stride is 1, the number of input channels is 4n, and the output channel is n for convolution. Among them, the model operation formula is: Y = Conv 3x3 ([X3,X2,X1,X]) Wherein, X represents the input, Xi represents the output of the intermediate operation, Y represents the final output, di represents the dilation rate, Conv represents the dilated convolution operation, [X i-1 , X i-2 ,..., X1] or [X3, X2, X1, X] represents the channel dimension connection.
5. The lower limb bone CT image segmentation algorithm based on the improved U-Net according to claim 1, characterized in that, For backpropagation training of the model, the SGD optimizer is used to train and optimize the model parameters, and the poly learning strategy is used to adjust the learning rate. The expression is: In the formula, lr is the initial learning rate, power is the degree of change in learning rate decay, total_epoch is the maximum number of training times, epoch is the current number of training times, and batch_size is the number of samples in each training during the training process.
6. The lower limb bone CT image segmentation algorithm based on the improved U-Net according to claim 1, wherein When importing the prediction map that meets the gold standard into the backpropagation training model for calculation to obtain the network training parameters, the backpropagation training model calculates the imported prediction map using a loss function. The expression is: L loss = L Dice + αL Focal Where, L loss is the total loss of the model, L Dice is the dice coefficient loss, L Focal is the focal loss, and α is the weight coefficient for neutralizing the two losses.
7. The lower limb bone CT image segmentation algorithm based on the improved U-Net according to claim 1, wherein, The U-Net network optimized and improved using the training parameters is utilized, and the test set data is imported into the optimized and improved U-Net network to obtain the segmented bone images for testing. The expressions for scoring the performance of the optimized and improved U-Net network using the segmented bone images passing the test are: Dice = 2TP / (2TP + FP + FN) IoU = TP / (TP + FP + FN) Recall = TP / (TP + FN) Precision = TP / (TP + FP) Wherein, TP (True Positives), TN (True Negatives), FP (False Positives), and FN (False Negatives) respectively represent the number of bone pixel points where both the prediction and the label are (true positives), the number of background pixel points where both the prediction and the label are (true negatives), the number of pixel points where the prediction is bone and the label is background (false positives), and the number of pixel points where the prediction is background and the label is bone (false negatives); Dice is the dice evaluation coefficient, IoU is the intersection over union of bone pixels, Recall is the recall rate of bone pixels, and Precision is the accuracy rate of bone pixels.
Citation Information
Patent Citations
Remote sensing image road segmentation method based on context information and attention mechanism
CN112183258A
Remote sensing image road extraction method of improved U-Net network
CN112418027A