A method, device and system for bone CT image segmentation
The UTrans-Net network with multi-size window self-attention and cross-attention modules improves bone CT image segmentation precision by effectively handling the unique challenges of bone CT imaging, such as low contrast and interference, through enhanced feature extraction.
Patent Information
- Application Number
- CN202311384909.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-10-24
AI Technical Summary
When the existing skeletal CT image segmentation algorithm faces problems such as low contrast between cortical bone and cancellous bone, close spatial location, small proportion of interfering objects and small proportion of areas to be segmented, the segmentation accuracy is difficult to improve. The deep learning segmentation algorithm lacks global modeling capabilities and effective utilization of global features, resulting in the loss of small target information and the amplification of segmented areas.
The Swin Transformer network with linear complexity is adopted, and the concept of multi-size window is introduced on it, and the multi-size window self-attention module and cross-attention module are proposed to improve the feature extraction ability of the UTrans-Net network for different targets, and to improve feature fusion through dual-coded structure and focus fusion module.
It improves the segmentation accuracy of bone CT images, alleviates weak induction bias, enhances the network's utilization of global and local features, improves the fusion ability of features at different stages, and improves the segmentation performance.
Smart Images

Figure CN117252891B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field, and specifically relates to a method, device and system for segmenting skeletal CT images. Background Technique
[0002] Skeletal CT image segmentation belongs to a type of medical image semantic segmentation. Its ultimate goal is to classify each pixel in the skeletal CT image, and then segment the bone region. According to the implementation principle of the algorithm, medical image segmentation algorithms can be divided into two categories: traditional segmentation algorithms and deep learning segmentation algorithms. Traditional segmentation algorithms mainly achieve image segmentation through image processing methods, and deep learning segmentation algorithms mainly achieve image segmentation by constructing deep neural network models.
[0003] However, skeletal CT images have the characteristics of low contrast between cortical bone and cancellous bone, close spatial positions, interference objects such as intramedullary nails and bone plates in some data, a small proportion of the area to be segmented, and small broken bone targets in some samples. These characteristics make the segmentation results based on traditional algorithms unsatisfactory. The mainstream deep learning segmentation algorithms are usually based on convolutional operations. Due to the limitations of convolutional operations, the segmentation algorithms based on convolutional neural networks lack the ability of global modeling and effective utilization of global features. In order to obtain a larger receptive field and reduce resource consumption, such algorithms usually use downsampling to reduce the feature map size, and then perform an upsampling operation to restore the original size after feature extraction. Although this approach can improve network performance and speed up network training, it also has problems such as loss of small target information and magnification of the segmented area. When facing these data characteristics, the feature extraction method of convolutional neural networks will further amplify the problems, resulting in the inability to improve the segmentation accuracy. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a method, device and system for segmenting skeletal CT images. By referring to the Swin Transformer network with linear complexity and introducing the concept of multi-size windows on this basis, a multi-size window self-attention module is proposed, which improves the feature extraction ability of the UTrans-Net network for different targets and can improve the segmentation accuracy of skeletal CT images.
[0005] In order to achieve the above technical objectives and reach the above technical effects, the present invention is realized through the following technical solutions:
[0006] In the first aspect, the present invention provides a method for segmenting skeletal CT images, including:
[0007] Obtain a skeletal CT image;
[0008] Input the bone CT image into the pre-trained UTrans-Net network, and the UTrans-Net network outputs a segmentation image; the UTrans-Net network includes a connected encoding unit and a decoding unit, and the encoding unit has a dual-encoding structure, including a first encoding structure and a second encoding structure. Among them, the second encoding structure is a Swin Transformer network introducing a multi-scale window self-attention module.
[0009] Optionally, the encoding unit includes a first encoding structure and a second encoding structure, and the encoding unit performs the following operations on the received bone CT image:
[0010] Input the bone CT image into the first encoding structure, and after downsampling through two groups of first CNN modules, generate a feature map X1. The feature map X1 is downsampled through a second CNN module to generate a feature map X2, and the feature map X2 is downsampled through a third CNN module to generate a feature map X3;
[0011] After performing a block division operation and linear encoding on the bone CT image in sequence, generate a feature map X-T. Input the feature map X-T into the second encoding structure, generate a feature map Y1 after passing through a first multi-scale window self-attention module, then send the feature map Y1 and the feature map X1 into a first cross-attention module to generate a feature map Z1. After performing a block merging operation on the feature map Z1, send it into a second multi-scale window self-attention module to generate a feature map Y2, then send the feature map Y2 and the feature map X2 into a second cross-attention module to generate a feature map Z2. After performing a block merging operation on the feature map Z2, send it into a third multi-scale window self-attention module to generate a feature map Y3, and then send the feature map Y3 and the feature map X3 through a third cross-attention module to generate a feature map Z3.
[0012] Optionally, the first CNN module, the second CNN module, and the third CNN module all perform the following steps on the received data:
[0013] A convolution operation with a convolution kernel size of 3x3, a stride of 1, an input channel number of n, and an output channel number of n;
[0014] Batch normalization operation;
[0015] Relu activation function operation;
[0016] A convolution operation with a convolution kernel size of 3x3, a stride of 2, an input channel number of n, and an output channel number of 2n;
[0017] Batch normalization operation;
[0018] Relu activation function operation.
[0019] Optionally, the first multi-size window self-attention module, the second multi-size window self-attention module, and the third multi-size window self-attention module all perform the following steps on the received data:
[0020] Perform the normalization operation LN and the multi-size window attention operation MS-SA on the input data Fin in sequence;
[0021] Perform the first residual connection on the data after the MS-SA operation and the original input data Fin. The output obtained here is used as one of the inputs for the next residual connection. At the same time, the output also serves as one of the inputs for the next residual connection after passing through LN and the multi-layer perceptron MLP;
[0022] The data after the second residual connection passes through LN and the shifted multi-size window attention operation SMS-SA as one of the inputs for the third residual connection;
[0023] The data after the third residual connection passes through LN and MLP and then serves as one of the inputs for the fourth residual connection. After performing the fourth residual connection, Fout is obtained;
[0024] The specific implementation of the MS-SA operation is as follows:
[0025] Divide the received input data into 4 groups on average along the input data channel dimension, namely feature map F1, feature map F2, feature map F3, and feature map F4;
[0026] Perform the sliding window self-attention operation with window sizes of [4, 8, 16, 32] on feature map F1, feature map F2, feature map F3, and feature map F4 respectively, and obtain feature maps A1, A2, A3, and A4 respectively;
[0027] Merge feature maps A1, A2, A3, and A4 along the input data channel dimension to obtain feature map F1;
[0028] Perform the channel attention operation of SE-Net on feature map F1.
[0029] Optionally, the first cross-attention module, the second cross-attention module, and the third cross-attention module all perform the following steps on the received data:
[0030] Perform a convolution operation with a convolution kernel size of 1x1, a stride of 1, an input channel number of n, and an output channel number of n on the received feature map X1, X2, or X3 to generate feature maps X1”, X2”, or X3”;
[0031] Perform the recombination operation R on the feature map X1″, feature map X2″, or feature map X3″, changing its size from HxWxC to (HxW)xC to be consistent with the feature map Y1, feature map Y2, or feature map Y3, and generate the feature map X1″′, feature map X2″′, or feature map X3″′;
[0032] Perform the layer normalization operation LayerNormalization on the feature map X1″′, feature map X2″′, or feature map X3″′ to generate the feature map X1″″, feature map X2″″, or feature map X3″″;
[0033] Perform the linear transformation Linear on the feature map X1″″, feature map X2″″, and feature map X3″″ to obtain Q1, Q2, and Q3;
[0034] Perform the linear transformation Linear on the feature map Y1, feature map Y2, or feature map Y3 to obtain K1, K2, K3 and V1, V2, V3;
[0035] Perform the cross-attention operation on the feature map X1″″, feature map X2″″, or feature map X3″″ and the corresponding feature map Y1, feature map Y2, or feature map Y3 to generate the feature map Z1′, feature map Z2′, and feature map Z3′.
[0036] Optionally, the expressions for the feature map Z1′, feature map Z2′, and feature map Z3′ are:
[0037]
[0038]
[0039]
[0040] Among them, D1, D2, and D3 respectively represent the number of channels of the feature map Y1, feature map Y2, and feature map Y3, and the superscript T represents the matrix transpose operation.
[0041] Optionally, the decoding unit receives the data output by the encoding unit and performs the following steps to obtain the segmented image;
[0042] Perform the block expansion operation on the feature map Z3′ to increase the resolution, and together with the feature map Z2, they are successively fed into the first focused fusion module and the fourth multi-scale window self-attention module to generate the feature map Z3″;
[0043] Perform the block expansion operation on the feature map Z3″ to increase the resolution, and together with the feature map Z1, they are successively fed into the second focused fusion module to generate the feature map Z3″′;
[0044] Perform a reorganization operation on the feature map Z3”' to change the resolution from (HxW)xC to HxWxC, and perform a linear projection operation to obtain a segmented image.
[0045] Optionally, both the first focus fusion module and the second focus fusion module perform the following steps on the received data:
[0046] Perform a linear transformation on one of the received input data F d to obtain the corresponding V d and K d ;
[0047] Perform a linear transformation on the other received input data F e to obtain the corresponding V e and Q e ;
[0048] Let Q d =K d , Q e =K e , to obtain and then obtain the intermediate generated feature AF d and AF e :
[0049]
[0050]
[0051] Perform a cross-attention operation on the intermediate generated features AF d and AF e to obtain the output feature map F out , and the calculation formula of the feature map F out is:
[0052]
[0053] where D is the number of feature channels.
[0054] In a second aspect, the present invention provides a bone CT image segmentation device, including:
[0055] An acquisition module for acquiring bone CT images;
[0056] A segmentation module for inputting the bone CT image into a pre-trained UTrans-Net network, and outputting a segmented image by the UTrans-Net network; the UTrans-Net network includes a connected encoding unit and a decoding unit, wherein the encoding unit has a dual-encoding structure.
[0057] In a third aspect, the present invention provides a bone CT image segmentation system, including a storage medium and a processor;
[0058] The storage medium is used for storing instructions;
[0059] The processor is configured to operate according to the instructions to execute the method according to any one of the first aspect.
[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0061] The present invention proposes a bone CT image segmentation method, device and system, which refers to the Swin Transformer network with linear complexity, and on this basis, introduces the concept of multi-size windows, and proposes a multi-size window self-attention module (MSwin Block), which improves the feature extraction ability of the UTrans-Net network for different targets and can improve the segmentation accuracy of bone CT images.
[0062] To improve the segmentation performance of the UTrans-Net network on small-scale datasets, the present invention proposes a cross-attention module (Cross Attention Module, CAM) applicable to the dual-encoding branch to realize the interaction of the dual branches, bringing global and local features to the network while alleviating the weak inductive bias.
[0063] In view of the characteristics of the bone dataset, the present invention proposes a Transformer Attention Fusion Module (TAFM) applicable to the Transformer architecture to improve the feature fusion at different stages and improve the utilization rate of features at different stages by the UTrans-Net network. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings, where:
[0065] Figure 1 It is a schematic diagram of the principle of a bone CT image segmentation method according to an embodiment of the present invention;
[0066] Figure 2 It is a structural diagram of a multi-size window self-attention module according to an embodiment of the present invention;
[0067] Figure 3 It is a structural diagram of a cross-attention module according to an embodiment of the present invention;
[0068] Figure 4 Structural diagram of a focusing fusion module according to an embodiment of the present invention;
[0069] Figure 5 Comparison diagram of the original skeletal CT image and the gold standard marked by a physician according to an embodiment of the present invention;
[0070] Figure 6 Schematic diagram of the experimental results of skeletal CT image segmentation according to an embodiment of the present invention. Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0072] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0073] Embodiment 1
[0074] A method for segmenting skeletal CT images is provided in an embodiment of the present invention, including:
[0075] Obtain a skeletal CT image;
[0076] Input the skeletal CT image into a pre-trained UTrans-Net network, and the UTrans-Net network outputs a segmented image; the UTrans-Net network includes a connected encoding unit and a decoding unit, and the encoding unit is a dual-encoding structure, including a first encoding structure and a second encoding structure. Among them, the second encoding structure is a Swin Transformer network introducing a multi-size window self-attention module.
[0077] In a specific implementation manner of the embodiment of the present invention, the encoding unit includes a first encoding structure and a second encoding structure, and the encoding unit performs the following operations on the received skeletal CT image:
[0078] Input the skeletal CT image into the first encoding structure, and generate a feature map X1 after downsampling through two groups of first CNN modules. The feature map X1 is downsampled through a second CNN module to generate a feature map X2, and the feature map X2 is downsampled through a third CNN module to generate a feature map X3;
[0079] After the skeletal CT image is sequentially subjected to block division operation and linear encoding, a feature map X-T is generated. The feature map X-T is input into the second encoding structure, and a feature map Y1 is generated after passing through the first multi-size window self-attention module. Then, the feature map Y1 and the feature map X1 are sent into the first cross-attention module to generate a feature map Z1. After performing a block merging operation on the feature map Z1, it is sent into the second multi-size window self-attention module to generate a feature map Y2. Then, the feature map Y2 and the feature map X2 are sent into the second cross-attention module to generate a feature map Z2. After performing a block merging operation on the feature map Z2, it is sent into the third multi-size window self-attention module to generate a feature map Y3. Then, the feature map Y3 and the feature map X3 are sent through the third cross-attention module to generate a feature map Z3.
[0080] In a specific implementation manner of the embodiment of the present invention, the first CNN module, the second CNN module, and the third CNN module all perform the following steps on the received data:
[0081] Convolution operation with a convolution kernel size of 3x3, a stride of 1, an input channel number of n, and an output channel number of n;
[0082] Batch normalization operation;
[0083] Relu activation function operation;
[0084] Convolution operation with a convolution kernel size of 3x3, a stride of 2, an input channel number of n, and an output channel number of 2n;
[0085] Batch normalization operation;
[0086] Relu activation function operation.
[0087] In a specific implementation manner of the embodiment of the present invention, as Figure 2 shown, the first multi-size window self-attention module, the second multi-size window self-attention module, and the third multi-size window self-attention module all perform the following steps on the received data:
[0088] The received input data is evenly divided into 4 groups in the dimension of the input data channel, namely feature map F1, feature map F2, feature map F3, and feature map F4;
[0089] The sliding window self-attention operation with window sizes of [4, 8, 16, 32] is respectively used for feature map F1, feature map F2, feature map F3, and feature map F4, and feature maps A1, A2, A3, and A4 are respectively obtained;
[0090] Feature maps A1, A2, A3, and A4 are merged in the dimension of the input data channel to obtain feature map F1;
[0091] The channel attention operation of SE-Net is performed on feature map F1.
[0092] In a specific implementation manner of the embodiment of the present invention, as Figure 3 shown, the first cross-attention module, the second cross-attention module, and the third cross-attention module all perform the following steps on the received data:
[0093] A convolution operation with a convolution kernel size of 1x1, a stride of 1, an input channel number of n, and an output channel number of n is performed on the received feature map X1, X2, or X3 to generate feature maps X1", X2", or X3";
[0094] A recombination operation is performed on feature maps X1", X2", or X3", changing their size from HxWxC to (HxW)xC, which is consistent with feature maps Y1, Y2, or Y3, to generate feature maps X1"', X2"', or X3"';
[0095] A layer normalization operation is performed on feature maps X1"', X2"', or X3"' to generate feature maps X1"", X2"", or X3"";
[0096] The cross-attention operation is performed on feature maps X1"", X2"", or X3"" and the corresponding feature maps Y1, Y2, or Y3 to generate feature maps Z1', Z2', Z3'.
[0097] Among them, the expressions of the feature maps Z1', Z2', Z3' are:
[0098]
[0099]
[0100]
[0101] Among them, D1, D2, and D3 respectively represent the number of channels of the feature maps Z1', Z2', and Z3'. The superscript T represents the matrix transpose operation.
[0102] In a specific implementation manner of the embodiment of the present invention, the decoding unit receives the data output by the encoding unit and performs the following steps to obtain the segmented image;
[0103] Perform a block expansion operation on the feature map Z3' to increase the resolution, and together with the feature map Z2, successively send them into the first focus fusion module and the fourth multi-scale window self-attention module to generate the feature map Z3'';
[0104] Perform a block expansion operation on the feature map Z3'' to increase the resolution, and together with the feature map Z1, successively send them into the second focus fusion module to generate the feature map Z3''';
[0105] Perform a recombination operation on the feature map Z3''' to change the resolution from (HxW)xC to HxWxC, and perform a linear projection operation to obtain the segmented image.
[0106] In a specific implementation manner of the embodiment of the present invention, as Figure 4 shown, both the first focus fusion module and the second focus fusion module include spatial refinement and channel refinement, and perform the following steps on the received data:
[0107] The two inputs of the focus fusion module are respectively F e , F d ; In the specific implementation process, for the first focus fusion module, its two inputs are respectively the feature map Z3' and the feature map Z2; for the second focus fusion module, its two inputs are respectively the feature map Z3'' and the feature map Z1;
[0108] Perform a linear transformation on one of the received input data F d to obtain the corresponding V d , K d ;
[0109] Perform a linear transformation on the other received input data F e to obtain the corresponding V e , Q e ;
[0110] Let Q d = K d , Q e = K e , to obtain Furthermore, obtain the intermediate generated feature AF d and AF e :
[0111]
[0112]
[0113] For AF d and AF e Perform cross-attention operation to obtain the output feature map F out The feature map F out The calculation formula of is as follows:
[0114]
[0115] Where D is the number of feature channels.
[0116] The method in the embodiments of the present invention will be described in detail below in conjunction with a specific embodiment.
[0117] In the actual application process, the bone CT image segmentation method in the embodiment of the present invention includes the following steps:
[0118] Step 1: Collect the lower limb bone CT images of the patient, and have professional doctors annotate the collected lower limb bone CT images of the patient to generate a gold standard image, and finally form a bone CT data set.
[0119] Step 2: Preprocess the bone CT data set, divide it according to a certain ratio, construct a training set and a test set, and implement an online data augmentation algorithm for the training set data. Specifically, it includes the following steps:
[0120] Step 2.1: Set the window level of the bone CT data set to 300 HU and the window width to 1250 HU.
[0121] Step 2.2: After slicing the bone CT data set along the transverse perspective, a total of 6000 two-dimensional CT images in dcm format are generated. Figure 5 For the comparison diagram of the original collected lower limb bone CT image of the patient and the gold standard image obtained after the doctor's annotation, see Figure 5 .
[0122] Step 2.3: Divide the data set into a training set and a validation set at a ratio of 8:2.
[0123] Step 2.4: Perform data augmentation on the training set. The data augmentation method uses a random online augmentation method, specifically: use data augmentation methods such as rotation, random cropping, contrast enhancement, and random occlusion when inputting into the network for training each time.
[0124] Step 3: Construct a UTrans-Net network. The network structure diagram of UTrans-Net is as shown in Figure 1As shown in the figure, the whole can be divided into an encoding unit and a decoding unit. The main operations are represented by 4 modules, namely: Convolution Module (CNN Block), Multi-Size Shift Window Block (MSwin Block), CrossAttention Module (CAM Block), and Transformer Attention Fusion Module (TAFM Block).
[0125] The encoding unit performs the following steps on the received skeletal CT image:
[0126] Step 3.1.1: The input image Input to be segmented with a size of 512x512x3 generates a feature map X1 with a size of 128x128x96 after passing through two groups of the first CNN modules.
[0127] Step 3.1.2: After the feature map X1 passes through the second CNN module, a feature map X2 with a size of 64x64x192 is generated.
[0128] Step 3.1.3: After the feature map X2 passes through the third CNN module, a feature map X3 with a size of 32x32x384 is generated.
[0129] Step 3.2.1: The input image Input to be segmented with a size of 512x512x3 is divided into blocks of size 4x4 to generate a feature map Input-T with a size of 128x128x48.
[0130] Step 3.2.2: Linearly encode the feature map Input-T to generate a feature map X-T. After passing the feature map X-T through the first multi-size window self-attention module (i.e., MSwin), a feature map Y1 with a size of 128x128x96 is generated. Then, the feature map Y1 and the feature map X1 pass through the first cross-attention module (i.e., CAM module) to generate a feature map Z1 with a size of 128x128x96.
[0131] Step 3.2.3: Perform a Patch Merging operation on Z1 to reduce the size of the feature map. Then, after passing through the second multi-size window self-attention module, a feature map Y2 with a size of 64x64x192 is generated. Then, the feature map Y2 and the feature map X2 pass through the second cross-attention module to generate a feature map Z2 with a size of 64x64x192.
[0132] Step 3.2.3: Perform a Patch Merging operation on the feature map Z2 to reduce the size of the feature map. Subsequently, after passing through the third multi-scale window self-attention module, a feature map Y3 with a size of 32x32x384 is generated. Then, the feature map Y3 and the feature map X3 pass through the third cross-attention module to generate a feature map Z3 with a size of 32x32x384.
[0133] Among them, the specific operations of the first CNN module, the second CNN module, and the third CNN module in steps 3.1.1 to 3.1.3 are as follows:
[0134] A convolution operation with a 3x3 convolution kernel, a stride of 1, an input channel number of n, and an output channel number of n;
[0135] Batch normalization operation;
[0136] Relu activation function operation;
[0137] A convolution operation with a 3x3 convolution kernel, a stride of 2, an input channel number of n, and an output channel number of 2n;
[0138] Batch normalization operation;
[0139] Relu activation function operation.
[0140] The structures of the multi-scale window self-attention modules in steps 3.2.1 to 3.2.3 are as Figure 2 shown, and the specific operations are as follows:
[0141] The implementation of this module is based on the Swin-transformer block structure. Among them, MS-SA is the multi-scale window attention operation proposed in the present invention, which replaces the W-MSA operation in the Swin-transformer block structure. SMS-SA is the mobile multi-scale window attention operation that replaces the SW-MSA operation in the Swin-transformer block structure. The specific operations of this module are as follows:
[0142] Perform normalization operation LN and multi-scale window attention operation MS-SA on the input data Fin of the multi-scale window self-attention module in sequence;
[0143] Perform the first residual connection on the data after the MS-SA operation and the original input data Fin. The output obtained here is used as one of the inputs for the next residual connection. At the same time, this output also serves as one of the inputs for the next residual connection after passing through LN and the multi-layer perceptron MLP;
[0144] The data after the second residual connection passes through LN and the mobile multi-scale window attention operation SMS-SA as one of the inputs for the third residual connection;
[0145] The data after the third residual connection is passed through LN and MLP and used as one of the inputs for the fourth residual connection. After performing the fourth residual connection, Fout is obtained:
[0146] The specific implementation of the MS-SA operation is as follows:
[0147] The input data channels are evenly divided into 4 groups, namely feature map F1, feature map F2, feature map F3, and feature map F4;
[0148] The sliding window self-attention operation with window sizes of [4, 8, 16, 32] is respectively applied to feature map F1, feature map F2, feature map F3, and feature map F4 to obtain feature maps A1, A2, A3, and A4;
[0149] Feature maps A1, A2, A3, and A4 are merged in the channel dimension to obtain feature map F1;
[0150] The channel attention operation of SE-Net is performed on feature map F1.
[0151] SMS-SA performs displacement and masking operations on the elements between windows based on MS-SA, and recalculates self-attention to achieve global modeling. The operation processes of the two are basically the same.
[0152] The structures of each cross-attention module in steps 3.2.1 to 3.2.3 are as Figure 3 shown, and the specific operation is:
[0153] A convolution operation with a convolution kernel size of 1x1, a stride of 1, an input channel number of n, and an output channel number of n is performed on the received feature map X1, X2, or X3 to generate feature maps X1″, X2″, or X3″;
[0154] A recombination operation R is performed on feature maps X1″, X2″, or X3″ to change their size from HxWxC to (HxW)xC, which is consistent with feature maps Y1, Y2, or Y3, to generate feature maps X1″′, X2″′, or X3″′;
[0155] A layer normalization operation LayerNormalization is performed on feature maps X1″′, X2″′, or X3″′ to generate feature maps X1″″, X2″″, or X3″″;
[0156] Linear transformations are performed on feature maps X1″″, X2″″, and X3″″ to obtain Q1, Q2, and Q3;
[0157] After performing a linear transformation Linear on the feature map Y1, feature map Y2, or feature map Y3, K1, K2, K3, and V1, V2, V3 are obtained;
[0158] Perform cross-attention operations on the feature map X1″″, feature map X2″″, or feature map X3″″ and the corresponding feature map Y1, feature map Y2, or feature map Y3 to generate the feature maps Z1′, Z2′, Z3′: The expressions for the feature maps Z1′, Z2′, Z3′ are:
[0159]
[0160]
[0161]
[0162] Among them, D1, D2, and D3 respectively represent the number of channels of the feature map Y1, feature map Y2, and feature map Y3, and the superscript T represents the matrix transpose operation.
[0163] The decoding unit performs the following steps on the received skeletal CT image:
[0164] Step 4.1.1: Perform a block expansion operation (Patch Expanding) on the feature map Z3' to increase the resolution. Together with the feature map Z2, they are successively fed into the first focused fusion module and the fourth multi-scale window self-attention module to generate a feature map Z3″ of 64x64x192;
[0165] Step 4.1.2: Perform a block expansion operation (Patch Expanding) on the feature map Z3″ to increase the resolution. Together with the feature map Z1, they are successively fed into the second focused fusion module to generate a feature map Z3″' of 128x128x96;
[0166] Step 4.1.3: Perform a recombination operation on the feature map Z3″' to change the resolution from (HxW)xC to HxWxC, and perform a linear projection operation (linear projection) to obtain a segmentation image Output of 512x512x2.
[0167] Among them, the specific operation of the focused fusion module is:
[0168] Based on the design idea of guiding with features at different stages as weights, the TAFM module calculates the weights by using cross-attention. The structure of the focused fusion module is as Figure 4 shown. The focused fusion module can be divided into two steps: spatial refinement and channel refinement. Spatial refinement and channel refinement respectively correspond to Figure 4As shown in the first and second dashed boxes, cross-attention operations are performed in the spatial dimension and channel dimension respectively to achieve feature fusion, where Q, K, and V represent the feature maps generated by linear transformation of the input features. The subscript d indicates that it is generated by Fd transformation, and the subscript e indicates that it is generated by Fe transformation. D is the number of feature channels, AF e and AF d is the intermediate generated feature. The specific implementation process can be expressed by the following formula:
[0169] Perform a linear transformation on one of the received input data F d to obtain the corresponding V d , K d ;
[0170] Perform a linear transformation on the other received input data F e to obtain the corresponding V e , Q e ;
[0171] Let Q d =K d , Q e =K e . Under this setting, calculating and operations can be achieved through a simple transpose operation, which can save two linear transformation and matrix multiplication operations and reduce the computational complexity. The specific derivation process is shown in the following formula:
[0172]
[0173] Substitute into the following formula:
[0174]
[0175]
[0176] to obtain the expressions of the final intermediate generated features AF d and AF e :
[0177]
[0178]
[0179] Perform cross-attention operations on AF d and AF e to obtain the output feature map F out , and the expression of the implementation process is:
[0180]
[0181] Finally, the original collected CT images of the patient's lower limb bones after data augmentation and the gold standard images obtained after the physician's annotation are input into the UTrans-Net network to perform corresponding policy updates and optimize the network model parameters. The specific policy is as follows: The open-source deep learning framework used is PyTorch 1.8.0, the Python version is 3.7, and the CUDA version is 11.0. The AdamW optimizer improved from the Adam optimizer is used to train the model parameters. The learning rate is adjusted using the cosine annealing algorithm, and the cosine period is 20 epochs. To avoid instability when training the hybrid architecture model, a warm-up stage is introduced in the learning rate adjustment strategy. The warm-up selects a linear growth method, starting from a learning rate of 0.0002 and growing for 20 epochs to reach the initial learning rate of 0.002, and then the cosine annealing strategy is executed. The maximum number of training times is set to 200, and the batch_size is set to 8. The loss function is implemented by combining the Dice loss function and the Focal loss function, as shown in the specific formula:
[0182] L = L Dice +βL Focal
[0183] where the hyperparameter β is the balancing factor. Through experiments in the present invention, it is found that when β takes the value of 0.5, the best performance can be obtained. The specific formula of the Dice loss function is as follows:
[0184]
[0185] The Focal loss function is as follows. In the selection of hyperparameters, α t takes the value of 0.75 and γ is 2.
[0186] L Focal =-α t (1 - ρ t ) γ log(ρ t ).
[0187] Comparative experiments:
[0188] To evaluate the effectiveness of the UTrans-Net network proposed in the present invention for the bone CT image segmentation task, advanced networks with similar improvement ideas to the present invention are selected as the control group, specifically including: the Baseline network of the present invention, the SwinTransformer Tiny version, denoted as SwinT, the SwinU-Net network improved using the Swin Transformer module, the TransU-Net network improved by serially mixing CNN and Transformer, and the U-Transformer ] network.
[36] Network, Transformer-Unet Improved by Parallel Non-Interactive Use of CNN and Transformer
[37] Network and the Excellent Network AFU-Net for Bone CT Image Segmentation
[0189] Use Dice, IoU, Recall, and Precision as evaluation metrics. The specific metrics are as follows:
[0190] Dice = 2TP / (2TP + FP + FN)
[0191] IoU = TP / (TP + FP + FN)
[0192] Recall = TP / (TP + FN)
[0193] Precision = TP / (TP + FP)
[0194] Among them, TP (True Positives), TN (True Negatives), FP (False Positives), and FN (False Negatives) represent the number of bone pixel points where both the prediction and the label are correct (true positives), the number of background pixel points where both the prediction and the label are correct (true negatives), the number of pixel points where the prediction is bone and the label is background (false positives), and the number of pixel points where the prediction is background and the label is bone (false negatives), respectively. The ranges of the above metrics are all between 0 and 1. The closer to 1, the stronger the prediction ability of the model.
[0195] All models use the same data configuration. For networks with a pure Transformer architecture, set the corresponding pre-trained models, and the remaining hybrid architecture networks are trained starting from random parameters. The results of the comparative experiments are shown in Table 1, where PreTrained indicates that the initial model is a pre-trained model.
[0196] Table 1 Results of Comparative Experiments
[0197]
[0198]
[0199] As can be seen from Table 1, in the task of bone CT image segmentation, SwinT and SwinU-Net without pre-trained weights perform poorly, and are lower than the other comparison networks in various metrics, which also verifies that the network with a pure Transformer architecture has certain limitations on small-scale medical data. After using the pre-trained model weights, the SwinT network has a significant improvement and surpasses the performance of the SwinU-Net network, but there is still a certain gap with the other networks. The serial hybrid architecture networks TransU-Net and U-Transformer obtained 86.56% and 88.12% in the Dice coefficient. The optimization logic of these two networks is to use Transformer as a plug-in module and embed it into U-Net. From the segmentation results, they are indeed better than U-Net, but there is still a certain gap with the algorithm proposed in the present invention.
[0200] The algorithm UTrans-Net of the present invention improves by 2.39% in the Dice coefficient compared with the Transformer-Unet network of the same parallel hybrid architecture, and is the best in various metrics. In terms of the Recall metric, UTrans-Net is 0.43% lower than AFU-Net, and the other metrics have a large increase. Through comprehensive analysis, the effect of UTrans-Net is better than that of the AFU-Net network, indicating that the hybrid architecture is more superior when facing bone CT images. Some experimental results are as Figure 6 shown.
[0201] Ablation experiment:
[0202] The present invention mainly realizes bone CT image segmentation based on the hybrid architecture of CNN and Transformer, and proposes the CAM and TAFM modules on this architecture. Therefore, it is necessary to set up ablation experiments to prove the contribution of the proposed architecture and modules to the overall algorithm. Select Swin Transformer Tiny as the benchmark network and initialize it with the corresponding pre-trained weights, and gradually add the modules proposed in this paper for ablation experiments. The original Swin Transformer uses the UperNet decoder head as the decoding module in the segmentation task. Considering the fairness of the comparative experiment, the following uniformly uses the U-Net-like decoding structure, that is, for the skip connection operation, the channel dimension splicing method and the Patch Merge method are used to reduce the number of channels. The specific settings of the experimental groups are as follows:
[0203] (1) SwinT+PreTrained. Use the Swin Transformer Tiny version network with the U-Net-like decoding structure to predict bone data segmentation, and the initial weight is the pre-trained weight.
[0204] (2) MSwin. Replace the Swin module with the MSwin module based on SwinT+PreTrained.
[0205] (3) SwinT+CAM. Use a dual-encoding branch for feature extraction based on SwinT+PreTrained.
[0206] (4) MSwin+CAM. Replace the Swin module with the MSwin module based on SwinT+CAM.
[0207] (5) UTrans-Net. Replace the feature fusion module with the TAFM module based on MSwin+CAM.
[0208] The experimental data selects the skeletal CT image dataset, and the experimental results are shown in Table 2. As can be seen from Table 2, in the skeletal CT image segmentation task, the pure Transformer architecture Swin Transformer has poor segmentation performance. The main reason is that the amount of skeletal CT image data is small and there are many data characteristics that affect the segmentation accuracy. The MSwin network introduces multi-size windows on the basis of Swin to improve the network's adaptability to different data and enhance the network's ability to extract target objects of different sizes, with a 1.08% increase in the Dice metric. Comparing the MSwin+CAM and Swin+CAM modules, it is found that the gap between having and not having the MSwin module is gradually narrowing, with a difference of 0.31%. And the performance decreases by 0.91% after using multi-size windows in the Precision metric. The main reason analyzed in this paper is that more feature information is introduced in the encoding stage, but the decoding stage cannot well summarize and predict from these features, which in turn causes the network to introduce more misjudgments and identify more false positive results. Moreover, more information will also make the network more difficult to train.
[0209] Table 2 Ablation Experiment Results Table
[0210]
[0211] Comparing SwinT+PreTrained with SwinT+CAM, MSwin with MSwin+CAM, it can be seen that after introducing the convolutional coding branch and the CAM module, the network has improved in multiple metrics. Taking the Dice coefficient as an example, the SwinT+CAM network has an improvement of 2.53% compared to Swin+PreTrained, and MSwin+CAM has an improvement of 1.76% compared to MSwin. Observing the data of the UTrans-Net proposed in this paper, it has an improvement of 4.69% in the Dice coefficient compared to the SwinT network and an improvement of 1.85% compared to the MSwin+CAM network, indicating that the TAFM module proposed in this chapter makes a certain contribution to the optimization of the decoding stage.
[0212] Example 2
[0213] Based on the same inventive concept as in Example 1, an apparatus for segmenting skeletal CT images is provided in an embodiment of the present invention, including:
[0214] An acquisition module, configured to acquire skeletal CT images;
[0215] A segmentation module, configured to input the skeletal CT images into a pre-trained UTrans-Net network, and output segmentation images by the UTrans-Net network; the UTrans-Net network includes a connected encoding unit and a decoding unit, wherein the encoding unit is a dual-encoding structure.
[0216] Example 2
[0217] Based on the same inventive concept as in Example 1, a system for segmenting skeletal CT images is provided in an embodiment of the present invention, including a storage medium and a processor;
[0218] The storage medium is used to store instructions;
[0219] The processor is configured to operate according to the instructions to execute the method according to any one of the embodiments in Example 1.
[0220] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0221] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0222] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0223] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0224] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope of the present invention as defined by the claims. All of these are within the protection scope of the present invention.
[0225] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all of these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for segmenting bone CT images, characterized in that, Including: Obtain a skeletal CT image; Input the skeletal CT image into a pre-trained UTrans-Net network, and the UTrans-Net network outputs a segmentation image; the UTrans-Net network includes a connected encoding unit and a decoding unit, the encoding unit is a dual-encoding structure, including a first encoding structure and a second encoding structure, wherein, the second encoding structure is a Swin Transformer network introducing a multi-size window self-attention module; The encoding unit includes a first encoding structure and a second encoding structure, and the encoding unit performs the following operations on the received skeletal CT image: Input the skeletal CT image into the first encoding structure, and generate a feature map X1 after downsampling through two groups of first CNN modules, generate a feature map X2 after the feature map X1 is downsampled through a second CNN module, and generate a feature map X3 after the feature map X2 is downsampled through a third CNN module; After the skeletal CT image is subjected to block division operation and linear encoding in sequence, a feature map X-T is generated, the feature map X-T is input into the second encoding structure, and a feature map Y1 is generated after passing through the first multi-size window self-attention module, then the feature map Y1 and the feature map X1 are sent into the first cross-attention module to generate a feature map Z1, after the block merging operation is performed on the feature map Z1, it is sent into the second multi-size window self-attention module to generate a feature map Y2, then the feature map Y2 and the feature map X2 are sent into the second cross-attention module to generate a feature map Z2, after the block merging operation is performed on the feature map Z2, it is sent into the third multi-size window self-attention module to generate a feature map Y3, then the feature map Y3 and the feature map X3 are sent through the third cross-attention module to generate a feature map Z3; The decoding unit receives the data output by the encoding unit and performs the following steps to obtain a segmentation image; Perform a block expansion operation on the feature map Z3' to enlarge the resolution, together with the feature map Z2, and are sequentially sent into the first focusing fusion module and the fourth multi-size window self-attention module to generate a feature map Z3”; Perform a block expansion operation on the feature map Z3” to enlarge the resolution, together with the feature map Z1, and are sequentially sent into the second focusing fusion module to generate a feature map Z3”'; Perform a recombination operation on the feature map Z3”' to change the resolution from (HxW)xC to HxWxC, and perform a linear projection operation to obtain a segmentation image.
2. The bone CT image segmentation method according to claim 1, wherein: The first CNN module, the second CNN module, and the third CNN module all perform the following steps on the received data: Convolution operation with a convolution kernel size of 3x3, a stride of 1, an input channel number of n, and an output channel number of n; Batch normalization operation; Relu activation function operation; Convolution operation with a convolution kernel size of 3x3, a stride of 2, an input channel number of n, and an output channel number of 2n; Batch normalization operation; Relu activation function operation.
3. A method for segmenting bone CT images according to claim 1, characterized in that: The first multi-size window self-attention module, the second multi-size window self-attention module, and the third multi-size window self-attention module all perform the following steps on the received data: Perform the normalization operation LN and the multi-scale window attention operation MS-SA on the input data Fin in sequence; Perform the first residual connection on the data after the MS-SA operation and the original input data Fin. The output obtained here is used as one of the inputs for the next residual connection. At the same time, the output also serves as one of the inputs for the next residual connection after passing through LN and the multi-layer sensor MLP; The data after the second residual connection passes through LN and the shifted multi-scale window attention operation SMS-SA and serves as one of the inputs for the third residual connection; The data after the third residual connection passes through LN and MLP and serves as one of the inputs for the fourth residual connection. After performing the fourth residual connection, Fout is obtained; The specific implementation of the MS-SA operation is as follows: Divide the received input data evenly into 4 groups along the input data channel dimension, namely feature map F1, feature map F2, feature map F3, and feature map F4; Perform sliding window self-attention operations with window sizes of [4, 8, 16, 32] on feature map F1, feature map F2, feature map F3, and feature map F4 respectively to obtain feature maps A1, A2, A3, and A4. Merge feature maps A1, A2, A3, and A4 along the input data channel dimension to obtain feature map F1; Perform channel attention operation of SE-Net on feature map F1.
4. A method for segmenting bone CT images according to claim 2, characterized in that: The first cross-attention module, the second cross-attention module, and the third cross-attention module all perform the following steps on the received data: Perform a convolution operation with a convolution kernel size of 1x1, a stride of 1, an input channel number of n, and an output channel number of n on the received feature maps X1, X2, and X3 to generate feature maps X1″, X2″, and X3″; Perform a reorganization operation R on feature maps X1″, X2″, and X3″ to change their size from HxWxC to (HxW)xC, which is consistent with feature maps Y1, Y2, and Y3, to generate feature maps X1″′, X2″′, and X3″′; Perform layer normalization operation Layer Normalization on feature maps X1″′, X2″′, and X3″′ to generate feature maps X1″″, X2″″, and X3″″; Perform a linear transformation Linear on feature maps X1″″, X2″″, and X3″″ to obtain Q1, Q2, and Q3; Perform a linear transformation Linear on feature maps Y1, Y2, and Y3 to obtain K1, K2, and K3 and V1, V2, and V3; Perform a cross-attention operation on feature maps X1″″, X2″″, and X3″″ and the corresponding feature maps Y1, Y2, and Y3 to generate feature maps Z1′, Z2′, and Z3′.
5. A method for segmenting bone CT images according to claim 4, characterized in that: The expressions of the feature maps Z1′, Z2′, and Z3′ are: where D1, D2, and D3 represent the number of channels of feature maps Y1, Y2, and Y3 respectively, and the superscript T represents the matrix transpose operation.
6. A method for segmenting bone CT images according to claim 5, characterized in that: Both the first focus fusion module and the second focus fusion module perform the following steps on the received data: For one of the received input data F d Perform linear transformation to obtain the corresponding V d , K d ; For another received input data F e perform a linear transformation to obtain the corresponding V e , Q e ; Let Q d = K d , Q e = K e , obtain Furthermore, obtain the intermediate generated feature AF d and AF e : For the intermediate generated feature AF d and AF e perform cross-attention operation to obtain the output feature map F out , the calculation formula of the feature map F out is as follows: where D is the number of feature channels.
7. A bone CT image segmentation device, characterized in that, It includes: An acquisition module for acquiring bone CT images; A segmentation module for inputting the bone CT image into a pre-trained UTrans-Net network, and the UTrans-Net network outputs a segmentation image; the UTrans-Net network includes a connected encoding unit and a decoding unit, where the encoding unit is a double-encoding structure; The encoding unit includes a first encoding structure and a second encoding structure, and the encoding unit performs the following operations on the received bone CT image: Input the bone CT image into the first encoding structure, and generate a feature map X1 after downsampling through two groups of first CNN modules. The feature map X1 is downsampled through a second CNN module to generate a feature map X2, and the feature map X2 is downsampled through a third CNN module to generate a feature map X3; After the bone CT image is sequentially subjected to a block division operation and linear encoding, a feature map X-T is generated. The feature map X-T is input into the second encoding structure, and a feature map Y1 is generated after passing through a first multi-size window self-attention module. Then, the feature map Y1 and the feature map X1 are sent into a first cross-attention module to generate a feature map Z1. After performing a block merging operation on the feature map Z1, it is sent into a second multi-size window self-attention module to generate a feature map Y2. Then, the feature map Y2 and the feature map X2 are sent into a second cross-attention module to generate a feature map Z2. After performing a block merging operation on the feature map Z2, it is sent into a third multi-size window self-attention module to generate a feature map Y3. Then, the feature map Y3 and the feature map X3 are sent into a third cross-attention module to generate a feature map Z3; The decoding unit receives the data output by the encoding unit and performs the following steps to obtain a segmentation image; Perform a block expansion operation on the feature map Z3' to enlarge the resolution, and together with the feature map Z2, they are sequentially sent into the first focus fusion module and a fourth multi-size window self-attention module to generate a feature map Z3”; Perform a block expansion operation on the feature map Z3” to enlarge the resolution, and together with the feature map Z1, they are sequentially sent into the second focus fusion module to generate a feature map Z3”'; Perform a recombination operation on the feature map Z3”' to change the resolution from (HxW)xC to HxWxC, and perform a linear projection operation to obtain a segmentation image.
8. A bone CT image segmentation system, characterized in that, It includes a storage medium and a processor; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Lower limb skeleton CT image segmentation algorithm based on improved U-Net
CN114913327A
Medical image segmentation model construction method based on multi-attention fusion
CN116309648A