A Thyroid Nodule Segmentation Method Based on Improved Unet Network
By introducing self-attention residual connection blocks, high and low frequency self-attention adaptive fusion modules and semantic enhancement modules in the Unnet network, the ultrasonic thyroid nodule segmentation method is improved, solving the problems of low segmentation accuracy and noise interference in the prior art, and achieving higher segmentation performance.
Patent Information
- Application Number
- CN202210723059.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-24
AI Technical Summary
The prior art is difficult to achieve high accuracy in ultrasonic thyroid nodule segmentation, and there are problems of noise interference and blurred edges, resulting in inaccurate diagnosis and treatment by doctors.
The improved Unet network is adopted to enhance the model's ability to capture global context and local information by adding self-attention residual connection blocks in the encoding stage, adding high and low frequency self-attention adaptive fusion module after downsampling convolution blocks, and connecting to the semantic enhancement module.
The accuracy of thyroid nodule segmentation is improved and the error is reduced. Compared with the Unet network, the Dice similarity coefficient and the average interaction ratio on the same data set are improved by 2.1% and 1.3%, respectively.
Smart Images

Figure CN114998296B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a thyroid nodule segmentation method based on an improved Unet network. Background Art
[0002] Currently, ultrasound imaging has become the preferred technology for thyroid diagnosis due to its advantages of low cost, non-radiation, and real-time. However, due to the low contrast of ultrasound images, the presence of a large amount of noise, and relatively blurred edges, conventional segmentation methods are difficult to achieve high precision in the segmentation of ultrasound thyroid nodules. Inaccurate segmentation is likely to cause misdiagnosis and mistreatment by doctors, with extremely serious consequences.
[0003] With the popularity of deep learning and the development of computer medical technology, people have increasingly focused on medical image segmentation technologies related to deep learning. CNN is widely used in the field of deep learning due to its powerful ability to handle scale invariance and modeling inductive bias. Among them, the Unet network proposed by Ronneberger still has good segmentation results in the field of medical images with scarce data volume.
[0004] Since the role of convolution is local, it is prone to the problem of limited receptive field, which poses a certain challenge to the capture of global context information, that is, the lack of the auxiliary effect of the surrounding tissues of the thyroid on nodule segmentation. At the same time, due to the lack of semantic prior information, it is easy to have the problem of boundary merging between smaller nodules and larger nodules, and the segmentation performance is difficult to reach an ideal state. The above problems can be solved by introducing a self-attention mechanism. However, the ordinary self-attention mechanism ignores the features in different frequency domains and pays too much attention to low-frequency global information, easily losing local details such as edges, resulting in inaccurate segmentation results. Summary of the Invention
[0005] The technical problem solved by the present invention is: to help doctors accurately segment thyroid nodules, precisely locate the nodule area, and at the same time enhance global context and local information, reducing the error of thyroid nodule segmentation.
[0006] The technical solution adopted by the present invention is: a thyroid nodule segmentation method based on an improved Unet network, including the following steps:
[0007] Step 1: Uniform the image size, randomly flip, rotate, and enhance the contrast of the public thyroid ultrasound image dataset, and divide it into a training set, a validation set, and a test set according to 8:1:1;
[0008] Step 2: Construct an improved Unet network. First, add a self-attention residual connection block after the first 3 downsampling convolutional blocks of the Unet network. Second, add a high-low frequency self-attention adaptive fusion module after the 4th downsampling convolutional block. Finally, connect a semantic enhancement module after the high-low frequency self-attention adaptive fusion module;
[0009] Adding a self-attention residual connection block in the encoding stage can capture long-range dependencies and global context information at multiple scales, and at the same time, initialize self-attention as convolution, thus making up for the defect that self-attention requires a large amount of data to learn position bias;
[0010] The high-low frequency self-attention adaptive fusion module is used to enhance the global context dependence of semantic features, and at the same time, enhance the accuracy of feature extraction by capturing local details at high frequencies and focusing on the global structure at low frequencies;
[0011] Step 3: Use argmax to determine whether the pixel value belongs to a nodule or the background.
[0012] Advantages of the present invention:
[0013] 1. Taking the Unet network as the main architecture of the network model of the present invention can help the model still achieve ideal results in the case of lack of data volume. Adding a self-attention residual connection block in the encoding stage can capture long-range dependencies and global context information at multiple scales, and at the same time, initialize self-attention as convolution, thus making up for the defect that self-attention requires a large amount of data to learn position bias;
[0014] 2. The high-low frequency self-attention adaptive fusion module endows self-attention with the ability to capture local edge details at high frequencies and global context structures at low frequencies simultaneously;
[0015] 3. The semantic enhancement module enhances the semantic features in the feature extraction stage by adding semantic prior information. The segmentation performance of the method of the present invention has been greatly improved on the thyroid dataset tn3k. The DSC has increased by 2.1% compared with Unet, and the mIoU has increased by 1.3% compared with Unet; for the problems of too many interference factors such as noise and unclear edges in the thyroid nodule segmentation task, it has been better solved. Description of the Drawings
[0016] Figure 1 is the flow chart of the thyroid nodule segmentation method based on the improved Unet network of the present invention;
[0017] Figure 2 is the image segmentation model diagram of the present invention;
[0018] Figure 3 is the structural diagram of the internal module of the segmentation model of the present invention;
[0019] Figure 4 is the high - low frequency self - attention adaptive fusion module of the present invention;
[0020] Figure 5 is the semantic enhancement module of the present invention;
[0021] Figure 6 is the segmentation effect diagram of the trained model of the present invention. Detailed implementation manners
[0022] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.
[0023] As Figure 1 shown, the thyroid nodule segmentation method based on the improved Unet network includes the following steps:
[0024] Step 1: Perform operations such as unifying the image size, randomly flipping, rotating, and enhancing the contrast on the public thyroid ultrasound image dataset, and divide it into a training set, a validation set, and a test set according to 8:1:1;
[0025] S1. Crop the public ultrasound thyroid nodule dataset tn3k into images with a size of 256×256. There are a total of 2879 training and validation images;
[0026] Step 2: Construct an improved Unet network. First, add self - attention residual connection blocks to the first 3 downsampling convolutions in the Unet network; secondly, add a high - low frequency self - attention adaptive fusion module after the 4th downsampling; finally, connect a semantic enhancement module after the high - low frequency self - attention adaptive fusion module;
[0027] Adding self - attention residual connection blocks in the encoding stage can capture long - range dependencies and global context information at multiple scales, and at the same time, initialize self - attention as convolution, thereby making up for the defect that self - attention requires a large amount of data to learn position bias;
[0028] The high - low frequency self - attention adaptive fusion module is used to enhance the global context dependence of semantic features. At the same time, by capturing local details with high frequency and paying attention to the global structure with low frequency, the accuracy of feature extraction is enhanced;
[0029] S2. Build a thyroid nodule segmentation network based on self - attention residual connection blocks, high - low frequency self - attention adaptive fusion modules, and semantic enhancement modules, that is, an improved Unet network, and send the thyroid image data read by pytorch into the network;
[0030] Furthermore, as Figure 2 is the diagram of the improved Unet network, which specifically includes:
[0031] S21. The Unet main structure includes a double convolution block, 4 downsampling convolution blocks, and 4 upsampling convolution blocks. The downsampling convolution block consists of max pooling and a double convolution block. The double convolution block is composed of two groups of 3×3 convolutions, a BN normalization layer, and a ReLU activation function. The downsampling convolution block consists of max pooling and a double convolution block. The double convolution block is composed of two groups of 3×3 convolutions, a BN normalization layer, and a ReLU activation function. The upsampling convolution block consists of bilinear upsampling and a double convolution block. The double convolution block is composed of two groups of 3×3 convolutions, a BN normalization layer, and a ReLU activation function.
[0032] A double convolution block changes the number of input channels to 64, and then passes through 4 downsampling convolution blocks. The number of channels becomes 128, 256, 512, 512 respectively. After that, it passes through 4 upsampling convolution blocks, and the number of channels becomes 256, 128, 64, 64. Finally, a 1×1 convolution turns it into the number of categories.
[0033] S22. As Figure 3 is the structure diagram of the self-attention residual connection block. The self-attention residual connection block is connected behind the first 3 downsampling convolution blocks, connecting the self-attention in a residual manner and initializing it as a convolution. The self-attention residual connection block is divided into two residual connection blocks. The first residual connection block is a BN normalization and an efficient self-attention mechanism (ESA) for multi-heads. In the efficient self-attention for multi-heads, the feature is first mapped into query (Q) through a Linear layer, then the feature is sent into the max pooling layer to become a feature map with a size of 8×8, and then through a Linear layer to calculate key (K) and value (V) in the self-attention to reduce the calculation. The formula is as follows:
[0034]
[0035] where Q is the query of the original feature mapping, and K, V are the key and value after being mapped through the max pooling layer and then mapped again; T is the transpose calculation; D h is the hidden dimension number of one head of the multi-head attention mechanism. In this embodiment is used to have a more stable gradient during training; Softmax is the activation function used for normalization.
[0036] The second residual connection block consists of BN normalization, a ReLU activation function, and a 3×3 convolution.
[0037] S23. As Figure 4It is the structural diagram of the high-low frequency self-attention adaptive fusion module. The high-low frequency self-attention adaptive fusion module is placed behind the 4th downsampling convolutional block in the form of residual connection, which is used to enhance the global context dependence of semantic features. At the same time, it captures local details through high frequency and focuses on the global structure through low frequency to enhance the accuracy of feature extraction. The high-low frequency self-attention adaptive fusion module consists of window multi-head attention and efficient multi-head self-attention in parallel. Among them, the window multi-head attention first divides the entire feature (512×16×16) into 4 small windows (512×8×8), calculates self-attention within each window respectively, and then integrates them. The query (Q) in the efficient multi-head self-attention is mapped by the Linear linear layer, while the key (K) and value (V) need to be respectively max-pooled and average-pooled first. The size of the pooled feature is 8×8. Then, the two pooled results are added together, and the key (K) and value (V) are calculated through the Linear linear layer. And two learnable parameters are set, initialized to 1, and the features of high-low frequency attention are adaptively added to the original feature in the form of residual connection. The formula is as follows:
[0038] y = x ori + a1×x gao + a2×x di (2)
[0039] Among them, x ori is the incoming feature, x gao , x di are the features processed by high-low frequency attention respectively, and a1 and a2 are two learnable weight parameters;
[0040] S24, such as Figure 5 is the semantic enhancement module. The semantic enhancement module is placed after the high-low frequency self-attention adaptive fusion module. The features after the high-low frequency self-attention adaptive fusion module are passed through the Linear linear layer to convert the number of channels into the number of categories 2 as the query (Q) and key (K). Similarly, it passes through the Linear linear layer without changing the number of channels as the value (V). Then, Q is converted into 512×16×16 and bilinearly upsampled by 8 times to the original image size, and supervised with the label. The generated loss is loss2. At the same time, the calculation of semantic attention (YSA) is carried out, and Q, K, and V are passed into the following formula:
[0041] YSA = Softmax(QK T )V (3)
[0042] Among them, Q, K, and V are the query, key, and value mapped by the Linear layer, T is the transpose, and Softmax is the activation function used for normalization.
[0043] Then, the features processed by semantic attention are passed through a Linear layer, and a learnable parameter λ is set to fine-tune the features. Then, they are added to the features connected by the residual connection to obtain the final semantically enhanced features. The loss generated by the semantic enhancement module is added to the loss generated by the final prediction map and enters the backpropagation process. The loss weight value of semantic enhancement is 0.4.
[0044] Step 3: Use argmax to determine whether the pixel value belongs to a nodule or the background.
[0045] For the loss function, the cross-entropy loss function and the Dice loss function are used for weighted summation, and each is given a weight value of 0.5.
[0046] Use the Adam optimizer for backpropagation. The initial learning rate is set to 0.0001. For learning rate adjustment, first warm up for 5 rounds, and then use the cosine annealing algorithm, that is, the learning rate is decreased through the cosine function. The training batch_size is 24. The device used in this experiment is a Tesla V100, and the software environment is python3.8 and pytorch1.7.0.
[0047] To verify the performance of the model of the present invention, the Dice similarity coefficient (dsc) and the mean intersection over union (miou) are used for experimental evaluation. The index calculation formulas are as follows:
[0048]
[0049]
[0050] Among them, P is the segmentation result predicted by the network, and T is the true label.
[0051] Finally, it is judged that the network has been stably trained on the validation set, and the model weights of the round with the highest dsc are selected for testing on the test set. It can be seen that the segmentation model of the present invention has improved by 2.1% in dsc and 1.3% in miou compared to Unet.
[0052] network structure dsc miou Unet 75.9% 84.8% the present invention 78.0% 86.1%
[0053] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A method for segmenting thyroid nodules based on an improved Unet network, characterized in that, It includes the following steps: Step 1: Perform operations such as unifying the image size, randomly flipping, rotating, and enhancing the contrast on the public thyroid ultrasound image dataset, and divide it into a training set, a validation set, and a test set according to 8:1:1; Step 2: Build an improved Unet network. First, add a self-attention residual connection block after the first 3 downsampling convolutional blocks of the Unet network; second, add a high-low frequency self-attention adaptive fusion module after the 4th downsampling convolutional block; finally, connect a semantic enhancement module after the high-low frequency self-attention adaptive fusion module; The detailed steps of adding a high-low frequency self-attention adaptive fusion module after the 4th downsampling convolutional block are as follows: The high-low frequency self-attention adaptive fusion module consists of window multi-head attention and efficient multi-head self-attention in parallel. Among them, the window multi-head attention first divides the entire feature 512×16×16 into 4 small windows 512×8×8, and calculates self-attention within each window respectively. The Q in the efficient multi-head self-attention is mapped by the Linear linear layer, while K and V are respectively subjected to max pooling and average pooling, and then the results of the two poolings are added, and then K and V are calculated through the Linear linear layer; and two learnable weight parameters are set to adaptively add the features of high-low frequency attention to the original features in a residual connection manner. The formula is as follows: y = x ori + a1 × x gao + a2 × x di (2) Among them, x ori is the incoming feature, and x gao , x di are the features processed by high and low frequency attention respectively, and a1 and a2 are two learnable weight parameters; The detailed steps of connecting a semantic enhancement module after the high-low frequency self-attention adaptive fusion module are as follows: The features after the high-low frequency self-attention adaptive fusion module are passed through a Linear linear layer to convert the number of channels into the number of categories 2 as Q and K. Similarly, it passes through a Linear linear layer without changing the number of channels as V. Then Q is converted into 512×16×16 and bilinearly upsampled 8 times to the original image size, and is supervised with the label. The generated loss is loss2. At the same time, semantic attention is calculated, and Q, K, and V are passed into the following formula: YSA = Softmax(QK T )V (3) Among them, Q, K, and V are query, key, and value mapped through the Linear layer, and T is the transpose; Then the features processed by semantic attention are passed through a Linear linear layer, and a learnable parameter λ is set to fine-tune the features, and then added to the features of the residual connection to obtain the finally semantically enhanced features. The loss generated by the semantic enhancement module is added to the loss generated by the prediction map and enters the backpropagation process; Step 3: Use argmax to determine whether the pixel value belongs to a nodule or the background.
2. The thyroid nodule segmentation method based on the improved Unet network according to claim 1, characterized in that: The Unet network includes a double convolutional block, 4 downsampling convolutional blocks, and 4 upsampling convolutional blocks; the downsampling convolutional block consists of max pooling and a double convolutional block. Among them, the double convolutional block is two groups of 3×3 convolutions, a BN normalization layer, and a ReLU activation function; the downsampling convolutional block consists of max pooling and a double convolutional block. Among them, the double convolutional block is two groups of 3×3 convolutions, a BN normalization layer, and a ReLU activation function; the upsampling convolutional block consists of bilinear upsampling and a double convolutional block. Among them, the double convolutional block is two groups of 3×3 convolutions, a BN normalization layer, and a ReLU activation function; A double convolutional block changes the number of input channels to 64, and then passes through 4 downsampling convolutional blocks, where the number of channels becomes 128, 256, 512, and 512 respectively. After that, it passes through 4 upsampling convolutional blocks, and the number of channels becomes 256, 128, 64, and 64. Finally, a 1×1 convolution is performed to obtain the number of classes.
3. The thyroid nodule segmentation method based on the improved Unet network according to claim 2, wherein, The detailed steps for adding self-attention residual connection blocks after the first 3 downsampling convolutional blocks in the Unet network are as follows: The self-attention residual connection block is connected behind the first 3 downsampling convolutional blocks. The self-attention residual connection block contains two residual connection blocks. The first residual connection block consists of BN normalization and an efficient multi-head self-attention mechanism. In the efficient multi-head self-attention, the features are first mapped to Q through a Linear layer, then the features are transformed into an 8×8 feature map through a max pooling layer, and then K and V in the self-attention are calculated through a Linear layer. The formula is as follows: Among them, Q is the query mapped through the Linear layer, K and V are the key and value mapped through the Linear layer after max pooling, T is the transpose, and D h is the hidden dimension number of one head of the multi-head attention mechanism, and Softmax is the activation function; The second residual connection block consists of BN normalization, a ReLU activation function, and a 3×3 convolution.
Citation Information
Patent Citations
Rapid space-time residual attention video super-resolution reconstruction method
CN111028150A
Cavity convolutional neural network image super-resolution reconstruction method based on attention mechanism
CN111047515A