Intelligent Classification Method for Thyroid Nodules with Dual Encoding Structure and Cascade Dilated Convolution
By using dual-coded structure and cascaded cavity convolution in the intelligent classification method of thyroid nodules, combined with the encoder of Transformer network and convolutional neural network, the problem of low classification accuracy of a small number of ultrasound images in the prior art is solved, and a higher classification accuracy of thyroid nodules and more efficient diagnostic assistance is achieved.
Patent Information
- Application Number
- CN202310574449.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-22
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-05-22
AI Technical Summary
When existing deep learning methods use a small number of ultrasound images to classify thyroid nodules, there are problems with risk of overfitting and low classification accuracy, which makes it difficult to effectively assist doctors in discrimination of benign and malignant nature.
An intelligent classification method for thyroid nodules based on double-coding structure and cascaded cavity convolution is proposed. Combined with the Encoder of Transformer Network and Convolutional Neural Network, high-level semantic information is extracted through the cascaded cavity convolution module to improve classification accuracy.
This method significantly improves the accuracy of thyroid nodule classification, reduces the doctor's time to read and diagnose, improves work efficiency, and provides support for the prognosis and improvement of treatment methods for patients with thyroid nodules.
Smart Images

Figure CN116612086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to an intelligent classification method for thyroid nodules with a double coding structure and cascaded dilated convolutions. Background Art
[0002] The thyroid gland is an important organ that regulates the hormone balance in the body. A thyroid nodule refers to a discrete lesion formed within the thyroid gland parenchyma, and its radiation characteristics are significantly different from those of the surrounding tissues. Thyroid nodules are very common in the general population, and the detection rate is about 19% to 68% using high-resolution ultrasound imaging. Usually, only nodules larger than 1 cm need to be further evaluated because they exhibit a higher potential for cancer. However, in rare cases, some nodules smaller than 1 cm may also affect future health conditions and mortality. Thyroid cancer accounts for 3% of all cancer incidences globally, with 586,000 new patients generated annually.
[0003] Since high-resolution ultrasound imaging provides a non-invasive, real-time, and low-cost examination method, it has become the best choice for diagnosing thyroid nodules. However, due to the influence of echo and speckle noise on the images, experienced radiologists usually diagnose nodules based on the shape, edge, and boundary of the ultrasonic features in the ultrasound image slices. This method is quite subjective and highly dependent on the clinical experience of the radiologists. Therefore, computer-aided classification systems are very important for automatically classifying and differentiating benign and malignant thyroid nodules using ultrasound images. This method can significantly reduce the workload of radiologists, improve the accuracy of nodule classification, and reduce the risk of diagnostic errors by inexperienced young radiologists.
[0004] In recent years, due to the better performance of deep learning methods than traditional learning methods, deep learning models have been widely applied to image classification tasks. An important advantage of deep learning is its ability to extract deep features from ultrasound images, which may be difficult for human radiologists to obtain through visual inspection. In addition, deep learning can also integrate feature extraction and classification into a unified framework, avoiding the complex process of manual feature extraction and classifier selection. Therefore, different deep learning methods have been proposed for different classification tasks.
[0005] In recent years, research on thyroid nodule classification based on deep neural networks has been carried out. However, since the number of natural images is much larger than that of medical images, it is a difficult task to achieve the same accuracy of thyroid nodule classification on the dataset. Different from natural images, it is difficult to obtain millions of ultrasound images in clinical practice. Therefore, it is a challenge to use a small number of ultrasound images to train a deep learning model for thyroid nodule classification. In clinical practice, experienced radiologists distinguish the benign or malignant nature of thyroid nodules by visually examining ultrasound images. However, this process is not only time-consuming and laborious but also extremely subjectively biased.
[0006] Data is an important factor in the performance of classification networks based on deep learning. Although classification networks with good performance in natural image classification, it is difficult to achieve the same high level of accuracy in medical image classification. Many existing thyroid nodule classification methods use natural image classification networks as the backbone network architecture. However, this kind of classification network is not fully adapted to medical images because the number of thyroid nodule images is much less than that of natural images. In the case of a small amount of data, there is a risk of overfitting when using deep networks suitable for natural images for thyroid nodule classification.
[0007] Therefore, the present invention proposes a new intelligent classification method for thyroid nodules based on a dual-encoding structure and cascaded dilated convolutions. Summary of the Invention
[0008] The present invention proposes an intelligent classification method for thyroid nodules with a dual-encoding structure and cascaded dilated convolutions, which can improve the accuracy of thyroid nodule classification.
[0009] The present invention adopts the following technical solutions.
[0010] An intelligent classification method for thyroid nodules with a dual-encoding structure and cascaded dilated convolutions, which is used for classifying medical images. The method is based on a classification network model, which fuses the encoders of both the Transformer network and the convolutional neural network to form an enhanced dual-encoder structure. The Swin Transformer of the Transformer network is used as the backbone of the first encoder to capture the long-range dependencies between the pixels of medical images; the Efficientnet-b3 module based on the convolutional neural network is used as the backbone of the second encoder, and the most suitable model parameters are found by scaling the three dimensions of the depth, width, and resolution of the network, realizing the simplification of the composite coefficient and the improvement of the efficiency of the classification network model.
[0011] The method includes the following steps;
[0012] Step S1: Perform medical image preprocessing operations; if the number of medical image datasets is limited, expand the dataset to reduce the risk of overfitting. The medical image preprocessing operation methods include image transformation and noise interference; image transformation includes vertical flipping, horizontal flipping, and rotation of the image, and the rotation direction angles include 45 / 90 / 135 / 180 / 225 / 270 / 315;
[0013] Step S2: Resize the medical images to a preset size and train the neural network. Use the Adam optimizer when training and validating the neural network, and terminate the training process after a preset number of epochs for the neural network;
[0014] Step S3: Input the medical images resized to the preset size into two encoders, namely Swin Transformer and convolutional neural network, for feature extraction, and then concatenate the outputs of the two encoders to fuse the two outputs; then perform cascaded dilated convolution CDI operations to extract the high-level semantic information existing in the images, and finally perform global average pooling GAP operations. After passing through the fully connected FC layer, give the classification prediction results of thyroid benign and malignant.
[0015] The classification network model consists of a dual-encoder structure, a cascaded dilated convolution module, a global average pooling GAP, and a fully connected FC layer; in the dual-encoder, the Transformer encoder of the Transformer network learns long-range dependencies, and the encoder based on the convolutional neural network captures high-level features; the outputs of the Transformer encoder and the encoder based on the convolutional neural network are both connected to the cascaded dilated convolution CDI module to further improve the classification performance.
[0016] The cascaded dilated convolution CDI module is used to extract high-level semantic information. It consists of four cascaded branches, including 3 × 3 max pooling, 1 × 1 convolutional layer, 1 × 1 and 5 × 5 convolutional layers, using dilated convolution with a dilation rate of 1, and also includes 1 × 1 and 3 × 3 convolutional layers, using dilated convolution with a dilation rate of 1.
[0017] The Swin Transformer replaces the standard multi-head self-attention (MSA) module with a shifted window module, which restricts self-attention calculations to non-overlapping local windows through a shifted window strategy to improve computational efficiency. The Swin Transformer module of the Swin Transformer includes a shifted window MSA and a two-layer MLP with a Gaussian error linear unit (GELU) non-linearity. A LayerNorm layer (LN) is placed before each MSA and MLP module to enable residual connections between two Swin Transformer modules.
[0018] The EfficientNet-b3 module based on the convolutional neural network is constructed based on the neural architecture search method and uses multiple MBConv blocks stacked together. The EfficientNet-b3 module finds the most suitable model parameters by scaling three dimensions of the neural network: depth, width, and resolution.
[0019] In order to reduce the significant amount of effort and time spent by doctors in classifying the benign and malignant nature of thyroid nodules during film reading, the present invention proposes an intelligent classification method for thyroid nodules based on a dual-encoding structure and cascaded dilated convolution. This method combines the dual-encoding structures of the Swin Transformer and the convolutional neural network, integrating the advantages of both, and taking into account the extraction of local and long-range information. The cascaded dilated convolution (CDI) module is used to extract high-level semantic information, further improving the accuracy of thyroid nodule classification. The present invention fills the gap in the discrimination of the benign and malignant nature of nodules in the field of thyroid image processing by current deep learning algorithms, laying a foundation for the application of artificial intelligence technology in the field of thyroid nodule assistance.
[0020] The present invention first proposes an algorithm based on a dual-encoder combined with cascaded dilated convolution to assist doctors in making judgments. Given a thyroid MRI image as input, the algorithm automatically outputs the probability values of the image being benign and malignant predicted by the algorithm. Automatically obtaining the preliminary judgment result of the benign and malignant nature of thyroid nodules can greatly reduce the film reading time and diagnostic difficulty of doctors, improve the work efficiency of doctors, and play a positive role in the prognosis and improvement of treatment methods for thyroid nodule patients, having very important significance and broad prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The following further details the present invention in conjunction with the drawings and specific embodiments:
[0022] Att Figure 1 is a schematic flow diagram of the present invention;
[0023] Att Figure 2 is a schematic diagram of the classification network model of the present invention;
[0024] Attached Figure 3 is a schematic diagram of the residual connection between two Swin Transformer modules;
[0025] Attached Figure 4 is a schematic diagram of the main architecture of the EfficientNet block;
[0026] Attached Figure 5 is a schematic diagram of the cascaded dilated convolution CDI module;
[0027] Attached Figure 6 is a schematic diagram of medical image preprocessing operations. Detailed implementation manners
[0028] As shown in the figure, a thyroid nodule intelligent classification method with a dual-encoding structure and cascaded dilated convolution is used for classifying medical images. The method is based on a classification network model. The classification network model fuses the encoders of both the Transformer network and the convolutional neural network to form an enhanced dual-encoder structure. The Swin Transformer of the Transformer network is used as the backbone of the first encoder to capture the long-range dependencies between the pixels of medical images; the Efficientnet-b3 module based on the convolutional neural network is used as the backbone of the second encoder, and the most suitable model parameters are found by scaling the three dimensions of the depth, width, and resolution of the network, realizing the simplification of the composite coefficient and the improvement of the efficiency of the classification network model.
[0029] The method includes the following steps;
[0030] Step S1, perform medical image preprocessing operations; if the number of the medical image dataset is limited, the dataset is augmented to reduce the risk of overfitting. The medical image preprocessing operation method includes image transformation and noise interference; the image transformation includes vertical flipping, horizontal flipping, and rotation of the image, and the rotation direction angles include 45 / 90 / 135 / 180 / 225 / 270 / 315;
[0031] Step S2, adjust the size of the medical image to a preset size and train the neural network. The Adam optimizer is used during the training and validation of the neural network, and the training process is terminated after a preset number of epochs for the neural network;
[0032] Step S3: The medical images adjusted to the preset size are respectively input into two encoders, namely Swin Transformer and convolutional neural network, for feature extraction. Then, the outputs of the two encoders are concatenated and fused. Next, a cascaded dilated convolution CDI operation is performed to extract the high-level semantic information existing in the image. Finally, a global average pooling GAP operation is carried out. After passing through the fully connected FC layer, the classification prediction result of thyroid benign and malignant is given.
[0033] The classification network model consists of a dual-encoder structure, a cascaded dilated convolution module, a global average pooling GAP, and a fully connected FC layer. In the dual-encoder, the Transformer encoder of the Transformer network learns long-range dependencies, and the encoder based on the convolutional neural network captures high-level features. The outputs of the Transformer encoder and the encoder based on the convolutional neural network are both connected to the cascaded dilated convolution CDI module to further improve the classification performance.
[0034] The cascaded dilated convolution CDI module is used to extract high-level semantic information. It consists of four cascaded branches, including 3×3 max pooling, 1×1 convolutional layer, 1×1 and 5×5 convolutional layers, using dilated convolution with a dilation rate of 1, and also includes 1×1 and 3×3 convolutional layers, using dilated convolution with a dilation rate of 1.
[0035] The Swin Transformer replaces the standard multi-head self-attention MSA module with a shifted window module. The shifted window module restricts the self-attention calculation to non-overlapping local windows through a shifted window strategy to improve the calculation efficiency. The Swin Transformer module of the Swin Transformer includes a shifted window MSA and a 2-layer MLP with a Gaussian error linear unit GELU non-linearity. There is a LayerNorm layer LN before each MSA and MLP module to achieve the residual connection between the two Swin Transformer modules.
[0036] The Efficientnet-b3 module based on the convolutional neural network is constructed based on the neural architecture search method and uses multiple MBConv blocks stacked. The EfficientNet-b3 module finds the most suitable model parameters by scaling three dimensions of the neural network: depth, width, and resolution.
[0037] Example:
[0038] In this example, image preprocessing operations are first carried out. Since the number of the medical image dataset is limited, the dataset is augmented to reduce the risk of overfitting.
[0039] For image preprocessing, in this example, several transformations are first selected, including vertical flipping, horizontal flipping, and rotation. The rotation directions include 45 / 90 / 135 / 180 / 225 / 270 / 315. In addition, noise interference is also adopted in this example, and Gaussian noise is selected. The visualization of the transformation is as Figure 6 shown.
[0040] Before performing classification prediction, in this example, the size of the image is adjusted to 224×224.
[0041] In the training and validation phases of this example, the Adam optimizer is used, the batch size is set to 16, the initial learning rate is set to 0.1, and the weight decay is set to 0.001. To avoid overfitting during training, the training process is terminated after 200 epochs.
[0042] In this example, classification prediction is performed on an Ubuntu system with a gpu above the Nvidia RTX 2080TI version to accelerate the prediction speed.
[0043] In the test phase, the images with the size adjusted to 224×224 are respectively input into two encoders, the Swin Transformer and the convolutional neural network, for feature extraction. Then, the outputs of the two encoders are concatenated and fused. Next, a cascaded dilated convolution (CDI) operation is performed to extract the high-level semantic information existing in the image. Finally, a global average pooling (GAP) operation is performed, and after passing through the fully connected (FC) layer, the classification prediction result of the benign and malignant thyroid is given, automatically assisting doctors and greatly reducing the doctors' reading time and diagnostic difficulty.
[0044] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, fall within the protection scope of the present invention.
Claims
1. An intelligent classification method for thyroid nodules with double - coding structure and cascaded dilated convolution, which is used for classifying medical images, Characterized in that: The method is based on a classification network model. The classification network model fuses the encoders of both the Transformer network and the convolutional neural network to form an enhanced double - encoder structure. It uses Swin Transformer of the Transformer network as the backbone of the first encoder to capture the long - range dependencies between the pixels of medical images; and uses the Efficientnet - b3 module based on the convolutional neural network as the backbone of the second encoder, and finds the most suitable model parameters by scaling the three dimensions of the depth, width, and resolution of the network, so as to simplify the composite coefficient and improve the efficiency of the classification network model; The method includes the following steps; Step S1: Perform pre - processing operations on medical images; If the number of the medical image dataset is limited, the dataset is augmented to reduce the risk of overfitting. The pre - processing operation methods of the images include image transformation and noise interference; the image transformation includes vertical flipping, horizontal flipping, or rotation of the image, and the rotation direction angles include 45 / 90 / 135 / 180 / 225 / 270 / 315; Step S2: Resize the medical images to a preset size for training the neural network. When training and validating the neural network, use the Adam optimizer, and terminate the training process after a preset number of epochs for the neural network; Step S3: Input the medical images resized to the preset size into the two encoders, namely Swin Transformer and the convolutional neural network, respectively, for feature extraction, and then concatenate the outputs of the two encoders, and splice and fuse the outputs of the two; then perform the cascaded dilated convolution CDI operation to extract the high - level semantic information existing in the image, and finally perform the global average pooling GAP operation. After passing through the fully - connected FC layer, give the classification prediction result of the thyroid benign and malignant; The classification network model consists of a double - encoder structure, a cascaded dilated convolution module, a global average pooling GAP, and a fully - connected FC layer; in the double - encoder, the Transformer encoder of the Transformer network learns long - range dependencies, and the encoder based on the convolutional neural network captures high - level features; The outputs of the Transformer encoder and the encoder based on the convolutional neural network are both connected to the cascaded dilated convolution CDI module to further improve the classification performance; The Efficientnet - b3 module based on the convolutional neural network is constructed based on the neural architecture search method and adopts multiple stacked MBConv blocks; the EfficientNet - b3 module finds the most suitable model parameters by scaling the three dimensions of the depth, width, and resolution of the neural network.
2. The intelligent classification method for thyroid nodules with double - coding structure and cascaded dilated convolution according to claim 1, Characterized in that: The cascaded dilated convolution CDI module is used to extract high-level semantic information. It consists of four cascaded branches, including 3×3 max pooling, 1×1 convolutional layer, 1×1 and 5×5 convolutional layers, using dilated convolution with a dilation rate of 1, and also including 1×1 and 3×3 convolutional layers, using dilated convolution with a dilation rate of 1.
3. The intelligent classification method for thyroid nodules with a dual coding structure and cascaded dilated convolution according to claim 1, characterized in that: The Swin Transformer uses a shifted window module to replace the standard multi-head self-attention MSA module. The shifted window module restricts the self-attention calculation to non-overlapping local windows through a shifted window strategy to improve computational efficiency; the Swin Transformer module of the Swin Transformer includes a shifted window MSA and a 2-layer MLP with a Gaussian error linear unit GELU non-linearity; there is a LayerNorm layer LN before each MSA and MLP module to achieve residual connections between two Swin Transformer modules.
Citation Information
Patent Citations
Image registration method based on Swin Transform and CNN double-branch coupling
CN115082293A