Transformer-Based Thyroid Nodule Detection Method

Through the Transformer-based nodule detection method, the problem of Anchor Box in the prior art has been solved, and efficient and accurate automatic detection of thyroid nodules is achieved.

CN114494215BActive Publication Date: 2025-07-22脉得智能科技(无锡)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210110296.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-29
Publication Date
2025-07-22
Estimated Expiration
2042-01-29

AI Technical Summary

Technical Problem

In the detection of thyroid nodules, the Anchor Box hyperparameters have problems such as large influence on Anchor Box, high computing resource consumption, and unbalanced positive and negative samples, resulting in low diagnostic efficiency and insufficient accuracy.

Method used

The nodule detection method based on Transformer is adopted, and the nodule detection model is constructed through image preprocessing and sample data set training. The Transformer network is used for feature extraction, encoding and decoding, and the automatic positioning and classification of nodules is realized, avoiding the use of Anchor Box and complex post-processing operations.

Benefits of technology

It realizes high automation and low computing resource requirements thyroid nodule detection, improves diagnostic accuracy and efficiency, and reduces computational volume and memory consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494215B_ABST
    Figure CN114494215B_ABST
Patent Text Reader

Abstract

The present invention discloses a thyroid nodule detection method based on Transformer, which relates to the technical field of image processing. After obtaining the ultrasonic image to be measured in the thyroid region and performing image preprocessing on the obtained ultrasonic image to be measured, it is input into a nodule detection model pre-trained based on the Transformer network. According to the output of the nodule detection model, the position and type of the nodule in the ultrasonic image to be measured are determined, and the detection of the nodule in the ultrasonic image to be measured is completed. The type of the nodule is used to indicate whether the nodule is a benign nodule or a malignant nodule. This method can automatically complete nodule positioning and classification, has a high degree of automation and good objectivity, does not need to construct a dense Anchor Box, does not need to use complex post-processing operations such as NMS, is easy to implement, and has low requirements for computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a thyroid nodule detection method based on Transformer. Background Art

[0002] Thyroid tumors are common and multiple tumors in the head and neck. In recent years, the incidence rate of thyroid cancer has been increasing year by year, which has received extensive attention from clinical staff and researchers. The malignancy of some thyroid nodules is relatively high. Therefore, the early diagnosis and treatment of thyroid nodules can effectively prevent thyroid cancer. In clinical practice, ultrasound is generally used to examine the thyroid gland. Ultrasound examination is a widely used imaging examination method in modern clinical practice, which can obtain information such as the boundary, shape and echo of the patient's thyroid nodules, and provide support for the further treatment of thyroid nodule patients. However, at present, medical resources in China are in short supply, the number of experienced ultrasound doctors is small, and doctors have heavy diagnosis and treatment tasks, which are prone to missed diagnosis and misdiagnosis. Therefore, how to assist doctors in the real-time diagnosis of thyroid nodules, identify malignant nodules from a large number of thyroid nodules, and improve the accuracy of doctors' diagnosis of the benign and malignant of thyroid nodules is of great significance and challenging for clinical practice.

[0003] At present, there are already many technologies that use deep learning methods for auxiliary diagnosis on medical images. For the auxiliary diagnosis of thyroid nodules in thyroid ultrasound images, deep learning-based algorithms are roughly divided into two categories: one is the Two Stage object detection algorithm based on Region Proposal (candidate box), such as RCNN, Faster RCNN, etc. The other is the One Stage object detection algorithm based on regression problems, such as YOLO, SSD, etc. The Two Stage algorithm needs to generate candidate boxes through the RPN network in advance, then classify the target through a convolutional neural network, and finally correct the position of the bounding box. The One Stage algorithm does not require the Region Proposal stage, directly generates the class probability and position coordinate values of the object, and can directly obtain the final detection result after a single detection. There are differences in their performances. The Two Stage algorithm has high accuracy but slow speed. The One Stage algorithm has an advantage in speed but slightly lower accuracy. With the development of research, the two types of algorithms are continuously improving their accuracy and speed.

[0004] In the existing technical solutions, both methods have the following disadvantages: (1) Different aspect ratios and sizes of Anchor Boxes need to be set before the experiment. However, due to the fact that the sizes of the input thyroid ultrasound images are not fixed and the sizes of nodule regions in the images vary in the actual situation, hyperparameters such as the size, aspect ratio, and number of Anchor Boxes will have a great impact on the experimental results, facing a huge test in the actual process of auxiliary diagnosis. (2) In order to achieve a high recall rate for the experimental results, a large number of dense Anchor Boxes are often set in one image. However, the number of nodules in each thyroid ultrasound image is extremely small. A large number of Anchor Boxes will bring a huge problem of imbalance between positive and negative sample categories during the class division in the training stage, and will increase the computational amount and consume memory resources when calculating IoU in the training and testing stages. Summary of the Invention

[0005] In view of the above problems and technical requirements, the inventor of the present invention proposes a thyroid nodule detection method based on Transformer. The technical solution of the present invention is as follows:

[0006] A thyroid nodule detection method based on Transformer, the method comprising:

[0007] Obtaining a to-be-detected ultrasound image of a thyroid region and performing image preprocessing on the obtained to-be-detected ultrasound image;

[0008] Inputting the to-be-detected ultrasound image after completing image preprocessing into a nodule detection model, where the nodule detection model is pre-trained based on a Transformer network;

[0009] Determining the position and type of nodules in the to-be-detected ultrasound image according to the output of the nodule detection model, and completing the detection of nodules in the to-be-detected ultrasound image, where the type of nodules is used to indicate whether the nodules are benign nodules or malignant nodules.

[0010] A further technical solution thereof is that performing image preprocessing on the obtained to-be-detected ultrasound image includes:

[0011] Performing image cropping on the to-be-detected ultrasound image by a threshold method, retaining the image of the ultrasound window region in the to-be-detected ultrasound image, and cropping off the image of the background region other than the ultrasound window region;

[0012] Performing histogram equalization on the image of the ultrasound window region to obtain the to-be-detected ultrasound image after completing image preprocessing.

[0013] A further technical solution thereof is that the method further includes:

[0014] Construct a sample data set, which includes a number of sample ultrasound images with preprocessed thyroid region images. Each sample ultrasound image includes a nodule annotation box, which is used to annotate the position and type of nodules in the sample ultrasound image. The sample ultrasound images in the sample data set include nodules of various different positions and / or types;

[0015] Use the sample data set to train a nodule detection model based on the Transformer network.

[0016] A further technical solution thereof is that using the sample data set to train a nodule detection model based on the Transformer network includes:

[0017] Perform pre-training based on the ImageNet data set using the Transformer network;

[0018] Transfer the pre-trained network parameters to the Transformer network and use the sample data set to perform network training to obtain a nodule detection model.

[0019] A further technical solution thereof is that using the sample data set to train a nodule detection model based on the Transformer network includes:

[0020] Divide the sample data set into a training set, a validation set and a test set. Use the sample ultrasound images in the training set to perform network training based on the Transformer network, use the sample ultrasound images in the validation set to optimize the training hyperparameters to obtain a nodule detection model, and use the sample ultrasound images in the test set to test the trained nodule detection model;

[0021] Among them, the difference in the number of benign nodules and malignant nodules included in the sample ultrasound images in the training set is within the first error range, and multiple sample ultrasound images belonging to the same patient are not simultaneously included in the training set and the validation set.

[0022] A further technical solution thereof is that the sample ultrasound images in the sample data set include nodules of at least two different size ranges, and the difference in the number of nodules of various size ranges included in the sample ultrasound images in the training set is within the second error range,

[0023] A further technical solution thereof is that the nodule detection model sequentially includes a feature extraction module, an encoding module, a decoding module and an FFN prediction module from input to output;

[0024] The feature extraction module extracts features from the input ultrasonic image to be measured that has completed image preprocessing and outputs a feature map; the encoding module is used to perform encoding processing on the feature map to obtain an encoding result; the decoding module is used to perform decoding processing on the encoding result to obtain a decoding result; the FFN prediction module includes a classification branch and a regression branch. The classification branch is used to classify the decoding result to determine the type of the nodule, and the regression branch is used to perform regression of the detection box on the decoding result to determine the position of the nodule.

[0025] Its further technical solution is that the encoding module includes an input unit and six encoding units connected in sequence from input to output. The input unit converts the feature map into serialized data and performs position encoding on the position information of the feature map. The serialized data and position encoding output by the input unit are added as input data and sequentially pass through six encoding units to obtain an encoding result.

[0026] The decoding result includes six decoding units connected in sequence from input to output. The input of the first decoding unit obtains N instance embedding sequences, and each instance embedding sequence corresponds to an object instance in the ultrasonic image to be measured; the encoding result output by the encoding module is respectively input into the multi-head cross-attention mechanism layer of the six decoding units. Each decoding unit aggregates the features of a predetermined object instance from the encoding result, and the last decoding unit outputs N feature vectors as the decoding result.

[0027] Its further technical solution is that the feature extraction module is constructed based on ResNet50 and performs feature extraction through five stages with 16-fold downsampling.

[0028] Its further technical solution is that the classification branch of the FFN prediction module includes a Linear layer with a hidden layer dimension of 512; the regression branch includes three Linear layers with hidden layer dimensions of 512 each.

[0029] The beneficial technical effects of the present invention are:

[0030] This application discloses a thyroid nodule detection method based on Transformer. This method locates and classifies nodules in the thyroid based on Transformer, has a high degree of automation, good objectivity, does not require constructing a dense Anchor Box, does not require using complex post-processing operations such as NMS, is easy to implement, and has low requirements for computing resources. Brief Description of the Drawings

[0031] Figure 1 is a flowchart of the thyroid nodule detection method in an embodiment.

[0032] Figure 2 is a schematic diagram of image preprocessing of the original ultrasonic image to be measured in an example.

[0033] Figure 3 It is a flowchart of a method for training a nodule detection model in an embodiment.

[0034] Figure 4 It is a model structure diagram of the trained nodule detection model. Specific Embodiments

[0035] The following further describes the specific embodiments of the present invention with reference to the accompanying drawings.

[0036] This application discloses a thyroid nodule detection method based on Transformer. Please refer to Figure 1 the flowchart shown below. The method includes the following steps:

[0037] Step 102: Obtain the ultrasonic image to be measured in the thyroid region and perform image preprocessing on the obtained ultrasonic image to be measured.

[0038] Since the original ultrasonic image to be measured collected contains some irrelevant information such as software interfaces, directly using it may have an adverse impact on model training. Moreover, considering the case where the texture features of isoechoic nodules are very similar to those of surrounding tissues and the boundaries are not obvious, after obtaining the original ultrasonic image to be measured, this step also performs image preprocessing, including:

[0039] (1) Crop the ultrasonic image to be measured by the threshold method, retain the image of the ultrasonic window area in the ultrasonic image to be measured, and crop the image of the background area outside the ultrasonic window area. The image of the ultrasonic window area is the image of the thyroid region, and the image of the background area outside the ultrasonic window area is the common software interface, etc. Specifically, scan the image row by row or column by column, calculate the pixel mean of each row or column, and filter out the background area outside the ultrasonic window area by setting a certain threshold, only retaining the ultrasonic window area. Generally, setting the threshold to 30 has the best effect according to experience.

[0040] (2) Perform histogram equalization on the image of the ultrasonic window area to enhance the image contrast and strengthen the boundary information, so as to improve the subsequent detection accuracy, thereby obtaining the ultrasonic image to be measured with completed image preprocessing. For the schematic diagram of image preprocessing of the ultrasonic image to be measured, please refer to Figure 2 .

[0041] Step 104: Input the ultrasonic image to be measured with completed image preprocessing into the nodule detection model, and the nodule detection model is pre-trained based on the Transformer network.

[0042] Step 106: Determine the position and type of nodules in the ultrasound image to be measured according to the output of the nodule detection model, and complete the detection of nodules in the ultrasound image to be measured. The type of nodules is used to indicate whether the nodules are benign nodules or malignant nodules. Specifically, a detection box will be displayed in the ultrasound image to be measured, and the area within the detection box corresponds to the detected nodules, thereby indicating the position of the nodules. At the same time, it will be indicated whether the nodules within the detection box are benign nodules or malignant nodules.

[0043] In the above step 104, before using the nodule detection model, it also includes the step of training the nodule detection model. Please refer to Figure 3 the flowchart shown below, which includes the following steps:

[0044] Step 302: Construct a sample data set.

[0045] Among them, the sample data set includes several pre-processed sample ultrasound images of the thyroid region. Each sample ultrasound image includes a nodule annotation box, which is used to annotate the position and type of nodules in the sample ultrasound image.

[0046] First, obtain multiple original ultrasound images of the thyroid region of multiple patients and perform image pre-processing. This step is similar to the above step 102 and will not be elaborated in this application. Then use the labelme annotation tool to select the nodule positions in the original ultrasound images and annotate the benign and malignant nature to obtain the nodule annotation boxes, and each nodule annotation box is confirmed by multiple doctors to ensure the accuracy of the content annotated in the nodule annotation boxes. For example, first, an experienced doctor annotates the nodule annotation box according to the diagnosis report, and then another doctor reviews and modifies it.

[0047] The sample ultrasound images in the constructed sample data set include nodules of various different positions and / or types. Generally, there is only one nodule in one sample ultrasound image, and each patient can have multiple sample ultrasound images. For example, in actual acquisition, about 10 sample ultrasound images are collected for each patient, and a total of 3000 sample ultrasound images are collected for 300 patients. Among them, there are about 1800 benign nodules and about 1200 malignant nodules.

[0048] In addition, the sample ultrasound images in the sample data set include nodules of at least two different size ranges. For example, all nodules are divided into two size ranges. Nodules with a size greater than 5 mm are defined as large nodules, and nodules with a size less than 5 mm are defined as small nodules. In the above example, there are about 1100 large nodules and about 1900 small nodules.

[0049] As Figure 3As shown, after constructing the sample data set, the sample data set is randomly divided into a training set, a validation set, and a test set, for example, randomly divided according to a ratio of 6:2:2. When dividing, ensure that multiple sample ultrasound images belonging to the same patient are not included in both the training set and the validation set at the same time, that is, they do not appear in both the training set and the validation set at the same time.

[0050] At the same time, in the later stage, the training set is mainly used for network training. Therefore, when dividing, the difference in the number of benign nodules and malignant nodules included in the sample ultrasound images in the training set is within the first error range, that is, the proportions of benign nodules and malignant nodules in the training set are close. The difference in the number of nodules in various size ranges included in the sample ultrasound images in the training set is within the second error range, that is, the proportions of nodules in various size ranges in the training set are close. For example, in the above example, the proportions of large nodules and small nodules are close.

[0051] Step 304, use the sample data set to train a nodule detection model based on the Transformer network. When performing network training, use the ImageNet data set to pre-train based on the Transformer network, and then transfer the pre-trained network parameters to the Transformer network and use the sample data set to perform network training to obtain the nodule detection model, which makes the network converge faster and has stronger generalization ability. When training the model, use a 24G, RTX3090 graphics card for training.

[0052] Specifically, when using the sample data set for network training, use the sample ultrasound images in the training set to perform network training based on the Transformer network. Use the sample ultrasound images in the validation set to optimize and fine-tune the hyperparameters of the training until convergence to obtain the nodule detection model. The hyperparameters are set as follows: use the AdamW optimizer, set the learning rate to 0.001, use CrossEntropy Loss for classification, and use Smooth L1Loss for the regression branch. Continuously iterate the training. After 500 epochs of iteration, the model reaches convergence. Use the sample ultrasound images in the test set to test the trained nodule detection model to ensure the accuracy and precision of the trained nodule detection model. Specifically, calculate the IoU between the detection box in the detection result of the nodule detection model and the nodule annotation box, and compare the type in the detection result with the type indicated by the nodule annotation box. When IoU > 0.5 and the type in the detection result is the same as the type indicated by the nodule annotation box, it is considered that the localization and recognition are accurate. If the test determines that the accuracy and precision meet the standards, the nodule detection model can be used to perform step 104 to localize and recognize the nodules in the ultrasound image to be measured. Otherwise, re-training is required.

[0053] Please refer to Figure 4, the trained nodule detection model sequentially includes a feature extraction module, an encoding module, a decoding module, and an FFN prediction module from input to output. In the model training stage, the image input to the feature extraction module is a sample ultrasound image. In the model usage stage, the image input to the feature extraction module is the ultrasound image to be measured. The processing of the image by each module in the model training stage and the model usage stage is similar. The following takes the processing of the ultrasound image to be measured that has completed image preprocessing in the model usage stage as an example for explanation:

[0054] The feature extraction module is used to extract features from the input ultrasound image to be measured that has completed image preprocessing and output a feature map. Specifically, the feature extraction module is built based on ResNet50 and undergoes four downsamplings through 5 stages, with a total of 16 times of downsampling for feature extraction. When the input ultrasound image to be measured that has completed image preprocessing is an image of 512x512x3, a feature map of 32x32x2048 is obtained.

[0055] The encoding module is used to perform encoding processing on the feature map to obtain an encoding result. The encoding module includes an input unit and 6 encoding units connected in sequence from input to output. Each encoding unit is a standard Transformer structure, and each encoding unit sequentially includes a Multi-Head Self-Attention Layer, a Normal Layer, and an FFN (feed-forward network). The input unit converts the feature map into serialized data, that is, stretches the features extracted by the feature extraction module in the spatial dimension and converts them into serialized data of 2048x1024. And perform position encoding on the position information of the feature map. Then, the serialized data output by the input unit and the position encoding are added together as input data and sequentially pass through 6 encoding units to obtain an encoding result.

[0056] The decoding module is used to decode the encoding result to obtain the decoding result. The decoding result includes 6 decoding units connected in sequence from input to output. Each decoding unit is a standard Tranformer structure. Each decoding unit includes Multi-Head Self-Attention Layer, Normal Layer, Multi-Head Cross-Attention Layer, FFN and NormalLayer in sequence. The input of the first decoding unit obtains N instance embedding sequences (Object Query), and each instance embedding sequence corresponds to an object instance in the ultrasound image to be tested. The encoding results output by the encoding module are respectively input to the multi-head cross-attention mechanism layers of the 6 decoding units. The multi-head cross-attention mechanism layer aggregates the features of the predetermined object instance from the encoding results output by the encoding module, and the multi-head self-attention mechanism layer models the relationship between the object instance and other object instances. The last decoding unit outputs N feature vectors (output Embedding) as the decoding result. During the model training phase, N sparse learnable object queries are used and updated as the network is trained, thus implicitly modeling the statistics of the entire training set.

[0057] The FFN prediction module includes a classification branch and a regression branch. The classification branch is used to classify the decoding results to determine the type of nodules, and the regression branch is used to regress the detection box of the decoding results to determine the location of the nodules. The classification branch of the FFN prediction module includes a Linear layer, the dimension of the hidden layer is 512, and the dimension of the output is the number of categories plus 1 (there is a background class). In this embodiment, the number of categories is 3. The regression branch includes three Linear layers, the dimension of the hidden layer is 512, and the dimension of the output layer is 4, indicating that the prediction box is the coordinate information of each vertex.

[0058] The above is only a preferred embodiment of the present application, and the present invention is not limited to the above embodiments. It is understood that other improvements and changes directly derived or associated by those skilled in the art without departing from the spirit and concept of the present invention should be considered to be included in the protection scope of the present invention.

Claims

1. A thyroid nodule detection method based on Transformer, characterized in that The method includes: Obtaining an ultrasonic image to be measured in the thyroid region and performing image preprocessing on the obtained ultrasonic image to be measured, including: performing image cropping on the ultrasonic image to be measured by a threshold method, retaining the image of the ultrasonic window region in the ultrasonic image to be measured, and cropping off the image of the background region other than the ultrasonic window region; performing histogram equalization on the image of the ultrasonic window region to obtain the ultrasonic image to be measured after completing image preprocessing; Inputting the ultrasonic image to be measured after completing image preprocessing into a nodule detection model, where the nodule detection model is pre-trained based on a Transformer network; Determining the position and type of nodules in the ultrasonic image to be measured according to the output of the nodule detection model, and completing the detection of nodules in the ultrasonic image to be measured. The type of nodules is used to indicate whether the nodules are benign nodules or malignant nodules; The nodule detection model sequentially includes a feature extraction module, an encoding module, a decoding module, and an FFN prediction module from input to output; the feature extraction module performs feature extraction on the input ultrasonic image to be measured after completing image preprocessing and outputs a feature map; the encoding module is used to perform encoding processing on the feature map to obtain an encoding result; the decoding module is used to perform decoding processing on the encoding result to obtain a decoding result; the FFN prediction module includes a classification branch and a regression branch. The classification branch is used to classify the decoding result to determine the type of nodules, and the regression branch is used to perform regression of the detection frame on the decoding result to determine the position of nodules; the encoding module includes an input unit and six encoding units connected in sequence from input to output. The input unit converts the feature map into serialized data and performs position encoding on the position information of the feature map. The serialized data and position encoding output by the input unit are added as input data and sequentially pass through six encoding units to obtain the encoding result; the decoding result includes six decoding units connected in sequence from input to output. The input of the first decoding unit obtains N instance embedding sequences, and each instance embedding sequence corresponds to an object instance in the ultrasonic image to be measured; the encoding result output by the encoding module is respectively input into the multi-head cross-attention mechanism layers of the six decoding units. Each decoding unit aggregates the features of a predetermined object instance from the encoding result, and the last decoding unit outputs N feature vectors as the decoding result.

2. The method according to claim 1, wherein The method further includes: Constructing a sample data set, where the sample data set includes a number of sample ultrasonic images in the thyroid region after completing image preprocessing. Each sample ultrasonic image includes a nodule annotation box, and the nodule annotation box is used to annotate the position and type of nodules in the sample ultrasonic image. The sample ultrasonic images in the sample data set include nodules of various different positions and / or types; Using the sample data set to perform network training based on a Transformer network to obtain the nodule detection model.

3. The method according to claim 2, wherein The using the sample data set to perform network training based on a Transformer network to obtain the nodule detection model includes: Pre-train based on the Transformer network using the ImageNet dataset; Transfer the pre-trained network parameters to the Transformer network and use the sample dataset for network training to obtain the nodule detection model.

4. The method according to claim 2, wherein The network training using the sample dataset based on the Transformer network to obtain the nodule detection model includes: Divide the sample dataset into a training set, a validation set, and a test set. Use the sample ultrasound images in the training set for network training based on the Transformer network, use the sample ultrasound images in the validation set to optimize the training hyperparameters to obtain the nodule detection model, and use the sample ultrasound images in the test set to test the trained nodule detection model; Among them, the difference in the number of benign nodules and malignant nodules included in the sample ultrasound images in the training set is within the first error range, and multiple sample ultrasound images belonging to the same patient are not simultaneously included in the training set and the validation set.

5. The method according to claim 4, wherein The sample ultrasound images in the sample dataset include nodules with at least two different size ranges, and the difference in the number of nodules with various size ranges included in the sample ultrasound images in the training set is within the second error range.

6. The method according to claim 1, wherein The feature extraction module is constructed based on ResNet50 and performs feature extraction through 5 stages with 16-fold downsampling.

7. The method according to claim 1, characterized in that The classification branch of the FFN prediction module includes a Linear layer with a hidden layer dimension of 512; the regression branch includes three Linear layers with hidden layer dimensions of 512 each.

Citation Information

Patent Citations

  • Swin Unet low-illumination image enhancement method

    CN113793275A

  • Pulmonary nodule benign and malignant auxiliary diagnosis system based on computed tomography

    CN113902702A