Automatic nodule benign and malignant scoring method and system based on thyroid ultrasound image
By adopting a deep learning model of encoder-decoder architecture and dynamic attention mechanism in ultrasonic image processing of thyroid nodules, combined with Swin Transformer and C-TIRADS standards, the problems of insufficient feature extraction and automated risk grading are solved, and the accurate detection and automatic grading of nodules are achieved, which improves diagnostic efficiency and clinical availability of the model.
Patent Information
- Application Number
- CN202510541657.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art lacks feature extraction in ultrasound images of thyroid nodules, is susceptible to interference from artifacts and background noise, and lacks an automated risk grading method combined with the C-TIRADS standard. The diagnostic results rely on doctor experience and are subjective. The existing deep learning models are complex and have poor real-time and applicability.
The deep learning model with an encoder-decoder architecture was adopted, combined with a dynamic attention mechanism for feature extraction, the thyroid nodule recognition results were graded using the Swin Transformer model, and the benign and malignant risk score was calculated in combination with the C-TIRADS standard.
It improves the precise positioning and feature extraction of nodules in ultrasound images, reduces dependence on artifacts and noise, realizes accurate detection and automatic grading of nodules, and improves diagnostic efficiency and clinical availability of the model.
Smart Images

Figure CN120070438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an automatic scoring method and system for the benign and malignant of nodules based on thyroid ultrasound images. Background Art
[0002] The present invention relates to the technical fields of medical image processing and artificial intelligence, and particularly to an automatic scoring method for the benign and malignant of thyroid nodules based on ultrasound images, aiming to provide an intelligent auxiliary diagnostic tool for the detection and classification of thyroid nodules through a deep learning model.
[0003] Thyroid nodules are a common clinical problem, and the accurate assessment of their benign and malignant is of great significance for the early diagnosis and treatment of thyroid cancer. Currently, ultrasound imaging is widely used in the detection and classification of thyroid nodules. However, due to the artifacts, noises and complex characteristics of nodules in ultrasound images, there is a certain degree of subjectivity and misdiagnosis risk in the diagnosis by doctors. Especially for doctors with insufficient experience or medical institutions with limited resources, it is a great challenge to accurately determine the benign and malignant of nodules.
[0004] To standardize the diagnostic criteria of thyroid nodules, the Chinese Thyroid Imaging Reporting and Data System (C-TIRADS) has proposed a risk grading system based on multiple features for quantifying the risk of nodules. However, in actual clinical practice, doctors need to conduct manual analysis by combining a large number of nodule features, which is not only time-consuming and laborious, but also may lead to inconsistent diagnostic results due to complexity. Therefore, how to realize automatic nodule detection and risk scoring based on ultrasound images by means of advanced artificial intelligence technology has become a research hotspot.
[0005] There is a literature (Yang D, Xia J, Li R, et al. Automatic thyroid nodule detectionin ultrasound imaging with improved YOLOv5 neural network[J]. IEEE Access,2024.) that proposed an improved YOLOv5 neural network to assist physicians in diagnosing thyroid cancer. It includes a Coordinate Attention (CA) module and a Label Smoothing Regularization (LSR) module. The literature (Jin Z, Zhu Y, Zhang S, et al.Ultrasound computer-aided diagnosis (CAD) based on the thyroid imagingreporting and data system (TI-RADS) to distinguish benign from malignantthyroid nodules and the diagnostic performance of radiologists with differentdiagnostic experience[J]. Medical Science Monitor: International MedicalJournal of Experimental and Clinical Research, 2020, 26: e918452-1.) proposed a computer-aided diagnosis (CAD) system based on the Thyroid Imaging Reporting and Data System (TI-RADS), and developed CAD using an improved TI-RADS based on a convolutional neural network (CNN). The literature (Yang Q, Geng C, Chen R, etal. DMU-Net: Dual-route mirroring U-Net with mutual learning for malignantthyroid nodule segmentation[J]. Biomedical Signal Processing and Control,2022, 77: 103805.) proposed a deep learning-based CAD (computer-aided diagnosis) method called Dual-route Mirroring U-Net (DMU-Net) to automatically segment malignant thyroid nodules.
[0006] With the development of technology, computer-aided diagnosis technology for thyroid nodules has gradually become mature. The current research mainly has the following problems in the assessment of the benign and malignant of thyroid nodules: insufficient extraction of key features of nodules in ultrasound images (such as boundaries, echo distribution, calcification features, etc.), being easily interfered by artifacts and background noise; lack of an automated risk grading method combined with the C-TIRADS standard, and the diagnostic results overly rely on doctors' experience, resulting in subjectivity; at the same time, the existing deep learning models have high complexity, poor real-time performance and applicability, and it is difficult to meet the needs of medical scenarios with limited resources. In view of the above problems, this paper proposes an automatic scoring method for the benign and malignant of thyroid nodules based on deep learning and the C-TIRADS standard, aiming to improve the accuracy of feature extraction, achieve accurate detection and automatic grading of nodules, and improve the diagnostic efficiency and clinical usability of the model. Summary of the Invention
[0007] The purpose of the present invention is to solve the problems in the prior art.
[0008] The technical solution adopted by the present invention to solve its technical problems is: to provide an automatic scoring method for the benign and malignant of nodules based on thyroid ultrasound images, including the following steps:
[0009] Obtain a thyroid ultrasound image dataset and perform annotation to construct a training dataset;
[0010] Construct a deep learning model including a nodule recognition network and a nodule classification network, and use the training dataset to train the model to obtain a trained deep learning model as an automatic scoring model;
[0011] Input the thyroid ultrasound image into the automatic scoring model to obtain the thyroid nodule recognition result and calculate the benign and malignant risk score;
[0012] The nodule recognition network adopts an encoder-decoder architecture. The encoder performs basic feature extraction on the input image and uses a dynamic attention mechanism for feature enhancement; the decoder gradually reduces the dimension of the enhanced features output by the encoder and finally outputs the thyroid nodule recognition result; the nodule classification network uses the Swin Transformer model to grade the thyroid nodule recognition result and calculate the benign and malignant risk score.
[0013] Preferably, the encoder performs basic feature extraction on the input image, expressed as:
[0014] ;
[0015] Where represents the basic feature, represents the input image, represents the feature extraction operation of ShuffleNetV2.
[0016] Preferably, the feature enhancement using the dynamic attention mechanism includes the following steps:
[0017] Enhance the attention to important feature channels in the basic features by using channel attention to obtain channel-enhanced features;
[0018] Enhance the key regions in the channel-enhanced features by using spatial attention to obtain enhanced features.
[0019] Preferably, the step of enhancing the attention to important feature channels in the basic features by using channel attention to obtain channel-enhanced features includes the following steps:
[0020] Extract the importance of each channel through global average pooling, expressed as:
[0021] ;
[0022] Wherein, represents the global description of channel-level features, and respectively represent the height and width of the basic feature , represents the basic feature at the coordinate ;
[0023] Calculate the channel attention weight through a fully connected layer, expressed as:
[0024] ;
[0025] Wherein, represents the first fully connected layer performing a linear transformation on the global pooling feature; represents the activation function; is the second fully connected layer further refining the attention information; σ represents the Sigmoid activation function; W represents the channel attention weight matrix;
[0026] Calculate the channel-enhanced features through channel weighting, expressed as:
[0027] ;
[0028] Wherein, represents the channel-enhanced features.
[0029] Preferably, the step of enhancing the key regions in the channel-enhanced features by using spatial attention to obtain enhanced features includes the following steps:
[0030] Use max pooling and average pooling to obtain the global information of the feature map, expressed as:
[0031] ;
[0032] Among them, is the maximum pooling result, represents maximum pooling, is the average pooling result, represents average pooling;
[0033] The spatial attention weights are calculated using convolution, expressed as:
[0034] ;
[0035] Among them, A represents the spatial attention weights, represents the convolution operation; σ represents the Sigmoid activation function;
[0036] The enhanced features are calculated through spatial weighting, expressed as:
[0037] ;
[0038] Among them, represents the enhanced features.
[0039] Preferably, the decoder gradually reduces the dimension of the enhanced features output by the encoder and finally outputs the thyroid nodule recognition result, including the following steps:
[0040] The feature dimension is gradually reduced through two layers of 3×3 convolution, expressed as:
[0041] ;
[0042] ;
[0043] Among them, is the feature map output by the first layer of 3×3 convolution, represents the feature map output by the second layer of 3×3 convolution; represents the convolution operation; represents the activation function;
[0044] The spatial features are converted into vectors using global average pooling, expressed as:
[0045] ;
[0046] Among them, represents the pixel point at coordinates in the feature map ;
[0047] Final classification is performed through a fully connected layer, expressed as:
[0048] ;
[0049] Among them, Y represents the output class probability, b represents the bias term, and W represents the channel attention weight matrix.
[0050] Preferably, the nodule classification network uses the Swin Transformer model to grade the thyroid nodule recognition result and calculate the benign and malignant risk score, including the following steps:
[0051] Crop the nodule region from the thyroid nodule recognition result output by the nodule recognition network and input it into the Swin Transformer model;
[0052] The Swin Transformer model uses the local window attention mechanism to extract features from the nodule region, and captures image information at different scales through the sliding window strategy to obtain key features including morphology, edge, structure, echo, and hyperechoic foci;
[0053] Use the C-TIRADS scoring system to score and sum the key features to obtain the final benign and malignant risk score result.
[0054] The present invention also provides an automatic scoring system for the benign and malignant of thyroid nodules based on thyroid ultrasound images, including:
[0055] The dataset acquisition module acquires the thyroid ultrasound image dataset and performs annotation to construct a training dataset;
[0056] The model training module constructs a deep learning model including a nodule recognition network and a nodule classification network, and uses the training dataset to train the model to obtain a trained deep learning model as an automatic scoring model;
[0057] The risk scoring module inputs the thyroid ultrasound image into the automatic scoring model to obtain the thyroid nodule recognition result and calculate the benign and malignant risk score;
[0058] The nodule recognition network adopts an encoder-decoder architecture. The encoder extracts basic features from the input image and uses a dynamic attention mechanism for feature enhancement; the decoder gradually reduces the dimension of the enhanced features output by the encoder and finally outputs the thyroid nodule recognition result; the nodule classification network uses the Swin Transformer model to grade the thyroid nodule recognition result and calculate the benign and malignant risk score.
[0059] The present invention has the following beneficial effects:
[0060] (1) Through the encoder and decoder structures in the deep learning model, the present invention introduces a dynamic attention mechanism combined with dynamic convolution to achieve precise localization and feature extraction of nodule regions in ultrasonic images. Specifically, dynamic channel attention is adopted to calculate the importance weights of each channel, enabling the model to adaptively adjust the contribution degree of channel information. Dynamic spatial attention is used to adaptively enhance the model's attention to the lesion region and improve the classification performance. The combination of dynamic attention and dynamic convolution can retain the low-level feature information of the original ultrasonic image while enhancing the important features related to the nodule region and improving the recognition accuracy.
[0061] (2) The present invention designs a nodule classification network based on the C-TIRADS standard to further refine the nodule features in multiple directions for the recognition results of the recognition network, thereby performing quantitative grading and generating a risk level assessment result.
[0062] The following further elaborates on the present invention in detail with reference to the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings
[0063] Figure 1 It is a method step diagram of the automatic scoring method for nodule benignity and malignancy based on thyroid ultrasonic images according to an embodiment of the present invention;
[0064] Figure 2 It is a structural schematic diagram of the automatic scoring model of the automatic scoring method for nodule benignity and malignancy based on thyroid ultrasonic images according to an embodiment of the present invention;
[0065] Figure 3 It is a structural schematic diagram of the dynamic attention unit of the automatic scoring method for nodule benignity and malignancy based on thyroid ultrasonic images according to an embodiment of the present invention, combining dynamic channel attention and dynamic spatial attention;
[0066] Figure 4 It is a schematic diagram of the input image and output result of the automatic scoring method for nodule benignity and malignancy based on thyroid ultrasonic images according to an embodiment of the present invention;
[0067] Figure 5 It is a structural diagram of the automatic scoring system for nodule benignity and malignancy based on thyroid ultrasonic images according to an embodiment of the present invention. Specific Embodiments
[0068] Refer to Figure 1 As shown, the automatic scoring method for nodule benignity and malignancy based on thyroid ultrasonic images according to an embodiment of the present invention includes the following steps:
[0069] S101, obtain a thyroid ultrasonic image dataset and perform annotation to construct a training dataset;
[0070] S102. Construct a deep learning model including a nodule recognition network and a nodule classification network, and use the training data set to train the model to obtain a trained deep learning model as an automatic scoring model;
[0071] S103. Input the thyroid ultrasound image into the automatic scoring model to obtain the thyroid nodule recognition result and calculate the benign and malignant risk score;
[0072] Specifically, the network structure of the automatic scoring model is shown in Figure 2 As shown, it includes a nodule recognition network and a nodule classification network. The nodule recognition network uses ShuffleNetV2 as the backbone network, and introduces a dynamic attention mechanism in its encoder to extract high-level features of the input image. These features contain spatial information and channel information, representing the texture, morphology, and boundary characteristics of thyroid nodules. ShuffleNetV2 is an efficient and lightweight convolutional neural network that uses grouped convolution and channel shuffle techniques to improve computational efficiency while maintaining strong feature extraction capabilities. Decoder part: The main task of the decoder is to convert the high-level features extracted by the encoder into the final classification result. Since the features output by the encoder usually have a high dimension and less spatial information, the decoder needs to gradually reduce the dimension and finally output the recognition result of the thyroid nodule. The decoder first performs feature fusion through two 3×3 convolutional layers. The first convolutional layer reduces the input high-dimensional features from 512 dimensions to 256 dimensions and introduces non-linearity through the ReLU activation function. The second convolutional layer further reduces the dimension to 128 dimensions and uses ReLU for activation again. Through these two layers of convolution, the decoder reduces the computational complexity while maintaining the feature expression ability.
[0073] The thyroid nodule ultrasound image is initially recognized by the nodule recognition network. The recognized nodules are then cropped and sent into the nodule classification network. The Swin Transformer model is used to grade the thyroid nodule recognition result and calculate the benign and malignant risk score. The Swin Transformer (Shifted Window Transformer) model is an improved vision Transformer (ViT) model, which is specifically used for computer vision tasks such as image classification, object detection, and medical image analysis. It improves computational efficiency through a hierarchical structure and shifted window attention, making the Transformer applicable to high-resolution image processing.
[0074] Specifically, see Figure 3As shown in the figure, it is a schematic structural diagram of the dynamic attention mechanism of the embodiment of the present invention, including dynamic channel attention and dynamic spatial attention.
[0075] Different channels of thyroid ultrasound images may contain different feature information, such as edges, morphology, echo distribution, etc. During the traditional CNN processing, the features of each channel are processed with equal weights, and the key feature channels cannot be highlighted. After extracting features using ShuffleNetV2 in the embodiment of the present invention, dynamic channel attention is also adopted to calculate the importance weights of each channel, enabling the model to adaptively adjust the contribution degree of channel information. The calculation process is as follows:
[0076] Global information aggregation, global average pooling (GAP) obtains the global information of each channel:
[0077] ;
[0078] Among them , represents the global feature of each channel.
[0079] Calculate the channel attention weight: adopt a two-layer fully connected (FC) network to calculate the dynamic attention weight:
[0080] ;
[0081] Among them: , are trainable parameters, σ is the Sigmoid activation function, ensuring that the attention weight is between 0 and 1.
[0082] Weighted channel features: Apply the calculated attention weight w to the original features:
[0083] ;
[0084] Here represents element-wise multiplication, that is, enhancing important channels and suppressing irrelevant channels.
[0085] In the ultrasound images of thyroid nodules, the importance of information contained in different spatial positions is different. For example, the lesion area is more important than the background area. During the traditional CNN processing, all spatial regions are treated equally, which may lead to the model paying too much attention to irrelevant regions and affecting the classification effect. In the embodiment of the present invention, dynamic spatial attention is adopted to adaptively enhance the model's attention to the lesion area and improve the classification performance. The calculation process is as follows:
[0086] Assume the input feature map with dimensions C×H×W, and the channel attention calculation process is as follows:
[0087] Spatial information extraction: Use max pooling (Max Pooling) and average pooling (Average Pooling) to extract the global information of the features:
[0088] ;
[0089] where represent the max pooling feature and the average pooling feature respectively.
[0090] Calculate the spatial attention weights: Use a 3×3 convolution to calculate the spatial attention distribution:
[0091] ;
[0092] where: .
[0093] Weighted spatial features: Calculate the final enhanced features:
[0094] ;
[0095] Here represents element-wise multiplication, that is, enhancing the features of the key spatial regions and suppressing the background information.
[0096] The overall calculation process of the dynamic attention is as follows: Through channel attention (Channel Attention), calculate the channel importance and perform weighted enhancement to obtain . Enter spatial attention (Spatial Attention), calculate the spatial attention distribution and perform weighted enhancement to obtain the final feature :
[0097] ;
[0098] This step can retain the low-level feature information of the original ultrasound image, while enhancing the important features related to the nodule region and improving the recognition accuracy.
[0099] Specifically, the nodule classification network deeply analyzes the cropped nodule region image, extracts and judges five major types of key features of the nodule, namely, morphology, margin, structure, echogenicity, and hyperechoic foci. Among them, the morphological features (such as vertical or horizontal) can be directly calculated based on the geometric shape of the detection box in the previous stage. Therefore, no additional data collection or calculation is required at this stage. Swin Transformer uses a local window attention mechanism for feature extraction and captures image information at different scales through a sliding window strategy.
[0100] The model uses the C-TIRADS scoring system to classify and train nodules based on five major types of features and predicts the corresponding C-TIRADS levels. The scoring criteria for the features are shown in Table 1.
[0101] Table 1 - Feature Scoring Criteria:
[0102]
[0103] The morphological features are calculated from the detection box in the previous stage and are divided into vertical (1 point) and horizontal (0 points); the margin features are divided into four categories: smooth, irregular, blurred, and extracapsular extension. Among them, except for smooth, the other three categories are all scored 1 point. After discussion with professional clinicians, we simplified the margin features into two categories: smooth and not smooth. The smooth margin feature gets 0 points, and the not smooth margin feature gets 1 point; for the structure feature, solid or mainly solid gets 1 point, cystic or mainly cystic gets 0 points, and the newly added mixed cystic-solid category gets 0 points; for the echogenicity feature, very low echogenicity gets 1 point, high echogenicity or isoechogenicity gets 0 points; for the hyperechoic foci feature, comet tail is a benign sign and gets -1 point, and punctate hyperecho is a malignant sign and gets 1 point; sum up the scores of the five major types of features for each nodule to calculate the final risk score. The C-TIRADS levels and malignancy rates corresponding to the risk scores are shown in Table 2:
[0104] Table 2 - Score Corresponding Risks:
[0105]
[0106] Specifically, a verification experiment was conducted on the embodiments of the present invention. The experimental platform and environment included: hardware configuration, with the operating system being Windows 10, the CPU being Intel(R) Core(TM) i7-11700, the memory being 16 GB, the GPU being NVIDIA GeForce GTX 3060, and equipped with 12 GB of video memory; software environment, with the programming language being Python 3.7 and the deep learning framework being PyTorch 1.10.1. The automatic scoring model finally outputs an evaluation report, and the report content includes the detection location of the nodule, morphological description, risk level, and recommended clinical treatment plan. The experimental results are shown in Figure 4 As shown, in the output image, the nodule was identified and relevant features were labeled: in the figure, the shape of the nodule was identified as wider-than-tall, the composition was solid, the echo characteristic was anechoic, there were no echogenic foci, the edge was extra-thyroidal extension, and the total score was 0. According to the C-TIRADS guidelines, this nodule was classified as benign (Score: 0 - Benign).
[0107] According to multiple verification experiments, the precision rate of the nodule recognition network of the embodiments of the present invention on the validation set was 0.887, the recall rate was 0.874, and the mean average precision (mAP@.5) reached 0.915. When the confidence threshold was set to 0.289, the F1 score of the model reached the highest, which was 0.89; the nodule classification network adopted the Swin Transformer model, and the accuracy rates on the categories of structure, echo, edge, and echogenic foci reached 0.985, 0.845, 0.826, and 0.867 respectively, and the average accuracy rate was 0.883.
[0108] As shown in Figure 5 is the system structure diagram of the embodiments of the present invention, including:
[0109] The dataset acquisition module 501 acquires the thyroid ultrasound image dataset and performs annotation to construct the training dataset;
[0110] The model training module 502 constructs a deep learning model including a nodule recognition network and a nodule classification network, and uses the training dataset to train the model to obtain a trained deep learning model as the automatic scoring model;
[0111] The risk scoring module 503 inputs the thyroid ultrasound image into the automatic scoring model to obtain the thyroid nodule recognition result and calculate the benign and malignant risk score.
[0112] The working processes of the modules of this system are the same as those of an automatic scoring method for the benign and malignant of thyroid nodules based on ultrasound images, and will not be described repeatedly here.
[0113] It can be seen that the present invention has high accuracy and real-time performance in the detection and classification of thyroid nodules, can assist doctors in improving the diagnosis efficiency, reducing the risk of misdiagnosis and missed diagnosis, and providing effective support for the early detection and treatment of thyroid diseases.
[0114] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for automatically scoring benign and malignant nodules based on thyroid ultrasound images, characterized in that: The following steps are involved: Obtain and annotate the thyroid ultrasound image dataset to construct a training dataset; Construct a deep learning model including a nodule recognition network and a nodule classification network, use the training data set to train the model, and obtain the trained deep learning model as an automatic scoring model; The thyroid ultrasound images were input into the automatic scoring model to obtain the thyroid nodule recognition results and calculate the benign and malignant risk scores; The nodule recognition network adopts an encoder-decoder architecture, the encoder extracts basic features from the input image, and uses a dynamic attention mechanism to enhance features; The decoder gradually reduces the dimension of the enhanced features output by the encoder and finally outputs the thyroid nodule recognition result; the nodule classification network uses the Swin Transformer model to classify the thyroid nodule recognition results and calculate the benign and malignant risk score.
2. The automatic scoring method for benign and malignant nodules based on thyroid ultrasound images according to claim 1 is characterized in that: The encoder extracts basic features from the input image, which is expressed as: ; in, Represents the basic features, represents the input image, Represents the feature extraction operation of ShuffleNetV2.
3. The automatic scoring method for benign and malignant nodules based on thyroid ultrasound images according to claim 1, characterized in that: The feature enhancement using the dynamic attention mechanism includes the following steps: Channel attention is used to enhance the attention on important feature channels in basic features to obtain channel enhanced features; Spatial attention is used to enhance the key areas in the channel features to obtain enhanced features.
4. The method for automatically scoring benign and malignant nodules based on thyroid ultrasound images according to claim 3, characterized in that: The method of using channel attention to enhance the attention to important feature channels in basic features to obtain channel enhanced features includes the following steps: The importance of each channel is extracted by global average pooling, expressed as: ; in, represents the global description of channel-level features, and Represents basic features The height and width of Represents basic features The median coordinate is Pixels of The channel attention weight is calculated through the fully connected layer and expressed as: ; in, Indicates that the first fully connected layer performs a linear transformation on the global pooling features; represents the activation function; Further refine the attention information for the second fully connected layer; σ represents the Sigmoid activation function; W represents the channel attention weight matrix; The channel enhancement feature is calculated by channel weighting, which is expressed as: ; in, Represents channel enhancement characteristics.
5. The method for automatically scoring benign and malignant nodules based on thyroid ultrasound images according to claim 3, characterized in that: The method of using spatial attention to enhance the key areas in the channel enhancement feature to obtain the enhanced feature includes the following steps: Use maximum pooling and average pooling to obtain the global information of the feature map, expressed as: ; in, is the maximum pooling result, represents the maximum pooling, is the average pooling result, represents average pooling; The spatial attention weight is calculated using convolution, expressed as: ; Among them, A represents the spatial attention weight, express Convolution operation; σ represents the Sigmoid activation function; The enhanced features are calculated by spatial weighting, expressed as: ; in, Indicates an enhanced feature.
6. The method for automatically scoring benign and malignant nodules based on thyroid ultrasound images according to claim 1, characterized in that: The decoder gradually reduces the dimension of the enhanced features output by the encoder and finally outputs the thyroid nodule recognition result, including the following steps: The feature dimension is gradually reduced through two layers of 3×3 convolution, expressed as: ; ; in, is the feature map output by the first layer 3×3 convolution. Represents the feature map of the second layer 3×3 convolution output; express Convolution operation; represents the activation function; Global average pooling is used to convert spatial features into vectors, expressed as: ; in, Representation feature map The median coordinate is Pixels of The final classification is performed through the fully connected layer, expressed as: ; Among them, Y represents the output category probability, b represents the bias term, and W represents the channel attention weight matrix.
7. The method for automatically scoring benign and malignant nodules based on thyroid ultrasound images according to claim 1, characterized in that: The nodule classification network uses the Swin Transformer model to classify the thyroid nodule recognition results and calculate the benign and malignant risk scores, including the following steps: The nodule area is cut out from the thyroid nodule recognition results output by the nodule recognition network and input into the SwinTransformer model; The Swin Transformer model uses a local window attention mechanism to extract features from nodule regions and captures image information of different scales through a sliding window strategy to obtain key features including morphology, edges, structures, echoes, and hyperechoic foci. The C-TIRADS scoring system was used to score the key features and sum them to obtain the final benign and malignant risk score results.
8. An automatic scoring system for benign and malignant nodules based on thyroid ultrasound images, characterized in that: include: The data set acquisition module acquires and annotates the thyroid ultrasound image data set to construct a training data set; Model training module, which builds a deep learning model including nodule recognition network and nodule classification network, uses the training data set to train the model, and obtains the trained deep learning model as the automatic scoring model; The risk scoring module inputs the thyroid ultrasound image into the automatic scoring model to obtain the thyroid nodule recognition results and calculate the benign and malignant risk scores; The nodule recognition network adopts an encoder-decoder architecture, the encoder extracts basic features from the input image, and uses a dynamic attention mechanism to enhance features; The decoder gradually reduces the dimension of the enhanced features output by the encoder and finally outputs the thyroid nodule recognition result; the nodule classification network uses the Swin Transformer model to classify the thyroid nodule recognition results and calculate the benign and malignant risk score.
Citation Information
Patent Citations
Thyroid nodule ultrasound image benign and malignant classification method based on deep learning
CN117197519A
Thyroid ultrasound contrast nodule benign and malignant classification method based on 3D ConvFormer
CN117237739A
Thyroid nodule ultrasonic image auxiliary diagnosis method fused with medical priori knowledge
CN117611895A
Thyroid nodule benign and malignant classification method based on segmented multi-feature information
CN117911772A
Benign and malignant automatic auxiliary identification system for follicular thyroid nodule ultrasonic image
CN118505589A
Cited By
Lightweight multi-feature fusion thyroid nodule classification method
CN120318613A
Malignant nodule detection image processing method based on benign thyroid
CN121190423A
Thyroid nodule ultrasound image multi-feature analysis system based on deep learning
CN121788523A