Ultrasonic thyroid nodule detection method and device based on deep learning feature fusion
Through a deep learning-based feature fusion method, the BDA-YOLO model and Attention C3K2 module are used to solve the problem of limited accuracy of existing thyroid nodule detection methods, and more efficient and accurate thyroid nodule detection is achieved.
Patent Information
- Application Number
- CN202510259588.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
The accuracy of the existing thyroid nodule detection methods is limited by the problem of image quality and algorithm performance, which is difficult to meet the actual needs of ultrasound examination.
Using a deep learning-based feature fusion method, the bounding box and lesion area of the thyroid nodule prediction model was extracted and fused ultrasound images were used to predict the bounding box and lesion area of the thyroid nodule through the BDA-YOLO thyroid nodule prediction model, combined with the Attention C3K2 module and the bidirectional feature fusion module.
It improves the accuracy and robustness of thyroid nodules detection, enhances the recognition ability of nodules of different scales, adapts to different ultrasound acquisition devices and image quality, and improves detection efficiency.
Smart Images

Figure CN120198756A_ABST
Abstract
Description
Technical Field:
[0001] The present invention relates to a method and device for ultrasonic thyroid nodule detection based on deep learning feature fusion, belonging to the field of computer artificial intelligence. Background Art:
[0002] Ultrasonic diagnosis is a non-invasive diagnostic method that uses ultrasonic waves to obtain the morphological, structural, and functional information of internal tissues and organs of the human body. It has the advantages of non-injury, non-intervention, economy, repeatability, and wide adaptability, and has become an indispensable important imaging examination method in clinical medicine. A thyroid nodule is a tissue mass presented in the thyroid tissue after abnormal proliferation of thyroid cells, which may compress the surrounding tissues, leading to physiological problems such as neck swelling, hoarseness, and difficulty in swallowing. Early detection and diagnosis of thyroid nodules are of great significance for the screening and treatment of thyroid cancer. Therefore, the accuracy and efficiency of thyroid nodule detection directly affect the quality and effect of disease diagnosis.
[0003] Currently, the detection methods of thyroid nodules are mainly divided into two categories: manual detection and automatic detection. Manual detection refers to manually marking the position of thyroid nodules by doctors or technicians. The advantage of this method is high accuracy, but the disadvantages are long time consumption, cumbersome operation, and the existence of subjective differences and human errors. Automatic detection refers to using methods of computer vision and machine learning to automatically detect and identify the position of thyroid nodules. The advantages of this method are high efficiency, simple operation, and the reduction of human interference. The disadvantage is that the accuracy is limited by image quality and algorithm performance. Currently, the automatic detection methods mainly include methods based on traditional machine learning and methods based on deep learning object detection. The methods based on traditional machine learning usually include two parts: feature extraction and classification model. By manually extracting image features such as texture, shape, and edge, and then using classification algorithms such as support vector machines and random forests to identify thyroid nodules. The advantage of this type of method lies in the relatively simple model and small amount of calculation, but it relies on manual feature extraction and is sensitive to image quality. The method based on deep learning object detection refers to automatically learning through a neural network model, extracting features from ultrasonic images, and directly identifying and locating nodules, with high precision and robustness. The advantage of this method is more efficient and accurate detection, but the disadvantages are the lack of fine-grained feature fusion, the neglect of local information attention, and the insufficient adaptability to complex scenarios.
[0004] Existing thyroid nodule detection methods all have certain limitations and are difficult to meet the actual needs of ultrasonic examinations. Therefore, there is an urgent need for a new general thyroid nodule detection method that can overcome the defects of existing methods, improve the accuracy and robustness of detection, and provide more reliable support for the treatment of thyroid diseases. Summary of the Invention:
[0005] In view of the above problems and difficulties of the prior art, the present invention proposes an ultrasonic thyroid nodule detection method and device based on deep learning feature fusion.
[0006] The first aspect of the present invention relates to an ultrasonic thyroid nodule detection method based on deep learning feature fusion, comprising the following steps:
[0007] S1: Collect ultrasonic videos;
[0008] S2: Extract the image sequence of the ultrasonic video, and perform standardization processing such as contrast enhancement on the images;
[0009] S3: Train the BDA-YOLO thyroid nodule prediction model to predict the bounding box and lesion area of the thyroid nodule;
[0010] S3-1: The ultrasonic thyroid nodule dataset contains a large number of ultrasonic images and is specifically used for the detection and segmentation tasks of thyroid nodules. Each image in the dataset is equipped with detailed annotation information, mainly including the segmentation contour of the nodule and the bounding box where the nodule is located. The annotation file of each image contains the contour coordinates and bounding box coordinates of the nodule. The contour of the nodule is annotated in the form of a polygon to accurately represent the boundary of the nodule, while the bounding box provides the circumscribed rectangle of the area where the nodule is located.
[0011] S3-2: Train the BDA-YOLO thyroid nodule prediction model. The BDA-YOLO thyroid nodule prediction model is based on the neural network BDA-YOLO and the ultrasonic thyroid nodule dataset, and is trained to predict the bounding box and segmentation contour of the thyroid nodule. BDA-YOLO mainly includes three parts: Backbone, Neck, and Head. The Backbone part is responsible for extracting the underlying features of the image, mainly using convolutional layer modules to process the input image layer by layer to obtain a series of hierarchical feature maps. The Neck part further processes these feature maps to enhance the information fusion of different scales and helps the network capture richer context information. The Head part uses the Attention C3K2 module introduced in the present invention for feature aggregation and processing. By combining the dynamic self-attention module and the bidirectional feature fusion module proposed in the present invention with the C3K2 module, it can more accurately extract and identify nodule features, strengthen the interaction between shallow and deep layer features, and improve the robustness to nodules of different scales. The Head part subsequently makes predictions based on the output features of the Attention C3K2 module to generate the final target box and class label. The BDA-YOLO thyroid nodule prediction model in the present invention uses a composite loss function, which combines the target box regression loss, classification loss, and bounding box confidence loss, and uses the precision and recall evaluation results of the model at different confidence scores;
[0012] S3-3: Input the image sequence of the ultrasound video obtained in step S2 into the BDA-YOLO thyroid nodule prediction model,
[0013] to predict the target bounding box and lesion area of the thyroid nodule;
[0014] S4: Visualize the target bounding box and lesion area of the thyroid nodule, showing the shape, size, position, and regional contour of the thyroid nodule, which can help doctors diagnose thyroid diseases more easily and provide support for subsequent patient disease treatment;
[0015] Preferably, the optimization objective of the BDA-YOLO thyroid nodule prediction model in step S3 can be formalized as formula (1):
[0016] L = λ cls ·L cls + λ box ·L box + λ obj ·L obj + λ seg ·L seg (1)
[0017] where L cls represents the classification loss, L box represents the localization loss, L obj represents the confidence loss, L seg represents the segmentation loss, and λ cls 、λ box 、λ obj 、λ seg represent the weight coefficients of each loss term;
[0018] L cls The classification loss can be formalized as formula (2):
[0019]
[0020] where S 2 represents the number of grids in the feature map, B represents the number of bounding boxes predicted by each grid, p ij (c) represents the true class label, represents the class probability predicted by the model, represents the indicator function, which is 1 when the j-th bounding box in the i-th grid is responsible for detecting the target, otherwise 0;
[0021] L box The localization loss can be formalized as formula (3):
[0022]
[0023] where b ijRepresent the coordinates of the true bounding box, represent the coordinates of the predicted bounding box, and CIoU represents an improved IoU metric that comprehensively considers the overlapping region, the distance between the center points, and the aspect ratio;
[0024] L obj The confidence loss can be formalized as in Equation (4):
[0025]
[0026] where, represents the object confidence predicted by the model, represents the indicator function, which is 1 when the j-th bounding box in the i-th grid is responsible for detecting the object, and 0 otherwise;
[0027] L seg The segmentation loss can be formalized as in Equation (5):
[0028] L seg = λ bce ·L bce + λ dice ·L dice (5)
[0029] where, L bce represents the binary cross-entropy loss, L dice represents the Dice loss, and λ bce 、λ dice represent the weight coefficients of each loss term;
[0030] L bce The binary cross-entropy loss can be formalized as in Equation (6):
[0031]
[0032] where, m ij represents the true segmentation mask, represents the segmentation mask probability predicted by the model;
[0033] L dice The Dice loss can be formalized as in Equation (7):
[0034]
[0035] Preferably, the dataset in step S3 is a dataset collected in reality or an existing dataset.
[0036] The second aspect of the present invention relates to an ultrasonic thyroid nodule detection device based on deep learning feature fusion, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement an ultrasonic thyroid nodule detection method based on deep learning feature fusion of the present invention.
[0037] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements an ultrasonic thyroid nodule detection method based on deep learning feature fusion of the present invention.
[0038] The present invention first collects ultrasonic videos; then extracts the image sequence of the ultrasonic videos; then trains the BDA-YOLO thyroid nodule prediction model to predict the target box and lesion area of the thyroid nodule; finally visualizes the thyroid nodule, including the shape, size, position and regional contour of the nodule. Through the intuitive visualization results, it helps doctors diagnose thyroid diseases more easily and provides support for the subsequent treatment of patients' diseases.
[0039] The present invention has the following advantages:
[0040] (1) By standardizing and preprocessing ultrasonic images, the feature expression of thyroid nodules is enhanced, and the image quality and processing efficiency are improved;
[0041] (2) By introducing a dynamic self-attention module into the prediction model, the attention degree of the model to different regions is dynamically adjusted. Especially in the case where the nodule boundary and texture information are blurred, the fine-grained feature capture ability of the model can be improved, thereby improving the localization and segmentation accuracy of the model for thyroid nodules;
[0042] (3) By introducing a bidirectional feature interaction module into the prediction model, by fusing feature maps of different levels, the representation ability of the model for the lesion edge and the discrimination ability for multiple lesions are enhanced. This enables the model to still maintain high detection ability when facing multiple thyroid nodules of different sizes;
[0043] (4) Through the method proposed by the present invention, good organ structure point localization effects can be achieved on ultrasonic image data of different ultrasonic acquisition devices, different organs, and different specifications, improving the accuracy and efficiency of ultrasonic thyroid nodule detection and being able to adapt to different thyroid nodule case scenarios. Description of the Drawings:
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0045] Figure 1 is the overall flowchart of the method of the present invention.
[0046] Figure 2 is the schematic diagram of the device of the present invention. Specific embodiments:
[0047] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on the present invention belong to the scope of protection of the present invention.
[0048] Embodiment 1
[0049] Refer to Figure 1 , this embodiment relates to an ultrasonic thyroid nodule detection method based on deep learning feature fusion, including the following steps:
[0050] S1: Obtain data and collect ultrasonic videos;
[0051] In this embodiment, an ultrasonic device is used to collect ultrasonic thyroid image data, and a series of video data files of ultrasonic thyroid images are obtained after the collection.
[0052] S2: Extract the image sequence around the prediction frame;
[0053] Read the image data obtained in step S1, and extract the image sequence from the video of the ultrasonic thyroid image area.
[0054] Perform operations such as grayscale conversion, median filtering, and histogram equalization on each frame of the ultrasonic thyroid image to enhance the image contrast and details.
[0055] S3: Train the algorithm model to predict the bounding box and lesion area of the thyroid nodule;
[0056] An algorithm model, the BDA - YOLO thyroid nodule instance segmentation model, is used in the present invention.
[0057] S3 - 1: Ultrasonic thyroid nodule dataset;
[0058] In this example, the ultrasound public dataset TN3K and DDTI of thyroid nodules are used. Each image in the dataset is accompanied by detailed annotation information, mainly including the segmentation contour of the nodule and the bounding box where the nodule is located. The annotation file of each image contains the contour coordinates and bounding box coordinates of the nodule. The contour of the nodule is annotated in the form of a polygon to accurately represent the boundary of the nodule, while the bounding box provides the circumscribed rectangle of the area where the nodule is located for the training of the deep learning model.
[0059] S3-2: Train the BDA-YOLO instance segmentation model for thyroid nodules;
[0060] The BDA-YOLO instance segmentation model for thyroid nodules is trained based on the neural network BDA-YOLO and the ultrasound thyroid nodule dataset until the preset number of iterations is reached or the stopping condition is met.
[0061] The initial input layer of this network receives a sequence of ultrasound images for training, with the specification of N*448*448, where N represents the number of ultrasound image channels. First, the underlying features of the images are extracted through the Backbone part. The convolutional layer module is mainly used to process the input images layer by layer to obtain a series of hierarchical feature maps. The Backbone part is based on the improved CSPDarknet architecture, and gradually extracts the deep features of the images through a series of convolutions and C3k2 modules. The input image first passes through an initial convolutional layer, which consists of a 3x3 convolutional kernel and a convolutional operation with a stride of 1, and is used to extract preliminary underlying features. Subsequently, the features are further extracted through the second convolutional layer to generate the initial feature map. Next, the input features are processed through multiple C3k2 modules. Each C3k2 module contains a 3x3 convolutional layer and a 1x1 convolutional layer, and realizes feature reuse and fusion through cross-stage partial connections. After each C3k2 module, there is a 3x3 convolutional layer, which is used to further extract features and adjust the number of channels. After continuous processing by four C3k2 modules, the size of the feature map gradually decreases, while the semantic information gradually increases. At the last layer of the Backbone, the SPPF module is introduced to enhance the model's perception ability of multi-scale features. The SPPF module performs fast pooling operations on the feature map through multiple max-pooling layers of different sizes, and splices the pooling results with the original feature map to form multi-scale feature fusion. Finally, the Backbone outputs multi-level feature maps for subsequent processing in the Neck and Head parts. The Neck part is responsible for further fusing and enhancing the multi-level features extracted by the Backbone to capture richer context information. The Neck part is based on the improved PANet architecture and realizes efficient interaction of multi-scale features through top-down and bottom-up bidirectional feature fusion paths. The Head part uses the Attention C3K2 module introduced in the present invention for feature aggregation and processing. By combining the dynamic self-attention module and the bidirectional feature fusion module proposed in the present invention with the C3K2 module, nodules are more accurately extracted and recognized. The dynamic self-attention module adaptively adjusts the model's attention degree to different regions through the self-attention mechanism. Especially in the case where the nodule boundary and texture information are blurred, it can improve the model's fine-grained feature capture ability. This module can effectively eliminate background interference and enhance the saliency of the nodule area, thus greatly improving the localization accuracy of thyroid nodules. The bidirectional feature fusion module performs bidirectional interactive fusion on the feature map, mines and fuses valuable information from different scales, and effectively improves the multi-scale information fusion ability. This enables the model to maintain high detection ability when facing thyroid nodules of different sizes, especially in the case of complex backgrounds in ultrasound images, showing excellent robustness.The Head part is subsequently predicted through multiple branches containing several convolutional layers based on the output features of the Attention C3K2 module to generate the final bounding boxes and segmentation contours.
[0062] S3-3: Predict thyroid nodules;
[0063] Input the image obtained in step S2 into the trained BDA-YOLO thyroid nodule instance segmentation model in step S3-1 for prediction to obtain the predicted bounding box and segmentation contour data.
[0064] S4: Visualize the coordinates of the organ structure points. Visualize the bounding box and segmentation contour data obtained in S3 on the frame sequence of the cardiac ultrasound video in S1 to show the shape, size, position, and regional contour of the thyroid nodules, which can help doctors diagnose thyroid diseases more easily and provide support for subsequent patient disease treatment.
[0065] Embodiment 2
[0066] Refer to Figure 2 , this embodiment relates to an ultrasound thyroid nodule detection device based on deep learning feature fusion, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for detecting ultrasound thyroid nodules based on deep learning feature fusion in Embodiment 1.
[0067] Embodiment 3
[0068] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for detecting ultrasound thyroid nodules based on deep learning feature fusion of the present invention.
[0069] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for detecting thyroid nodules by ultrasound based on deep learning feature fusion, characterized in that: The following steps are involved: S1: Acquisition of ultrasound video; S2: extract the image sequence of the ultrasound video and perform standardization processing such as contrast enhancement on the image; S3: Train the BDA-YOLO thyroid nodule prediction model to predict the bounding box and lesion area of the thyroid nodule; S4: Visualize the target box and lesion area of the thyroid nodule, showing the shape, size, location and regional outline of the thyroid nodule.
2. The ultrasonic thyroid nodule detection method based on deep learning feature fusion according to claim 1, characterized in that: The step S3 specifically includes: S3-1: The ultrasound thyroid nodule dataset contains ultrasound images for thyroid nodule detection and segmentation tasks. Each image in the dataset is annotated with the segmentation contour of the nodule and the bounding box where the nodule is located. The annotation file of each image contains the contour coordinates and bounding box coordinates of the nodule. The contour of the nodule is annotated in the form of a polygon, which accurately represents the boundary of the nodule, while the bounding box provides the circumscribed rectangular frame of the area where the nodule is located. S3-2: Training the BDA-YOLO thyroid nodule prediction model; The BDA-YOLO thyroid nodule prediction model is based on the neural network BDA-YOLO and the ultrasound thyroid nodule dataset, and is trained to predict the bounding box and segmentation contour of the thyroid nodule. BDA-YOLO consists of three parts: Backbone, Neck, and Head; The Backbone part is responsible for extracting the underlying features of the image, and uses the convolutional layer module to process the input image layer by layer to obtain a series of hierarchical feature maps; The Neck part further processes these feature maps to enhance the fusion of information at different scales and help the network capture richer contextual information; The Head part uses the Attention C3K2 module for feature aggregation and processing. By combining the dynamic self-attention module and the bidirectional feature fusion module with the C3K2 module, the nodule features are extracted and identified, and the interaction between deep and shallow features is strengthened, thereby improving the robustness to nodules of different scales; The Head part subsequently predicts based on the output features of the Attention C3K2 module to generate the final target box and category label; S3-3: Input the image sequence of the ultrasound video obtained in step S2 into the BDA-YOLO thyroid nodule prediction model to predict the target box and lesion area of the thyroid nodule.
3. The method for locating ultrasonic organ structure points based on deep learning according to claim 1, characterized in that: The BDA-YOLO thyroid nodule prediction model described in step S3-2 adopts a composite loss function Loss, combining the target box regression loss, classification loss, and bounding box confidence loss to minimize the Loss. The calculation formula of the Loss is: L=λ cls ·L cls +λ box ·L box +λ obj ·L obj +λ seg ·L seg (1) Among them, L cls represents the classification loss, L box represents the positioning loss, L obj represents the confidence loss, L seg represents the segmentation loss, λ cls , box , obj , seg Represents the weight coefficient of each loss item; L cls The classification loss is calculated as: Among them, S 2 represents the number of grids in the feature map, B represents the number of bounding boxes predicted by each grid, and p ij (c) represents the true category label, represents the class probability predicted by the model, represents the indicator function, which is 1 when the j-th bounding box of the i-th grid is responsible for detecting the target, otherwise it is 0; L box The calculation formula of positioning loss is: Among them, b ij represents the coordinates of the true bounding box, Represents the coordinates of the predicted bounding box, CIoU represents the improved IoU index that comprehensively considers the overlapping area, center point distance and aspect ratio; L obj The confidence loss is calculated as: in, represents the target confidence of the model prediction, represents the indicator function, which is 1 when the j-th bounding box of the i-th grid is responsible for detecting the target, otherwise it is 0; L seg The calculation formula of segmentation loss is: L seg =λ bce ·L bce +λ dice ·L dice (5) Among them, L bce represents the binary cross entropy loss, L dice represents the Dice loss, λ bce , dice Represents the weight coefficient of each loss item; L bce The binary cross entropy loss is calculated as: Among them, m ij represents the true segmentation mask, represents the segmentation mask probability predicted by the model; L dice The calculation formula of Dice loss is:
4. The method for ultrasonic organ structure point positioning based on deep learning according to claim 1, characterized in that: The step S4 specifically includes: Read the image data obtained in step S1, determine the prediction frame, and extract the image sequence adjacent to the prediction frame from the cardiac ultrasound image area video as the image sequence around the prediction frame; for multiple available prediction frames, screen and select image sequences with less noise and clear images for subsequent analysis and evaluation; or, according to actual needs, select the systole or diastole of the heart for analysis.
5. An ultrasonic organ structure point positioning device based on deep learning, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement an ultrasonic thyroid nodule detection method based on deep learning feature fusion as described in any one of claims 1-4.
6. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, an ultrasonic thyroid nodule detection method based on deep learning feature fusion as described in any one of claims 1-4 is implemented.