Thyroid follicular lesion segmentation model training method, segmentation method and equipment
By enhancing contrast and smoothing the edge of the thyroid follicular grayscale ultrasound images, and training the Unet model with the Transformer module, the problem of poor segmentation effect caused by low ultrasound image quality is solved, and more efficient and reliable segmentation of thyroid follicular lesions is achieved.
Patent Information
- Application Number
- CN202511044811.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-07-29
AI Technical Summary
In the existing thyroid follicle lesion segmentation technology, the low quality of ultrasound image results in poor robustness to noise and artifacts, making it difficult to achieve efficient and stable segmentation effect.
Through enhanced contrast-limiting and smooth edge transitions on thyroid follicular grayscale ultrasound images, the Unet model is trained in combination with the cross-fusion feature context Transformer module and the regional attention module to improve image clarity and reliability of data samples.
The training effectiveness and reliability of the thyroid follicle lesion segmentation model is improved, the accuracy of the segmentation results is enhanced, the impact of noise is reduced, and the overall visual consistency of the data samples is maintained.
Smart Images

Figure CN120543979A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image data processing technology, and in particular to a thyroid follicular lesion segmentation model training method, segmentation method, and device. Background Art
[0002] Follicular thyroid tumors are a type of tumor that develops in the thyroid gland, specifically within the follicular cells of the thyroid gland. Lesion segmentation, a machine learning task, accurately segments thyroid lesions, automatically extracting key information such as their morphology, size, boundaries, and structure. This provides a powerful basis for classifying benign and malignant diseases. Segmentation technology is particularly important when distinguishing between follicular thyroid cancer and follicular adenomas, which share very similar pathological structures.
[0003] Currently, in the existing thyroid follicular lesion segmentation technology, most algorithms only focus on optimizing the network structure, and insufficiently explore image preprocessing and image enhancement technologies, thus ignoring the low quality of the ultrasound image itself. This results in poor robustness of the model to noise and artifacts, making it difficult to achieve efficient and stable segmentation effects in a real clinical environment. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a thyroid follicular lesion segmentation model training method, segmentation method and device to eliminate or improve one or more defects in the prior art.
[0005] A first aspect of the present application provides a thyroid follicular lesion segmentation model training method, comprising: performing contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image, so as to obtain a contrast enhanced block for each image block; Merging the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image and performing edge smoothing to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images; A target Unet model is trained using each of the enhanced image samples and the label group corresponding to each of the enhanced image samples, so as to train the target Unet model into a thyroid follicular lesion segmentation model for outputting the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the label group includes the lesion mask label and the classification label of the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; and the classification label includes a benign label and a malignant label.
[0006] In some embodiments of the present application, performing contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block includes: Divide each thyroid follicular grayscale ultrasound image into multiple image blocks; Calculating the cumulative distribution value of the histogram corresponding to each of the image blocks respectively; If the cumulative distribution value of the histogram corresponding to the image block is greater than a contrast limit threshold, the target cumulative distribution value corresponding to the image block is set as the contrast limit threshold; if the cumulative distribution value of the histogram corresponding to the image block is less than or equal to the contrast limit threshold, the target cumulative distribution value corresponding to the image block is set as the cumulative distribution value of the histogram corresponding to the image block; Contrast enhancement processing is performed on each of the image blocks according to the target cumulative distribution value corresponding to each of the image blocks, so as to obtain a contrast enhanced block for each of the image blocks.
[0007] In some embodiments of the present application, merging the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image and performing edge smoothing to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images includes: Merging the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image to obtain merged images corresponding to the respective thyroid follicular grayscale ultrasound images; In a bilinear interpolation manner, edge smoothing transition processing is performed between the contrast enhancement blocks belonging to the same merged image to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images.
[0008] In some embodiments of the present application, a cross-fusion feature context Transformer module is provided between the output end of the encoder and the input end of the decoder in the target Unet model; The cross-fusion feature context Transformer module is used to perform feature aggregation, multi-scale feature extraction, group convolution and feature interaction processing based on a multi-head attention mechanism on the downsampled feature data corresponding to the enhanced image sample output by the encoder in sequence to obtain a target feature map corresponding to the enhanced image sample, and input the target feature map into the decoder so that the decoder outputs the thyroid follicular lesion segmentation result data corresponding to the enhanced image sample according to the target feature map; wherein the thyroid follicular lesion segmentation result data is used to represent the thyroid follicular lesion area segmented from the enhanced image sample and the classification label corresponding to the thyroid follicular lesion area.
[0009] In some embodiments of the present application, the cross-fusion feature context Transformer module includes: a feature aggregation layer, a multi-scale feature extraction layer, a grouped convolution layer based on depthwise separable convolution, a Transformer module based on a multi-head attention mechanism, and an output layer; The feature aggregation layer is used to perform initial feature aggregation processing on the downsampled feature data corresponding to the enhanced image samples output by the encoder using a convolution kernel, and perform layer normalization processing on the downsampled feature data after the initial feature aggregation processing to obtain aggregated feature data corresponding to the enhanced image samples; The multi-scale feature extraction layer is used to perform convolution feature extraction on the aggregated feature data using convolution kernels of different sizes to obtain extracted feature maps of different scales corresponding to the aggregated feature data, and to fuse the extracted feature maps to obtain convolution feature data corresponding to the enhanced image sample; The grouped convolution layer based on depthwise separable convolution is used to divide the convolution-extracted feature data into feature image blocks of fixed size; flatten each feature image block and map each feature image block into an embedding vector by linear projection; perform depthwise convolution and pointwise convolution on each embedding vector using depthwise separable convolution to obtain an injection feature map of each channel corresponding to the convolution-extracted feature data, and then use a convolution kernel to aggregate the injection feature maps of each channel to obtain a dynamic feature vector corresponding to the enhanced image sample; The Transformer module based on the multi-head attention mechanism is used to perform feature interaction on the dynamic feature vector based on the multi-head attention mechanism to obtain convolution injection feature data corresponding to the enhanced image sample; The output layer is used to fuse the convolution extraction feature data and the convolution injection feature data corresponding to the enhanced image sample to obtain a target feature map corresponding to the enhanced image sample.
[0010] In some embodiments of the present application, a regional attention module is provided in the encoder of the target Unet model; The regional attention module is used to extract features from the input data using a spatial attention mechanism and a channel attention mechanism in sequence to obtain regional attention feature data corresponding to the enhanced image sample; and input the regional attention feature data into a downsampling layer in the encoder described in itself, so that the downsampling layer outputs downsampled feature data corresponding to the regional attention feature data; Among them, if the encoder where the regional attention module is located is the first encoder in the target Unet model, the input data is the enhanced image sample; if the encoder where the regional attention module is located is not the first encoder in the target Unet model, the input data is the downsampled feature data output by another encoder connected to the input side of the encoder.
[0011] In some embodiments of the present application, the region attention module includes: a spatial attention layer and a channel attention layer; The spatial attention layer is used to obtain a spatial attention map corresponding to the input data using sequentially connected convolutional layers and ReLU activation functions, and multiply the spatial attention map by the input data element-by-element to obtain spatial attention feature data corresponding to the enhanced image sample; The channel attention layer is used to use a global pooling layer, a one-dimensional convolution layer and a ReLU activation function connected in sequence to obtain the spatial and channel attention maps corresponding to the spatial attention feature data, and multiply the spatial and channel attention maps with the input data element by element to obtain the regional attention feature data corresponding to the enhanced image sample.
[0012] A second aspect of the present application provides a thyroid follicular lesion segmentation method, comprising: performing contrast enhancement processing based on contrast limitation on each image block corresponding to the target thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block; Merging the contrast enhancement blocks and performing edge smoothing to obtain an enhanced image sample corresponding to the target thyroid follicular grayscale ultrasound image; The enhanced image sample is input into a thyroid follicular lesion segmentation model so that the thyroid follicular lesion segmentation model outputs the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein, the thyroid follicular lesion segmentation model is pre-trained based on the thyroid follicular lesion segmentation model training method described in the first aspect above.
[0013] The third aspect of the present application provides a thyroid follicular lesion segmentation model training device, comprising: a first contrast enhancement module, configured to perform contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image, so as to obtain a contrast enhanced block for each image block; a first smooth transition module, configured to merge and perform edge smoothing on the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image, so as to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images; A model training module is used to train a target Unet model using each of the enhanced image samples and the label group corresponding to each of the enhanced image samples, so as to train the target Unet model into a thyroid follicular lesion segmentation model for outputting the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the label group includes the lesion mask label and the classification label of the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; and the classification label includes a benign label and a malignant label.
[0014] A fourth aspect of the present application provides a thyroid follicular lesion segmentation device, comprising: a second contrast enhancement module, configured to perform contrast enhancement processing based on contrast limitation on each image block corresponding to the target thyroid follicular grayscale ultrasound image, so as to obtain a contrast enhanced block for each image block; a second smooth transition module, configured to merge the contrast enhancement blocks and perform edge smooth transition processing to obtain an enhanced image sample corresponding to the target thyroid follicular grayscale ultrasound image; A model segmentation module is used to input the enhanced image sample into a thyroid follicular lesion segmentation model so that the thyroid follicular lesion segmentation model outputs the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the thyroid follicular lesion segmentation model is pre-trained based on the thyroid follicular lesion segmentation model training method described in the first aspect above.
[0015] The fifth aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the thyroid follicular lesion segmentation model training method provided in the first aspect is implemented, and / or the thyroid follicular lesion segmentation method provided in the second aspect is implemented.
[0016] The sixth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the thyroid follicular lesion segmentation model training method provided in the first aspect, and / or implements the thyroid follicular lesion segmentation method provided in the second aspect.
[0017] The seventh aspect of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the thyroid follicular lesion segmentation model training method provided in the first aspect, and / or implements the thyroid follicular lesion segmentation method provided in the second aspect.
[0018] The thyroid follicular lesion segmentation model training method provided in the present application performs contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block; merges and performs edge smoothing transition processing on each contrast enhanced block belonging to the same thyroid follicular grayscale ultrasound image to obtain an enhanced image sample corresponding to each thyroid follicular grayscale ultrasound image; uses each enhanced image sample and each corresponding label group of the enhanced image sample to train a target Unet model to train the target Unet model into a thyroid follicular lesion segmentation model for outputting the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the label group includes the lesion mask label and the classification label of the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; the classification label includes a benign label and a malignant label. That is to say, by independently performing contrast enhancement for each image block, the present application can selectively adjust the contrast for different image areas, thereby reducing the risk of overall distortion; by performing contrast enhancement processing based on contrast limitation, it can prevent excessive enhancement of noise and artifacts, thereby effectively reducing noise while improving image clarity; by merging and edge smoothing the various contrast enhanced blocks belonging to the same thyroid follicular grayscale ultrasound image, it can eliminate the grayscale differences between adjacent image blocks to maintain the overall visual consistency of the data sample; thereby, it can improve the application reliability of the data samples used to train the thyroid follicular lesion segmentation model, effectively improve the training effectiveness and reliability of the thyroid follicular lesion segmentation model, and thereby improve the accuracy of the sampled high thyroid follicular lesion segmentation results.
[0019] Additional advantages, purposes, and features of the present application will be described in part in the following description and will become apparent to those skilled in the art upon study of the following or may be learned from practice of the present application. The purposes and other advantages of the present application may be achieved and obtained by the structures specifically pointed out in the specification and drawings.
[0020] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are intended to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger than other components in the exemplary device actually manufactured according to the present application. In the drawings: Figure 1 This is a schematic diagram of the first flow chart of the thyroid follicular lesion segmentation model training method in one embodiment of the present application.
[0022] Figure 2 This is a second flow chart of the thyroid follicular lesion segmentation model training method in one embodiment of the present application.
[0023] Figure 3 Schematic diagram of the overall architecture of the target Unet model in one embodiment of the present application.
[0024] Figure 4 Schematic diagram of the architecture of the cross-fusion feature context Transformer module in one embodiment of the present application.
[0025] Figure 5 FIG. 1 is a schematic diagram of the architecture of a CFCT module in an application example of the present application.
[0026] Figure 6 Schematic diagram of the specific architecture of the target Unet model in one embodiment of the present application.
[0027] Figure 7 Schematic diagram of the architecture of the regional attention module in one embodiment of the present application.
[0028] Figure 8 Schematic diagram of the flow of a thyroid follicular lesion segmentation method in one embodiment of the present application.
[0029] FIG9( a ) is a schematic diagram showing comparison results of an original image of a follicular thyroid carcinoma (FTC) and an image pre-processed with CLAHE in an application example of the present application.
[0030] FIG9( b ) is a schematic diagram showing the comparison results of the original image of a thyroid follicular adenoma (FA) and the image after CLAHE preprocessing in an application example of the present application.
[0031] Figure 10 A schematic diagram of the data set construction process in an application example of this application.
[0032] FIG11( a ) is a schematic diagram showing a comparison of the segmentation results of thyroid follicular lesion segmentation model in an application example of the present application and other models for follicular thyroid carcinoma (FTC) images.
[0033] FIG11( b ) is a schematic diagram showing a comparison of the segmentation results of thyroid follicular lesion segmentation model in an application example of the present application and other models for thyroid adenoma (FA) images.
[0034] Figure 12 Schematic diagram of the structure of a thyroid follicular lesion segmentation model training device in one embodiment of the present application.
[0035] Figure 13 Schematic diagram of the structure of a thyroid follicular lesion segmentation device in one embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail in conjunction with the embodiments and drawings. Here, the illustrative embodiments of this application and their descriptions are used to explain this application, but are not intended to limit this application.
[0037] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, the accompanying drawings only show structures and / or processing steps that are closely related to the scheme according to the present application, while other details that are not closely related to the present application are omitted.
[0038] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.
[0039] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0040] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0041] Artificial intelligence technology is becoming increasingly integrated with medical imaging. Using AI to process medical images is becoming a major research hotspot. In the medical process, doctors spend a significant portion of their time manually segmenting regions of interest to calculate organ function indicators. Applying AI to image segmentation can not only significantly reduce time costs and accelerate the medical diagnosis process, but can also serve as a precursor to other tasks such as auxiliary diagnosis. Ultrasound image segmentation, due to its low resolution and high signal-to-noise ratio, is currently a research hotspot.
[0042] The thyroid gland is the largest endocrine gland in the human body, located below the thyroid cartilage in the neck, on either side of the trachea. It is a crucial organ for controlling energy use, producing proteins, and regulating sensitivity to other endocrine hormones.
[0043] Thyroid cancer is a malignant tumor that originates from the thyroid follicular or parafollicular epithelium and is the most common malignant tumor of the head and neck. Based on tumor origin and differentiation, thyroid cancer is further categorized into papillary thyroid carcinoma (PTC) and follicular thyroid carcinoma (FTC). Follicular thyroid carcinoma is the second most common pathological type of thyroid cancer, accounting for 10% to 15% of all thyroid cancers. Follicular carcinomas are typically encapsulated and invasive, with hematogenous metastasis to the lungs and bone. Their prognosis is generally worse than that of papillary carcinoma. Thyroid follicular tumors are categorized into benign follicular adenomas and malignant follicular carcinomas. Approximately 10% of follicular tumors are susceptible to malignant transformation. The cellular morphology of these two conditions is similar, and preoperative cytological examination cannot confirm the presence of a capsule or vascular invasion. Therefore, preoperative diagnosis of follicular thyroid carcinoma and follicular adenoma is relatively difficult.
[0044] With the continuous advancement of science and technology, techniques such as magnetic resonance imaging, computed tomography, and ultrasound have become widely used in medical examinations. Ultrasound, due to its advantages such as affordability, convenience, non-invasiveness, and real-time performance, has become a common diagnostic method for thyroid nodules. It has achieved remarkable results in locating the size and number of nodules, combining it with fine needle aspiration cytology (FNAC) to diagnose malignancies, differentiating solid from cystic lesions, and improving the performance of ultrasound in the diagnosis of thyroid nodules using the ultrasound risk stratification system (US-RSS). However, ultrasound has limitations in distinguishing between follicular thyroid carcinoma and follicular adenoma: Follicular thyroid carcinoma and follicular adenoma often exhibit similar features on ultrasound imaging, making differentiation difficult. The uncertainty of some follicular thyroid carcinomas and the presence of benign lesions with atypical features further complicate ultrasound differentiation. Capsular invasion and vascular invasion are key diagnostic "gold standard" features of follicular thyroid carcinoma, but these features are difficult to accurately observe on ultrasound. Preoperative diagnosis of follicular carcinoma and benign adenoma is difficult, and only postoperative pathological examination can effectively distinguish between the two. However, postoperative pathological diagnosis will affect the normal function of the thyroid gland.
[0045] In the current era of big data, artificial intelligence (AI) has become a research hotspot at the intersection of medicine and engineering, and its application is rapidly driving the development of medical diagnosis and disease management. Currently, the application of AI in disease diagnosis covers multiple areas, such as data analysis and real-time early warning of critical illnesses, precise lesion localization and prediction based on imaging findings, aiding in differential diagnosis of diseases, risk prediction of benign and malignant lesions, and evaluating the effectiveness of treatment options. However, research in the interdisciplinary field of AI for follicular thyroid carcinoma remains significantly underdeveloped. Most current research focuses on screening and risk prediction for papillary thyroid carcinoma, while research on follicular thyroid carcinoma is relatively scarce. The similarities in pathological and imaging findings between follicular carcinoma and follicular adenoma significantly increase the difficulty of diagnosis. Furthermore, most existing studies rely on single-modality imaging data, grayscale ultrasound. Follicular carcinoma often appears highly similar to adenoma in grayscale ultrasound images, making differentiation between the two challenging using only a single modality.
[0046] Furthermore, grayscale ultrasound has inherent limitations in medical imaging. Compared with other imaging techniques such as computed tomography (CT), grayscale ultrasound images typically have a lower signal-to-noise ratio, lower resolution, and less contrast, resulting in insufficient image detail. Furthermore, the acquisition of ultrasound images relies on operator experience and equipment performance, and these human factors and equipment differences can lead to significant fluctuations in image quality. These unfavorable conditions combine to severely restrict the accuracy and robustness of ultrasound-based artificial intelligence models in classifying and diagnosing follicular thyroid tumors.
[0047] Therefore, it is crucial to develop an intelligent screening and malignancy risk prediction method for follicular thyroid tumors that can overcome existing technological bottlenecks. Such an approach should fully leverage artificial intelligence technologies to achieve efficient fusion and deep learning of multimodal data, automatically segmenting and extracting features from ultrasound images, and overcoming the limitations of traditional grayscale ultrasound imaging. Furthermore, by combining ultrasound data with data from other modalities, diagnostic accuracy and stability can be effectively improved, providing new technical support for the early screening and precise diagnosis of follicular thyroid tumors.
[0048] Ultrasound is the first-line examination for thyroid nodules, and has high sensitivity and specificity in distinguishing benign and malignant nodules. However, existing studies on ultrasound in thyroid malignant transformation are all based on papillary thyroid carcinoma. Because thyroid follicular carcinoma and thyroid follicular adenoma often show similar characteristics on ultrasound imaging, the controversy over whether ultrasound can be directly used for thyroid follicular carcinoma is still unresolved. In medicine, the World Health Organization has proposed a diagnostic decision tree for follicular thyroid cancer, which requires a 7-step decision-making process to complete the definitive diagnosis of thyroid follicular carcinoma. And because fine needle aspiration biopsy (FNAB) cannot distinguish between thyroid follicular carcinoma and adenoma, in summary, the preoperative diagnosis of thyroid follicular carcinoma in clinical practice is relatively difficult.
[0049] Deep learning is currently being widely applied in medical diagnosis, and AI-based medical diagnostic methods are gradually emerging in applications such as disease risk prediction and the classification of difficult diseases. In the classification and diagnosis of thyroid nodules, attempts have been made to use computer-assisted diagnosis (CAD) and deep learning methods, with promising results. Some scholars have applied artificial intelligence technology to thyroid cell pathology, which helps to distinguish papillary carcinoma from benign lesions, distinguish follicular adenomas from carcinomas, and identify non-invasive follicular thyroid tumors with papillary nuclear characteristics; some scholars use the deep separable convolutional neural network Xception as the main classification network to perform multi-scale convolution on ultrasound images to extract feature information, and predict and classify various thyroid nodular lesions including thyroid cancer. In the comparative experiment of multiple groups of networks, the best results were obtained, with a classification accuracy of up to 97.2%; some scholars use a two-stage convolutional neural network (CNN) to use the results of tissue section staining to predict the benign or malignant results of follicular nodules, with a sensitivity of 92.0% and a specificity of 90.5%; some scholars use the residual network ResNet, or an improved network with other network structure improvements based on ResNet, as well as the classic convolutional neural network VGG network model for multi-classification tasks of thyroid nodular lesions, and have also achieved remarkable results, obtaining diagnostic results that are not inferior to those of professional physicians.
[0050] However, existing research has the following problems: First, due to the low resolution and high signal-to-noise ratio of ultrasound thyroid nodule images, they may not be able to fully represent all the characteristics of the lesions. The potential information loss problem may significantly affect the performance of the deep learning model. In addition, since ultrasound images need to be manually captured, there is no definite standard for image capture methods, and the accuracy of data annotation is also a major factor in the performance of the imaging model. Most existing studies focus on major categories of diseases (such as inflammation, cancer, etc.) without detailed research, especially in the research of thyroid follicular tumors, where there is still a large gap. Thyroid follicular carcinoma and thyroid adenoma have no obvious distinguishing features on ultrasound images, and it is difficult to distinguish between the two based on ultrasound images alone.
[0051] Thanks to the increased computing power, algorithm optimization, and inherent flexibility of GPUs, deep learning methods have surpassed traditional machine learning approaches in object classification, localization, and semantic segmentation. Based on the U-net model, some studies have achieved segmentation of nasopharyngeal tumor magnetic resonance imaging (MRI) images. This approach utilizes a contracting path to acquire surrounding information and, based on this, expands the path to achieve precise localization. Some researchers have combined the results of the previous dilated convolution algorithm with those of traditional convolution. The stitched image is then transferred to the next dilated convolutional layer, and a dense atrous spatial pyramid pooling (DenseASPP) model is proposed. Some scholars have combined tracking three-dimensional ultrasound with CNN segmentation, significantly reducing the inter-observer variability of thyroid volume measurement and improving measurement accuracy through shorter acquisition times. Some scholars have proposed a new hybrid Transformer-UNet (H-TUNet) to segment the thyroid gland in ultrasound sequences. The proposed method outperforms other state-of-the-art methods in the ultrasound dataset TSUD and the medical imaging dataset TG3k. Transformer-UNet is a deep learning architecture that combines the Transformer, a deep learning model based on an attention mechanism, with the U-shaped fully convolutional neural network (UNet), primarily used in fields such as medical image segmentation. Its core idea is to combine the global context processing capabilities of the Transformer with the local feature extraction advantages of the UNet to improve segmentation accuracy. Based on the densely connected convolutional network DenseNet-121 network structure model, combined with the core module ASPP (Atrous Spatial Pyramid Pooling) for dense prediction tasks (especially semantic segmentation), some scholars proposed a new thyroid nodule ultrasound image segmentation model, which greatly improved the segmentation effect of thyroid nodule ultrasound images. Some scholars also proposed a thyroid region prior guided feature enhancement network TRFE-Net for thyroid nodule segmentation. This framework uses inferred thyroid region priors to enhance the feature representation of thyroid nodule segmentation. Other scholars proposed an FCN architecture called IVUS-Net, followed by a post-processing contour extraction step, to automatically segment the internal (lumen) and external (media and adventitia) regions of human arteries.
[0052] However, existing research has the following problems: First, due to the large amount of noise, fuzzy boundaries and complex background problems in ultrasound images, the traditional segmentation model has low segmentation accuracy in ultrasound images; in order to overcome the shortcomings of ultrasound images, the complexity of the segmentation model increases, which is not ideal for accelerating training and reducing optimization difficulty; due to the heterogeneity of the appearance of thyroid nodules and the possibility that thyroid tissue may be confused with the edge effects of other tissues under ultrasound, how to accurately segment and classify them is a challenge.
[0053] Follicular thyroid tumors can be divided into benign follicular adenomas (FA) and malignant follicular thyroid carcinomas (FTC). Benign adenomas are typically encapsulated and noninvasive, whereas positive follicular thyroid carcinomas may invade blood vessels and potentially metastasize, with common metastatic sites including the bones or lungs. Because adenomas and follicular carcinomas exhibit similar cytologic features, reliably distinguishing them using fine-needle aspiration biopsy is difficult. Distinguishing follicular adenomas from thyroid follicular carcinomas by ultrasound remains challenging due to the lack of specific ultrasound features.
[0054] In recent years, deep learning-based methods have been widely used in the classification of thyroid nodules due to their ability to analyze complex features in medical imaging and pathology data. Lesion segmentation, as a machine learning task, can accurately segment thyroid lesions and automatically extract key information such as morphology, size, boundaries, and structure, providing a powerful basis for benign and malignant classification. Segmentation technology is particularly important when distinguishing between follicular thyroid carcinoma and follicular adenoma, which have very similar pathological structures. Compared to manual image analysis, automated segmentation using artificial intelligence is not only faster but also reduces subjective errors caused by differences in physician experience. Therefore, it can better guide thyroid cancer staging, surgical scope planning, and radiotherapy dose calculation, thereby optimizing treatment strategies. However, some existing segmentation models often focus solely on optimizing network structure, insufficiently exploring image preprocessing and image enhancement techniques, and thus neglecting the inherent low quality of ultrasound images. This results in poor robustness to noise and artifacts, making it difficult to achieve efficient and stable segmentation results in real-world clinical settings.
[0055] Based on this, in order to improve the clarity and application reliability of data samples used to train the thyroid follicular lesion segmentation model, the embodiments of the present application respectively provide a thyroid follicular lesion segmentation model training method, a thyroid follicular lesion segmentation model training device for executing the thyroid follicular lesion segmentation model training method, a thyroid follicular lesion segmentation method, a thyroid follicular lesion segmentation device, a physical device, a computer-readable storage medium and a computer program product, which can reduce the risk of overall distortion of data samples, reduce noise and maintain the overall visual consistency of data samples, effectively improve the training effectiveness and reliability of the thyroid follicular lesion segmentation model, and thus improve the accuracy of the sampled high thyroid follicular lesion segmentation results.
[0056] The details are described in detail through the following examples.
[0057] Based on this, the embodiment of the present application provides a thyroid follicular lesion segmentation model training method that can be implemented by a thyroid follicular lesion segmentation model training device, see Figure 1 The thyroid follicular lesion segmentation model training method specifically includes the following contents: Step 100: performing contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block.
[0058] In one or more embodiments of the present application, the grayscale ultrasound image of thyroid follicles may be referred to as a thyroid ultrasound image or ultrasound image, both of which refer to ultrasound images showing thyroid follicular adenoma or thyroid follicular carcinoma. The image blocks refer to sub-blocks obtained by cutting or dividing a thyroid follicular grayscale ultrasound image, which can also be called windows or grid areas. In one example, the thyroid follicular grayscale ultrasound image can be divided into multiple image blocks of the same size.
[0059] In step 100, contrast enhancement processing is performed independently for each image block. This localized approach enables selective contrast adjustment for different image regions, thereby reducing the risk of overall distortion of grayscale ultrasound images of thyroid follicles. To further reduce the risk of overall data sample distortion and noise, in step 100, contrast enhancement processing based on contrast limiting is performed independently for each image block. By setting contrast limits, excessive enhancement of image block noise and artifacts is prevented during the contrast enhancement process.
[0060] In an example, the contrast limitation may be implemented by using a contrast limitation threshold E, which will be described in detail in subsequent embodiments.
[0061] Step 200: merging the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image and performing edge smoothing to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images.
[0062] In step 200, for each thyroid follicular grayscale ultrasound image, each of the contrast enhancement blocks corresponding to the thyroid follicular grayscale ultrasound image is merged and edge smoothing transition is performed to obtain an enhanced image sample corresponding to the thyroid follicular grayscale ultrasound image. By smoothing the transition between adjacent sub-blocks, the grayscale difference between the edges of adjacent image blocks is eliminated to maintain the overall visual consistency of the data sample.
[0063] Specifically, because the principle of ultrasound imaging is based on the propagation, reflection, and scattering characteristics of high-frequency sound waves, the ultrasound probe acts as both a transmitter and a receiver. The internal signal processing system generates ultrasound images by calculating the depth of the reflection interface and the echo intensity. If the acoustic impedance of two adjacent tissues—that is, the ability of sound waves to pass through a medium—is slightly different, the reflected sound wave signal will be weak, resulting in poor image contrast. Ultrasound waves scatter and absorb when propagating through human tissue, a phenomenon that is particularly pronounced in soft tissue. The scattered reflected signal is distributed chaotically, weakening the echo signal received by the receiver and reducing image contrast. Furthermore, common artifacts in ultrasound images, such as acoustic shadowing, enhancement effects, and mirroring artifacts, can also affect contrast. Therefore, poor ultrasound image contrast is the result of a combination of factors, posing a significant challenge to the accurate identification of lesions.
[0064] Based on this, in order to improve the quality of thyroid ultrasound images before training, the technical means provided by the above steps 100 and 200 in the embodiment of the present application can be referred to as a preprocessing method based on contrast limited adaptive histogram equalization (CLAHE), which can be referred to as CLAHE for short. CLAHE is an advanced adaptive histogram equalization technology that aims to enhance the contrast of local areas of the image while avoiding excessive amplification of noise and artifacts. It maintains visual consistency by applying contrast enhancement in local areas and smoothing the transition between adjacent areas.
[0065] In addition, while acquiring each of the thyroid follicular grayscale ultrasound images or before executing step 300, the thyroid follicular lesion segmentation model training device also receives a label group corresponding to each of the thyroid follicular grayscale ultrasound images. The label group includes a lesion mask label and a classification label for the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; and the classification label includes a benign label and a malignant label. The lesion mask label and classification label can be extracted from the results of a pathological examination of a slice of the thyroid follicular grayscale ultrasound image.
[0066] The label group of the thyroid follicular grayscale ultrasound image is then used as the label group of the enhanced image sample corresponding to the thyroid follicular grayscale ultrasound image to perform the following step 300. It should also be noted that the thyroid follicular grayscale ultrasound images used to train the target Unet model mentioned in this application are all ultrasound images authorized by the patient.
[0067] Step 300: Use each of the enhanced image samples and the label group corresponding to each of the enhanced image samples to train a target Unet model to train the target Unet model into a thyroid follicular lesion segmentation model for outputting the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the label group includes the lesion mask label and the classification label of the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; the classification label includes a benign label and a malignant label.
[0068] In step 300, the target Unet model refers to the Unet model currently serving as the target model framework. The Unet model may adopt a general Unet model architecture, or may be further improved to improve the accuracy and robustness of thyroid ultrasound image segmentation, as will be described in detail in subsequent embodiments.
[0069] The benign label is used to indicate the probability value that the thyroid follicle corresponding to the enhanced image sample is a benign thyroid follicular adenoma; the malignant label is used to indicate the probability value that the thyroid follicular adenoma corresponding to the enhanced image sample is a malignant thyroid follicular carcinoma.
[0070] From the above description, it can be seen that the thyroid follicular lesion segmentation model training method provided in the embodiment of the present application can selectively adjust the contrast for different image areas by independently performing contrast enhancement for each image block, thereby reducing the risk of overall distortion; by performing contrast enhancement processing based on contrast limitation, it can prevent excessive enhancement of noise and artifacts, thereby effectively reducing noise while improving image clarity; by merging and smoothing the edges of the contrast enhanced blocks belonging to the same thyroid follicular grayscale ultrasound image, it can eliminate the grayscale difference between adjacent image blocks to maintain the overall visual consistency of the data sample; thereby, it can improve the application reliability of the data samples used to train the thyroid follicular lesion segmentation model, effectively improve the training effectiveness and reliability of the thyroid follicular lesion segmentation model, and thereby improve the accuracy of the high-sampling thyroid follicular lesion segmentation results.
[0071] In order to further improve the clarity and application reliability of the data samples used to train the thyroid follicular lesion segmentation model, and reduce the risk of overall distortion of the data samples and reduce noise, in a thyroid follicular lesion segmentation model training method provided in an embodiment of the present application, see Figure 2 Step 100 in the thyroid follicular lesion segmentation model training method specifically includes the following contents: Step 110: Divide each thyroid follicular grayscale ultrasound image into a plurality of image blocks; Step 120: Calculating the cumulative distribution value of the histogram corresponding to each of the image blocks respectively; Step 130: If the cumulative distribution value of the histogram corresponding to the image block is greater than the contrast limit threshold, setting the target cumulative distribution value corresponding to the image block as the contrast limit threshold; if the cumulative distribution value of the histogram corresponding to the image block is less than or equal to the contrast limit threshold, setting the target cumulative distribution value corresponding to the image block as the cumulative distribution value of the histogram corresponding to the image block; In step 130, the target cumulative distribution value of the image block can be calculated using the following formula: in, Grayscale ultrasound image of thyroid follicles The target cumulative distribution value of the image patch in ; Indicates the size (dimensions) of the image block; represents the cumulative distribution value of the histogram corresponding to the image block; is the contrast limit threshold; Indicates that if the cumulative distribution value of the histogram corresponding to the image block is less than or equal to the ratio limit threshold ; Indicates that if the cumulative distribution value of the histogram corresponding to the image block is greater than the ratio limit threshold .
[0072] Step 140: performing contrast enhancement processing on each of the image blocks according to the target cumulative distribution value corresponding to each of the image blocks, so as to obtain a contrast enhanced block for each of the image blocks.
[0073] In order to further improve the clarity and application reliability of the data samples used to train the thyroid follicular lesion segmentation model and maintain the overall visual consistency of the data samples, in a thyroid follicular lesion segmentation model training method provided in an embodiment of the present application, see Figure 2 Step 200 in the thyroid follicular lesion segmentation model training method specifically includes the following contents: Step 210: Merge the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image to obtain merged images corresponding to the respective thyroid follicular grayscale ultrasound images.
[0074] Step 220: performing edge smoothing transition processing between the contrast enhancement blocks belonging to the same merged image in a bilinear interpolation manner to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images.
[0075] Specifically, since each image block is processed independently of its grayscale histogram, in order to avoid discontinuity at the boundaries between image blocks, an interpolation method can be used to smooth the transition between adjacent image blocks. The output grayscale value of each pixel is calculated as follows: in, is the position of the input thyroid follicle in the grayscale ultrasound image The pixel value of is the position of the input thyroid follicle in the grayscale ultrasound image The corresponding output pixel value at and Represent the minimum and maximum values of the gray level respectively; Indicates the location of thyroid follicles in grayscale ultrasound images The target cumulative distribution value of the image block corresponding to .
[0076] Subsequently, bilinear interpolation is applied to smooth the transition between image patches.
[0077] In solving the above-mentioned existing thyroid follicular lesion segmentation technology, most algorithms only focus on optimizing the network structure, and insufficiently explore image preprocessing and image enhancement technologies, thereby ignoring the low quality of ultrasound images themselves. In addition, the existing technology also has the following technical problems: Thyroid cancers vary greatly in morphology and size, and some tumors have irregular shapes. In the actual clinical diagnosis of thyroid cancer, physicians must not only identify long-range features such as the anatomical structure of the thyroid gland and surrounding tissues and conduct a comprehensive analysis of the overall morphology of the image, but also analyze short-range features such as smaller lesion details or tissue boundaries. Long-range features can help the model capture global anatomical structures and long-range dependencies, while short-range features play a key role in distinguishing local details and lesion boundaries. However, most models only focus on the extraction of local information and ignore the necessity of coordinated processing of global and local information. This results in ordinary models not being able to segment tumors with complex morphology or lesions with blurred boundaries in an ideal way.
[0078] Based on this, based on the embodiment of the above-mentioned thyroid follicular lesion segmentation model training method, in view of the large differences in the morphology and size of thyroid cancer, and the problem that some tumors have irregular shapes, in a thyroid follicular lesion segmentation model training method provided in the embodiment of the present application, see Figure 3, a cross-fusion feature context Transformer module is provided between the output end of the encoder and the input end of the decoder in the target Unet model in the thyroid follicular lesion segmentation model training method; The cross-fusion feature context Transformer module is used to perform feature aggregation, multi-scale feature extraction, group convolution and feature interaction processing based on a multi-head attention mechanism on the downsampled feature data corresponding to the enhanced image sample output by the encoder in sequence to obtain a target feature map corresponding to the enhanced image sample, and input the target feature map into the decoder so that the decoder outputs the thyroid follicular lesion segmentation result data corresponding to the enhanced image sample according to the target feature map; wherein the thyroid follicular lesion segmentation result data is used to represent the thyroid follicular lesion area segmented from the enhanced image sample and the classification label corresponding to the thyroid follicular lesion area.
[0079] The Transformer module (or model) is a deep learning model architecture primarily designed for sequence-to-sequence tasks, with particular strength in natural language processing (NLP). The core of the Transformer model is its attention mechanism, which transforms input data by calculating the relationships between vectors within a matrix. Its purpose is to remove invalid information and enhance valid information, thereby better solving the final mathematical problem and calculating the optimal solution.
[0080] Among them, the cross-fusion feature contextual Transformer module can be referred to as the CFCT (Cross-fusion Feature Contextual Transformer) module.
[0081] In order to further solve the problems of Unet's ability to capture global context information, ignoring the fusion of global features, and insufficient extraction of fine-grained short-range features, in a thyroid follicular lesion segmentation model training method provided in an embodiment of the present application, see Figure 4 The cross-fusion feature context Transformer module in the thyroid follicular lesion segmentation model training method includes: a feature aggregation layer, a multi-scale feature extraction layer, a grouped convolution layer based on depthwise separable convolution, a Transformer module based on a multi-head attention mechanism, and an output layer; The feature aggregation layer is used to perform initial feature aggregation processing on the downsampled feature data corresponding to the enhanced image samples output by the encoder using a convolution kernel, and perform layer normalization processing on the downsampled feature data after the initial feature aggregation processing to obtain aggregated feature data corresponding to the enhanced image samples; wherein, see Figure 5, the convolution kernel in the feature aggregation layer can adopt a 3 × 3 convolution kernel.
[0082] The multi-scale feature extraction layer is used to perform convolution feature extraction on the aggregated feature data using convolution kernels of different sizes to obtain extracted feature maps of different scales corresponding to the aggregated feature data, and to fuse the extracted feature maps to obtain convolution feature data corresponding to the enhanced image sample; wherein, see Figure 5 , the convolution kernels of different sizes in the multi-scale feature extraction layer can adopt a 1×1 convolution kernel and a 3×3 convolution kernel.
[0083] The grouped convolution layer based on depthwise separable convolution is used to divide the convolution-extracted feature data into feature image blocks of fixed size; flatten each feature image block and map each feature image block into an embedding vector by linear projection; perform depthwise convolution and pointwise convolution on each embedding vector using depthwise separable convolution to obtain an injection feature map of each channel corresponding to the convolution-extracted feature data, and then use a convolution kernel to aggregate the injection feature maps of each channel to obtain a dynamic feature vector corresponding to the enhanced image sample; The Transformer module based on the multi-head attention mechanism is used to perform feature interaction on the dynamic feature vector based on the multi-head attention mechanism to obtain the convolution injection feature data corresponding to the enhanced image sample; wherein, see Figure 5 The dynamic feature vector fed into the Transformer module based on the multi-head attention mechanism consists of Q (Query), K (Key), and V (Value). This is the core component of self-attention, which dynamically calculates the correlation between different positions in the input sequence to achieve contextual modeling. Q represents the query vector, which is used to match the key vectors at other positions; K represents the key vector, which provides semantic information about the query; and V represents the value vector, which carries the actual features being transmitted.
[0084] The output layer is used to fuse the convolution extraction feature data and the convolution injection feature data corresponding to the enhanced image sample to obtain a target feature map corresponding to the enhanced image sample.
[0085] In addressing the above-mentioned existing thyroid follicular lesion segmentation technologies, most algorithms focus only on optimizing the network structure, with insufficient exploration of image preprocessing and image enhancement technologies, thus ignoring the low quality of ultrasound images themselves. In addition, most models focus only on extracting local information, ignoring the necessity of collaborative processing of global and local information. This results in the common models not being able to segment tumors with complex morphology or lesions with blurred boundaries in an ideal manner. Furthermore, existing technologies also have the following technical problems: Thyroid ultrasound images typically include not only the thyroid gland and the tumor itself, but also many irrelevant anatomical structures, such as muscle tissue, trachea, arteries, and veins, which are unhelpful or even disruptive to diagnosis. The presence of these irrelevant elements increases the complexity of the algorithm in extracting effective information related to the lesion. For example, the high-contrast echoes of the trachea and blood vessels may mask the signal returned by the thyroid nodule, leading to misjudgment of the model or the inability to accurately capture the boundary of the lesion. Existing algorithms lack a focus on the target area when extracting features. This not only reduces the algorithm's segmentation accuracy but can also lead to unstable performance of the model in real clinical scenarios.
[0086] Based on this, based on the embodiment of the thyroid follicle lesion segmentation model training method described above, thyroid ultrasound images contain many irrelevant anatomical structures, such as muscle tissue, trachea, arteries and veins. This information is not helpful for diagnosis and even interferes with the irrelevant content, which increases the complexity and effectiveness of the algorithm in extracting effective information related to the lesion. In the thyroid follicle lesion segmentation model training method provided in the embodiment of the present application, see Figure 6 , the encoder in the target Unet model in the thyroid follicular lesion segmentation model training method is provided with a regional attention module; The regional attention module is used to extract features from the input data using a spatial attention mechanism and a channel attention mechanism in sequence to obtain regional attention feature data corresponding to the enhanced image sample; and input the regional attention feature data into a downsampling layer in the encoder described in itself, so that the downsampling layer outputs downsampled feature data corresponding to the regional attention feature data; Among them, if the encoder where the regional attention module is located is the first encoder in the target Unet model, the input data is the enhanced image sample; if the encoder where the regional attention module is located is not the first encoder in the target Unet model, the input data is the downsampled feature data output by another encoder connected to the input side of the encoder.
[0087] The target Unet model contains four encoders and four decoders, with each decoder corresponding to each encoder. Each encoder includes a regional attention module and a downsampling layer, and each decoder includes an upsampling layer. The input (output) and the four corresponding decoder and encoder groups form five stages, with the number of channels in each stage being 16, 32, 64, 128, and 160, respectively, from top to bottom.
[0088] Among them, the regional attention module can be abbreviated as RA (Regional Attention) module.
[0089] In order to accurately extract features related to the target area while effectively suppressing information from irrelevant areas, in a thyroid follicular lesion segmentation model training method provided in an embodiment of the present application, see Figure 7 , the regional attention module in the thyroid follicular lesion segmentation model training method includes: a spatial attention layer and a channel attention layer; The spatial attention layer is used to obtain a spatial attention map corresponding to the input data using sequentially connected convolutional layers and ReLU activation functions, and multiply the spatial attention map by the input data element-by-element to obtain spatial attention feature data corresponding to the enhanced image sample; The channel attention layer is used to use a global pooling layer, a one-dimensional convolution layer and a ReLU activation function connected in sequence to obtain the spatial and channel attention maps corresponding to the spatial attention feature data, and multiply the spatial and channel attention maps with the input data element by element to obtain the regional attention feature data corresponding to the enhanced image sample.
[0090] Based on the above embodiment of the thyroid follicular lesion segmentation model training method, the present application also provides an embodiment of a thyroid follicular lesion segmentation method, see Figure 8 The thyroid follicular lesion segmentation method specifically includes the following contents: Step 400: performing contrast enhancement processing based on contrast limitation on each image block corresponding to the target thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block.
[0091] Step 500: Merge the contrast enhancement blocks and perform edge smoothing to obtain an enhanced image sample corresponding to the target thyroid follicular grayscale ultrasound image.
[0092] Step 600: Input the enhanced image sample into a thyroid follicular lesion segmentation model, so that the thyroid follicular lesion segmentation model outputs a thyroid follicular lesion region segmented from the enhanced image sample and a classification label of the thyroid follicular lesion region; wherein the thyroid follicular lesion segmentation model is pre-trained based on the thyroid follicular lesion segmentation model training method.
[0093] In the thyroid follicular lesion segmentation method, the thyroid follicular lesion segmentation model training method adopted in step 600 can be implemented based on the thyroid follicular lesion segmentation model training method mentioned in the above embodiment. The specific steps refer to the thyroid follicular lesion segmentation model training method mentioned in the above embodiment and are not repeated here.
[0094] From the above description, it can be seen that the thyroid follicular lesion segmentation method provided in the embodiment of the present application can improve the accuracy and application reliability of the high-sampling thyroid follicular lesion segmentation results, and provide a more accurate and effective auxiliary diagnosis data basis for clinical practice.
[0095] In order to further illustrate the above-mentioned thyroid follicular lesion segmentation model training method and the embodiment of the thyroid follicular lesion segmentation method, this application also provides a specific application example of the thyroid follicular lesion segmentation model training method, that is, a thyroid follicular lesion segmentation algorithm based on Transformer, by deploying an integrated image enhancement and segmentation pipeline, and designing a segmentation model based on the Unet model, Transformer model, channel attention and spatial attention to solve the above problems, it can effectively improve the quality of thyroid follicular lesion segmentation and achieve the purpose of assisting doctors in diagnosis and treatment. The application examples of this application are specifically described as follows: (1) Image preprocessing method based on CLAHE For the CLAHE algorithm, it mainly includes three steps: (1) Image Blocking: The CLAHE algorithm divides the ultrasound image into smaller blocks or grid regions. For example, the ultrasound image can be divided into a total of 8×8=64 sub-blocks, called windows, and contrast enhancement is performed independently within each window. This localized approach enables selective contrast adjustment of different image regions, thereby reducing the risk of overall distortion.
[0096] (2) Contrast Limitation: Within each image block, the CLAHE algorithm limits the magnitude of contrast enhancement. By setting a limiting threshold, it constrains the values of certain grayscale histograms to prevent excessive enhancement of noise and artifacts. Excess histogram values exceeding the threshold are redistributed to other grayscale levels, thereby effectively reducing noise while improving image clarity.
[0097] (3) Grayscale mapping: After applying contrast-limited histogram equalization to each image block, CLAHE merges the results into an overall image and eliminates the grayscale differences between image blocks and image block edges by smoothing the transitions between adjacent image blocks.
[0098] Specifically, if we now have a grayscale ultrasound image of a thyroid follicle ,set up Grayscale The number of occurrences, is the total number of pixels. Then the gray level The probability of occurrence of thyroid follicles in grayscale ultrasound images It can be expressed as formula (3-1): (3-1) The Cumulative Distribution Function (CDF) of a histogram is a function of the grayscale level of an image, describing the number of pixels in the image that are not greater than that grayscale level. The definition of is given by formula (3-2): (3-2) In the CLAHE algorithm, each image is divided into During the histogram equalization process of each image block, CLAHE limits the contrast limit threshold of the cumulative distribution function , by preventing over-enhancement of high contrast areas, the risk of amplifying noise is reduced, where E = 40. The target cumulative distribution value for each image block is It can be calculated by formula (3-3): (3-3) Since each image block in CLAHE processes its grayscale histogram independently, in order to avoid discontinuity at the boundaries between image blocks, CLAHE uses interpolation to smooth the transition between adjacent image blocks. The output grayscale value of each pixel is calculated as shown in formula (3-4). is the position of the input thyroid follicle in the grayscale ultrasound image The pixel value of is the position of the input thyroid follicle in the grayscale ultrasound image The corresponding output pixel value at and Represent the minimum and maximum values of the gray level respectively; Indicates the location of thyroid follicles in grayscale ultrasound images The target cumulative distribution value of the image block corresponding to . Then, bilinear interpolation is applied to smooth the boundary transition between image blocks.
[0099] (3-4) Figure 9(a) shows a comparison of the original and CLAHE-preprocessed images of follicular thyroid carcinoma (FTC); Figure 9(b) shows a comparison of the original and CLAHE-preprocessed images of follicular thyroid adenoma (FA). Compared with the original FA and FTC images, the method of this application enhances local details and reduces the effects of noise. Furthermore, this method improves the problem of uneven brightness without significantly increasing computational complexity. Therefore, this method is suitable for real-time diagnostic applications, providing clearer visibility of subtle structures, thereby improving diagnostic quality in ultrasound images with poor visibility.
[0100] (II) Thyroid follicular carcinoma segmentation lesion network 1. Network Architecture Designing thyroid tumor-based segmentation models has become a hot research topic. Initially designed to address cellular-level segmentation tasks, UNet is now a widely used benchmark model for medical image segmentation. Despite its widespread adoption, UNet suffers from shortcomings in modeling long-range dependencies, limiting its ability to fully capture global context. Because skip connections in UNet directly transmit local features, global feature fusion is often overlooked. Furthermore, high-resolution spatial and channel information is partially lost during multiple downsampling steps, hindering the extraction of fine-grained features.
[0101] To address the above limitations, the application example of this application proposes integrating a novel cross-fusion feature contextual transformer (CFCT) module and a regional attention (RA) module into the two-dimensional Unet model. The introduction of these two modules is to solve the aforementioned problems of long- and short-distance fusion and target area focusing.
[0102] The architecture of the target UNet model is as follows Figure 6 The encoder of the baseline 2D UNet model consists of five stages, with the number of channels in stages 1 to 5 being 16, 32, 64, 128, and 160, respectively. The size of both the input and output ultrasound images is set to 224×224.
[0103] (1) Cross-fusion feature context Tranformer module Because thyroid grayscale hyperimages require attention not only to long-range anatomical features but also to fine-grained, short-range features like glandular texture, this application proposes CFCT to address these issues, given that UNet's ability to capture global contextual information overlaps, neglects the fusion of global features, and fails to extract fine-grained, short-range features.
[0104] The architecture of the CFCT module is as follows Figure 5 As shown in Figure 2, the CFCT module begins with initial feature aggregation. The input feature map is first passed through a 3×3 convolution kernel for initial feature aggregation. This step not only extracts surface features but also smoothes the input features, effectively suppressing noise and irrelevant details. Layer normalization is then applied to normalize the feature distribution, stabilizing the training process and enhancing the model's generalization capabilities. The layer normalization process is shown in Table 1.
[0105] Table 1 Multi-scale feature extraction is then performed. This is achieved by combining 1×1 and 3×3 convolutional layers. The 1×1 convolutional layer focuses on extracting fine-grained local features while reducing channel redundancy. The 3×3 convolutional layer, on the other hand, leverages its larger receptive field to capture richer contextual information, complementing local operations to model long-range dependencies. This multi-scale strategy simultaneously extracts both short-range and long-range features, effectively fusing local details with global context to generate a comprehensive feature representation.
[0106] Later, the application example of this application introduces the convolution-injected operation. At this stage, the input image is divided into feature image blocks of fixed size, which are then flattened and mapped into embedding vectors through linear projection. However, this operation destroys the spatial relationship in the image, which may lead to problems such as the loss of local details and poor multi-scale feature modeling. To alleviate these problems, the application example of this application introduces depth-wise separable convolutions in the Transformer module. The depth-wise separable convolution consists of depth-wise convolution and point-wise convolution. The role of depth-wise convolution is to extract spatial features, while the role of point-wise convolution is to extract channel features. It can be understood that the depth-wise separable convolution groups the convolutions in the feature dimension, performs independent depth-wise convolution on each channel, and aggregates all channels using a 3×3 convolution before output.
[0107] The advantage of using depthwise separable convolution is that it naturally embeds position information through local computation in the spatial dimension, thereby eliminating the need for explicit position encoding in the original Transformer. Moreover, using depthwise separable convolution instead of the original linear mapping allows the input features to interact directly with the spatial context information of the image through the convolution operation, achieving efficient modeling of local and global features. The scalability of the depthwise separable convolution operation also allows the application examples of this application to flexibly adapt to feature extraction requirements of different scales by adjusting the convolution kernel size.
[0108] If the information is extracted after the Unet encoder, the obtained encoding feature data set is , if the operation of the entire CFCT module can be described as , then the output set of the CFCT module designed in this application example is If here, the input enhanced image sample is , the target feature map corresponding to the enhanced image sample output after processing is , respectively and Representing the processing of the two parts of convolution extraction and convolution injection Transformer, the overall output can be expressed as formula (3-5): (3-5) in, express; Indicates that a CFCT operation is performed on the enhanced image sample; Representing the convolution-extracted feature data corresponding to the enhanced image sample; Represents the convolution injection feature data.
[0109] In summary, the design of CFCT ensures that the scale of the input and output feature maps remains unchanged, thus enabling in-depth feature re-extraction and multi-level feature fusion. By dynamically capturing global context and local details, the CFCT module significantly enhances the ability to model long-range dependencies and integrate complex features, improving its performance in various medical imaging tasks.
[0110] (2) Regional Attention Mechanism Since thyroid ultrasound images contain many irrelevant anatomical structures, such as muscle tissue, trachea, arteries, and veins, these irrelevant contents are not helpful for diagnosis and even interfere with the diagnosis, which increases the complexity and effectiveness of the algorithm in extracting effective information related to the lesion. In order to solve this problem, this application example proposes a regional attention (RA) module designed specifically for ultrasound image segmentation. The architecture of the regional attention (RA) module is as follows: Figure 7As shown. In the design of the application example of this application, the spatial attention mechanism (SAM) can dynamically highlight the areas in the image corresponding to follicular carcinoma and adenoma. By enhancing the features of these spatial regions, the spatial attention mechanism can reduce the interference of irrelevant structures and enhance the spatial positioning of the target. The channel attention mechanism (CAM) assigns different weights to feature channels, emphasizing features related to the target region, while also suppressing redundant features caused by irrelevant structures. This ensures that the model focuses on diagnostically valuable information and improves segmentation accuracy. The combination of spatial attention and channel attention makes it possible to accurately extract features related to the target region while effectively suppressing information in irrelevant regions.
[0111] The spatial attention mechanism reduces the dimensionality of the input feature map to 1 through convolution operations, thereby achieving spatial weighting. This operation generates a spatial attention map that can be used to highlight relevant areas such as lesions while effectively suppressing interference from anatomical structures such as the trachea and arteries. In this way, the spatial attention module helps the model focus on more important areas and reduces the influence of irrelevant structures on segmentation. This is particularly true in thyroid ultrasound images, where there are multiple complex anatomical structures. This spatial weighting mechanism of spatial attention can effectively improve the spatial positioning accuracy of the target area, thereby enhancing the segmentation effect.
[0112] The channel attention mechanism adopts the lightweight architecture of the network model ECANe for learning context-dependent vector representations of nodes, but improves it by replacing the fully connected layer with a 1×1 convolution operation. The main purpose of this design is to improve computational efficiency by reducing the number of model parameters while retaining sufficient expressive power of the model. The channel attention module calculates the weight of each channel through a global average pooling operation and adjusts the channels of the input feature map according to the channel weight. In this way, channel attention can not only enhance features related to the target area, but also suppress redundant features unrelated to the target. This mechanism enables the model to focus on key diagnostic information and improve segmentation accuracy, which is particularly important when dealing with complex thyroid structures.
[0113] In addition, the application example of the present application also adds a point-by-point jump connection to the output part of each attention module. This design allows the features of different attention modules to be directly fused, which helps to maintain multi-scale information and avoid the loss of detail information during the feature extraction process. Especially when dealing with multi-scale thyroid regions, the application example of the present application can effectively avoid over-simplification or loss of details, thereby improving the robustness of the model. The final output is processed by the ReLU activation function. The use of ReLU as an attention gating mechanism can further enhance the nonlinear representation of features, not only helping the model to add nonlinear relationships, but also filtering out useless feature information, so that the model can better adapt to complex medical images in the case of various pathological morphologies and complex structures. Through the combination of the above series of mechanisms, spatial attention and channel attention work together to significantly improve the accuracy and robustness of thyroid ultrasound image segmentation.
[0114] (3) Dataset construction The dataset used in this application example contains grayscale ultrasound images of the thyroid gland collected from actual clinical practice, covering two different types of thyroid tumors: benign follicular adenoma (FA) and malignant follicular thyroid carcinoma (FTC). Each case was pathologically confirmed through subsequent histopathological examination, further enhancing the reliability of the dataset. The overall dataset construction process is as follows: Figure 10 shown.
[0115] After that, the collected thyroid ultrasound images need to be enhanced and masked. During the image enhancement process, all images are enhanced by the introduced CLAHE algorithm, and further data enhancement is performed by random inversion and random cropping. During the experiment, a single NVIDIA RTX A6000 graphics card is used, which is equipped with 48GB of video memory. The high-performance GPU provides sufficient computing power support for intensive training and testing tasks. The entire experiment is based on Pytorch, an open source deep learning framework for machine learning and deep learning. In terms of training configuration, the batch size is set to 16, which strikes a balance between video memory usage and training efficiency. The initial learning rate that controls the step size of the model weight update is set to 10 (-4), a smaller initial value ensures stable convergence of the training process. The model was trained for 300 complete cycles, and cross entropy was selected as the loss function. The cross entropy function shows significant advantages in classification tasks by calculating the difference between the predicted probability distribution and the true label. In terms of parameter optimization, the Adam optimizer widely used in the field of deep learning is adopted. This algorithm combines the advantages of the adaptive gradient algorithm AdaGrad and the adaptive learning rate optimization algorithm RMSProp, and can achieve adaptive and efficient updates of model weights. In order to further improve the training effect, the application example of this application also introduces a cosine annealing learning rate decay strategy: as the training process progresses, the learning rate is dynamically adjusted according to the cosine curve. Such a design can not only achieve more precise parameter fine-tuning in the later stage of training, but also effectively suppress overfitting during training.
[0116] The comparison of the effects of the thyroid follicular lesion segmentation model used in the application example of this application and other models is shown in Figure 11 (a) and Figure 11 (b). Figure 11 (a) includes the original image of thyroid follicular cancer (FTC Original), label (Ground Truth), U-shaped convolutional network (Unet), U-shaped convolutional network based on attention gating mechanism (AttentionUnet), U-shaped convolutional network ++ (Unet++), U-shaped convolutional network based on converter architecture (TransUnet), U-shaped convolutional network + contrast limited adaptive histogram equalization (Unet+CLAHE), U-shaped convolutional network + cross-fusion feature context Transformer module (Unet+CFCT), U-shaped convolutional network + regional attention (Unet+RA) and the segmentation model provided in this application. Figure 11 (b) includes the original image of thyroid follicular adenoma (FA Original), label (GroundTruth), U-shaped convolutional network (Unet), U-shaped convolutional network based on attention gating mechanism (AttentionUnet), U-shaped convolutional network ++ (Unet++), U-shaped convolutional network based on converter architecture (TransUnet), U-shaped convolutional network + contrast limited adaptive histogram equalization (Unet+CLAHE), U-shaped convolutional network + cross fusion feature context Transformer module (Unet+CFCT), U-shaped convolutional network + regional attention (Unet+RA) and the segmentation model provided in this application.
[0117] From the software level, the present application also provides a thyroid follicular lesion segmentation model training device for executing all or part of the thyroid follicular lesion segmentation model training method, see Figure 12 The thyroid follicular lesion segmentation model training device specifically includes the following contents: The first contrast enhancement module 10 is configured to perform contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image, so as to obtain a contrast enhanced block for each image block.
[0118] The first smooth transition module 20 is configured to merge the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image and perform edge smoothing to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images.
[0119] The model training module 30 is used to train a target Unet model using each of the enhanced image samples and the label group corresponding to each of the enhanced image samples, so as to train the target Unet model into a thyroid follicular lesion segmentation model for outputting the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the label group includes the lesion mask label and the classification label of the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; and the classification label includes a benign label and a malignant label.
[0120] The embodiment of the thyroid follicular lesion segmentation model training device provided in this application can be specifically used to execute the processing flow of the embodiment of the thyroid follicular lesion segmentation model training method in the above-mentioned embodiment. Its functions will not be repeated here, and reference can be made to the detailed description of the above-mentioned thyroid follicular lesion segmentation model training method embodiment.
[0121] The portion of the thyroid follicular lesion segmentation model training performed by the thyroid follicular lesion segmentation model training device can be completed in the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor for the specific processing of the thyroid follicular lesion segmentation model training.
[0122] The aforementioned client device may include a communication module (i.e., a communication unit) capable of establishing a communication connection with a remote server to facilitate data transmission with the server. The server may include a server at the task scheduling center or, in other implementation scenarios, a server on an intermediate platform, such as a server on a third-party server platform that is communicatively linked to the task scheduling center server. The server may comprise a single computer device, a server cluster consisting of multiple servers, or a distributed server configuration.
[0123] The server and the client device may communicate using any suitable network protocol, including network protocols that have not yet been developed as of the filing date of this application. Examples of such network protocols include TCP / IP, UDP / IP, HTTP, and HTTPS. Furthermore, examples of such network protocols include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols, which are used on top of the aforementioned protocols.
[0124] From the above description, it can be seen that the thyroid follicular lesion segmentation model training device provided in the embodiment of the present application can selectively adjust the contrast for different image areas by independently performing contrast enhancement for each image block, thereby reducing the risk of overall distortion; by performing contrast enhancement processing based on contrast limitation, it can prevent excessive enhancement of noise and artifacts, thereby effectively reducing noise while improving image clarity; by merging and smoothing the edges of the various contrast enhanced blocks belonging to the same thyroid follicular adenoma grayscale ultrasound image, it can eliminate the grayscale differences between adjacent image blocks to maintain the overall visual consistency of the data sample; thereby, it can improve the application reliability of the data samples used to train the thyroid follicular cancer lesion segmentation model, effectively improve the training effectiveness and reliability of the thyroid follicular cancer lesion segmentation model, and thereby improve the accuracy of the sampled thyroid follicular cancer lesion segmentation results.
[0125] From the software level, the present application also provides a thyroid follicular lesion segmentation device for executing all or part of the thyroid follicular lesion segmentation method, see Figure 13 The thyroid follicular lesion segmentation device specifically includes the following contents: a second contrast enhancement module 40 for performing contrast enhancement processing based on contrast limitation on each image block corresponding to the target thyroid follicular grayscale ultrasound image, so as to obtain a contrast enhanced block for each image block; A second smooth transition module 50 is configured to merge the contrast enhancement blocks and perform edge smooth transition processing to obtain an enhanced image sample corresponding to the target thyroid follicular grayscale ultrasound image; The model segmentation module 60 is used to input the enhanced image sample into the thyroid follicular lesion segmentation model so that the thyroid follicular lesion segmentation model outputs the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the thyroid follicular lesion segmentation model is pre-trained based on the thyroid follicular lesion segmentation model training method described in the first aspect above.
[0126] The embodiment of the thyroid follicular lesion segmentation device provided in this application can be specifically used to execute the processing flow of the embodiment of the thyroid follicular lesion segmentation method in the above-mentioned embodiment. Its functions will not be described in detail here, and reference can be made to the detailed description of the above-mentioned thyroid follicular lesion segmentation method embodiment.
[0127] An embodiment of the present application further provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is configured to execute the thyroid follicular lesion segmentation model training method described in the above embodiment, and / or the thyroid follicular lesion segmentation method provided in the second aspect. The processor and the memory may be connected via a bus or other means, with bus connection being an example. The receiver may be connected to the processor and the memory via a wired or wireless manner.
[0128] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0129] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs, and modules, such as the program instructions / modules corresponding to the thyroid follicular lesion segmentation model training method in the embodiments of the present application and / or the thyroid follicular lesion segmentation method provided in the second aspect. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, implementing the thyroid follicular lesion segmentation model training method in the above-mentioned method embodiment and / or the thyroid follicular lesion segmentation method provided in the second aspect.
[0130] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0131] The one or more modules are stored in the memory, and when executed by the processor, perform the thyroid follicular lesion segmentation model training method in the embodiment and / or the thyroid follicular lesion segmentation method provided in the second aspect.
[0132] In some embodiments of the present application, the user equipment may include a processor, a memory and a transceiver unit, and the transceiver unit may include a receiver and a transmitter. The processor, memory, receiver and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.
[0133] As an implementation method, the functions of the receiver and transmitter in this application can be considered to be implemented through a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general-purpose chip.
[0134] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program code for implementing the functions of the processor, receiver, and transmitter is stored in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.
[0135] The present application also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the steps of the aforementioned thyroid follicular lesion segmentation model training method and / or the thyroid follicular lesion segmentation method provided in the second aspect. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0136] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the aforementioned thyroid follicular lesion segmentation model training method and / or the thyroid follicular lesion segmentation method provided in the aforementioned second aspect.
[0137] It should be understood by those skilled in the art that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether it is implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted on a transmission medium or communication link via a data signal carried in a carrier.
[0138] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0139] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0140] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art will appreciate that various modifications and variations of the present embodiment are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A thyroid follicular lesion segmentation model training method, characterized in that: include: performing contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image, so as to obtain a contrast enhanced block for each image block; Merging the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image and performing edge smoothing to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images; A target Unet model is trained using each of the enhanced image samples and the label group corresponding to each of the enhanced image samples, so as to train the target Unet model into a thyroid follicular lesion segmentation model for outputting the thyroid follicular lesion area segmented from the enhanced image sample and the classification label of the thyroid follicular lesion area; wherein the label group includes the lesion mask label and the classification label of the thyroid follicular grayscale ultrasound image corresponding to the enhanced image sample; the lesion mask label is used to display the thyroid follicular lesion area; and the classification label includes a benign label and a malignant label.
2. The thyroid follicular lesion segmentation model training method according to claim 1, characterized in that: The step of performing contrast enhancement processing based on contrast limitation on each image block corresponding to each thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block includes: Divide each thyroid follicular grayscale ultrasound image into multiple image blocks; Calculating the cumulative distribution value of the histogram corresponding to each of the image blocks respectively; If the cumulative distribution value of the histogram corresponding to the image block is greater than a contrast limit threshold, the target cumulative distribution value corresponding to the image block is set as the contrast limit threshold; if the cumulative distribution value of the histogram corresponding to the image block is less than or equal to the contrast limit threshold, the target cumulative distribution value corresponding to the image block is set as the cumulative distribution value of the histogram corresponding to the image block; Contrast enhancement processing is performed on each of the image blocks according to the target cumulative distribution value corresponding to each of the image blocks, so as to obtain a contrast enhanced block for each of the image blocks.
3. The thyroid follicular lesion segmentation model training method according to claim 1, characterized in that: The merging and edge smoothing processing of the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images includes: Merging the contrast enhancement blocks belonging to the same thyroid follicular grayscale ultrasound image to obtain merged images corresponding to the respective thyroid follicular grayscale ultrasound images; In a bilinear interpolation manner, edge smoothing transition processing is performed between the contrast enhancement blocks belonging to the same merged image to obtain enhanced image samples corresponding to the respective thyroid follicular grayscale ultrasound images.
4. The thyroid follicular lesion segmentation model training method according to claim 1, characterized in that: A cross-fusion feature context Transformer module is provided between the output end of the encoder and the input end of the decoder in the target Unet model; The cross-fusion feature context Transformer module is used to perform feature aggregation, multi-scale feature extraction, group convolution and feature interaction processing based on a multi-head attention mechanism on the downsampled feature data corresponding to the enhanced image sample output by the encoder in sequence to obtain a target feature map corresponding to the enhanced image sample, and input the target feature map into the decoder so that the decoder outputs the thyroid follicular lesion segmentation result data corresponding to the enhanced image sample according to the target feature map; wherein the thyroid follicular lesion segmentation result data is used to represent the thyroid follicular lesion area segmented from the enhanced image sample and the classification label corresponding to the thyroid follicular lesion area.
5. The thyroid follicular lesion segmentation model training method according to claim 4, characterized in that: The cross-fusion feature context Transformer module includes: a feature aggregation layer, a multi-scale feature extraction layer, a grouped convolution layer based on depthwise separable convolution, a Transformer module based on a multi-head attention mechanism, and an output layer; The feature aggregation layer is used to perform initial feature aggregation processing on the downsampled feature data corresponding to the enhanced image samples output by the encoder using a convolution kernel, and perform layer normalization processing on the downsampled feature data after the initial feature aggregation processing to obtain aggregated feature data corresponding to the enhanced image samples; The multi-scale feature extraction layer is used to perform convolution feature extraction on the aggregated feature data using convolution kernels of different sizes to obtain extracted feature maps of different scales corresponding to the aggregated feature data, and to fuse the extracted feature maps to obtain convolution feature data corresponding to the enhanced image sample; The grouped convolution layer based on depthwise separable convolution is used to divide the convolution-extracted feature data into feature image blocks of fixed size; flatten each feature image block and map each feature image block into an embedding vector by linear projection; perform depthwise convolution and pointwise convolution on each embedding vector using depthwise separable convolution to obtain an injection feature map of each channel corresponding to the convolution-extracted feature data, and then use a convolution kernel to aggregate the injection feature maps of each channel to obtain a dynamic feature vector corresponding to the enhanced image sample; The Transformer module based on the multi-head attention mechanism is used to perform feature interaction on the dynamic feature vector based on the multi-head attention mechanism to obtain convolution injection feature data corresponding to the enhanced image sample; The output layer is used to fuse the convolution extraction feature data and the convolution injection feature data corresponding to the enhanced image sample to obtain a target feature map corresponding to the enhanced image sample.
6. The thyroid follicular lesion segmentation model training method according to claim 1, characterized in that: The encoder in the target Unet model is provided with a regional attention module; The regional attention module is used to extract features from the input data using a spatial attention mechanism and a channel attention mechanism in sequence to obtain regional attention feature data corresponding to the enhanced image sample; And input the regional attention feature data into the downsampling layer of the encoder itself, so that the downsampling layer outputs the downsampling feature data corresponding to the regional attention feature data; Among them, if the encoder where the regional attention module is located is the first encoder in the target Unet model, the input data is the enhanced image sample; if the encoder where the regional attention module is located is not the first encoder in the target Unet model, the input data is the downsampled feature data output by another encoder connected to the input side of the encoder.
7. The thyroid follicular lesion segmentation model training method according to claim 6, characterized in that: The regional attention module includes: a spatial attention layer and a channel attention layer; The spatial attention layer is used to obtain a spatial attention map corresponding to the input data using sequentially connected convolutional layers and ReLU activation functions, and multiply the spatial attention map by the input data element-by-element to obtain spatial attention feature data corresponding to the enhanced image sample; The channel attention layer is used to use a global pooling layer, a one-dimensional convolution layer and a ReLU activation function connected in sequence to obtain the spatial and channel attention maps corresponding to the spatial attention feature data, and multiply the spatial and channel attention maps with the input data element by element to obtain the regional attention feature data corresponding to the enhanced image sample.
8. A method for segmenting thyroid follicular lesions, characterized in that: include: performing contrast enhancement processing based on contrast limitation on each image block corresponding to the target thyroid follicular grayscale ultrasound image to obtain a contrast enhanced block for each image block; Merging the contrast enhancement blocks and performing edge smoothing to obtain an enhanced image sample corresponding to the target thyroid follicular grayscale ultrasound image; The enhanced image sample is input into a thyroid follicular lesion segmentation model so that the thyroid follicular lesion segmentation model outputs a thyroid follicular lesion area segmented from the enhanced image sample and a classification label of the thyroid follicular lesion area; wherein the thyroid follicular lesion segmentation model is pre-trained based on the thyroid follicular lesion segmentation model training method according to any one of claims 1 to 7.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the thyroid follicular lesion segmentation model training method according to any one of claims 1 to 7, and / or implements the thyroid follicular lesion segmentation method according to claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the thyroid follicular lesion segmentation model training method according to any one of claims 1 to 7, and / or implements the thyroid follicular lesion segmentation method according to claim 8.
Citation Information
Patent Citations
Bone tumor focus positioning method and device based on precise recognition
CN117218200A
Thyroid follicle classification model training device, classification device and equipment
CN118094362A
Method for automatically judging benign and malignant thyroid superficial nodules based on bimodal ultrasound
CN118864999A
Cancer image contrast enhancement method and system based on deep learning
CN119599925A
Method for assessing the risk of recurrent differentiated thyroid cancer after radioiodine therapy
RU2743275C1