Thyroid ultrasound image segmentation method and device and readable storage medium thereof
By introducing CMFEM and CFIFM modules into the TransUNet network, the bottleneck layer feature expression and the decoder cross-layer feature fusion are optimized, and the problems of boundary blurring and noise interference in thyroid ultrasound image segmentation are solved, achieving higher segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202510918834.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The prior art has problems such as blurred boundary, strong noise interference, low segmentation accuracy and poor robustness in thyroid ultrasound image segmentation, especially the lack of bottleneck layer feature expression capabilities of deep learning models such as TransUNet and insufficient fusion of cross-layer feature of decoder.
The collaborative multipath feature enhancement module (CMFEM) and cross-layer feature interaction fusion module (CFIFM) are introduced into the TransUNet network, and the bottleneck layer feature expression is optimized through dual-dimensional attention coupling, group feature interaction and dynamic fusion gate, and the deep fusion of high and low-level features is achieved through multi-domain feature perception and spatial interaction weights.
The accuracy and robustness of thyroid ultrasound image segmentation are significantly improved, the Dice coefficient is increased by 1.49% to 2.26%, and the 95% Hausdorff distance is reduced by 16.14 to 25.68 pixels, enhancing noise immunity and stability, adapting to different samples and complex scenarios.
Smart Images

Figure CN120411528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing technology, and in particular to an automated segmentation method and system for thyroid ultrasound images, which achieves accurate segmentation of thyroid nodules through deep learning network optimization. Background Art
[0002] Early and accurate diagnosis of thyroid nodules is crucial for improving survival and effectively managing the disease. Ultrasound imaging is the preferred method due to its radiation-free, real-time, and non-invasive nature. However, ultrasound images are often affected by speckle noise and acoustic artifacts, resulting in blurred nodule boundaries and irregular morphology. This makes manual segmentation time-consuming and highly subjective.
[0003] Traditional image segmentation methods (such as those based on contours, regions, or traditional machine learning) have limitations such as sensitivity to image quality and reliance on artificial feature design. Although deep learning models (such as U-Net and TransUNet) improve performance through encoding-decoding structures and Transformer, the bottleneck layer of TransUNet has problems such as insufficient channel feature discrimination, limited spatial detail recovery, and weak noise resistance. In addition, the traditional decoder feature fusion mechanism (simple splicing / addition) has difficulty in mining the correlation between high- and low-level features and has poor noise suppression capabilities, which restricts the accuracy and clinical applicability of thyroid ultrasound image segmentation.
[0004] Therefore, there is an urgent need for a thyroid ultrasound image segmentation method, device and readable storage medium thereof to solve the problems existing in the prior art. Summary of the Invention
[0005] The embodiments of the present invention provide a thyroid ultrasound image segmentation method, device, and readable storage medium thereof. These methods address the problems of traditional image segmentation methods that rely on artificial features and have poor adaptability, and deep learning models (such as TransUNet) that have insufficient bottleneck layer feature expression capabilities and insufficient decoder cross-layer feature fusion, resulting in low segmentation accuracy and poor robustness for nodules with blurred boundaries and noise interference in thyroid ultrasound images.
[0006] The core technology of this invention is mainly achieved by introducing the collaborative multipath feature enhancement module (CMFEM) and the cross-layer feature interaction fusion module (CFIFM) into TransUNet. The former optimizes the bottleneck layer feature expression through two-dimensional attention coupling, group feature interaction and dynamic fusion gate, while the latter realizes the deep fusion of high and low layer features of the decoder through multi-domain feature perception, channel and spatial interaction weights, thereby improving the accuracy and robustness of thyroid ultrasound image segmentation.
[0007] In a first aspect, the present invention provides a method for segmenting a thyroid ultrasound image, the method comprising the following steps: Obtain ultrasound images of the thyroid gland; Input the thyroid ultrasound image into the C2F-TransUNet network, which includes an encoder, a decoder, a collaborative multi-path feature enhancement module set between the encoder and the decoder, and a cross-layer feature interaction and fusion module set at the cross-layer connection position in the decoder; Through the collaborative multi-path feature enhancement module, perform collaborative enhancement in the channel and spatial dimensions and group feature interaction on the bottleneck layer features output by the encoder to generate enhanced bottleneck layer features; Through the cross-layer feature interaction and fusion module, perform multi-domain feature perception and dynamic weight fusion in the channel and spatial dimensions on the high-level and low-level features of the upsampling and downsampling paths in the decoder to generate interaction and fusion features; Based on the enhanced bottleneck layer features and the interaction and fusion features, output the segmentation result of the thyroid nodule through the decoder.
[0008] Furthermore, the collaborative multi-path feature enhancement module includes: A two-dimensional attention coupling module for co-optimizing channel attention and spatial attention on the input features. The channel attention generates a channel attention vector through global average pooling, layer normalization, and 1×1 convolution for dimensionality reduction and upsampling. The spatial attention generates a spatial attention matrix through parallel processing of a 7×7 depthwise separable convolution and a 3×3 dilated convolution with a dilation rate of 3, and the element-wise product of the two forms an adaptive enhancement feature; A grouped feature interaction layer for performing grouped normalization, grouped convolution on the input features after grouping them along the channels, and realizing cross-group feature interaction through a channel shuffle operation; A dynamic feature fusion gate for multiplying the output features of the two-dimensional attention coupling module with the original input features element-wise, then concatenating them with the output features of the grouped feature interaction layer, and generating the final enhanced features after 1×1 convolution, batch normalization, and GELU activation.
[0009] Furthermore, the process of the two-dimensional attention coupling module generating the channel attention vector includes: performing global average pooling on the input feature map to obtain global semantic statistics; After processing through layer normalization and the GELU activation function, perform 1×1 convolution for dimensionality reduction and upsampling to generate the channel attention vector where C is the number of channels.
[0010] Furthermore, the process of the two-dimensional attention coupling module generating the spatial attention matrix includes: Extract local fine structure features through a 7×7 depthwise separable convolution and extract long-range spatial dependence features through a 3×3 dilated convolution with a dilation rate of 3; Perform 1×1 convolution for dimensionality reduction on the two-way features respectively and then add them, and generate the spatial attention matrix after Sigmoid activation where H and W are the spatial height and width.
[0011] Further, the processing process of the grouped feature interaction layer includes: dividing the input features into g subgroups along the channels, and performing grouped normalization on each subgroup; Performing 1×1 grouped convolution on the normalized subgroups respectively to generate intra-group structured features; Recombining the channels of different subgroups through channel shuffling operation to restore to the original feature dimension .
[0012] Further, the cross-layer feature interaction and fusion module includes: A multi-domain feature perception unit that respectively uses 3×3, 5×5, and 7×7 convolutions to extract multi-scale features from the input high-level and low-level features in parallel, and generates multi-domain enhanced features after fusion; A channel interaction regulation unit that performs global average pooling and max pooling on the multi-domain enhanced features in the spatial dimension, and after splicing, generates channel fusion weights through 1×1 convolution and Softmax to dynamically adjust the channel contribution degrees of the high-level and low-level features; A spatial interaction focusing unit that performs global average pooling and max pooling on the multi-domain enhanced features in the channel dimension, and after splicing, generates spatial fusion weights through 2D convolution and Softmax to focus on key spatial regions; A residual fusion unit that adds 1 after adding the channel and spatial fusion weights, multiplies element-wise with the multi-domain enhanced features and then adds them to generate interaction and fusion features.
[0013] Further, the process of the channel interaction regulation unit generating channel fusion weights includes: Performing global average pooling and global max pooling on the multi-domain enhanced features in the spatial dimension to obtain channel statistical features; Splicing the statistical features along the spatial width, reducing the dimension through 1×1 convolution and then normalizing through Softmax to generate channel fusion weights .
[0014] Further, the process of the spatial interaction focusing unit generating spatial fusion weights includes: performing global average pooling and global max pooling on the multi-domain enhanced features in the channel dimension to obtain spatial statistical features; Splicing the statistical features along the channels, processing through 2D convolution and then normalizing through Softmax to generate spatial fusion weights .
[0015] In a second aspect, the present invention provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned thyroid ultrasound image segmentation method.
[0016] In a third aspect, the present invention provides a readable storage medium storing a computer program, the computer program including program codes for controlling a process to execute the process, the process including the thyroid ultrasound image segmentation method described above.
[0017] The main contributions and innovations of the present invention are as follows: 1. Significantly improved segmentation accuracy: On the public dataset DDTI and the self-built dataset Dataset A, the Dice coefficients reach 84.62% and 80.03% respectively, an increase of 1.49% and 2.26% compared to TransUNet; the 95% Hausdorff distance is reduced to 16.14 pixels and 25.68 pixels, and the boundary positioning is more accurate.
[0018] 2. Enhanced noise resistance and stability: The module design effectively suppresses ultrasound noise and artifacts, and the standard deviations of the Dice coefficient and the 95% Hausdorff distance are the lowest among the comparison models, adapting to different samples and complex scenarios.
[0019] 3. Optimized feature expression and fusion: CMFEM enhances the discriminative power of the bottleneck layer channels and spatial details, and CFIFM realizes deep fusion of high-level and low-level features through dynamic weights, solving the problems of feature homogenization and insufficient fusion in traditional models.
[0020] 4. Outstanding clinical applicability: It has strong training stability in small-sample scenarios, and has excellent segmentation effects on nodules with small volumes and fuzzy boundaries, providing quantitative support for the intelligent diagnosis of thyroid nodules.
[0021] Details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flowchart of the thyroid ultrasound image segmentation method according to an embodiment of the present invention; Figure 2 is a schematic diagram of the CMFEM architecture according to an embodiment of the present invention; Figure 3 is a schematic diagram of the CFIFM architecture according to an embodiment of the present invention; Figure 4 is a schematic diagram of the qualitative comparison of C2F-TransUNet and SOTA segmentation models on Dataset A and DDTI datasets according to an embodiment of the present invention; Figure 5Schematic diagram of the C2F-TransUNet architecture according to an embodiment of the present invention; Figure 6 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0023] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0024] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0025] Currently, there are speckle noises and acoustic artifacts in thyroid nodule ultrasound images, resulting in blurred nodule boundaries and irregular shapes. Traditional segmentation methods (such as U-Net, TransUNet) have problems such as insufficient global context modeling, loss of local details, and weak anti-noise ability when processing low-contrast and high-noise images. The bottleneck layer and decoder feature fusion mechanism of the existing TransUNet cannot meet the high-precision requirements for thyroid ultrasound image segmentation.
[0026] Based on this, the present invention is based on the C2F-TransUNet network architecture to solve the problems existing in the prior art.
[0027] Embodiment 1 The present invention aims to propose a thyroid ultrasound image segmentation method and system (C2F-TransUNet) based on deep learning. By introducing a Coupled Multi-path Feature Enhancement Module (CMFEM) and a Cross-layer Feature Interactive Fusion Module (CFIFM) into the traditional TransUNet network, it solves the problem of insufficient segmentation accuracy caused by blurred nodule boundaries and strong noise interference in thyroid ultrasound images. This method optimizes the feature expression of the bottleneck layer and cross-layer feature fusion of the decoder through module design, significantly improving the segmentation accuracy, noise resistance, and clinical applicability.
[0028] Specifically, an embodiment of the present invention provides a thyroid ultrasound image segmentation method, which can refer to Figure 1 , and the method includes the following steps: Step 1: Obtain a thyroid ultrasound image; Step 2: Input the thyroid ultrasound image into the C2F-TransUNet (Coupled&Cross-layer Fusion TransUNet) network, which includes an encoder, a decoder, a coupled multi-path feature enhancement module arranged between the encoder and the decoder, and a cross-layer feature interactive fusion module arranged at the cross-layer connection position in the decoder; In this embodiment, as Figure 2 shown, the coupled multi-path feature enhancement module is innovatively integrated after the output features of the bottleneck layer of the TransUNet network, aiming to deeply optimize and dynamically enhance the high-dimensional features output by the bottleneck layer, so as to significantly improve the expression ability of the features in thyroid ultrasound image segmentation.
[0029] Among them, as Figure 5 shown, the C2F-TransUNet network architecture is as follows: Based on the encoder-decoder architecture, the CMFEM is embedded between the encoder and the decoder, and the CFIFM is integrated at the cross-layer connection position of the decoder. After inputting the thyroid ultrasound image, features are extracted by the encoder, the bottleneck layer features are enhanced by the CMFEM, and the high and low layer features of the decoder are fused by the CFIFM, and finally the segmentation result is output.
[0030] Specifically, Figure 5 the core process in is "encoding-enhancing-decoding-fusing-segmenting":
[0031] 2. Encoder (left half): ResNet50 backbone: First, extract the basic features of the image (such as nodule contours, textures) through ResNet50 to generate multi-scale "Hidden Feature".
[0032] Transformer layer: Further encode the features output by ResNet50, and use multiple layers (n = 12) of Transformer to mine global semantic associations (such as long-range dependencies between nodules and the background), and output high-dimensional features (n patch ×768 dimensions).
[0033] 3. Feature Enhancement (CMFEM Module): For the high-dimensional features output by the Transformer ( ), optimize through CMFEM (Collaborative Multi-Path Feature Enhancement Module): Principle: Integrate channel attention and spatial attention to strengthen key features (such as highlighting nodule regions and suppressing noise artifacts), and then use grouped convolution and channel shuffling to enhance feature diversity, and finally output the enhanced bottleneck layer features ( ).
[0034] 4. Decoder (right half): Upsampling and Cross-Layer Fusion (CFIFM Module): The enhanced features are gradually restored to the spatial resolution through upsampling ( → etc.), and at the same time, fuse the skip connection features (SkipConnection) of different stages of the encoder through CFIFM (Cross-Layer Feature Interaction and Fusion Module): Principle: Use multi-scale convolution to capture local-global information, and combine channel + spatial attention for dynamic weighting (strengthening nodule boundaries and suppressing background interference) to make the high-level and low-level features complement each other deeply.
[0035] 5. Segmentation Head: The final features are processed by convolution (Conv3×3 + ReLU) and upsampling (Upsample) to output the segmentation result (H×W×16 → binary or multi-class segmentation map) to distinguish nodules from the background.
[0036] It can be seen that Figure 5 clearly shows the entire process of "from ultrasound image input, to feature encoding - enhancement - cross-layer fusion, and finally outputting the nodule segmentation result". The core relies on two modules, CMFEM and CFIFM, to break through the traditional segmentation bottleneck and adapt to the needs of medical imaging.
[0037] Step 3: Through the collaborative multi-path feature enhancement module, perform collaborative enhancement in the channel and spatial dimensions and group feature interaction on the bottleneck layer features output by the encoder to generate enhanced bottleneck layer features; In this embodiment, receive the bottleneck layer features output by the encoder-Transformer, that is, the input feature map , where C represents the number of channels, H and W are the spatial height and width respectively. The design of the collaborative multi-path feature enhancement module includes three key modules: the two-dimensional attention coupling module, the grouped feature interaction layer, and the dynamic feature fusion gate: 1. Two-dimensional attention coupling module (Bidirectional Attention Coupling Block, BACB): BACB aims to perform adaptive feature recalibration and enhancement on the input features through the collaborative optimization and deep fusion of the channel dimension and the spatial dimension. As Figure 2 shown, the structure and process of BACB include the following key steps: First, adopt a bottleneck compression-expansion structure for channel dimension adaptive feature enhancement. Use global average pooling (GAP) to aggregate global semantic statistics, and integrate layer normalization (LN) and the GELU activation function to significantly improve the stability of the feature distribution and the non-linear expression ability. Implement the dimension reduction-dimension increase transformation through 1×1 convolution, efficiently compress information redundancy, and dynamically strengthen the response of key semantic channels, and finally dynamically generate the channel attention vector:
[0038] where represents global average pooling, and represent the 1×1 convolutional layers for dimension reduction and dimension increase. represents layer normalization, represents the GELU activation function.
[0039] Secondly, adopt a heterogeneous parallel convolution architecture for heterogeneous perception and fusion in the spatial dimension. Path one uses a 7×7 depthwise separable convolution to efficiently focus on local fine structures; path two uses a 3×3 dilated convolution with a dilation rate of 3 to significantly expand the receptive field to model long-range spatial dependencies. After the outputs of the two paths are integrated through 1×1 convolution for dimension reduction, a spatial attention matrix is generated through the Sigmoid activation function:
[0040] Among them, and are corresponding 1×1 convolution operations; and respectively represent depthwise separable convolution with a convolution kernel size of 7×7 and dilated convolution with a dilation rate of 3 and a convolution kernel size of 3×3, represents the Sigmoid activation function.
[0041] Finally, through the cross-dimensional tensor broadcast product, the channel attention vector and the spatial attention matrix are element-wise multiplied and fused to obtain an adaptively enhanced feature representation:
[0042] Among them, represents the broadcast multiplication. The BACB enhanced feature is obtained.
[0043] Through the above innovative design, BACB realizes deep interaction and collaborative optimization in both the channel and spatial dimensions, significantly enhancing the discriminative response of key semantic channels, suppressing homogeneous expressions, precisely focusing on the spatially significant regions crucial for the segmentation task, and effectively strengthening the robustness of features under complex backgrounds and noise interference. This module significantly improves the expressive ability of bottleneck layer features, laying a solid discriminative feature foundation for subsequent processing.
[0044] 2. Grouped Feature Interaction Layer (GFIL) GFIL aims to promote intra-group consistency and inter-group interactivity of features by fusing group normalization, grouped convolution, and channel shuffle operations. Its processing flow can be described as:
[0045] Among them, represents group normalization, is grouped convolution, represents the channel shuffle operation.
[0046] First, intra-group feature normalization is performed. The input feature map is divided into g independent subgroups along the channel dimension, and group normalization (GN) is applied to the features of each subgroup. This strategy effectively alleviates the statistical bias problem in the mini-batch training scenario, significantly improves the consistency of intra-group feature distributions, and thus enhances the training stability of the model on small-sample medical image datasets.
[0047] Subsequently, perform intra-group structural transformation. Independently perform grouped convolution (Grouped Convolution, ) operations on each normalized feature subgroup, and use 1×1 convolution kernels to focus on efficient structural representation learning of intra-group features. Such a grouped processing mechanism effectively avoids the computational redundancy of traditional full-channel convolution while retaining the prior of the intra-group feature structure, and strengthens the internal correlation of local features.
[0048] Finally, perform cross-group feature interaction and recombination. On the basis of intra-group processing, introduce the channel shuffle operation (Channel Shuffle, ). This operation realizes the full interaction and fusion of cross-subgroup feature information by systematically reorganizing the channels of different subgroups. Specifically, as Figure 2 shown, for the feature representation after grouped convolution, the definition of the channel shuffle operation is as follows:
[0049] This operation rearranges and combines the features of different groups and finally restores them to the original shape .
[0050] Through the above hierarchical processing, GFIL not only ensures the distribution consistency of intra-group features, improves the robustness of the model under small sample conditions, but also greatly promotes cross-group information flow and interaction, and significantly improves the diversity and discriminative ability of feature expression.
[0051] 3. Dynamic Feature Fusion Gate (DFFG) DFFG adaptively fuses the enhanced features output by the BACB module and the structured features output by the GFIL module through a gating mechanism, and its fusion process is expressed as:
[0052] where , represents element-wise multiplication, represents concatenation along the channel dimension, is the corresponding 1×1 convolution, represents batch normalization.
[0053] First, perform element-wise multiplication ( ) on the enhanced features and the original input features to strengthen the key feature regions and enhance the discriminative feature responses of the key regions. Subsequently, concatenate the above enhanced features and the structured features <s along the channel dimension ( ), a comprehensive feature representation containing complementary information is formed. Next, a 1×1 convolutional kernel is used to perform a linear transformation on the concatenated features in the channel dimension, aiming to further adjust the relative contribution weights of the enhanced features and the structured features and achieve the adaptive fusion of the two-way features. Finally, batch normalization (BatchNormalization, ) and the GELU activation function are used for feature normalization and non-linear activation. The finally output optimized enhanced features are obtained to achieve dynamic weighted fusion of features.
[0054] Through the above design, the DFFG mechanism efficiently fuses the complementary information from different paths, and also dynamically optimizes the contribution ratio of the features of each path in the final output through gate-driven weight learning, maximizing the fusion benefit, thereby significantly improving the model's ability to represent complex medical image features and segmentation performance. It can be seen that Figure 2 details the whole process of "how the bottleneck layer features go from input → two-dimensional enhancement → grouped interaction → dynamic fusion output". The core is to use multiple modules to collaboratively solve the problem of insufficient feature expression in traditional networks and adapt to the requirements of thyroid ultrasound image segmentation.
[0055] Step 4: Perform multi-domain feature perception and dynamic weight fusion in the channel and spatial dimensions on the high-level and low-level features of the upsampling and downsampling paths in the decoder through a cross-layer feature interaction fusion module to generate interaction fusion features; In this embodiment, aiming at the core defects such as insufficient feature correlation mining and weak noise suppression ability in the traditional decoder feature fusion method, the present invention proposes a cross-layer feature interactive fusion module (Cross-layer Feature Interactive Fusion Module, CFIFM) and integrates it into the key bridging position connecting the upsampling path and the downsampling path in the TransUNet network, aiming to solve the problem of insufficient cross-layer feature fusion in the encoder-decoder architecture. As Figure 3 shown, for the input image , where represents the low-level features extracted during the downsampling process, represents the high-level features restored during the upsampling process, C is the number of channels, H and W are the spatial height and width respectively. The module design includes the following key modules: CFIFM first performs multi-domain feature perception on the input high-level and low-level features respectively. Specifically, for each path of features, the module simultaneously uses multiple groups of convolutional kernels with different receptive fields (such as 3×3, 5×5, 7×7) to perform parallel convolution operations, and the formal expression is as follows:
[0056] Among them, represents convolution operations of different sizes. Through this multi-scale collaborative coding method, the module can fully capture rich information from local details to global context dependencies. The convolution results of each scale are fused after non-linear activation to form multi-domain enhanced feature representations of high-level features and low-level features respectively and . This design significantly improves the adaptability of features to spatial structures of different scales and provides rich context information for subsequent interactive fusion.
[0057] Subsequently, global context-aware interactive fusion is performed on the two groups of enhanced features and in both the channel and spatial dimensions. First, CFIFM adopts a channel collaborative regulation mechanism to achieve dynamic optimization of features from different sources in the channel dimension. Specifically, global average pooling ( and ) and global max pooling ( ) are performed on the multi-domain enhanced features in the spatial dimension to obtain two groups of global statistical features. Subsequently, through the operation, the above statistical features are concatenated in the spatial width direction to form a joint feature description. After being processed by one-dimensional convolution ( ), two groups of feature representations are generated, and two groups of channel fusion weights are obtained through Softmax normalization ( ), which is formally expressed as follows:
[0058] Among them, and are used to represent global average pooling and global max pooling in the spatial dimension respectively, represents dimension reshaping, represents one-dimensional convolution, represents Softmax normalization.
[0059] This weight can dynamically adjust the contribution degrees of upsampled and downsampled features in each channel to achieve efficient interaction and optimization in the channel dimension.
[0060] Then, CFIFM adopts a spatial collaborative focusing mechanism to further improve the discriminative ability in the spatial dimension. Specifically, global average pooling ( and ) and global max pooling ( ) are performed on the multi-domain enhanced features in the channel dimension to obtain spatial statistical features. Subsequently, through Operate to splice the above statistical features in the channel direction to form joint spatial features, and use Remove the spatial height to form a compact joint spatial feature description. After the compact feature is processed by two-dimensional convolution ( ), two sets of spatial feature representations are generated, and then normalized by Softmax into two sets of spatial fusion weights , and the formal expression is as follows:
[0061] Among them, and respectively represent global average pooling and global max pooling in the channel dimension, represents two-dimensional convolution.
[0062] This weight can dynamically focus on key spatial regions, thereby enhancing the discriminative ability of the model in the spatial dimension and the fineness of feature expression.
[0063] Subsequently, add the two sets of channel fusion weights and the two sets of spatial fusion weights respectively, and introduce an identity mapping (i.e., adding 1 operation) to ensure the effective retention of the original feature information and the stability of gradient flow.
[0064] It can be seen that the core innovation of the CFIFM module lies in constructing a fusion framework that is organically coordinated by three major mechanisms: multi-domain feature perception, channel interaction regulation, and spatial interaction focusing, Figure 3 which details the whole process of "how high-level and low-level features are extracted from input → multi-scale → dynamically weighted → precisely fused". The core is to use intelligent weights to make different-level features complement each other's advantages and solve the problem of ultrasonic image segmentation.
[0065] Multi-domain feature perception mechanism: Through multi-scale convolution paths, it efficiently captures and fuses multi-level spatial information from local fine structures to global contexts, significantly enriching the spatial representation dimension of features.
[0066] Channel interaction regulation mechanism: Based on the global statistical characteristics of cross-path features, it dynamically generates and applies channel-level interaction weights to achieve adaptive integration of the feature channel dimension, significantly strengthening the expression of key semantic channels and suppressing redundancy.
[0067] Spatial interaction focusing mechanism: Through dynamic spatial attention modeling, it guides the model to focus on spatial regions crucial for the current fusion task, such as complex boundaries and suspected nodule regions, significantly improving the capture accuracy of fuzzy boundaries and fine structures.
[0068] In addition, the module adopts a residual design of fusion weights, which effectively guarantees the reliable transmission of the original feature information and the stability of gradient flow while introducing deep interaction fusion.
[0069] Through the synergistic effect of the above innovative mechanism, CFIFM effectively overcomes the problems of information redundancy, insufficient discriminative power, and noise sensitivity caused by traditional simple splicing and fusion. It significantly improves the performance in model structure analysis, precise positioning of fuzzy boundaries, and robustness against strong noise interference, laying a crucial feature fusion foundation for achieving high-precision and high-robustness medical image segmentation.
[0070] Step Five: Based on the enhanced bottleneck layer features and interactive fusion features, output the segmentation result of the thyroid nodule through the decoder.
[0071] In this embodiment, finally, the fusion weights are multiplied element-wise with the multi-domain enhanced features and then added together to obtain the interactive fusion features, which are formally expressed as follows:
[0072]
[0073] The fused features As the output of CFIFM, it combines rich global context information and fine local details, providing highly expressive features for subsequent network layers. This module is flexibly designed and can be seamlessly integrated into the feature bridging positions of various encoder-decoder structures (such as UNet, TransUNet, etc.). This module effectively overcomes the deficiencies of traditional feature fusion methods in integrating global and local information, significantly improving the model's ability to capture complex structures and details, and is particularly suitable for scenarios with extremely high requirements for cross-layer feature fusion such as medical image segmentation.
[0074] Preferably, for feasibility verification, this embodiment also provides a comparative experiment: To comprehensively verify the effectiveness of the method proposed in the present invention, the present invention conducts experiments on the publicly available thyroid ultrasound image segmentation dataset (DDTI) and the self-built thyroid ultrasound image segmentation dataset (Dataset A) respectively. In the comparative experiment, a variety of mainstream models representative in the current medical image segmentation field are selected, including U-Net, DeepLabV3+, TransUNet, DA-TransUNet, U-KAN, and C2F-TransUNet proposed in the present invention as the comparison objects. The experiment uses two commonly used segmentation performance evaluation indicators, the Dice coefficient (Dice Similarity Coefficient, DSC) and the 95% Hausdorff distance (HD95), to measure the overlap degree and boundary localization accuracy of the segmentation results respectively. Through systematic comparison on different datasets and various mainstream models, the accuracy, robustness of the method of the present invention in the thyroid ultrasound image segmentation task and its promotion value in practical applications can be fully verified.
[0075] Table 1 Dataset division method
[0076] As shown in Table 1, we split both datasets into training and test sets in an 8:2 ratio. Specifically, in the publicly available DDTI dataset, 509 samples were used for training and 128 samples were used for testing; in our self-built Dataset A, 900 samples were selected as the training set and 226 samples were selected as the test set.
[0077] Table 2 Segmentation performance of C2F-TransUNet and SOTA segmentation models on the DDTI dataset
[0078] Experiments on the thyroid ultrasound image segmentation dataset DDTI demonstrate that the proposed C2F-TransUNet demonstrates superior segmentation performance. As shown in Table 2, in terms of segmentation accuracy, the average Dice coefficient (DSC) of C2F-TransUNet reaches 84.62±14.11, a significant improvement over mainstream models. This performance is significantly higher than the classic medical image segmentation network U-Net (79.34±18.16), significantly better than the traditional TransUNet (83.13±15.14), and significantly ahead of the convolutional network DeepLabV3+ (70.48±25.84), surpassing the improved DA-TransUNet (78.99±18.52), and outperforming the recently proposed U-KAN (79.26±16.18). In terms of boundary localization accuracy, the HD95 of C2F-TransUNet is 16.14±14.63, which is better than U-Net (25.18±22.42), DeepLabV3+ (29.75±24.74), DA-TransUNet (22.93±22.76), and U-KAN (28.50±23.36), and also surpasses TransUNet (18.39±15.39) in this indicator, showing that this method has stronger robustness and precision in dealing with complex boundaries and detailed structures.
[0079] Furthermore, C2F-TransUNet achieved the lowest standard deviations of all models in both DSC and HD95 metrics, demonstrating its more stable segmentation performance and its ability to adapt to changes in different samples and complex scenarios. In summary, the experimental results of C2F-TransUNet on the DDTI dataset demonstrate its effectiveness and advancement in thyroid ultrasound image segmentation tasks, particularly in improving segmentation accuracy, boundary discrimination, and model stability, outperforming existing mainstream segmentation methods.
[0080] Table 3 Segmentation Performance of C2F-TransUNet and SOTA Segmentation Models on Dataset A
[0081] On the thyroid ultrasound image segmentation dataset Dataset A, C2F-TransUNet proposed by the present invention shows significant advantages in two key evaluation indicators, the Dice coefficient and HD95. As shown in Table 3, the Dice coefficient of C2F-TransUNet reaches the highest accuracy of 80.03±17.89, exceeding U-Net (72.37±27.06), DeepLabV3+ (62.42±31.48), TransUNet (77.77±21.90), DA-TransUNet (77.76±22.13), and U-KAN (73.49±22.08), fully demonstrating the advantages of this method in terms of segmentation accuracy. At the same time, C2F-TransUNet also performs outstandingly in the HD95 index, achieving 25.68±36.76, which is better than U-Net (58.49±80.10), DeepLabV3+ (52.33±59.39), TransUNet (29.81±39.21), DA-TransUNet (30.13±39.50), and comparable to U-KAN (25.55±29.61), indicating that this method also has a leading level in terms of boundary localization and robustness to abnormal samples.
[0082] It is worth noting that the standard deviations of C2F-TransUNet in the Dice and HD95 indices are lower than those of all comparison models, significantly lower than that of U-Net (80.10). This shows that its segmentation performance is more stable and can adapt to the changes in different samples and complex scenarios. Generally speaking, the experimental results of C2F-TransUNet on Dataset A fully verify its effectiveness and advancement in the thyroid ultrasound image segmentation task, especially in terms of improving segmentation accuracy, boundary discrimination ability, and model stability, being superior to existing mainstream segmentation methods.
[0083] Preferably, the present invention also conducts a visual comparison of the segmentation results of typical samples on the self-built Dataset A thyroid ultrasound image segmentation dataset of the publicly disclosed DDTI. As Figure 4As shown, when traditional U-Net and DeepLabV3+ face nodules with blurred boundaries, complex structures, or small targets, the segmentation results often show phenomena such as incomplete target contours, missed detections, or missegmentations, making it difficult to accurately restore the true shape of the nodules. TransUNet and DA-TransUNet have improved in the overall segmentation effect and can better capture the target area, but there are still problems such as uneven boundaries, detail loss, or sensitivity to noise in some samples. U-KAN can segment the target well in some samples, but there are still cases of discontinuous segmentation or missegmentation when dealing with small targets or complex backgrounds.
[0084] Compared with several major mainstream segmentation methods, the C2F-TransUNet proposed in the present invention shows good adaptability to nodules of different shapes and boundary discrimination ability in the segmentation results on two datasets. Regardless of the size and boundary-blurred nodules, the segmentation contour of C2F-TransUNet is highly consistent with the gold standard annotation, and it has a stronger ability to suppress background noise and irrelevant regions. Overall, C2F-TransUNet not only improves the segmentation accuracy and detail restoration ability, but also significantly enhances the generalization ability and robustness of the model across datasets, further verifying the advancement and practical value of the method of the present invention in the thyroid ultrasound image segmentation task.
[0085] Preferably, to further verify the effectiveness of each innovative module proposed in the present invention in the thyroid ultrasound image segmentation task and its contribution to the overall model performance, the present invention designs and conducts a systematic ablation experiment on two datasets. By integrating the collaborative multi-path feature enhancement module (CMFEM), cross-layer feature interaction and fusion module (CFIFM), and their joint application on the basis of TransUNet respectively, the impact of each module on the segmentation accuracy, boundary discrimination ability, and model robustness can be comprehensively evaluated. The ablation experiment not only helps to reveal the independent roles and collaborative gain effects of each module, but also provides a strong empirical basis for the scientific nature and advancement of the method of the present invention.
[0086] Table 4 Ablation experiment results of different improvement modules on Dataset A and DDTI
[0087] As shown in Table 4, the ablation experiment results on the publicly available thyroid ultrasound image DDTI dataset show that the Dice coefficient and HD95 of TransUNet are 83.13±15.14 and 18.39±15.39, respectively. After introducing CMFEM, the Dice coefficient is increased to 84.09±14.47, and HD95 is decreased to 17.87±16.28, showing the positive effect of this module in enhancing feature expression and improving segmentation accuracy. When integrating CFIFM, the Dice coefficient is 84.15±14.18, and HD95 is 17.05±15.86, indicating that CFIFM has significant advantages in feature fusion and detail restoration. When the two are applied jointly, the model performance is further improved, the Dice coefficient reaches 84.62±14.11, and HD95 is decreased to 16.14±14.63. The above results show that the two modules of CMFEM and CFIFM can effectively improve the segmentation accuracy and boundary discrimination ability on this dataset, and the joint application of the two can further enhance the robustness and generalization ability of the model, fully verifying the effectiveness and advancement of the method of the present invention.
[0088] The ablation experiment on the self-built Dataset A thyroid ultrasound image segmentation dataset shows the same trend. The Dice coefficient of the TransUNet model is 77.77±21.90, and HD95 is 29.81±39.21. When CMFEM is introduced, the Dice coefficient is increased to 78.87±20.63, and HD95 is decreased to 28.39±38.13. When CFIFM is integrated alone, the Dice coefficient is increased to 79.71±19.74, and HD95 is decreased to 26.49±37.19. When the two are combined, the model performance reaches the optimal, the Dice coefficient is increased to 80.03±17.89, and HD95 is further decreased to 25.68±36.76, fully demonstrating the synergistic gain effect of the two major modules and being able to maximize the performance of the model in complex ultrasound image segmentation tasks.
[0089] Embodiment 2 This embodiment also provides an electronic device, referring to Figure 6 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0090] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC for short), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.
[0091] Among them, the memory 404 may include a mass memory 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0092] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0093] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the thyroid ultrasound image segmentation methods in the above embodiments.
[0094] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0095] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0096] The input / output device 408 is used to input or output information.
[0097] Embodiment III This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the thyroid ultrasound image segmentation method according to Embodiment I.
[0098] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated here.
[0099] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers, or other computing devices, or some combination thereof.
[0100] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to perform the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, at this point, it should be noted that any box in the logical flow in the figure can represent a program step, or interconnected logical circuits, boxes, and functions, or a combination of program steps and logical circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media are non-transitory media.
[0101] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.
[0102] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it cannot be understood as a limitation to the scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A method for segmenting thyroid ultrasound images, characterized in that, It includes the following steps: Obtain thyroid ultrasound images; Input the thyroid ultrasound images into the C2F-TransUNet network, which includes an encoder, a decoder, a collaborative multi-path feature enhancement module arranged between the encoder and the decoder, and a cross-layer feature interaction and fusion module arranged at the cross-layer connection position in the decoder; Through the collaborative multi-path feature enhancement module, perform collaborative enhancement in the channel and spatial dimensions and group feature interaction on the bottleneck layer features output by the encoder to generate enhanced bottleneck layer features; Through the cross-layer feature interaction and fusion module, perform multi-domain feature perception and dynamic weight fusion in the channel and spatial dimensions on the high-level and low-level features of the upsampling and downsampling paths in the decoder to generate interaction and fusion features; Based on the enhanced bottleneck layer features and the interaction and fusion features, output the segmentation result of thyroid nodules through the decoder.
2. The thyroid ultrasound image segmentation method according to claim 1, characterized in that, The collaborative multi-path feature enhancement module includes: A two-dimensional attention coupling module for co-optimizing channel attention and spatial attention on the input features. The channel attention generates a channel attention vector through global average pooling, layer normalization, and 1×1 convolution for dimension reduction and upsampling. The spatial attention generates a spatial attention matrix through parallel processing of 7×7 depthwise separable convolution and 3×3 dilated convolution with a dilation rate of 3, and the element-wise product of the two forms an adaptive enhancement feature; A group feature interaction layer for performing group normalization and group convolution on the input features after grouping them along the channel, and realizing cross-group feature interaction through channel shuffling operation; A dynamic feature fusion gate for multiplying the output features of the two-dimensional attention coupling module with the original input features element-wise, then splicing with the output features of the group feature interaction layer, and generating final enhanced features after 1×1 convolution, batch normalization, and GELU activation.
3. The thyroid ultrasound image segmentation method according to claim 2, wherein, The process of the two-dimensional attention coupling module generating the channel attention vector includes: performing global average pooling on the input feature map to obtain global semantic statistics; After being processed by layer normalization and the GELU activation function, it is dimension-reduced and then dimension-increased through a 1×1 convolution to generate a channel attention vector , where C is the number of channels.
4. The thyroid ultrasound image segmentation method according to claim 2, characterized in that, The process of the two-dimensional attention coupling module generating the spatial attention matrix includes: Extracting local fine-grained structure features through 7×7 depthwise separable convolution, and extracting long-range spatial dependence features through 3×3 dilated convolution with a dilation rate of 3; After performing 1×1 convolution for dimensionality reduction on the two-way features respectively and then adding them together, a spatial attention matrix is generated through Sigmoid activation , where H and W are the spatial height and width respectively.
5. The thyroid ultrasound image segmentation method according to claim 2, wherein, The processing process of the group feature interaction layer includes: dividing the input features into g subgroups along the channel, and performing group normalization on each subgroup; Performing 1×1 group convolution on the normalized subgroups respectively to generate intra-group structured features; Recombine the channels of different subgroups through the channel shuffle operation and restore them to the original feature dimension 。 6. A method for segmenting thyroid ultrasound images according to claim 1, characterized in that, The cross-layer feature interaction and fusion module includes: A multi-domain feature perception unit that respectively uses 3×3, 5×5, and 7×7 convolutions to extract multi-scale features from the input high-level and low-level features in parallel, and generates multi-domain enhanced features after fusion; A channel interaction regulation unit that performs global average pooling and max pooling in the spatial dimension on the multi-domain enhanced features, splices them, and then generates channel fusion weights through 1×1 convolution and Softmax to dynamically adjust the channel contribution degrees of the high-level and low-level features; A spatial interaction focusing unit that performs global average pooling and max pooling in the channel dimension on the multi-domain enhanced features, splices them, and then generates spatial fusion weights through 2D convolution and Softmax to focus on key spatial regions; The residual fusion unit adds 1 after adding the channel and spatial fusion weights, multiplies element-wise with the multi-domain enhanced features, and then adds them up to generate the interactive fusion features.
7. The method for segmenting thyroid ultrasound images according to claim 6, wherein The process of the channel interaction regulation unit generating the channel fusion weight includes: Performing spatial global average pooling and global max pooling on the multi-domain enhanced features to obtain channel statistical features; Concatenate the statistical features along the spatial width, reduce the dimension through 1×1 convolution, and then normalize through Softmax to generate the channel fusion weights .
8. The thyroid ultrasound image segmentation method according to claim 6, wherein, The process of the spatial interaction focusing unit generating the spatial fusion weight includes: performing channel global average pooling and global max pooling on the multi-domain enhanced features to obtain spatial statistical features; Concatenate the statistical features along the channels, perform 2D convolution processing, and then normalize through Softmax to generate the spatial fusion weights .
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the thyroid ultrasound image segmentation method according to any one of claims 1 to 8.
10. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium, and the computer program includes program codes for controlling a process to execute the process, and the process includes the thyroid ultrasound image segmentation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Large receptive field attention-enhanced breast ultrasound image segmentation method
CN118261925A
Ultrasonic image segmentation method and system based on multistage feature extraction
CN119693383A
Deep learning-based ultrasonic image segmentation method and system, device and medium
WO2024182997A1