Neck transparent layer segmentation method of U-shaped coding and decoding structure based on improved Transform
Through the improved Transformer structure of the U-shaped encoder-decoder network TransNTSeg, combined with efficient self-attention and DDMix-FFN modules, the problems of insufficient accuracy and high computational complexity in NT ultrasound image segmentation are solved, and efficient and accurate segmentation of the NT area is achieved.
Patent Information
- Application Number
- CN202510769403.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-30
AI Technical Summary
Existing technologies for fetal nuchal translucency (NT) ultrasound image segmentation have problems such as insufficient accuracy, poor robustness, and high computational complexity of traditional Transformer structures, making it difficult to achieve accurate NT region segmentation.
TransNTSeg, a U-shaped encoder-decoder network with an improved Transformer structure, is combined with an efficient self-attention mechanism and DDMix-FFN module. Through multi-scale feature fusion and cross-layer feature fusion modules, the model improves the segmentation accuracy and robustness of the NT area while reducing the computational complexity.
Efficient and accurate segmentation of NT ultrasound images was achieved, which significantly improved the recognition accuracy of the NT region and the practicality of the model and reduced the computational complexity.
Smart Images

Figure CN120726319A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image segmentation, and in particular relates to a nuchal translucency (NT) segmentation method based on an improved Transformer U-type encoding and decoding structure. Background Art
[0002] Nuchal translucency (NT) refers to the accumulation of subcutaneous fluid at the back of the fetal neck, manifesting as an echo-free area on ultrasound images. NT thickness is a key indicator in early pregnancy ultrasound screening. Abnormally increased NT thickness may indicate an increased risk of fetal chromosomal abnormalities (such as Down syndrome and Edwards syndrome) and structural malformations. Therefore, accurate NT measurement is crucial for risk assessment and subsequent diagnosis.
[0003] However, NT measurement faces many challenges. First, ultrasound images have inherent limitations of low resolution and poor contrast, and are susceptible to noise and artifacts, making it difficult to clearly distinguish NT areas. Second, NT areas are usually small, with blurred boundaries, making them difficult to accurately locate. In addition, standard NT measurements need to be performed on a strictly defined fetal midsagittal section, and fetal mobility also increases the difficulty of obtaining standard sections.
[0004] To address the above issues, researchers have explored a variety of NT segmentation methods. Early semi-automatic methods relied on users to manually identify the membrane structure at the boundary of the transparent layer and guide the edge tracking algorithm for segmentation. Chudhary et al. proposed an automated mid-sagittal plane NT detection technology based on SIFT key points and GRNN. Deng et al. proposed a NT measurement method that requires manual assistance, using morphological filtering, empirical threshold segmentation, and gradient vector flow serpentine model for segmentation. In addition, some researchers proposed a hierarchical structure model that uses Gaussian pyramid representation to train discriminant classifiers to identify NT, head, and body regions.
[0005] In recent years, deep learning technology has made significant progress in the field of medical image segmentation. Thomas et al. built the SegNet model based on CNN and used transfer learning to accelerate training to achieve NT segmentation. Some researchers have also introduced the Transformer structure into the medical image segmentation task to enhance the model's ability to model global contextual information, such as combining the Transformer as an encoder with the U-Net decoder. However, directly applying the Transformer to high-resolution images will result in excessive computational complexity. Improved Transformer structures, such as SwinTransformer, reduce computational complexity by introducing a sliding window mechanism and hierarchical feature representation. However, the window-based partitioning method may not be suitable for processing irregularly shaped targets, and may not be as effective as the global attention mechanism for modeling dependencies between distant regions in the image.
[0006] In view of the importance of NT region segmentation and the shortcomings of the existing technology, the present invention proposes a TransNTSeg model, which adopts a U-shaped codec network architecture with an improved Transformer structure, specifically for NT segmentation. The encoder and decoder of TransNTSeg both adopt a hierarchical structure. Each stage contains an overlapping block embedding module and an improved Transformer module, which are responsible for extracting local and global feature information of NT ultrasound images, respectively. At the same time, the jump connection structure adopts a cross-layer feature fusion module based on DDMix-FFNTransformer, aiming to more effectively fuse multi-scale information from the encoder. Compared with the existing technology, the present invention not only achieves excellent performance in NT segmentation, but also significantly reduces the trainable parameters required for the model, thereby greatly reducing the computational complexity and making it more practical. Summary of the Invention
[0007] In response to the problems of insufficient accuracy, poor robustness and high computational complexity of the existing NT ultrasound image segmentation technology, this paper proposes a U-type codec network TransNTSeg based on the improved Transformer structure, aiming to achieve accurate NT ultrasound image segmentation and provide a more accurate basis for the segmentation results.
[0008] Specific content of the invention is as follows:
[0009] Step 1: Obtain a NT ultrasound image dataset, consisting of NT ultrasound images and corresponding labeled images. Divide the dataset into a training set and a test set. To address issues such as low contrast and high noise in NT ultrasound images, perform data preprocessing on the images. This includes contrast enhancement, horizontal flipping, random cropping, random rotation, and random noise addition. This data augmentation technique simulates low-quality, low-contrast ultrasound images, increases data diversity, and improves model robustness. Furthermore, during data preprocessing, resize all NT ultrasound images to 224×224 pixels.
[0010] Step 2: Construct a NT segmentation network model based on TransNTSeg, which is an encoder-decoder image segmentation network based on a U-shaped structure. The encoder is used to extract effective features from NT ultrasound images and suppress the interference of noise and artifacts. Multi-scale feature fusion is used to enhance the model's perception of small targets, solving the problem of small NT regions in NT ultrasound images that are difficult to segment. To address the high computational complexity of the traditional Transformer structure, an improved Transformer structure is introduced. By designing an efficient self-attention mechanism and a DDMix-FFN structure, the segmentation capability is enhanced while reducing the model's computational complexity and solving the problems of vanishing and exploding gradients.
[0011] Step three: Construct the skip connection structure in the network model. Given the small size and blurred boundaries of the NT region in NT ultrasound images, a cross-layer feature fusion module based on the DDMix-FFNTransformer was designed between the encoder and decoder of the TransNTSeg network to effectively utilize the multi-scale features extracted by the encoder and compensate for the loss of segmentation details.
[0012] Step 4: Network training. The collected NT ultrasound image dataset is input into the designed TransNTSeg model. Training is performed using a combination of binary cross entropy and Dice loss function. The AdamW optimizer is used to optimize model parameters and improve model generalization capabilities, ultimately obtaining segmentation results.
[0013] Preferably, step 2 is specifically as follows: To better extract features from NT ultrasound images and fuse multi-scale information, TransNTSeg adopts a U-shaped encoder-decoder structure. The encoder adopts a hierarchical structure consisting of four stages, extracting features from NT ultrasound images layer by layer. Each stage includes an overlapping block embedding module and an improved Transformer module, respectively used to extract local features from NT ultrasound images, particularly edges and texture information from NT regions, while preserving local continuity through an overlapping design. The improved Transformer structure is used to capture global contextual information and utilizes an efficient self-attention mechanism to focus on important regions, suppressing interference information in NT ultrasound images. Corresponding to the encoder, the decoder also adopts a hierarchical structure to gradually restore the resolution and detail information of the NT ultrasound images. In addition to the overlapping block embedding module and the improved Transformer module, a convolutional block attention module (CBAM) is added before each Transformer module. By extracting features from both the channel and spatial dimensions and adaptively adjusting weights, it addresses the problem of high noise and unclear features in NT ultrasound images. Furthermore, the encoder and decoder are connected via a cross-layer feature fusion module based on the DDMix-FFN Transformer. This module fuses the encoder's multi-scale features to compensate for the loss of segmentation detail. This fusion of multi-scale information is particularly important in NT ultrasound images, where the NT region is small and has blurred edges, helping the model to more accurately identify the target. Finally, in the final decoder layer, a segmentation head is used to map the decoder's output feature maps into pixel-level segmentation results, ultimately achieving precise segmentation.
[0014] The TransNTSeg network utilizes an improved Transformer architecture designed to improve the accuracy and efficiency of NT ultrasound image segmentation. Considering the low resolution, small size, and susceptibility to noise in NT ultrasound images, this architecture incorporates several improvements over the traditional Transformer architecture. First, to process NT ultrasound image feature maps with limited computational resources while maintaining global modeling capabilities, the present invention employs a highly efficient self-attention mechanism, significantly reducing computational complexity through spatial reduction techniques. This enables it to effectively capture long-range dependencies between the NT region and other tissues in NT ultrasound images, improving NT segmentation accuracy. Secondly, in response to the small size and blurred edges of the NT region in NT ultrasound images, the designed DDMix-FFN module innovatively employs dilated depthwise separable convolutions to expand the receptive field, effectively extracting local information with fewer parameters. Multi-layer normalization and residual connections are then used to accelerate model convergence and alleviate the vanishing gradient problem, thereby enhancing the expressive power of features and enabling the model to more accurately identify and segment small target NT regions. The improved Transformer structure contains key components such as layer normalization, efficient self-attention mechanism, residual connection, DDMix-FFN module, etc. By working together, these components can effectively suppress noise and artifacts in NT ultrasound images and fully extract the features of smaller NT areas in NT ultrasound images, thereby improving segmentation performance.
[0015] Preferably, step three specifically includes the following: Considering that NT ultrasound images are small and susceptible to noise, making it difficult to extract clear segmentation details, the TransNTSeg network designs a cross-layer feature fusion module based on the DDMix-FFN Transformer between the encoder and decoder. This module aims to leverage multi-scale features extracted by different encoder stages to offset the shortcomings of single-scale features, thereby improving the model's ability to perceive small NT regions. This module adaptively reshapes and splices feature maps from different encoder stages, ensuring that the extracted NT feature maps have the same size and number of channels, facilitating subsequent processing. Furthermore, the reshaping process highlights important information in the feature maps at different scales that is helpful for NT region segmentation, such as edges and texture. An efficient self-attention mechanism extracts long-term dependencies between features, thereby suppressing noise and artifacts in NT ultrasound images and focusing on the correlation between the NT region and other tissues. Subsequently, the spliced feature maps are split into multiple scales and fed into different DDMix-FFN modules for feature extraction and cross-scale information exchange. The DDMix-FFN modules effectively extract local information and enhance the expressiveness of features, thereby improving the model's recognition accuracy for NT regions. Finally, the aggregated features are fused with the input features through residual connections to enhance feature representation, and the fused, rich feature information is passed to the decoder. Through the above design, the module can fully utilize the multi-scale information of the encoder to improve the segmentation accuracy and robustness of NT ultrasound images.
[0016] Preferably, the step four is specifically as follows: In response to the segmentation accuracy and category imbalance problems in the NT ultrasound image segmentation task, the present invention adopts a combined loss function to optimize the segmentation performance of the TransNTSeg network, ensuring that the model can effectively learn and accurately segment the NT region. Since the NT region in the NT ultrasound image usually accounts for a very small proportion, it is a typical category imbalance problem. Therefore, the use of binary cross entropy loss (BCELoss) alone can easily cause the model to be biased towards the segmentation of the background region. Therefore, the present invention combines DiceLoss to effectively balance the segmentation accuracy and category imbalance problems, so that the model can accurately segment the NT region of small targets while paying attention to the overall segmentation accuracy. By reasonably setting the weight factors (α and β) of the two loss functions, the network learning can be effectively guided, and ultimately more accurate segmentation results can be obtained. In the experiment, α is set to 0.4 and β is set to 0.6.
[0017] This invention offers significant benefits: it proposes a novel U-shaped codec network based on an improved Transformer architecture for accurate NT segmentation. The improved Transformer, the core of this model, is cleverly embedded in both the encoder and decoder. Two key components of its architecture—an efficient self-attention mechanism and a DDMix-FFN module—ensure that rich global information is extracted while effectively reducing computational complexity, thereby enhancing the model's practicality. By combining the U-shaped codec with the Transformer architecture, this module combines the advantages of both, focusing on local details while also emphasizing global context, ultimately achieving efficient and accurate NT segmentation. Furthermore, the skip connections utilize a cross-layer feature fusion module based on the DDMix-FFN Transformer. This unique architecture maximizes the fusion of multi-scale features from the encoder, promoting complementarity and enhancement between feature information. The decoder cleverly employs a convolutional block attention module to extract and weight features from both the channel and spatial dimensions, highlighting important features and suppressing noise. A carefully designed segmentation head is incorporated into the final decoder layer, ultimately achieving precise segmentation of the NT region. Compared with other network structures, the present invention has obvious advantages in performance, efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a structural diagram of the TransNTSeg network of the present invention.
[0019] Figure 2 This is the improved Transformer structure diagram.
[0020] Figure 3 This is a schematic diagram of the structure of DDMix-FFN in Transformer.
[0021] Figure 4 This is the structural diagram of the cross-layer feature fusion module based on DDMix-FFNTransformer.
[0022] Figure 5 It is a structural diagram of the convolutional block attention module. DETAILED DESCRIPTION
[0023] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can more easily understand its advantages and features and define the scope of protection more clearly and precisely accordingly.
[0024] In this example, we used a NT ultrasound dataset for pregnancies ranging from 11 to 13 weeks, provided by a hospital. This dataset contains 732 images, and the training and test sets are divided into 80% and 20% ratios.
[0025] A U-type codec network TransNTSeg based on an improved Transformer structure is used for NT segmentation. The specific steps are as follows:
[0026] Step 1: Obtain a NT ultrasound image dataset, consisting of NT ultrasound images and corresponding labeled images. Divide the dataset into a training set and a test set. To address issues such as low contrast and high noise in NT ultrasound images, perform data preprocessing on the images. This includes contrast enhancement, horizontal flipping, random cropping, random rotation, and random noise addition. This data augmentation technique simulates low-quality, low-contrast ultrasound images, increases data diversity, and improves model robustness. Furthermore, during data preprocessing, resize all NT ultrasound images to 224×224 pixels.
[0027] Step 2: Construct a NT segmentation network model based on TransNTSeg, which is an encoder-decoder image segmentation network based on a U-shaped structure. The encoder is used to extract effective features from NT ultrasound images and suppress the interference of noise and artifacts. Multi-scale feature fusion is used to enhance the model's perception of small targets, solving the problem of small NT regions in NT ultrasound images that are difficult to segment. To address the high computational complexity of the traditional Transformer structure, an improved Transformer structure is introduced. By designing an efficient self-attention mechanism and a DDMix-FFN structure, the segmentation capability is enhanced while reducing the model's computational complexity and solving the problems of vanishing and exploding gradients.
[0028] Step three: Construct the skip connection structure in the network model. Given the small size and blurred boundaries of the NT region in NT ultrasound images, a cross-layer feature fusion module based on the DDMix-FFNTransformer was designed between the encoder and decoder of the TransNTSeg network to effectively utilize the multi-scale features extracted by the encoder and compensate for the loss of segmentation details.
[0029] Step 4: Network training. The collected NT ultrasound image dataset is input into the designed TransNTSeg model. Training is performed using a combination of binary cross entropy and Dice loss function. The AdamW optimizer is used to optimize model parameters and improve model generalization capabilities, ultimately obtaining segmentation results.
[0030] In the step 2, a NT segmentation network model based on TransNTSeg is constructed. This model adopts a U-shaped encoder-decoder structure and introduces an improved Transformer module for accurate segmentation of NT ultrasound images. Figure 1The specific implementation steps are as follows: For the encoder, a hierarchical structure is adopted, consisting of four stages. Each stage contains an overlapping block embedding module and an improved Transformer module. The two work together to extract local features and global context information from NT ultrasound images. The overlapping block embedding is a convolutional layer with a specific 7×7 convolution kernel, a stride of 4, and 3 pixel padding. It is used to extract local features such as edges and textures of the NT region in the NT ultrasound image. The overlapping pixel blocks are divided to preserve local continuity, which is conducive to the subsequent capture of the subtle structure of the NT region. Then, the feature map output by the convolutional layer is converted into a one-dimensional feature vector sequence through a deformation operation. The feature vector sequence is then mapped to a preset high-dimensional feature space using a linear layer to generate a series of high-order embedding vectors. The resulting vectors are input into the improved Transformer module to learn the global context information of the NT ultrasound image. Given the susceptibility of NT ultrasound images to noise interference, the improved Transformer module can adaptively focus on key areas in the image, prevent the interference of noise and artifacts, and improve the feature expression capability. The encoder performs downsampling at each stage, gradually reducing the resolution of the feature map and increasing the number of channels while extracting deep semantic information. For example, if the input image is 224×224 pixels, the feature map size after each encoder stage changes to 56×56, 28×28, 14×14, and 7×7, respectively. The decoder is symmetrical to the encoder and also adopts a layered structure consisting of four stages. Each stage includes an overlapping block embedding module for upsampling, expanding the feature map size and extracting local features, gradually restoring detailed information in the NT ultrasound image, and an improved Transformer module to supplement global dependencies. In addition, a convolutional attention module is added before each Transformer layer. To address the small size and blurred edges of NT regions in NT ultrasound images, this module comprises channel-attention and spatial-attention submodules, respectively extracting features from the channel and spatial dimensions. The channel-attention submodule uses global average pooling and global max pooling to obtain channel statistics from the feature map. It then uses a multi-layer perceptron to learn inter-channel dependencies and adaptively adjust channel weights. The spatial-attention submodule, based on the weighted channel feature map, uses convolution operations to extract spatial weight information, thereby emphasizing key regions and suppressing noise interference. This spatial attention mechanism enables the model to focus on key locations in the NT region while ignoring background noise interference, improving segmentation accuracy and robustness. Finally, in the final decoder layer, a segmentation head is used to map the decoder output feature map into pixel-level segmentation results. This head, typically composed of convolutional layers, normalization layers, and activation functions, converts high-dimensional feature information into pixel-level segmentation probability maps, thereby achieving accurate NT segmentation.
[0031] One of the core modules of the TransNTSeg network is the improved Transformer module, such as Figure 2 As shown. This module is optimized based on the traditional Transformer, aiming to more effectively capture the global context information in NT ultrasound images and maintain high efficiency on the NT feature map, thereby improving the segmentation performance of the model. Taking into account the characteristics of NT ultrasound images, the Transformer module of TransNTSeg is mainly improved in the following aspects: Considering the high noise of NT images and the high computational complexity of traditional Transformer, efficient self-attention is adopted. Given the input feature map X, the dimensions of the Query (Q), Key (K) and Value (V) feature maps are first changed from H×W×C to Where R is the spatial reduction rate, and then a linear projection layer is used to restore the number of channels to C. Finally, the scaled dot product attention mechanism is used to calculate the self-attention weight and weight the Value feature map. This method reduces the computational complexity from O(N 2 ) is reduced to This allows it to be applied to NT ultrasound image feature maps. This enables the model to effectively capture the long-range dependencies between the NT region and other tissue structures, adaptively focus on key areas in the image, suppress the interference of noise and artifacts, and improve the model's robustness to NT ultrasound images. This can be described in formula language as follows:
[0032]
[0033] Another key module of TransNTSeg is DDMix-FFN, such as Figure 3 As shown. In response to the problems that the NT area in NT ultrasound images is small, has blurred edges, and is easily disturbed by noise, the DDMix-FFN module aims to enhance the ability to extract local information, accelerate model convergence, alleviate the gradient vanishing problem, and thus improve the accuracy and robustness of segmentation. To achieve this goal, the DDMix-FFN module uses expanded depthwise separable convolution to replace the traditional convolution operation, expand the receptive field with fewer parameters, and capture pixel dependencies at longer distances, which is especially important for the recognition of small targets in NT ultrasound images. Multi-layer normalization operations are used to accelerate model convergence, improve training stability, and introduce residual connections to alleviate the gradient vanishing problem, so that the model can better learn the complex features of NT ultrasound images. Specifically, for the input feature map x, a fully connected layer FC is first used to map it to a higher-dimensional feature space to obtain y1=FC(x in ), and then use the dilated depth separable convolution to extract local information to get y2 = Conv 3×3(y1), and then perform layer normalization to obtain y3 = LN(y2), and then perform residual connection operation to obtain y4 = y3 + y1. This process is repeated to enhance the expressive power of the feature. Finally, the GELU activation function is used for nonlinear transformation, and a fully connected layer FC is used to map the feature map back to the original dimension to obtain the output x out =FC(GELU(DDMix))+x in . It can be described in formula language as:
[0034] y1=LN((Conv 3×3 (FC(x in ))+FC(x in )))
[0035] DDMix=LN(LN(y1+FC(x in ))+FC(x in ))
[0036] x out =FC(GELU(DDMix))+x in
[0037] Among them, LN refers to the normalization layer, FC refers to the fully connected layer, and GELU is the activation function. Through this structure, the expressive ability of the model can be enhanced.
[0038] The improved Transformer module includes layer normalization, efficient self-attention, residual connections, layer normalization, and DDMix-FFN modules. These components work together to enhance the model's ability to capture global contextual information while maintaining high efficiency on high-resolution feature maps.
[0039] In the step 3, a cross-layer feature fusion module is implemented to effectively utilize the multi-scale features extracted by the encoder and compensate for the loss of segmentation details. In view of the characteristics of the NT area in the NT ultrasound image being small, with blurred edges and susceptible to interference, the present invention designs a cross-layer feature fusion module based on DDMix-FFNTransformer, such as Figure 4 As shown in the figure, it aims to improve the model's ability to perceive small target NT areas and reduce the impact of noise on segmentation results. The specific implementation method is: First, feature reshaping is performed, that is, the feature maps from different stages of the encoder are adjusted to a uniform number of channels C using convolutional layers, for example, the size is The features will be reshaped to a size of This operation not only unifies the representations of feature maps at different scales but also highlights key features that are helpful for NT region segmentation through convolution operations. Next, feature concatenation is performed: the feature maps from the four stages are spatially concatenated to generate a fused feature map containing multi-scale information. This concatenation operation fully utilizes the multi-scale information extracted by the encoder, thereby improving the model's perception of the NT region and compensating for the shortcomings of single-scale features. Multi-scale information fusion is particularly important when segmenting small NT regions. Next, the global modeling capabilities of the efficient self-attention module are utilized to extract long-term dependencies between locations in the fused feature map. Furthermore, the efficient self-attention mechanism can suppress noise interference in NT ultrasound images, allowing the model to focus more on key areas in the image. Finally, the feature information at different scales is refined. The fused feature map processed by the efficient self-attention module is split into four specific scales according to the size of the original feature map and then fed into different DDMix-FFN modules. This allows for refined processing of information at different scales, extracting feature information at different scales that is helpful for NT region segmentation. The features processed by different DDMix-FFN modules are then aggregated and, finally, a residual connection is performed to fuse the aggregated features with the input features. Finally, the feature maps processed by the DDMix-FFN Transformer cross-layer feature fusion module are passed to the decoder as input. This design effectively fuses the multi-scale features extracted by the encoder, compensating for the loss of segmentation details and providing the decoder with richer contextual information, thereby improving the segmentation accuracy and robustness of the model for NT ultrasound images.
[0040] In step 4, in order to optimize the segmentation performance of the TransNTSeg network for NT ultrasound images and solve the problems of segmentation accuracy and class imbalance, the combined loss function of the network binary cross entropy and Dice is used. The final loss function can be described in formula language as follows:
[0041]
[0042] Where Loss is the final combined loss, Represents the final predicted segmentation result. y represents the actual segmentation result. BCE represents the binary cross entropy loss function, which is used to measure the classification accuracy at the pixel level and encourages the model to make correct predictions about the category of each pixel. Especially for NT ultrasound images, noise and artifacts may cause pixel-level classification difficulties, so BCELoss is needed to ensure basic classification accuracy. Dice represents the Dice loss function, which is used to measure the degree of overlap of the segmented regions and encourage the segmentation results generated by the model to be as consistent as possible with the true labels, with particular attention paid to the shape and size of the NT region. Since the NT region in NT ultrasound images usually accounts for a very small proportion, it is a typical small target segmentation problem. DiceLoss can encourage the model to pay more attention to the segmentation of the NT region. α and β represent the weight factors of the loss function. In the experiment, we set α to 0.4 and β to 0.6. By reasonably setting the combined loss function and weight factor, the TransNTSeg network can better learn the characteristics of NT ultrasound images, improve segmentation accuracy and robustness, and thus achieve more accurate NT measurement.
[0043] This example is implemented in PyTorch and trained on a 16GB 4070Ti Super GPU. Data augmentation operations, such as horizontal flipping, random rotation, and random noise addition, are used to improve generalization. The TransNTSeg model training in this example is divided into two stages. In the first stage, the model is trained without pretrained weights. In the second stage, the model is fine-tuned using the optimal weights obtained in the first stage. For the NT ultrasound image dataset, the batch_size is set to 4, and all images are resized to 224×224 for consistent input. The model is trained using the AdamW optimizer with an initial learning rate of 0.0001. The loss function is a weighted combination of cross-entropy loss and Dice loss. All baseline models follow the same training configuration and loss function. Segmentation results are quantified using the Dice coefficient, Intersection over Union (IoU), precision, and sensitivity metrics.
[0044] In this study, we used the Dice coefficient and IoU as the main evaluation indicators. This is because in the NT segmentation task, the accuracy of segmentation is crucial, and the Dice coefficient and IoU can directly measure the degree of overlap between the predicted segmentation results and the true label. In addition, we also used sensitivity and precision as auxiliary evaluation indicators. Sensitivity is used to evaluate the model's ability to correctly identify the NT region among all real NT regions, that is, to avoid misjudging the NT region as the background. Precision is used to evaluate the proportion of pixels predicted by the model as NT regions that actually belong to the NT region, that is, to avoid misjudging the background region as the NT region. The calculation formula is as follows:
[0045]
[0046]
[0047] Where X is the true area and Y is the predicted area. TP is when the model predicts the positive class and the prediction is correct, FP is when the model predicts the positive class but the prediction is incorrect, and FN is when the model predicts the negative class but the prediction is incorrect. In evaluation, we want these values to be as close to 1 as possible.
[0048] On the test set, the performance of the proposed TransNTSeg model was compared with eight other deep neural networks: 1. U-Net (Ronneberger, Fischer, Brox, 2015), a powerful image segmentation model whose U-shaped structure and skip connections enable it to effectively utilize contextual and spatial information to achieve accurate segmentation. 2. Brau-net++ (Lan LB, Cai PZ, Jiang L, et al. 2024), a U-shaped hybrid CNN-Transformer network for medical image segmentation. It combines the advantages of CNN and Transformer to more effectively extract image features and achieve higher segmentation accuracy. 3. Unet++ (Zhou Z, Rahman Siddiquee MM, Tajbakhss N, et al. 2018), a segmentation model that introduces more upsampling nodes and skip connections, thereby improving the model's ability to model long-range dependencies. 4. DeepLabV3+ (LiuY, BaiX, WangJ, et al. 2024), a model with powerful context modeling capabilities and feature enhancement capabilities of the attention mechanism, which can achieve higher segmentation accuracy. 5. ResUNet-a (DiakogiannisFI, WaldnerF, CaccettaP, et al. 2020), a model based on the U-Net architecture and using the ResNet module as the building block of the encoder. This model also integrates an attention mechanism to improve segmentation accuracy. 6. Transunet (ChenJ, LuY, YuQ, et al. 2021), the design is inspired by the Transformer model in the field of natural language processing, which can automatically learn features in images and convert them into semantic information. 7. R2u-net (Alom MZ, Hasan M, Yakopcic C, et al. 2018), a method based on the U-Net architecture that uses R2CNN modules as encoder and decoder building blocks, combines the advantages of RNN and ResNet to extract richer image features and capture long-range dependencies. 8. MT-Unet (Wang H, Xie S, Lin L, et al. 2022), based on the U-Net architecture, uses a hybrid Transformer module to extract features and capture long-range dependencies. This method combines the global modeling capabilities of the Transformer with the local feature extraction capabilities of the CNN to improve segmentation accuracy.
[0049] Table 1 summarizes the segmentation performance comparison results of the TransNTSeg model and existing mainstream methods on the NT dataset. Experimental results show that the TransNTSeg model has achieved significant performance improvements in the four key indicators of Dice coefficient, IoU, precision, and sensitivity, reaching 90.7%, 83.6%, 89.4%, and 92.7%, respectively, which is better than other comparison methods. In addition, although the TransNTSeg model does not have the least trainable parameters, compared with the model that also uses the U-Net and Transformer hybrid architecture, the number of parameters required by the present invention is significantly smaller, which shows that the TransNTSeg model can achieve higher segmentation performance with fewer parameters and has higher parameter utilization efficiency.
[0050] In summary, the TransNTSeg segmentation network proposed in this paper achieves excellent performance in the NT segmentation task. This network fully leverages the advantages of the DDMix-FFNTransformer structure, efficiently capturing global contextual information and, through a cross-layer feature fusion module, fully utilizing the multi-scale features extracted by the encoder to generate more expressive feature representations. Furthermore, by introducing an efficient self-attention mechanism to reduce computational complexity, utilizing a convolutional block attention module to adaptively extract spatial and channel features, and employing a segmentation head for precise segmentation, TransNTSeg achieves significant improvements in both the performance and efficiency of NT segmentation.
[0051] Table 1
[0052]
Claims
1. A method for segmenting the nuchal translucency layer based on an improved Transformer U-shaped codec structure, characterized in that: The implementation steps of this method are as follows: Step 1: Obtain a NT ultrasound image dataset, including NT ultrasound images and corresponding labeled images; divide the NT ultrasound image dataset into a training set and a test set, and perform data preprocessing; Step 2: Build a NT segmentation network model based on TransNTSeg. This NT segmentation network model is an encoder-decoder image segmentation network based on a U-shaped structure. The encoder is used to extract effective features from NT ultrasound images and suppress the interference of noise and artifacts. The model's perception of small targets is enhanced through multi-scale feature fusion. The improved Transformer structure is introduced. By designing an efficient self-attention mechanism and DDMix-FFN structure, the segmentation capability is enhanced while reducing the model's computational complexity and solving the problems of gradient vanishing and exploding. Step 3: Construct a skip connection structure in the network model. Considering the small size and blurred boundaries of the NT region in NT ultrasound images, a cross-layer feature fusion module based on DDMix-FFNTransformer is designed between the encoder and decoder of the TransNTSeg network to utilize the multi-scale features extracted by the encoder and compensate for the loss of segmentation details. Step 4: Network training: The collected NT ultrasound image dataset is input into the designed TransNTSeg model, and the combined loss function of binary cross entropy and Dice is used for training. The AdamW optimizer is used to optimize the model parameters, and finally the nuchal translucency segmentation result of the NT ultrasound image is obtained.
2. The nuchal translucency segmentation method based on the improved Transformer U-shaped codec structure according to claim 1 is characterized in that: The step 1 is specifically as follows: the data preprocessing includes contrast enhancement, horizontal flipping, random cropping, random rotation, and adding random noise data enhancement technology to adjust all NT ultrasound images to NT ultrasound images of 224×224 size.
3. The nuchal translucency segmentation method based on the improved Transformer U-shaped codec structure according to claim 1, characterized in that: The second step is as follows: to extract the features of the NT ultrasound image and fuse multi-scale information, TransNTSeg adopts a U-shaped encoder-decoder structure; the encoder adopts a hierarchical structure, including four stages, to extract the features of the NT ultrasound image layer by layer; Each stage contains an overlapping block embedding module and an improved Transformer module, which are used to extract local features of NT ultrasound images, especially the edges and texture information of the NT area, and retain local continuity through overlapping design; the improved Transformer structure is used to capture global context information, and use an efficient self-attention mechanism to focus on important areas and suppress interference information in the NT ultrasound image; corresponding to the encoder, the decoder adopts a hierarchical structure to gradually restore the resolution and detail information of the NT ultrasound image; in addition to the overlapping block embedding module and the improved Transformer module, a convolutional block attention module CBAM is added before each Transformer module to adaptively adjust the weights by extracting features from two dimensions: channel and space; the encoder and decoder are connected through a cross-layer feature fusion module based on DDMix-FFNTransformer, which is used to fuse the multi-scale features of the encoder and compensate for the loss of segmentation details; finally, in the last layer of the decoder, the segmentation head is used to map the feature map output by the decoder into pixel-level segmentation results, ultimately achieving accurate segmentation.
4. The nuchal translucency segmentation method based on the improved Transformer U-shaped codec structure according to claim 3 is characterized in that: The TransNTSeg network adopts an improved Transformer structure and an efficient self-attention mechanism. Secondly, in view of the characteristics of small NT areas and blurred edges in NT ultrasound images, the DDMix-FFN module is designed to use expanded depthwise separable convolution to expand the receptive field and enhance the expressive ability of features. The improved Transformer structure successively includes layer normalization, efficient self-attention mechanism, residual connection, and DDMix-FFN module to effectively suppress noise and artifacts in NT ultrasound images, and fully extract the features of smaller NT areas in NT ultrasound images, thereby improving segmentation performance.
5. The method for segmenting the nuchal translucency layer based on the improved Transformer U-shaped codec structure according to claim 1, characterized in that: The specific step three is: the TransNTSeg network designs a cross-layer feature fusion module based on DDMix-FFNTransformer between the encoder and decoder; The cross-layer feature fusion module can adaptively reshape and splice feature maps from different stages of the encoder so that the extracted NT feature maps have the same size and number of channels; the spliced feature maps are split into multiple scales and sent to different DDMix-FFN modules for feature extraction and cross-scale information exchange; the DDMix-FFN module can effectively extract local information, enhance the expressiveness of features and improve the recognition accuracy of NT regions; finally, the aggregated features are fused with the input features through residual connections to enhance feature representation, and the fused rich feature information is passed to the decoder.
6. The nuchal translucency segmentation method based on the improved Transformer U-shaped codec structure according to claim 1, characterized in that: The step four is specifically as follows: to address the segmentation accuracy and category imbalance problems in the NT ultrasound image segmentation task, a combined loss function is used to optimize the segmentation performance of the TransNTSeg network; and DiceLoss is combined to effectively balance the segmentation accuracy and category imbalance problems.
7. The nuchal translucency segmentation method based on the improved Transformer U-shaped codec structure according to claim 1, characterized in that: In the step 2: construct a NT segmentation network model based on TransNTSeg; the model adopts a U-shaped encoder-decoder structure and introduces an improved Transformer module for accurately segmenting NT ultrasound images; the specific implementation steps are as follows: in terms of the encoder, a hierarchical structure is adopted, consisting of four stages, each stage contains an overlapping block embedding module and an improved Transformer module, the two work together to extract local features and global context information of the NT ultrasound image; wherein, the overlapping block embedding is a convolutional layer with a specific 7×7 convolution kernel, a step size of 4 and 3 pixel padding, which is used to extract the edge and texture local features of the NT area in the NT ultrasound image, and retains local continuity by dividing overlapping pixel blocks, which is conducive to the subsequent capture of the subtle structure of the NT area; then, the feature map output by the convolution layer is converted into a one-dimensional feature vector sequence through a deformation operation, and then the feature vector sequence is mapped to a preset high-dimensional feature space using a linear layer to generate a series of high-order embedding vectors, and the obtained vectors are input into the improved Transformer module for learning the global context information of the NT ultrasound image; the encoder performs a downsampling operation at each stage, while extracting deep semantic information The decoder is symmetrical with the encoder and also adopts a hierarchical structure, consisting of four stages. Each stage contains an overlapping block embedding module for upsampling, which expands the size of the feature map and extracts local features, gradually restoring the detailed information of the NT ultrasound image, and an improved Transformer module to supplement the global dependency. A convolutional block attention module is added before each layer of the Transformer module, which includes a channel attention submodule and a spatial attention submodule to extract features from the channel and spatial dimensions respectively. The channel attention submodule obtains channel statistics of the feature map through global average pooling and global maximum pooling operations, and then uses a multi-layer perceptron to learn the dependencies between channels, thereby adaptively adjusting the channel weights. The spatial attention submodule uses convolution operations based on the channel-weighted feature map to extract spatial weight information, thereby emphasizing key areas and suppressing noise interference. Finally, in the last layer of the decoder, a segmentation head is used to map the feature map output by the decoder into a pixel-level segmentation result. The segmentation head consists of a convolutional layer, a normalization layer, and an activation function component, which is used to convert high-dimensional feature information into a pixel-level segmentation probability map, thereby achieving accurate NT segmentation.
8. The method for segmenting the nuchal translucency layer based on the improved Transformer U-shaped codec structure according to claim 7, characterized in that: One of the core modules of the TransNTSeg network is the improved Transformer module. The Transformer module of TransNTSeg has been improved in the following aspects: using efficient self-attention, given the input feature map X, firstly, the dimensions of the Query (Q), Key (K) and Value (V) feature maps are changed from H×W×C to , where R is the spatial reduction rate, and then a linear projection layer is used to restore the number of channels to C. Finally, the scaled dot product attention mechanism is used to calculate the self-attention weight and weight the Value feature map; the computational complexity is reduced from O(N²) to O( ), applied to the NT ultrasound image feature map; described in formula language: The DDMix-FFN module in TransNTSeg uses dilated depth-separable convolution to replace the convolution operation; for the input feature map x, a fully connected layer FC is first used to map it to a higher-dimensional feature space to obtain = , and then use the dilated depth separable convolution to extract local information to get = ( ), and then perform layer normalization to obtain =LN( ), and then perform residual connection operation to obtain = + , the process is repeated to enhance the expressive power of the features. Finally, the GELU activation function is used for nonlinear transformation, and a fully connected layer FC is used to map the feature map back to the original dimension to obtain the output =FC(GELU(DDMix))+ , described in formula language as: ))) Among them, LN refers to the normalization layer, FC refers to the fully connected layer, and GELU is the activation function; the improved Transformer module contains layer normalization, efficient self-attention, residual connection, layer normalization and DDMix-FFN modules in sequence.
9. The nuchal translucency segmentation method based on the improved Transformer U-shaped codec structure according to claim 1, characterized in that: In the step three: a cross-layer feature fusion module is implemented to effectively utilize the multi-scale features extracted by the encoder and compensate for the loss of segmentation details; a cross-layer feature fusion module based on DDMix-FFN Transformer is designed, and the specific implementation method is as follows: first, feature reshaping is performed, that is, the feature maps from different stages of the encoder are adjusted to a unified channel number C using a convolutional layer; then, feature splicing is performed, that is, the feature maps of the four stages are spliced in the spatial dimension to generate a fused feature map containing multi-scale information; feature information of different scales is refined more finely, that is, the fused feature map processed by the efficient self-attention module is split into four specific scales according to the size of the original feature map, and then sent to different DDMix-FFN modules respectively; the features processed by different DDMix-FFN modules are aggregated, and finally, a residual connection is performed to fuse the aggregated features with the input features; Finally, the feature map processed by the DDMix-FFNTransformer cross-layer feature fusion module is passed to the decoder as the input of the decoder.
10. The method for segmenting the nuchal translucency layer based on the improved Transformer U-shaped codec structure according to claim 1, characterized in that: In step 4, in order to optimize the segmentation performance of the TransNTSeg network for NT ultrasound images and solve the problems of segmentation accuracy and class imbalance, the network binary cross entropy and Dice combined loss function are used. The final loss function is described in formula language as follows: Where Loss is the final combined loss, represents the final predicted segmentation result; y represents the actual segmentation result, and BCE represents the binary cross entropy loss function.