Unet-based longitudinal carotid plaque segmentation method
Through the multi-scale edge enhancement and adaptive feature reconstruction of the EdgeWaveNet network, the problem of insufficient segmentation accuracy in carotid plaque segmentation is solved, efficient segmentation of complex boundaries and small targets is achieved, and segmentation performance is improved.
Patent Information
- Application Number
- CN202510451706.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art has problems such as insufficient segmentation accuracy and weak generalization ability of single-modal ultrasound images in carotid plaque segmentation, especially in complex boundary and small-objective segmentation tasks.
The EdgeWaveNet network is adopted, combining multi-scale edge enhancement module, edge focus attention module and wavelet frequency domain attention module, and improve segmentation performance through multi-scale feature extraction and adaptive feature reconstruction.
The accuracy and robustness of carotid plaque segmentation were significantly improved, especially in complex boundary and small-objective segmentation tasks, with segmentation performance improvements of more than 10%.
Smart Images

Figure CN120374975A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image detection, and particularly relates to a method for segmenting longitudinal carotid artery plaques based on UNet. Background Art
[0002] Carotid artery plaques are important inducements for ischemic stroke, and accurate identification and segmentation thereof are crucial for the prevention and treatment of stroke. At present, the segmentation of carotid artery plaques mainly relies on manual operations by doctors, which is time-consuming and highly subjective. The segmentation results are easily affected by doctors' experience, state, and the contact pressure of the ultrasound probe, resulting in inconsistent manifestations of plaques in the images. With the development of computer science and automation technology, image processing has been widely applied in the field of medical image segmentation. Deep learning-based methods can quickly and accurately segment the target area, thus providing a faster and more objective assistance in the segmentation of carotid artery plaque ultrasound images.
[0003] In recent years, with the rapid development of computer vision and deep learning technologies, many researchers have begun to explore how to achieve automatic segmentation of carotid artery plaques. Multiple methods have been proposed in the prior art to improve the accuracy and efficiency of segmentation. For example, a depth-guided network based on U-Net introduces an image filtering module to restore structural information. However, due to its high dependence on the quality of the guiding image, when there is significant noise in the image, the segmentation accuracy will decrease significantly. For example, the Swin-UNet architecture performs excellently in global feature extraction, but the local characteristics of convolutional operations limit its ability to capture details and small targets. For example, a multi-modal segmentation network that combines ultrasound B-mode images and color Doppler images has achieved remarkable results in multi-modal information fusion, but its generalization ability on single-modal ultrasound images is weak.
[0004] In summary, there are problems of insufficient segmentation accuracy and weak generalization ability on single-modal ultrasound images in the prior art. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method for segmenting longitudinal carotid artery plaques based on UNet to solve the problems existing in the above prior art.
[0006] To achieve the above object, the present invention provides a method for segmenting longitudinal carotid artery plaques based on UNet, including:
[0007] Obtain longitudinal carotid artery ultrasound images, and segment the longitudinal carotid artery ultrasound images through a segmentation model based on the UNet network structure to obtain the segmentation result of carotid artery plaques;
[0008] The segmentation model is the EdgeWaveNet network. The segmentation model includes a multi-scale edge enhancement module for downsampling, an edge focus attention module for upsampling, and a wavelet frequency domain attention module for skip connection between the multi-scale edge enhancement module and the corresponding edge focus attention module. The multi-scale edge enhancement module is used to enhance and extract the edge features of the ultrasonic image. The wavelet frequency domain attention module is used to optimize the features after enhanced extraction. The edge focus attention module is used to reconstruct the features after feature optimization according to the assigned weights and further perform per-pixel annotation to segment the carotid artery plaque.
[0009] Optionally, the segmentation model further includes an initial convolutional layer connected to the multi-scale edge enhancement module. The initial convolutional layer is used to perform initial feature extraction on the input ultrasonic image to meet the input requirements of the multi-scale edge enhancement module. After the edge focus attention module, feature segmentation is performed through a convolutional layer and a sigmoid function to obtain the segmentation result.
[0010] Optionally, several layers of the multi-scale edge enhancement module are connected in sequence, and several layers of the edge focus attention module are connected in sequence. Among the several layers of the multi-scale edge enhancement module, the last layer of the multi-scale edge enhancement module is connected to the first layer of the edge focus attention module.
[0011] Optionally, in the multi-scale edge enhancement module, the process of enhancing and extracting the edge features of the ultrasonic image includes:
[0012] The feature map of the ultrasonic image is downsampled through max pooling and average pooling. Feature extraction is performed on the fused downsampled structure through PSConv. The extracted features are normalized and processed by an activation function to obtain the processed features. The processed features are weighted through a spatial attention mechanism to obtain spatial attention features. The spatial attention features are weighted through a channel attention mechanism to obtain channel attention features. The spatial attention features, channel attention features, and processed features are fused, and the feature map is fused with the fused result to obtain the enhanced extracted features, which are the output results of the multi-scale edge enhancement module.
[0013] Optionally, in the wavelet frequency domain attention module, the process of optimizing the features after enhanced extraction includes:
[0014] The features after enhanced extraction are subjected to wavelet transform to obtain high-frequency features and low-frequency features. The high-frequency features are processed through a Manhattan self-attention mechanism, convolution, normalization, and an activation function to obtain the activated high-frequency features. An interpolation operation is performed on the activated high-frequency features to obtain the processed high-frequency features;
[0015] The low-frequency features are processed through a linear embedding layer and a Transformer structure, and the processed low-frequency features and the processed high-frequency features are fused and reconstructed to obtain a reconstructed signal, that is, the features after feature optimization. The reconstructed signal is transmitted to the edge-focus attention module through a skip connection.
[0016] Optionally, in the edge-focus attention module, the process of feature reconstruction for the features after feature optimization according to the assigned weights includes:
[0017] The features after feature optimization are upsampled through a transposed convolution to obtain a high-resolution feature map. The high-resolution feature map is concatenated with the output features of the previous layer and feature extraction is performed to obtain local features. The local features are subjected to feature extraction through local convolution, and the local features are subjected to weight assignment processing through global convolution to obtain local information and global information. The local information and global information are concatenated and feature extraction is performed to obtain the feature reconstruction result.
[0018] Optionally, the process of feature extraction through local convolution includes: performing convolution, normalization, and activation function processing on the local features to obtain local information.
[0019] Optionally, the process of weight assignment processing for the local features through global convolution includes:
[0020] Performing convolution, normalization, and activation function processing on the local features several times, and performing convolution processing on the results of the activation function processing again to obtain the first branch information; performing normalization processing on the local features, and performing weight assignment processing on the results of the normalization processing through the SS2DwithSCAM module, and regularizing the results of the SS2DwithSCAM module processing. The regularized results are fused with the results of the SS2DwithSCAM module processing to obtain the second branch information. The first branch information and the second branch information are subjected to convolution processing to obtain global information. The SS2DwithSCAM module is an SS2D module with a spatial channel attention module added.
[0021] Optionally, the process of processing through the SS2DwithSCAM module includes:
[0022] Performing global modeling on the results of the normalization processing through the SS2D module to obtain the modeling results. The modeling results are processed through the spatial channel attention module, and further weight assignment and transformation are performed on the results of the spatial channel attention module processing to obtain the results of the SS2DwithSCAM module processing;
[0023] Among them, the process of processing through the spatial-channel attention module includes: calculating the spatial weight and channel weight of the modeling result through the spatial attention mechanism and the channel attention mechanism respectively, and multiplying the modeling result with the spatial weight and the channel weight element by element to obtain the processing result of the spatial-channel attention module.
[0024] Compared with the prior art, the present invention has the following advantages and technical effects:
[0025] The present invention proposes an ultrasound image segmentation method based on hybrid edge-aware attention and multi-scale feature enhancement. By introducing a wavelet frequency-domain attention module, a multi-scale edge enhancement module, and an edge-focus attention module, the deficiencies of traditional UNet in boundary blurring and noise interference are effectively improved. This method has achieved significant improvement in segmentation performance on the self-built carotid plaque dataset, Breast Ultrasound Images Dataset, and DDTI thyroid ultrasound image dataset, and is particularly excellent in complex boundary and small target segmentation tasks.
[0026] Specifically, the wavelet frequency-domain attention module significantly enhances the significant edge information in the image through frequency-domain feature processing, thereby improving the segmentation accuracy of low-quality images; the multi-scale edge enhancement module effectively captures complex boundary information through upsampling operations and the fusion of multi-scale features, enhancing the model's segmentation ability for details; the edge-focus attention module realizes the efficient expression of multi-scale information through downsampling operations and the adaptive integration of wavelet features, further optimizing the segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0028] Figure 1 is the overall structure diagram of the EdgeWaveNet network according to the embodiment of the present invention;
[0029] Figure 2 is the structure diagram of the WFA module according to the embodiment of the present invention;
[0030] Figure 3 is the structure diagram of the MSEE module according to the embodiment of the present invention;
[0031] Figure 4 is the structure diagram of the EFA module according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0033] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0034] Aiming at the deficiencies of traditional ultrasonic image segmentation methods in dealing with blurred boundaries, noise interference, and multi-scale feature extraction, the present invention proposes a lightweight ultrasonic image segmentation network EdgeWaveNet. The network first realizes frequency-domain decomposition through a wavelet frequency-domain attention module (WFA), effectively extracts high-frequency edge information and low-frequency structure information, and combines an attention mechanism to enhance feature expression; secondly, a multi-scale edge enhancement module (MSEE) is adopted, combined with a combined structure of MaxPool, AvgPool, and PSConv, and cooperated with an asymmetric padding strategy to achieve precise capture of blurred edges, and at the same time uses a spatial-channel dual attention mechanism to optimize the feature extraction efficiency; finally, an edge focus attention module (EFA) is introduced, which innovatively combines ConvTranspose with local-global feature fusion, significantly improving the quality of edge detail reconstruction. Experimental results show that on three ultrasonic image datasets of thyroid, breast cancer, and carotid artery plaque, the mIoU index of EdgeWaveNet has increased by 11.63%, 8.86%, and 2.55% respectively compared with existing methods. In particular, the network improves the segmentation performance through an efficient module design and demonstrates excellent robustness in challenging tasks such as complex noise environments and small target detection.
[0035] The present invention proposes innovative solutions from three aspects of feature extraction, noise suppression, and computational efficiency for the problems existing in the above method:
[0036] (1) Design an edge focus attention module (Edge Focus Attention Block, EFA). This module innovatively combines ConvTranspose with a local-global feature fusion mechanism, and significantly enhances the model's ability to reconstruct edge details through adaptive feature weight assignment. Experiments show that this module significantly improves the segmentation accuracy in complex boundary and small target segmentation tasks, and the average mIoU on the three datasets has increased by more than 10%.
[0037] (2) A Wavelet-Frequency Attention Block (WFA) is proposed. Compared with the traditional Fourier transform, the wavelet transform has unique advantages in time-frequency localization and can more accurately separate high-frequency edge information (such as tissue boundaries) and low-frequency structural information (such as background noise). Combining the channel-space dual attention mechanism, the performance degradation of this module in a noisy environment is controlled within 5%, which is significantly better than existing methods.
[0038] (3) A Multi-Scale Edge Enhancement Block (MSEE) is constructed. This module combines MaxPool (strengthening the edge mutation response), AvgPool (preserving texture details), and PSConv (asymmetric padding strategy to adapt to the ultrasound edge diffusion pattern), and enhances the extraction effect of edge features through the space-channel dual attention mechanism. Experiments show that the introduction of this module significantly improves the segmentation performance of the model on three datasets of carotid artery plaques, breast, and thyroid. When dealing with fuzzy boundaries and small targets, the mIoU is improved by more than 8% compared with the baseline model.
[0039] Regarding the above technical content, the technical solution of the present invention will be described in detail:
[0040] The technical solution of the present invention mainly provides a method for longitudinal carotid artery plaque segmentation of an EdgeWaveNet network based on the unet network structure. The longitudinal carotid artery ultrasound images are recognized by the EdgeWaveNet network to segment the carotid artery plaques.
[0041] The EdgeWaveNet network, as a segmentation model, includes a multi-scale edge enhancement module for downsampling, an edge focus attention module for upsampling, and a wavelet-frequency attention module for skip connection between the multi-scale edge enhancement module and the corresponding edge focus attention module.
[0042] The overall network architecture of the EdgeWaveNet proposed by the present invention is as Figure 1As shown, it includes an initial convolution and an Edge Focus Attention Block (EFA), a Wavelet-Frequency Attention Block (WFA), and a Multi-Scale Edge Enhancement Block (MSEE). The network first extracts features from the input image (224×224×3) through the initial convolution and the downsampling module. The MSEE module enhances the extraction effect of the edge features of the ultrasound image by combining MaxPool (strengthening the edge mutation response) and AvgPool (retaining texture details) for parallel processing, and cooperating with the asymmetric padding strategy of PSConv. The WFA module utilizes the time-frequency localization advantage of wavelet transform to accurately separate high-frequency edge information and low-frequency structure information, and optimizes features through a channel-spatial dual attention mechanism to improve the ability to capture multi-scale information. The EFA module innovatively combines ConvTranspose with a local-global feature fusion mechanism, and enhances the model's ability to reconstruct edge details through adaptive feature weight allocation, which is especially suitable for processing segmentation tasks of complex boundaries and small targets. Finally, the spatial resolution is gradually restored through the upsampling module to generate the carotid plaque segmentation result. The loss function L adopts the Dice loss L Dice combined with the cross-entropy loss L CE : L = 0.5L Dice + 0.5L CE .
[0043] WFA module (Wavelet-Frequency Attention Block)
[0044] The edge information of the anatomical structure in the ultrasound image is mainly contained in the high-frequency components, while the low-frequency components contain large-scale tissue features and noise. Traditional spatial domain convolution is prone to spectral confusion when separating high-frequency edges and low-frequency noise, which limits the segmentation accuracy. To solve this problem, the present invention proposes a frequency processing module (WFA), the core idea of which is to separately process high-frequency and low-frequency components through frequency domain decomposition to improve the segmentation performance.
[0045] First, WFA uses the discrete wavelet transform (DWT) to decompose the input feature map into high-frequency and low-frequency components, initially suppressing the noise in the ultrasound image. Considering the real-time requirements of ultrasound image processing and the importance of preserving edge features, the present invention selects the Haar wavelet basis as the basis function of DWT. The Haar wavelet basis has the characteristics of high computational efficiency and simple structure. It can not only effectively retain the edge jump features of the ultrasound image but is also particularly suitable for dealing with speckle noise in the ultrasound image. In the specific implementation, a four-layer decomposition structure is adopted, and the'reflect' mode is used for boundary processing, effectively extracting multi-scale features while ensuring computational efficiency. The high-frequency components contain edge detail information, while the low-frequency components retain the large-scale structural features.
[0046] For the high-frequency components, WFA uses a Transformer structure for feature enhancement and combines the Manhattan Self-Attention (MaSA) mechanism for secondary processing. Since MaSA can effectively suppress the interference of high-frequency noise while retaining key edge details by modeling spatial relationships. For the low-frequency components, WFA focuses on retaining the local structural information contained in them and only performs preliminary denoising to reduce computational complexity.
[0047] Finally, WFA fuses the processed high-frequency and low-frequency components with the original feature map and passes the information processed by WFA at each layer to the decoder through skip connections similar to those in U-Net. As Figure 2 shown, this inter-layer information transfer mechanism and the structural design of WFA jointly promote the effective utilization of features and the improvement of segmentation performance, while reducing the interference of noise in the ultrasound image on the model performance.
[0048] Compared with traditional spatial domain convolution, the WFA module can effectively avoid spectral aliasing, thus improving the segmentation accuracy. When traditional spatial convolution processes ultrasound images, since it cannot effectively separate frequency domain information, the high-frequency edge features and low-frequency structural noise overlap in the spatial domain. Especially during the downsampling process, high-frequency information will be aliased into the low-frequency region, generating false signals. This spectral aliasing will cause the edge region to become blurred. During the upsampling process, due to the spectral aliasing caused by the auxiliary information provided by the downsampling, it will affect the accurate positioning of the carotid plaque boundary, resulting in an increase in measurement error and a reduction in the segmentation result. However, the WFA module can accurately separate frequency components at different scales through the multi-resolution analysis of wavelet transform, effectively avoiding this aliasing problem and improving the segmentation accuracy. Experimental results show that the WFA module performs better than existing methods on public datasets, especially having significant advantages in dealing with complex edges and background noise.
[0049] A detailed description of the data processing flow of the above WFA module is as follows:
[0050] First, the input feature map F in ∈R H×W×C is decomposed by the Haar wavelet transform (using the'reflect' mode), which can be expressed as obtaining a low-frequency component and three high-frequency components.
[0051] LL, LH, HL, HH = DWT(F in ) (1)
[0052] where LL is the low-frequency component, and LH (horizontal high-frequency), HL (vertical high-frequency), and HH (diagonal high-frequency) are the three high-frequency components. These three high-frequency components are combined into a unified high-frequency feature layer:
[0053] F low = LL (2)
[0054] F high = Concat(LH, HL, HH) (3)
[0055] In this way, the input feature map is decomposed into a low-frequency feature F low and a combined high-frequency feature F high , preparing for subsequent processing.
[0056] The high-frequency feature first passes through the Manhattan Self-Attention (MaSA) module and then through a 3×3 convolution operation:
[0057] (I high , I low ) = WaveletDecomp(F low ) (4)
[0058] H processed = Conv(MaSA(I high )) (5)
[0059] Next, the processed high-frequency information passes through batch normalization BatchNorm and the activation function GELU:
[0060] H norm = BatchNorm(H processed ) (6)
[0061] H activated = GELU(H norm ) (7)
[0062] The interpolating operation Interp is performed on the high-frequency part Hactivated after activation:
[0063] H interpolated = Interp(H activated ) (8)
[0064] The low-frequency part is processed through a linear embedding layer Embl and a Transformer structure:
[0065] L processed = Transformer(Emb(I low )) (9)
[0066] The reconstructed module is used to reconstruct the fused features and is passed to the decoder through skip connections, expressed as:
[0067] I reconstructed = Reconstruction(L processed ) (10)
[0068] Finally, the reconstructed module is used to reconstruct the fused feature I fused for reconstruction Reconstruction and is passed to the decoder through skip connections:
[0069] I fused = Concat(H interpolated ,I reconstructed ) (11)
[0070] Through this process, the high-frequency and low-frequency information of the input image is processed, fused, and reconstructed through a series of operations, thereby obtaining the final output features.
[0071] MSEE module (Multi-Scale Edge EnhancementBlock)
[0072] In the downsampling module of the ultrasound image, this module combines MaxPool, AvgPool, and PSConv (PartialSeparable Convolution) as well as spatial attention and channel attention mechanisms, aiming to improve the performance of ultrasound image segmentation. Through multi-scale feature extraction and attention mechanisms, the module can better process complex boundaries and detailed regions in ultrasound images, improving segmentation accuracy and robustness.
[0073] First, MaxPool enhances the edge mutation response in the image through max-pooling operation, strengthening the ability to extract edge features; while AvgPool preserves tissue texture information through average-pooling operation, capturing more detailed features and reducing information loss. The combination of the two can enhance the extraction of edge features while retaining details. Second, PSConv, through the design of asymmetric padding and multi-directional diffusion convolutional kernels, has a receptive field distribution that highly matches the spatial characteristics of ultrasonic targets. Specifically, asymmetric padding uses different padding methods for different regions, enhancing the ability to extract local features; the multi-directional diffusion convolutional kernel captures the diffusion pattern of blurred edges through the diffusion of convolutional kernels in the horizontal and vertical directions, which is particularly suitable for processing Gaussian-distributed edges in ultrasonic images. Compared with standard convolution, PSConv can more efficiently capture the diffusion pattern of blurred edges in the underlying feature extraction stage, improving the feature expression ability.
[0074] In addition, the introduction of spatial attention and channel attention mechanisms further optimizes the feature extraction process. Spatial attention weights the spatial dimension of the feature map to enhance the feature expression of important regions; channel attention weights the channel dimension of the feature map to enhance the feature expression of important channels. The outputs of spatial attention and channel attention are fused to obtain the final feature map, thereby further enhancing the model's attention to key regions and improving the segmentation accuracy. It should be noted that the introduction of PSConv complements the traditional attention mechanism: the traditional attention mechanism mainly enhances the feature expression of important regions by weighting the spatial and channel dimensions of the feature map, while PSConv, through the design of asymmetric padding and multi-directional diffusion convolutional kernels, more efficiently captures the blurred edges and local details in ultrasonic images. This combination enables the model to perform better when dealing with complex backgrounds and local information limitations, further improving the segmentation accuracy and robustness.
[0075] In summary, this downsampling module effectively improves the performance of ultrasonic image segmentation by combining MaxPool, AvgPool, PSConv, and spatial attention and channel attention mechanisms, especially performing excellently when dealing with local information limitations and complex backgrounds. Through multi-scale feature extraction and attention mechanisms, the module can better handle complex boundaries and detail regions in ultrasonic images, improving the segmentation accuracy and robustness. The structure diagram of the MSEE module is as Figure 3 shown.
[0076] The specific data processing flow of this module is as follows:
[0077] The input feature map is first downsampled by MaxPool and AvgPool respectively for the input feature map x:
[0078] X max = MaxPool(X), Xavg = AvgPool(X) (12)
[0079] Where X max , X avg ∈ R C×2H×2W , X max , X avg represent the corresponding processed feature maps;
[0080] PSConv extracts features through asymmetric padding and multi-directional diffusion convolutional kernels:
[0081] X psconv = PSConv(X max ⊕ X avg , k = 3) (13)
[0082] Where ⊕ represents the feature concatenation operation and k = 3 is the convolutional kernel size.
[0083] Perform batch normalization and activation on the output X psconv of PSConv:
[0084] X bn = BatchNorm(X psconv ), X relu = ReLU(X bn ) (14)
[0085] X psconv = PSConv(X max ⊕ X avg , k = 3) (15)
[0086] A spatial = Sigmoid(PSConv(Concat(AdaptiveMaxPool(X psconv ), AdaptiveAvgPool(X psconv ))), k = 7) (16)
[0087] Where AdaptiveMaxPool is the adaptive max pooling layer, AdaptiveAvgPool is the adaptive average pooling layer, and A spatial represents the spatial attention weight;
[0088] Spatial attention weights the spatial dimension of the feature map:
[0089]
[0090] Where represents element-wise multiplication.
[0091] Channel attention weights the channel dimension of the feature map:
[0092] A channel = Sigmoid(Concat(Conv(ReLU(Conv(AdaptiveMaxPool(X psconv) ), k = 1), k = 1), Conv(ReLU((AdaptiveAvgPool(X psconv ), k = 1)), k = 1))) (17)
[0093]
[0094] Where, A channel represents the channel attention weight, X spatial represents the spatial attention output feature, X channel l represents the channel attention output feature.
[0095] If the number of channels of the input and output is the same (i.e., judge In_channels == Out_channels), a residual connection is added, and at the same time, the outputs of spatial attention and channel attention are combined to obtain the final feature map X out :
[0096] X o = X shortcut + X channel (19)
[0097] X out = Convs(X o ) (20)
[0098] Where X shortcut is the result of the added residual connection, and is the result of the original input processed by PSconv and conv
[0099] This module realizes efficient feature extraction and downsampling by combining MaxPool, AvgPool, PSConv, spatial attention and channel attention mechanisms.
[0100] EFA module (Edge Focus Attention Block)
[0101] Ultrasound images usually have low resolution and less detailed information. Directly using traditional interpolation methods (such as bilinear interpolation) will lead to information loss or image blurring. In addition, the tissue structure in ultrasound images is complex, and local information and global context information are crucial for accurately understanding the image content. To solve these problems, such as Figure 4As shown, in the upsampling process, this module uses ConvTranspose (transposed convolution) to replace the traditional interpolation method to reduce information loss. ConvTranspose can generate a smoother high-resolution image, effectively addressing the low-resolution problem of ultrasound images.
[0102] Meanwhile, this module achieves a more complete upsampling by fusing local information (ConvPath) and global information (ConvSSM). Specifically, local information captures local features (such as edges, textures, etc.) in the image through convolution operations, while ConvSSM helps the model understand the complex relationships between different regions in the image by modeling global context information, which is particularly important for identifying lesion regions.
[0103] In ConvSSM, this module introduces an improved SS2DwithSCAM structure. Based on SS2D, this structure adds SCAM (Spatial Channel Attention Module) after each convolutional layer. SCAM enhances the model's understanding of global context information in each patch through spatial and channel attention mechanisms, thus more accurately capturing the relationship between the background and the target. In addition, SCAM can effectively suppress noise interference in ultrasound images and improve the model's recognition ability for complex structures, further enhancing the upsampling effect.
[0104] The specific data processing flow of this module is as follows:
[0105] Let the input low-resolution ultrasound image be X ∈ R H×W×C , where H, W, and C represent the height, width, and number of channels of the image respectively.
[0106] Use ConvTranspose to upsample the input image to generate a high-resolution feature map F up :
[0107] F up = ConvTranspose(X; θ up ), F up ∈ R 2H×2W×C (19)
[0108] where θ up represents the parameters of ConvTranspose, and C is the number of channels after upsampling.
[0109] Extract local features F ConvPath ∈ R 2H×2W×C
[0110] F ConvPath = Conv(F up ; θ conv ) (20)
[0111] Among them, θ conv represents the parameters of the convolutional layer.
[0112] For the above local features, they are processed in parallel through the local convolution ConvPath and the global convolution ConvSSM respectively. Among them, in the local convolution ConvPath, it is set to be processed sequentially through several layers of Conv, BN, and ReLu to obtain local information. In the global convolution ConvSSM, the local features are sequentially processed through Conv, BN, ReLu, Conv, BN, ReLu, and Conv to form the first branch feature. The local features are also sequentially processed through BN, the SS2DwithSCAM module, and a drop path judgment is made to determine whether to skip this process to obtain the second branch. The first branch and the second branch are processed through CONV to obtain global information. The global information is concatenated with the features processed by the local convolution through concat, and after the concatenation, it is processed through CONV to obtain the final output of this module.
[0113] In ConvSSM, the SS2DwithSCAM module is used to extract the global context information F global ∈R 2H×2W×C . The SS2D module performs global modeling on the input feature map F up :
[0114] F SS2D = SS2D(F up ; θ SS2D ) (21)
[0115] Among them, F SS2D represents the features processed by the SS2D module. In the SS2D module, linear processing, conv, and Silu are sequentially performed and the residual connection of expansion and convolution of another branch is passed through;
[0116] On the basis of SS2D, SCAM (Spatial Channel Attention Module) is added to enhance the output of each convolutional layer:
[0117] F SCAM = SCAM(F SS2D ; θ SCAM )(22)
[0118] Among them, θ SCAM represents the parameters of SCAM.
[0119] The calculation process of SCAM includes: spatial attention: calculating the spatial weight W spatial :
[0120] W spatial = Sigmoid(Conv(FSS2D ; θ spatial ))(23),
[0121] Channel attention: Calculate the channel weight W channel :
[0122] W channel = Sigmoid(Conv(F SS2D ; θ channel ))(24)
[0123] Feature enhancement: Apply the spatial and channel weights to the feature map:
[0124]
[0125] where represents element-wise multiplication.
[0126] And through the forward function forword core, merging, normalization, linear processing of SS2D in mamba, and finally obtaining the output features of the final SS2DwithSCAM module through dropout processing.
[0127] Fuse the local information F local and the global information F global to generate the final high-resolution feature map F final ∈ R 2H×2W×C :
[0128] F final = F local + F global (26)
[0129] Perform pixel-level annotation on the final EFA module to achieve the final segmentation of carotid artery plaques.
[0130] In addition, to enhance the generalization of the model, we also used publicly available ultrasound image datasets, including the malignant dataset part in BreastUltrasound Images Dataset and the samples in DDTI:Thyroid UltrasoundImages dataset.
[0131] Perform relevant experimental settings and result explanations on the above content:
[0132] Dataset
[0133] There is no publicly available standard dataset in the field of carotid plaque segmentation. In this study, ultrasound images from a hospital were collected and a new dataset was created. This study has been approved by the Ethics Committee of the Affiliated Hospital of Inner Mongolia Medical University. All data collection and use follow the hospital's data security management system, and patient information is de-identified. For ease of model training, the original ultrasound images collected were cropped to RGB images of 224x224. The cropped images were annotated using the open-source tool Roboflow.
[0134] Training Environment and Training Parameters
[0135] The experiment was implemented using Python under the ubuntu system based on the Pytorch 1.8 framework. When training, the initial learning rate was set to 0.0001, the batch size was set to 8, and 100 epochs were trained. Under this setting, accelerated training was completed using NVIDIA GeForce RTX3090, and mixed precision was used for accelerated training. The dataset was randomly divided into a training set, a test set, and a validation set in the ratio of 8:1:1. The loss function used was 0.5 Dice + 0.5 Cross Entropy
[0136] Performance Metrics
[0137] To evaluate the performance of the proposed method in the ultrasound image segmentation task, several commonly used evaluation metrics were adopted, including the mean intersection over union (mIoU), Dice coefficient (mDice), recall, and precision. These metrics help to comprehensively measure the performance of the model in handling segmentation tasks, especially in complex boundaries and small object detection.
[0138] Mean Intersection over Union (mIoU)
[0139] The mean intersection over union (mIoU) is a metric that measures the degree of overlap between the model's prediction results and the ground truth labels. For each class, the formula for calculating the intersection over union (IoU) is as follows:
[0140] IoU = TP / (TP + FP + FN) (27)
[0141] Where TP, FP, and FN represent the number of samples that are truly positive and predicted to be positive, the number of samples that are predicted to be positive but are truly negative, and the number of samples that are truly positive but predicted to be negative, respectively. The mIoU metric evaluates the overall segmentation performance of the model by averaging the IoU scores of all classes. For binary segmentation tasks, mIoU is equivalent to IoU.
[0142] Dice Coefficient (mDice)
[0143] The Dice coefficient (mDice, Mean Dice Coefficient) is another commonly used metric for measuring the overlap between the predicted results and the ground truth labels. Its formula is as follows:
[0144] Dice = 2×TP / (2×TP + FP + FN) (28)
[0145] The Dice coefficient measures the similarity between the two by comparing twice the overlapping area with the sum of the predicted and ground truth results. For multi-class tasks, mDice is the average of the Dice coefficients for each class and is used to evaluate the performance of the model in the overall segmentation task.
[0146] Recall
[0147] Recall reflects the model's ability to identify positive samples, and its calculation formula is:
[0148] Recall = TP / (TP + FN) (29)
[0149] Recall measures the proportion of positive samples that the model can correctly identify among all positive samples. In the ultrasound image segmentation task, a higher Recall indicates that the model can capture more target regions, especially in cases where the boundaries are blurred or the targets are small.
[0150] Precision
[0151] Precision is used to measure the accuracy of the model's predictions of positive samples, and its calculation formula is:
[0152] Precision = TP / (TP + FP) (30)
[0153] Precision reflects the proportion of samples that are actually positive among the samples predicted as positive by the model. A higher Precision indicates that the model's predictions of the target regions are more accurate and can effectively reduce the occurrence of false positives.
[0154] Comparative experiments and visualization
[0155] To verify the effectiveness of the proposed method in the ultrasound image segmentation task, we conducted experiments on the self-built carotid plaque dataset, the malignant subset of the Breast Ultrasound Images Dataset, and the DDTI:Thyroid Ultrasound Images dataset, and compared the performance with classical segmentation networks, including UNet, UNet++, UCTransNet, Attention UNet, ResUNet, MutiresUNet, and TransNet, etc.
[0156] Performance on the carotid plaque dataset
[0157] To verify the performance of the proposed method in the ultrasound image segmentation task, experiments were conducted on the self-built carotid plaque dataset and compared with various classical and state-of-the-art segmentation networks. The experimental results show that the proposed method exhibits significant advantages in multiple key metrics. Specifically, the proposed method achieved the best performance in mIoU (61.38%), mDice (76.07%), Recall (70.15%), and Precision (83.07%), which are 2.55%, 1.99%, 4.90%, and -2.58% higher than the second-best method PDF-UNet (2023), respectively.
[0158] In the comparison method, in Table 1, PDF-UNet and Attunet performed outstandingly, approaching the proposed method in mIoU and Precision respectively. However, although the Recall and Precision of PDF-UNet are relatively high, its mIoU and mDice are still lower than those of the proposed method, indicating that there is still room for improvement in its overall segmentation performance. Attunet performed best in Precision (85.04%), but its Recall (61.71%) and mIoU (55.66%) are relatively low, indicating its insufficient ability to recognize some targets. Other methods such as Net, Resunet, and DCSAU-UNet performed acceptably in some metrics, but there is still a large gap between their overall performance and the proposed method. For example, Net has a relatively high Precision (76.60%), but its mIoU (47.62%) and Recall (55.73%) are relatively low, indicating that its segmentation results are not comprehensive enough. Mutiresunet and UltraLight_VMUnet(2024) performed the worst in all metrics, indicating that their model designs may not be suitable for the current task. In addition, VMUnet(2024) performed acceptably in Precision (78.75%), but its mIoU (45.09%) and Recall (51.34%) are relatively low, indicating that its overall segmentation performance still needs to be improved. UNetv2 performed poorly in all metrics, especially its mIoU (36.94%) and Recall (41.99%) are significantly lower than other methods, indicating that its model design may not be suitable for the current task or its generalization ability on the dataset is poor.
[0159] In summary, the proposed method shows excellent performance in the ultrasonic image segmentation task, especially having significant advantages in mIoU and Recall, reflecting its comprehensive recognition ability of the target area. At the same time, the proposed method also maintains a high level in Precision, indicating the accuracy of its segmentation results. These results show that the proposed method has high practicality and promotion value in the ultrasonic image segmentation task. Table 1 shows the comparison test results on the carotid plaque dataset.
[0160] Table 1
[0161]
[0162] Performance on the malignant breast ultrasound dataset
[0163] To further verify the performance of the proposed method in breast ultrasound image segmentation tasks, we conducted experiments on a malignant breast ultrasound image dataset and compared it with various classical and state-of-the-art segmentation networks. The experimental results show that the proposed method exhibits significant advantages in multiple key metrics. Specifically, the proposed method achieved the best performance in mIoU (55.15%), mDice (71.09%), Recall (64.55%), and Precision (79.11%), which are 8.86%, 7.80%, 6.28%, and 6.59% higher than the second-best method, DCSAU-UNet (2023), respectively.
[0164] Among the comparison methods, as shown in Table 2, DCSAU-UNet and Attunet performed relatively prominently, approaching the proposed method in mIoU and Precision respectively. However, although the Recall (58.27%) and Precision (69.25%) of DCSAU-UNet are relatively high, its mIoU (46.29%) and mDice (63.29%) are still lower than those of the proposed method, indicating that there is still room for improvement in its overall segmentation performance. Attunet performed well in Precision (71.86%), but its Recall (54.99%) and mIoU (45.25%) are relatively low, indicating its insufficient ability to identify some targets. Other methods such as UNet, Resunet, and Mutiresunet performed acceptably in some metrics, but there is still a large gap between their overall performance and the proposed method. For example, UNet has a relatively high Recall (60.84%), but its mIoU (39.75%) and Precision (53.43%) are relatively low, indicating that its segmentation results are not comprehensive enough. Mutiresunet and UltraLight_VMUnet (2024) performed poorly in all metrics, indicating that their model designs may not be suitable for the current task.
[0165] In summary, the proposed method exhibits excellent performance in breast ultrasound image segmentation tasks, especially having significant advantages in mIoU and Recall, which reflects its comprehensive recognition ability for target regions. At the same time, the proposed method also maintains a relatively high level in Precision, indicating the accuracy of its segmentation results. These results show that the proposed method has high practicality and promotion value in breast ultrasound image segmentation tasks. Table 2 shows the comparison test results on the malignant breast ultrasound dataset.
[0166] Table 2
[0167]
[0168]
[0169] Performance on Thyroid Ultrasound Dataset
[0170] To verify the performance of the proposed method on the DDTI dataset in terms of generalization ability, we conducted experiments on this dataset and compared it with multiple classic and state-of-the-art segmentation networks. The experimental results show that the proposed method exhibits significant advantages in multiple key metrics. Specifically, the proposed method achieved the best performance in mIoU (49.17%), mDice (65.93%), Recall (60.79%), and Precision (72.01%), which are 11.63%, 11.35%, 15.47%, and 3.41% higher than the second-best method, PDFUNet, respectively.
[0171] Among the comparison methods, in Table 3, PDFUNet and DCSAU-UNet (2023) performed relatively prominently, being close to the proposed method in mIoU and mDice respectively. However, although the Recall (45.32%) and Precision (68.60%) of PDFUNet are relatively high, its mIoU (37.54%) and mDice (54.58%) are still lower than the proposed method, indicating that there is still room for improvement in its overall segmentation performance. DCSAU-UNet (2023) performed well in mDice (52.86%) and Recall (45.16%), but its mIoU (35.93%) and Precision (63.74%) are relatively low, indicating its insufficient ability to recognize some targets. Other methods such as UNet, Attunet, and Mutiresunet performed okay in some metrics, but there is still a large gap between their overall performance and the proposed method. For example, Attunet has a relatively high Precision (72.47%), but its Recall (34.06%) and mIoU (30.16%) are relatively low, indicating that its segmentation results are not comprehensive enough. Transnet and UltraLight_VMUnet (2024) performed poorly in all metrics, indicating that their model designs may not be suitable for the current task.
[0172] In summary, the proposed method exhibits excellent performance on the DDTI dataset, especially having significant advantages in mIoU and Recall, reflecting its comprehensive recognition ability of target regions. At the same time, the proposed method also maintains a relatively high level in Precision, indicating the accuracy of its segmentation results. These results show that the proposed method has high practicality and promotion value in the DDTI dataset segmentation task.
[0173] Table 3
[0174]
[0175] Ablation Experiment
[0176] To verify the effectiveness of each module, ablation experiments were conducted on three datasets (Mydata, DDTI, and Malignant). From Tables 4, 5, and 6, on the Mydata dataset, DDTI dataset, and Malignant dataset, after adding the WFA module to the basic model, the mIoU and mDice were respectively increased to 54.91% and 70.89%, proving that the WFA module significantly enhanced the feature extraction ability of the model. After further adding the EFA module, the mIoU and mDice were respectively increased to 60.72% and 75.56%, proving that the EFA module plays an important role in multi-scale feature fusion. Finally, after adding the MSEE module, the mIoU and mDice reached 61.38% and 76.07% respectively, and the Rec. was increased to 70.15%, indicating that the MSEE module further optimized the feature weight allocation. The effectiveness of the EFA, WFA, and MSEE modules was verified through ablation experiments. The gradual addition of each module significantly improved the model performance, reflecting the generalization and reliability of the proposed method. Table 4 is the ablation experiment on the carotid artery plaque ultrasound dataset, Table 5 is the ablation experiment on the malignant breast ultrasound dataset, and Table 6 is the ablation experiment on the thyroid ultrasound dataset.
[0177] Table 4
[0178]
[0179] Table 5
[0180]
[0181] Table 6
[0182]
[0183] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. The longitudinal carotid plaque segmentation method based on UNet, characterized in that, Including: Obtain longitudinal carotid artery ultrasound images, and segment the longitudinal carotid artery ultrasound images through a segmentation model based on the unet network structure to obtain carotid plaque segmentation results; The segmentation model is the EdgeWaveNet network. The segmentation model includes a multi-scale edge enhancement module for downsampling, an edge focusing attention module for upsampling, and a wavelet frequency domain attention module for skip connection between the multi-scale edge enhancement module and the corresponding edge focusing attention module. The multi-scale edge enhancement module is used to enhance and extract the edge features of the ultrasound image, the wavelet frequency domain attention module is used to optimize the enhanced and extracted features, and the edge focusing attention module is used to reconstruct the features optimized by the features according to the assigned weights and further perform per-pixel annotation to segment the carotid plaque.
2. The method according to claim 1, wherein In the segmentation model, an initial convolutional layer connected to the multi-scale edge enhancement module is further included, and the initial convolutional layer is used to perform initial feature extraction on the input ultrasound image to meet the input requirements of the multi-scale edge enhancement module. After the edge focusing attention module, the features are segmented through a convolutional layer and a sigmoid function to obtain a segmentation result.
3. The method according to claim 1, wherein A plurality of layers of the multi-scale edge enhancement modules are sequentially connected, a plurality of layers of the edge focusing attention modules are sequentially connected, and among the plurality of layers of the multi-scale edge enhancement modules, the last multi-scale edge enhancement module is connected to the first edge focusing attention module.
4. The method according to claim 1, wherein In the multi-scale edge enhancement module, the process of enhancing and extracting the edge features of the ultrasound image includes: Downsample the feature map of the ultrasound image through max pooling and average pooling, extract features from the fused downsampling structure through PSConv, perform normalization and activation function processing on the extracted features to obtain processed features, and weight the processed features through a spatial attention mechanism to obtain spatial attention features. Weight the spatial attention features through a channel attention mechanism to obtain channel attention features, fuse the spatial attention features, channel attention features, and processed features, and fuse the feature map with the fused result to obtain the enhanced and extracted features, which are the output results of the multi-scale edge enhancement module.
5. The method according to claim 1, wherein In the wavelet frequency domain attention module, the process of optimizing the enhanced and extracted features includes: Perform wavelet transform on the enhanced and extracted features to obtain high-frequency features and low-frequency features. The high-frequency features are processed through a Manhattan self-attention mechanism, convolution, normalization, and activation function to obtain activated high-frequency features, and interpolation operation is performed on the activated high-frequency features to obtain processed high-frequency features; The low-frequency features are processed through a linear embedding layer and a Transformer structure, and the processed low-frequency features and the processed high-frequency features are fused and reconstructed to obtain a reconstructed signal, that is, the features after feature optimization. The reconstructed signal is transmitted to the edge-focus attention module through a skip connection.
6. The method according to claim 1, wherein in the edge-focus attention module, the process of feature reconstruction of the features after feature optimization according to the assigned weights includes: Upsampling the features after feature optimization through a transposed convolution to obtain a high-resolution feature map, splicing and feature extraction of the output features of the previous layer on the high-resolution feature map to obtain local features, performing feature extraction on the local features through a local convolution, performing weight assignment processing on the local features through a global convolution to obtain local information and global information, splicing the local information and the global information and performing feature extraction to obtain a feature reconstruction result.
7. The method according to claim 6, wherein the process of feature extraction through a local convolution includes: performing convolution, normalization, and activation function processing on the local features to obtain local information.
8. The method according to claim 6, wherein the process of performing weight assignment processing on the local features through a global convolution includes: performing convolution, normalization, and activation function processing on the local features for several times, and performing convolution processing on the result of the activation function processing again to obtain the first branch information; performing normalization processing on the local features, and performing weight assignment processing on the result of the normalization processing through the SS2DwithSCAM module, and regularizing the result of the SS2DwithSCAM module processing, fusing the regularization result with the result of the SS2DwithSCAM module processing to obtain the second branch information, and performing convolution processing on the first branch information and the second branch information to obtain global information. The SS2DwithSCAM module is an SS2D module with a spatial channel attention module added.
9. The method according to claim 8, wherein the process of processing through the SS2DwithSCAM module includes: performing global modeling on the result of the normalization processing through the SS2D module to obtain a modeling result, processing the modeling result through the spatial channel attention module, and performing further weight assignment and transformation on the result of the spatial channel attention module processing to obtain the result of the SS2DwithSCAM module processing; wherein, the process of processing through the spatial channel attention module includes: calculating the spatial weight and the channel weight of the modeling result through the spatial attention mechanism and the channel attention mechanism respectively, and multiplying the modeling result element-wise with the spatial weight and the channel weight to obtain the result of the spatial channel attention module processing.
Citation Information
Patent Citations
Coronary artery calcification plaque image segmentation system based on improved UNet network
CN116823853A
Bladder tumor image segmentation method and system based on detail enhancement reverse attention network
CN117975011A
Ultrasonic image plaque segmentation method, system and equipment based on deep learning
CN118397015A
Cited By
Lightweight target detection method suitable for low-light space environment
CN121074866A
一种适用于低光空间环境的轻量级目标检测方法
CN121074866B