Ultrasonic mammary gland image segmentation method based on fine-grained guidance
Through the fine-grained condition-guided diffusion model and the adaptive detail-oriented attention module, the problems of semantic correlation and feature imbalance in medical ultrasound image segmentation are solved, and high-precision lesion region segmentation is achieved.
Patent Information
- Application Number
- CN202510205461.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-04
AI Technical Summary
The existing medical ultrasound image segmentation method has insufficient semantic correlation between lesion areas and normal tissues, making it difficult to capture fine-grained features, and the traditional model is unbalanced between global and local features, resulting in inaccurate segmentation results.
A fine-grained condition-based diffusion model is adopted, combined with a denoising diffusion probability model, a fine-grained condition-guided network and an adaptive detail-oriented attention module, to enhance the semantic correlation between the foreground and the background, and capture the complex correlation between global information and channels through the context decoding cross attention layer.
It improves the accuracy and efficiency of medical ultrasound image segmentation, especially in fine-grained feature capture and global feature extraction in the lesion area, significantly improving the accuracy and robustness of the segmentation results.
Smart Images

Figure CN120259645A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image segmentation, and in particular to an ultrasonic breast image segmentation method based on fine-grained guidance. Background Art
[0002] Medical image segmentation is the process of segmenting medical images into meaningful regions. It is a basic step in many medical image analysis applications (assisted diagnosis, treatment planning, disease detection, etc.). Medical ultrasound imaging is safe, non-invasive, and low-cost. Compared with computed tomography (CT) and magnetic resonance imaging (MRI), it is more commonly used and has become an effective initial diagnosis method for disease diagnosis and treatment. However, ultrasound images are often affected by factors such as speckle noise, blurred edges, insufficient contrast, and complex lesion shapes, which invisibly increases the workload of doctors and increases the risk of missed diagnosis and misdiagnosis.
[0003] With the development of deep learning technology, more and more studies have successfully applied neural network-based models to medical image segmentation tasks, especially convolutional neural networks (CNNs), which have been widely used and achieved remarkable results. Among them, UNet has achieved pixel-level prediction through end-to-end training, making a major breakthrough in the field of medical image segmentation. UNet embeds low-resolution features into high-resolution features by introducing jump connections between the encoder and decoder, thereby significantly improving the image segmentation ability. Inspired by the success of UNet, many mainstream models have been proposed based on UNet, including ResUNet, DenseUNet, UNet++, and models based on ensemble learning. However, these methods mainly focus on the overall area of medical objects, show low sensitivity in the segmentation of medical objects in small areas, and lack the ability to capture fine-grained features.
[0004] Since the successful application of the attention mechanism in the Transformer model, the technology has quickly attracted widespread attention. The attention mechanism helps the model focus on key content and significantly improves the ability to detect important features. The attention mechanism has also been introduced into the medical image segmentation task, further improving the model segmentation performance. In addition, the denoising diffusion probability model (DDPM), as a powerful generation model, has also been widely used in medical image segmentation in recent years. DDPM has attracted much attention for its high diversity and high-quality generation capabilities. DDPM was originally applied to the generation field without absolute true values. Recent studies have shown that it is also effective for medical image segmentation with unique true labels. However, these methods still face three main problems in performing medical image segmentation processing tasks:
[0005] 1: Existing diffusion-based segmentation methods ignore the semantic correlation between the lesion area (foreground) and normal tissues (background). Specifically, diffusion-based segmentation methods add Gaussian noise to the ground truth labels through a diffusion process, and then perform reverse diffusion on the ground truth labels of pure noise to recover the ground truth label regions. During this process, the segmentation results obtained by predicting the added noise distribution exhibit various randomnesses, causing the segmentation results to be distorted.
[0006] 2: The lesion areas in medical ultrasound images are often blurred and difficult to separate from the background. Adaptive calibration is the key to obtaining accurate segmentation results. Also, due to the inherent inductive bias in traditional convolutional architectures, each convolutional kernel only focuses on a local subset of pixels in the entire image and forces the network to focus on local patterns rather than global context, making it difficult for them to capture long-range dependencies in the image.
[0007] 3: Common attention mechanisms (such as self-attention) usually ignore capturing fine-grained and coarse-grained information, resulting in an imbalance between global and local features, losing the ability of context awareness and causing overfitting.
[0008] Therefore, we propose a fine-grained-guided ultrasound breast image segmentation method. We use a denoising diffusion probability model to obtain a clear denoised image; based on a fine-grained conditional guidance network and the included adaptive detail-guided attention module, we learn and fuse the prior information of the image to enhance the semantic correlation between the foreground and the background; we add a context decoding cross-attention layer to effectively capture the global information of the input features and the complex correlations between channels, overall improving the accuracy and efficiency of image segmentation. Summary of the Invention
[0009] The object of the present invention is to solve the above problems and provide a medical ultrasound image segmentation method based on a fine-grained conditional guidance diffusion model.
[0010] To achieve the above object, the technical solution adopted by the present invention is as follows: According to one aspect of the present invention, there is provided a medical ultrasound image segmentation method based on a fine-grained conditional guidance diffusion model, including using a multi-scale conditional guidance diffusion model, using a denoising diffusion probability model to obtain a clear denoised image, based on a fine-grained conditional guidance network and the included adaptive detail-guided attention module, learning and fusing the prior information of the image to enhance the semantic correlation between the foreground and the background; adding a context decoding cross-attention layer to effectively capture the global information of the input features and the complex correlations between channels, overall improving the accuracy and efficiency of image segmentation. The specific steps are as follows: Step 1: Use a denoising diffusion probability model to effectively remove image noise The denoising diffusion probability model includes a forward diffusion stage and the reverse denoising stage , in the forward stage, the segmentation label x0~q(x) is gradually added Gaussian noise through a series of time steps T, and the formula is as follows:
[0011] (1) In the formula is the hyperparameter of the variance of the Gaussian distribution, which is a fixed value. N is the added noise at each step of the model. The denoising comes from the normal distribution (Gaussian distribution). I is an identity matrix, which means that the random variables in each dimension are independent and the variance is 1. Through the known x 0 and β t (the hyperparameter of the variance of the Gaussian distribution is a fixed value), directly obtain x t , , , ,…, The constant is similar to the hyperparameter, is the noise added in the forward diffusion process. As T increases, it becomes smaller and smaller. Therefore, at any time satisfies:
[0012] (2) In the reverse process, reverse the above process and sample from , and restore the original image from the Gaussian noise N(0, I), and use a deep learning model to predict a reverse distribution : : (3) where the variance and mean are respectively: (4) (5) Repeat the above process until the original image is restored. The step estimation function can be adjusted by the prior information of the original image, which can be expressed as: (6) In the formula is the global conditional feature embedding, is the local block conditional feature embedding, is the segmentation map feature embedding at the current step. The sum of these three components is sent to the UNet decoder D for feature extraction. The step index t is integrated with the added embedding and decoder features; Step 2: Capture features from different conditional sources based on the fine-grained conditional guidance network The multi-scale conditional guidance network contains a global feature extraction network and a local block feature extraction network. The encoding layer of the global feature extraction network consists of Transformer Layer1 and Transformer Layer2, and the encoding layer of the local block feature extraction network consists of Transformer Layer1_p, Transformer Layer2_p, Transformer Layer3_p, and Transformer Layer4_p. The settings of the encoding modules are all the same; And a context decoding cross-attention module is introduced in the decoder module to help the model capture and fuse features from different input sources and context-related features, improving the ability to extract global and local features; Detail attention and positional encoding are added to the input. At the same time, the two-dimensional self-attention is decomposed into two one-dimensional self-attentions: height self-attention and width self-attention, which can be expressed as: (7) Among them, the parameter operations in formula (7) refer to Figure 2 the connection method in, where is the query value (Query) of the input tensor, is the key value (Key) of the input tensor, is the value (Value) of the input tensor, , , and are four gates to control the amount of information provided by the positional embedding to the key, query, and value. , and are positional embedding vectors. Each layer of the encoder extracts features from the Height dimension and the Weight dimension, and then obtains the global conditional feature and the local conditional feature of the current time step through the decoding layer. Finally, these two components and the segmentation map of the current time step are added together and sent to the UNet encoder D for reconstruction;
[0013] In each iteration process, the original image and the segmentation label are randomly sampled for training. The number of iterations is randomly drawn from a uniform distribution, comes from a Gaussian distribution. Finally, we train the parametric model , and the training objective is : (8) Among them, .
[0014] Step 3: Use the adaptive detail-oriented attention module to enhance the interaction ability between global and local features The adaptive detail-oriented attention module is used to perform sigmoid gating and convolution operations, and the specific content is as follows: In each convolutional block of the deep network, the adaptive detail-oriented attention module adaptively optimizes the intermediate feature map. The adaptive detail-oriented attention module consists of three main parts: an adaptive pooling attention module, a self-attention module, and a spatial attention module; The adaptive detail-oriented attention module does not change the shape of the data, but highlights important content and suppresses unimportant content in the spatial and channel dimensions. It has three paths: an efficient channel attention layer, a multi-head attention mechanism module based on relative position embedding, and emphasizing important spatial regions in the input feature map to perform attention processing on the intermediate feature map in different dimensions; Step 4: Add a context decoding cross-attention layer to further improve the image segmentation accuracy The context decoding cross-attention layer is placed in the decoder. By effectively capturing the global information of the input features and the complex correlations between channels, it improves the image segmentation accuracy and efficiency; Containing feature tensors from different input sources , , pass through the decoding layer of the denoising diffusion probability model layer by layer. At the decoder1 layer, the context decoding cross-attention layer extracts information of different dimensions of different feature maps through average pooling and max pooling operations. The average pooling is shown in Equation 15, and the max pooling is shown in Equation 16: (15) In the formula, , H is the height of the feature tensor , and W is the width of the feature tensor .
[0015] (16) Subsequently, through the channel concatenation operation , the two pooling features are fused into . Next, two fully connected layers are used to perform dimensionality reduction and dimensionality increase operations on the concatenated features. The first fully connected layer reduces the channel dimension from to ; (17)(17) In the formula, , is the reduction rate, represents the activation function, and b1 is the offset of the tensor. The second fully connected layer restores the feature dimension to C; (18) In the formula, , finally, the sigmoid activation function is used to generate the channel attention weight coefficient : (19) And multiply this weight with the input features channel by channel to obtain the enhanced feature map: (20) In the formula, represents the channel-by-channel multiplication operation. After the above process, the residual connection method is used to add the enhanced feature map to perform an add operation to obtain a new feature tensor containing multi-source features, and finally obtain the segmentation result through the above steps.
[0016] Furthermore, among the three paths in the adaptive detail-oriented attention module, the upper path is the efficient channel attention layer. By performing global average pooling to take the average value of each channel in the spatial dimension, a pooled feature map with the shape of (B, C, 1, 1) is generated; then, 1D convolution operation is used to extract the fine-grained information between channels, and the attention weight is generated through the sigmoid activation function. Finally, the channel weight is scaled by the learnable parameter γ and applied to the input features, and the output tensor is as shown in Equation (9).
[0017] (9) In Equation (9): (10) In Equations (9) and (10), is the one-dimensional convolution, is the activation function, is the global average pooling, is the input tensor; The middle path is a multi-head attention mechanism module based on relative position embedding, which is used to model the context relationship of input features and enhance the expression ability of spatial information. By introducing four gates to control the amount of information provided by the position embedding keys, queries, and values, the two-dimensional self-attention is decomposed into two one-dimensional self-attention mechanisms. The inner products of the query and the position embedding are calculated as shown in Equation (11), the inner products of the key and the position embedding are calculated as shown in Equation (12), and the inner products between the query and the key are calculated as shown in Equation (13), so as to capture the local and global dependencies in the input features. The formulas are as follows: (11) (12) (13) In Equation (11), Q is the query, is the position embedding of the query. In Equation (12), is the key, is the position embedding of the key. After the similarity matrix is concatenated and scaled, the softmax function is used to generate a normalized attention weight matrix. Subsequently, the weight matrix performs weighted summation on the value features and the position embedding features, and the output features are obtained through concatenation and batch normalization. Finally, the pooling operation is combined to achieve the extraction of multi-scale features, which can efficiently integrate spatial positions and context information;
[0018] The lower path mainly emphasizes the important spatial regions in the input feature map, suppresses the unimportant regions, performs average pooling and max pooling on the input feature map in the channel dimension to generate two attention maps, then concatenates these two attention maps into a feature map with two channels, and then performs convolution on the concatenated feature map through a 3×3 convolutional kernel to extract coarse-grained information and generate a spatial attention map. The sigmoid activation function is used to generate spatial attention weights to weight the input feature map. The overall output of the adaptive detail-oriented attention module is shown in Equation (14).
[0019] (14) The parameters in Equation (14) are respectively the upper path, the middle path, the final feature extraction results of the three paths of the lower path.
[0020] Compared with the prior art, the present invention has the following beneficial effects: (1) An innovative medical image segmentation model based on DDPM, namely MSCD-DiffNet, is designed. This model extracts features of different dimensions from the prior image in a multi-scale manner at the feature level through a multi-scale conditional guidance network (FG-ADDN), and then fuses the current-step segmentation module with the prior image. During the iterative sampling process, MSCD-DiffNet learns the prior information of the image. The current-step segmentation labels with added noise interference dynamically enhance the conditional features, thereby improving the reconstruction accuracy.
[0021] (2) The self-designed adaptive detail-oriented attention module (AODA) and the Transformer network are used as conditional networks. The attention mechanism and Transformer are good at capturing the global relationships in the input data and can solve the long-distance relationships that are difficult for CNNs to handle. The AODA module can balance the attention to global context and fine-grained features, and complete the tasks of comprehensive feature extraction and accurate image segmentation.
[0022] (3) A context decoding cross-attention layer (CACD) is added according to specific tasks and embedded in the DDPM decoder to help the model solve the problem of feature interaction between different conditional sources and enhance the feature interaction ability of the network. Through different pooling methods, residual connections, and multi-layer fully connected layers, it can achieve a balance between global and local features. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is the flowchart of the technical solution of the present invention; Figure 2 It is the illustration of the adaptive detail-oriented attention (AODA) module of the present invention; Figure 3 It is the feature obtained by FG-ADDN of the present invention and the segmentation map of the current time step; Figure 4 It is the segmentation result map of the mainstream method and the method of the present invention on three different ultrasonic image part datasets; Figure 5 It is two images of small lesion areas and two images of irregular lesions randomly selected from the BUSI public dataset for the comparative experiment of the present invention; Figure 6 It is the ultrasonic image of gallbladder stones in the self-built library of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0024] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0025] The present invention provides a method for ultrasonic breast image segmentation based on fine-grained guidance. The present invention attempts to use a multi-scale conditional guided diffusion model (MSCD-DiffNet) to make up for the deficiencies of traditional segmentation methods, such as weak semantic association between foreground and background, imbalance between global and local features, and lack of fine-grained feature capture ability. Its core design is as follows: using a denoising diffusion probability model to obtain a clear denoised image; based on a fine-grained conditional guidance network (FG-ADDN) and an embedded adaptive detail-oriented attention module (AODA), learning and fusing image prior information to enhance the semantic correlation between foreground and background; adding a context decoding cross-attention layer (CACD) to effectively capture the global information of input features and the complex correlation between channels, and overall improving the accuracy and efficiency of image segmentation. The present invention selects ultrasonic images of gallbladder stones as experimental images to carry out the specific application and verification of the MSCD-DiffNet method. The comparative experimental results on relevant datasets show that the method of the present invention obtains better image segmentation results than similar popular methods.
[0026] Implementation steps and key algorithms of the technical solution of the present invention Step 1: Use a denoising diffusion probability model (DDPM) to effectively remove image noise This model is a generative model based on probability modeling, including a forward diffusion stage and a reverse denoising stage . In the forward stage, the segmentation label x0~q(x) is gradually added Gaussian noise through a series of time steps T. Since each moment t in the forward stage
[0027] is only related to the moment t-1, it can also be regarded as a Markov process:
[0028] (1) In the formula is a hyperparameter of the variance of the Gaussian distribution, which is a fixed value, N is the noise added in each step of the model, and the denoising comes from a normal distribution (Gaussian distribution). I is an identity matrix, which means that the random variables in each dimension are independent and the variance is 1. Through the known x 0 and β t (the hyperparameter of the variance of the Gaussian distribution is a fixed value), directly obtain x t , , , ,…, The constant is similar to the hyperparameter, is the noise added in the forward diffusion process. As T increases, it becomes smaller and smaller. Therefore, at any moment satisfies:
[0029] (2) During the reverse process, if the above process can be reversed and sampling can be performed from sampling, the original image can be restored from the Gaussian noise N(0, I). However, relevant literature has proven that it is not possible to simply infer , so a deep learning model (with parameters and a main architecture of U-Net + attention structure) needs to be used to predict an inverse distribution :
[0030] (3) where the variance and mean are respectively: (4) (5) It will predict the noise based on the current input, then subtract the predicted noise from the current image, and finally obtain the denoised image. Repeat the above process until the original image is restored until. As Figure 1 shown, the step estimation function can be adjusted through the prior information of the original image, which can be expressed as:
[0031] (6) In the formula is the global conditional feature embedding, is the local block conditional feature embedding, is the segmentation map feature embedding at the current step. These three components are added together and sent to the UNet decoder D for feature extraction. The step index t is integrated with the added embeddings and decoder features.
[0032] Step 2: Capture features from different conditional sources based on the fine-grained conditional guidance network (FG-ADDN) FG-ADDN contains a global feature extraction network and a local block (patches) feature extraction network, which can improve the overall understanding of the image. The encoding layer of the global feature extraction network consists of Transformer Layer1 and TransformerLayer2. The encoding layer of the local block feature extraction network consists of Transformer Layer1_p, TransformerLayer2_p, Transformer Layer3_p, and Transformer Layer4_p. The settings of the encoding modules are all the same. In the design of FG-ADDN of the present invention, different resolutions of the original image ( and ) are used as the inputs of the network respectively. In the local block feature extraction network, four patches of the image with sizes of I / 4 × I / 4 are created. Among them, I is the size of the original image. Each patch is fed forward through the network, and the output feature map is resampled based on its position to obtain the output feature map. Aiming at the problems that the self-attention layer has high computational complexity and does not utilize any position information when calculating the non-local context, we add detail attention and position encoding to the input so that it can not only extract fine features but also distinguish elements at different positions. At the same time, the two-dimensional self-attention is decomposed into two one-dimensional self-attentions: height self-attention and width self-attention. It is called the axial attention mechanism and can be expressed as:
[0033] (7) Among them, the parameter operations in formula (7) refer to the connection method in Figure 2 , where is the query value (Query) of the input tensor, is the key value (Key) of the input tensor, is the value (Value) of the input tensor, , , and are four gates to control the amount of information provided by the position embedding to the key, query, and value. , and are the position embedding vectors. Each layer of the encoder extracts features from the Height dimension and the Weight dimension, and then obtains the global conditional feature of the current time step and the local conditional feature of the current time step through the decoding layer. Finally, these two components and the segmentation map of the current time step are added together and sent to the UNet encoder D for reconstruction;
[0034] The formula in Equation 2 follows the attention model proposed in. Each layer of the encoder extracts features from the Height dimension and the Weight dimension. Then, the global conditional feature of the current time step and the local conditional feature of the current time step are obtained through the decoding layer. Finally, these two components and the segmentation map of the current time step are added together and sent to the UNet encoder D for reconstruction. MSCD-DiffNet follows the standard training process of DDPM. In each iteration, the original image and the segmentation label Random sampling is performed for training, and the number of iterations is randomly drawn from a uniform distribution. It is derived from a Gaussian distribution. Finally, we train the parametric model , and the training objective is :
[0035] (8) where .
[0036] Since the encoding module carries rich information of the conditional module, the present invention introduces a context decoding cross-attention module in the decoder module to help the model capture and fuse features from different input sources and context-related features, which cover image information of different scales, thereby improving the ability to extract global and local features.
[0037] Step 3: Use the Adaptive Detail-Oriented Attention Module (AODA) to enhance the interaction ability between global and local features Regarding the traditional attention mechanism (such as self-attention) which often fails to fully capture fine-grained and coarse-grained information, lacks context awareness ability, and is prone to overfitting problems. The present invention designs a new feed-forward convolutional neural network attention module, namely the Adaptive Detail-Oriented Attention Module (AODA). This module can retain more detailed information in the feature map and, by combining multi-scale pooling and position embedding, enables the model to have stronger context awareness ability and achieve a balance between global and local features. AODA dynamically adjusts the focus of attention according to the complexity of the input features through sigmoid gating and convolutional operations, making it more flexible than existing attention mechanisms.
[0038] Its specific design is as follows: In each convolutional block of the deep network, the AODA module adaptively optimizes the intermediate feature map. This module consists of three main parts: an adaptive pooling attention module, a self-attention module, and a spatial attention module. As an attention mechanism, AODA does not change the shape of the data but highlights important content and suppresses unimportant content in the spatial and channel dimensions, and can be applied to all levels of the network. Figure 2 Shows the overall structure of the AODA module, where the three paths respectively perform attention processing on the intermediate feature map in different dimensions.
[0039] As Figure 2 shown, after passing through three paths with different functions, the feature structure obtains rich and diverse information, and finally through the add operation, the information is fused and output as a new feature H×W×Cout.
[0040] Figure 2The upper-middle path is the efficient channel attention layer. By performing global average pooling to take the average value of each channel in the spatial dimension, a pooled feature map with the shape of (B, C, 1, 1) is generated. Then, 1D convolution operation is used to extract the fine-grained information between channels, and sigmoid activation function is used to generate the attention weights. Finally, the channel weights are scaled by the learnable parameter γ and applied to the input features. The output tensor As shown in Equation (9).
[0041] (9) In Equation (9): (10) In Equations (9) and (10), is a one-dimensional convolution,[[]] is an activation function,[[]] is global average pooling,[[]] is the input tensor.[[]]
[0042] The middle path is a multi-head attention mechanism module based on relative position embedding, which is used to model the context relationship of the input features and enhance the expression ability of spatial information. By introducing four gates to control the amount of information provided by the position embedding keys, queries, and values, the two-dimensional self-attention is decomposed into two one-dimensional self-attention mechanisms: the height dimension and the width dimension, to overcome the disadvantages of large computational complexity and long distance in calculating the attention matrix in traditional transformers. Specifically, the query (Q), key (K), and value (V) representations are generated through linear transformation, and the features are decomposed using a grouping strategy to reduce the computational complexity. At the same time, a predefined relative position embedding matrix is combined. We calculate the inner product of the query and the position embedding as shown in Equation (11), the inner product of the key and the position embedding as shown in Equation (12), and the inner product between the query and the key as shown in Equation (13), respectively, so as to capture the local and global dependencies in the input features.[[]]
[0043] (11) (12) (13) In Equation (11), Q is the query,[[]] is the position embedding of the query. In Equation (12), is the key,[[]] is the position embedding of the key. After the similarity matrix is concatenated and scaled, the softmax function is used to generate the normalized attention weight matrix. Subsequently, the weight matrix performs weighted summation on the value features and the position embedding features, and the output features are obtained through concatenation and batch normalization. Finally, multi-scale feature extraction is achieved by combining pooling operations, which can efficiently integrate spatial position and context information;[[]]
[0044] The lower path mainly performs average pooling and max pooling on the input feature map in the channel dimension by emphasizing important spatial regions in the input feature map and suppressing unimportant regions, generating two attention maps. Then, these two attention maps are concatenated into a feature map with two channels, and a 3×3 convolutional kernel is used to perform convolution on the concatenated feature map to extract coarse-grained information, generating a spatial attention map. The sigmoid activation function is used to generate spatial attention weights to weight the input feature map. This design enables the AODA module to enhance the attention to details, dynamically adjust the attention focus according to the complexity of the input, and balance the attention to global context and fine-grained features. The overall output of the AODA module is shown in Equation (14).
[0045] (14) The parameters in Equation (14) are respectively the upper path, the middle path, the final feature extraction results in the three paths of the lower path.
[0046] Step 4: Add a Context-Aware Cross Decoding Attention layer (CACD) to further improve the image segmentation accuracy.
[0047] In the diffusion model, there are requirements for modeling long-range dependent features in the medical image generation process and problems of insufficient global context information and interaction between features in multi-input source feature fusion. Therefore, the present invention designs a Context-Aware Cross Decoding Attention layer (CACD). As Figure 3 shown, the CACD module is placed in the decoder, and by effectively capturing the global information of the input features and the complex correlations between channels, it improves the image segmentation accuracy and efficiency.
[0048] In Figure 3 , the feature obtained by FG-ADDN and the segmentation map at the current time step The fused H×W×Cin is fed into the Decoder Block in DDPM. In the Decoder Block, layer-by-layer CCA layers and Decoder layers are used to achieve the fusion of global context information of multi-input source features.
[0049] Containing feature tensors from different input sources , , of Pass through the decoding layer of DDPM layer by layer. In the decoder1 layer, the CACD module extracts information of different dimensions of different feature maps through average pooling (shown in Equation 15) and max pooling (shown in Equation 16) operations.
[0050] (15) In the formula, , H is the height of the feature tensor , and W is the width of the feature tensor .
[0051] (16) Subsequently, through the channel concatenation operation , the two pooling features are fused into . Next, two fully connected layers are used to perform dimensionality reduction and dimensionality increase operations on the concatenated features. The first fully connected layer reduces the channel dimension from to .
[0052] (17) In the formula, , is the reduction rate, represents the activation function. The second fully connected layer restores the feature dimension to C.
[0053] (18) In the formula, , and finally, the sigmoid activation function is used to generate the channel attention weight coefficient : (19) And multiply this weight with the input features channel by channel to obtain the enhanced feature map: (20) In the formula, represents the channel-by-channel multiplication operation. After the above process (CCA1), the residual connection method is used to perform the add operation on the enhanced feature map and to obtain a new feature tensor containing multi-source features.
[0054] Through the above steps, the segmentation result is finally obtained. The CACD module realizes long-range dependence modeling and efficient interaction of multi-source features by strengthening important features and suppressing irrelevant information, improves the accuracy of medical image segmentation, and reduces the computational complexity.
[0055] The experimental images used in this invention are from the public BUSI dataset, the public TN3K dataset, and the gallstone ultrasound image dataset we built ourselves. The public datasets contain 780 and 3,493 samples respectively and have publicly available segmentation labels. The self-built dataset contains 123 original images, which are extended to 738 images using mirroring and rotation methods. 80% of the images are used for training and 20% for validation.
[0056] To evaluate the effectiveness of the method of this invention, several commonly used image segmentation evaluation metrics are adopted, namely IoU, Dice, Precision, Recall, and Accuracy. Among them, IoU (Intersection over Union) is called the intersection over union ratio, which represents the proportion of the intersection of the segmented region and the true region to the union, reflecting the accuracy of the segmentation; the Dice similarity coefficient (DSC) measures the overlapping degree of the segmented region and the true annotation region, with a value range between 0 and 1, and 1 indicating complete overlap. Precision represents how many of the positive examples segmented are truly positive examples, evaluating the accuracy of the model in predicting the lesion region. A higher precision means fewer false positives (FP). Recall represents the proportion of the true lesion region that is correctly segmented, measuring the model's ability to identify positive examples. A higher recall means fewer false negatives (FN). Accuracy is suitable for tasks with higher requirements for the overall segmentation effect. Especially on datasets where the distribution of lesions and normal regions is balanced, it can give an overall evaluation of the model's segmentation ability. These metrics combined can more comprehensively reflect the performance of the medical image segmentation model (↑ indicates the higher the better. Underline indicates sub-optimal results, and the best result is in bold).
[0057] Table 1 Quantitative comparison results of the method of this invention and common image segmentation methods.
[0058]
[0059] The quantitative comparison results of the method of the present invention and eight comparison methods on two public datasets are shown in Table 1. For BUSI images, the IoU, Dice, Precision, Recall, and Accuracy of the method of the present invention are 84.68%, 85.71%, 87.37%, 88.25%, and 97.08% respectively, which are 1.15%, 0.79%, 1.62%, and 0.86% higher than the Dice, Precision, Recall, and Accuracy of the second-best segmentation method MedsegDiff in the table. For TN3K images, the IoU, Dice, Precision, Recall, and Accuracy of the method of the present invention are 75.36%, 77.11%, 89.73%, 87.28%, and 96.86% respectively. It is 0.43%, 0.50%, 1.09%, 2.2%, and 1.43% higher than the IoU, Dice, Precision, Recall, and Accuracy of the second-best segmentation method MedsegDiff in the table.
[0060] The method of the present invention captures structural information at different scales through a multi-scale conditional guidance network (FG-ADDN). Especially in fuzzy and complex lesion areas, the noise influence is reduced by dynamically adjusting conditional features, significantly improving the segmentation accuracy. This makes the improvement of IoU particularly significant. The adaptive detail-oriented attention module (AODA) achieves a balance between global and local feature modeling, has a good performance in segmenting edge-blurred and low-contrast regions, and strengthens the attention to small-scale and complex-shaped regions, so that indicators such as Dice, Precision, and Recall are significantly better than MedSegDiff. The context decoding cross-attention layer (CACD) solves the problem of inconsistency between features through a multi-source feature interaction and detail information retention mechanism, and realizes a balanced expression of global information and local details. It helps to improve indicators such as Precision and Recall. The present invention introduces a dynamic learning mechanism in the framework of the denoising diffusion probability model (DDPM) to gradually generate high-quality segmentation images, and has stronger robustness when dealing with noise and fuzzy textures. The progressive optimization sampling process makes the segmentation result tend to be accurate. The present invention integrates model modules such as denoising diffusion, multi-scale conditions, and context cross-attention to form a comprehensive modeling framework from global semantics to local details. Compared with methods such as TransUNet and MedSegDiff, the method of the present invention can better perform complex tasks of medical ultrasound image processing through joint optimization in multiple dimensions.
[0061] The quantitative comparison results of the method of the present invention with eight comparison methods on the self-built ultrasonic image dataset of gallbladder stones are shown in Table 2 (↑ indicates the higher the better. The best results are in bold). The method of the present invention has better performance in the processing of details such as edge contours, and improves by 0.51%, 1.24%, 2.18%, 3.39% and 0.82% respectively compared with the second-best IoU, Dice, Precision, Recall and Accuracy in the table.
[0062] Table 2: Quantitative comparison results of the method of the present invention and common image segmentation methods on the self-built dataset
[0063] Figure 4 The qualitative comparison results of the method of the present invention and eight current mainstream and classical medical image segmentation methods on two public datasets are presented. All methods are implemented using their default settings. nnUNet, AttentionUNet, TransUNet, SwinUNet, CIMD, MDFNet, MedUniSeg, MedSegDiff and the method of the present invention can accurately segment the location area where the lesion is located. However, nnUNet performs excellently in segmenting large targets, but its performance is not ideal when dealing with small lesion areas. Small lesion areas will be ignored or confused with the background, resulting in information loss; AttentionUNet has a certain improvement in segmentation accuracy and flexibility due to the introduction of the attention mechanism, but it still performs poorly in the segmentation of small targets (such as tiny lesions). Small target areas may be ignored during multiple downsampling processes, and it is difficult to fully focus on these detailed areas even with regional weighting through the attention mechanism; TransUNet has limited performance in dealing with local details of fine boundaries, and it is unable to balance between global and local, resulting in insufficient segmentation accuracy; SwinUNet is vulnerable to noise, and in the processing of low-quality images or images with artifacts, it may lead to inaccurate feature extraction and affect the segmentation performance; The feature extraction of dilated convolution in CIMD and MDFNet is limited by a fixed scale, and the ability to recognize some local structures with blurred edges or complexity is limited. MedUniSeg is a general medical segmentation model that attempts to adapt to multiple segmentation tasks in the same model, but this leads to insufficient performance optimization for specific tasks and it is difficult to achieve the best results of dedicated models for each task. MedSegDiff has limited effect in generating small lesions or lesion samples with complex structures, and it may be difficult to truly reproduce the details of these areas, thus affecting the segmentation performance of the model in small target or detail-sensitive tasks.
[0064] Figure 4Among them, (a)-(d) represent the BUSI public dataset; (e)-(h) represent the TN3K public dataset; where the blue squares mark the edge details differences of the lesion areas, and the red curves are the boundaries of the lesions.
[0065] From Figure 4 the experimental results, it can be seen that the above eight comparison methods all have problems of insufficient or over-segmentation of the tumor area to varying degrees, especially in the processing of the edges and detail noises. In contrast, the method of the present invention retains more segmentation details of the tumor edge and minimizes information loss. To more intuitively display the segmentation effect of the model, we randomly selected four typical images with small lesion areas and irregular edges, and the segmentation results are as Figure 5 shown. It can be clearly seen from it the advantages of the method of the present invention in detail processing, especially its superiority in edge sharpness and lesion shape integrity.
[0066] In Figure 5 the red square represents the tumor position.
[0067] Finally, we conducted a comparative experiment on the self-built gallbladder ultrasound image dataset (as Figure 6 shown). This experiment aims to evaluate the segmentation effect of the method of the present invention in the scenario of small-sample medical data and its adaptability in dealing with special lesion morphologies. Compared with the public dataset, the self-built gallbladder dataset contains a smaller number of images, only 123 original images, and only reaches 738 after amplification. In order to obtain reliable and credible evaluation results, we introduced the GSB dataset as external verification data. Since the sample size in the GSB dataset is small and the scenario is complex, directly using it for training may lead to overfitting, so we only use it as external verification data. Through testing on this dataset, we can observe the segmentation performance of the network when facing complex images that have never been seen before, so as to test whether the model can accurately distinguish the lesion area and process detail features.
[0068] In Figure 6 the red curve represents the gallbladder organ area, and the blue square marks the segmentation boundary difference.
[0069] The experimental results show that the method of the present invention has excellent performance on both the self-built gallbladder ultrasound dataset (GSBD) and the external verification dataset. On the GSBD dataset, the model proposed by the present invention can accurately segment the edges of gallbladder stones and completely retain the shape and structural features of the lesion area, significantly superior to the comparative methods. Despite the differences between the image scenarios and the training data, the model still shows good generalization performance and can stably process complex details and irregular boundaries.
[0070] The above comparative experiments verify the effectiveness of the method of the present invention in medical ultrasound image segmentation, and at the same time demonstrate its robustness and generalization ability in dealing with different data distributions.
[0071] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0072] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for ultrasonic breast image segmentation based on fine-grained guidance, characterized in that: Including the use of a multi-scale conditional guided diffusion model, obtaining a clear denoised image using a denoising diffusion probability model, learning and fusing image prior information based on a fine-grained conditional guidance network and an embedded adaptive detail-oriented attention module to enhance the semantic correlation between the foreground and the background; adding a context decoding cross-attention layer to effectively capture the global information of the input features and the complex correlation between channels, overall improving the accuracy and efficiency of image segmentation. The specific steps are as follows: Step 1: Use a denoising diffusion probability model to effectively remove image noise The denoising diffusion probability model includes a forward diffusion stage as shown in Equation (1) and a reverse denoising stage As shown in Equation (3), in the forward stage, the segmentation label x0~q(x) is gradually added Gaussian noise through a series of time steps T, and the formula is as follows: (1) where is the hyperparameter of the variance of the Gaussian distribution, which is a fixed value. N is the added noise at each step of the model, and the denoising comes from a normal distribution (Gaussian distribution). I is an identity matrix, meaning that the random variables in each dimension are independent and have a variance of 1. Given the known x0 and β t (the hyperparameter of the variance of the Gaussian distribution is a fixed value), x can be directly obtained t , , , ,…, Constants are similar to hyperparameters, is the noise added in the forward diffusion process; It becomes smaller and smaller as T increases, so at any moment all satisfy: (2) During the reverse process, reverse the above process and start from Sampling, restore the original image from the Gaussian noise N(0, I), and use a deep learning model to predict an inverse distribution : (3) where the variance and mean are respectively: (4) (5) Repeat the above process until the original image is restored Up to this point, the step estimation function can be adjusted by the prior information of the original image, which can be expressed as: (6) where is the global conditional feature embedding, is the local block conditional feature embedding, is the segmentation map feature embedding at the current step. These three components are added together and sent to the UNet decoder D for feature extraction. The step index t is integrated with the added embeddings and decoder features; Step 2: Capture features from different conditional sources based on a fine-grained conditional guidance network The multi-scale conditional guidance network includes a global feature extraction network and a local block feature extraction network. The encoding layer of the global feature extraction network consists of Transformer Layer1 and Transformer Layer2, and the encoding layer of the local block feature extraction network consists of Transformer Layer1_p, Transformer Layer2_p, Transformer Layer3_p, and Transformer Layer4_p. The settings of the encoding modules are all the same; And a context decoding cross-attention module is introduced in the decoder module to help the model capture and fuse features from different input sources and context-related features, improving the ability to extract global and local features; Add detail attention and position encoding to the input, and at the same time decompose the two-dimensional self-attention into two one-dimensional self-attentions: height self-attention and width self-attention, which can be expressed as: (7) Among them, the parameter operations in formula (7) refer to the connection method in Figure 2. represents the query value (Query) of the input tensor. is the key value (Key) of the input tensor. is the value (Value) of the input tensor. , , and are four gates to control the amount of information provided by the position embedding to the key, query, and value. , and are the position embedding vectors; each layer of the encoder will extract features from the Height dimension and the Weight dimension, and then obtain the global conditional features at the current time step through the decoding layer. and the local conditional features at the current time step . Finally, these two components and the segmentation map at the current time step are added together and sent to the UNet encoder D for reconstruction. In each iteration, the original image and the segmentation label are randomly sampled for training, and the number of iterations is randomly drawn from a uniform distribution, which is from a Gaussian distribution. Finally, we train the parametric model , and the training objective is : (8) Among them, ; Step 3: Use an adaptive detail-oriented attention module to enhance the interaction ability between global and local features Use an adaptive detail-oriented attention module through sigmoid gating and convolution operations. The specific content is as follows: In each convolutional block of the deep network, the adaptive detail-oriented attention module adaptively optimizes the intermediate feature map. The adaptive detail-oriented attention module includes three main parts: an adaptive pooling attention module, a self-attention module, and a spatial attention module; The adaptive detail-oriented attention module does not change the shape of the data, but highlights important content in the spatial and channel dimensions and suppresses unimportant content. It has an efficient channel attention layer, a multi-head attention mechanism module based on relative position embedding, and three paths that respectively perform attention processing on the intermediate feature map in different dimensions by emphasizing important spatial regions of the input feature map; Step 4: Add a context decoding cross-attention layer to further improve the accuracy of image segmentation The context decoding cross-attention layer is placed in the decoder. By effectively capturing the global information of the input features and the complex correlation between channels, it improves the accuracy and efficiency of image segmentation; Containing feature tensors from different input sources , , are passed layer by layer through the decoding layer of the denoising diffusion probability model. In the decoder1 layer, the context decoding cross-attention layer extracts information of different dimensions of different feature maps through average pooling and max pooling operations. The average pooling is shown in Equation 15, and the max pooling is shown in Equation 16: (15) wherein, , H is the height of the feature tensor , and W is the width of the feature tensor . (16) Subsequently, through the channel concatenation operation , the two pooling features are fused into . Next, two fully connected layers are used to perform dimensionality reduction and dimensionality increase operations on the concatenated features. The first fully connected layer reduces the channel dimension from to ; (17) In the formula, , is the reduction rate, represents the activation function, b1 is the offset of the tensor, and the second fully connected layer restores the feature dimension to C; (18) wherein, , b2 is the offset of the tensor. Finally, the sigmoid activation function is used to generate the channel attention weight coefficient : (19) And multiply this weight with the input features channel by channel to obtain an enhanced feature map: (20) In the formula, represents the per-channel multiplication operation. After the above process, the residual connection method is used to combine the enhanced feature map with to obtain a new feature tensor containing multi-source features through an add operation. The segmentation result is finally obtained through the above steps.
2. The method for ultrasonic breast image segmentation based on fine-grained guidance according to claim 1, wherein: Among the three paths in the adaptive detail-oriented attention module, the upper path is an efficient channel attention layer. It takes the average value of each channel in the spatial dimension through global average pooling to generate a pooled feature map with the shape of (B, C, 1, 1). Then, 1D convolution operation extracts the fine-grained information between channels, and generates attention weights through the sigmoid activation function. Finally, the channel weights are scaled by the learnable parameter γ and applied to the input features to output a tensor As shown in Equation (9): (9) (10) In equations (9) and (10), is a one-dimensional convolution,[ is an activation function,[ is global average pooling,[ is the input tensor; The intermediate path is a multi-head attention mechanism module based on relative position embedding, which is used to model the context relationship of input features and enhance the expression ability of spatial information. By introducing four gates to control the amount of information provided by the position embedding keys, queries, and values, the two-dimensional self-attention is decomposed into two one-dimensional self-attention mechanisms. The inner products of the query and the position embedding as shown in Equation (11), the inner product of the key and the position embedding as shown in Equation (12), and the inner product between the query and the key as shown in Equation (13) are calculated respectively, so as to capture the local and global dependencies in the input features. The formula is as follows: (11) (12) (13) In formula (11), Q is the query, is the position embedding of the query. In formula (12), is the key, is the position embedding of the key; After the similarity matrix is spliced and scaled, the softmax function is used to generate a normalized attention weight matrix. Subsequently, the weight matrix performs weighted summation on the value feature and the position embedding feature, and the output feature is obtained through splicing and batch normalization. Finally, the pooling operation is combined to extract multi-scale features, which can efficiently integrate spatial position and context information; The lower path mainly emphasizes the important spatial regions in the input feature map, suppresses the unimportant regions, performs average pooling and max pooling on the input feature map in the channel dimension to generate two attention maps, then concatenates these two attention maps into a feature map with two channels, and then performs convolution on the concatenated feature map through a 3×3 convolutional kernel to extract coarse-grained information, generating a spatial attention map. The sigmoid activation function is used to generate the spatial attention weight to weight the input feature map. The overall output of the adaptive detail-oriented attention module is as shown in Equation (14): (14) The parameters in Equation (14) are respectively the final feature extraction results of the upper path the middle path and the lower path among the three paths
Citation Information
Cited By
Vertebral bone tissue segmentation method combined with structure-guided conditional diffusion network
CN121121125A
A method for vertebra bone tissue segmentation combining structure guided conditional diffusion network
CN121121125B
Image classification method capable of interpreting breast cancer pathology based on dual-granularity attention and potential prototype alignment
CN121305248A
Two-dimensional engineering drawing element identification method and system based on time sequence-frequency spectrum joint diffusion
CN121686508A
A two-dimensional engineering drawing element identification method and system based on time-spectrum joint diffusion
CN121686508B