Medical ultrasound image segmentation method based on diffusion model, multi-scale and attention module

By adopting an encoder-decoder architecture in medical ultrasound image segmentation, combining the denoising diffusion probability model, multi-scale dynamic condition module and Gaussian cross-fusion attention module, the limitations of existing methods when processing complex backgrounds and rich images are solved, and a more efficient image segmentation effect is achieved.

CN119180826BActive Publication Date: 2025-06-06LANZHOU JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411203045.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-06-06
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

Existing medical ultrasound image segmentation methods have limitations in processing complex backgrounds and rich details, especially convolutional neural networks have limitations in modeling explicit remote relationships, and pure Transformer methods also have the problem of feature loss.

Method used

It adopts an encoder-decoder architecture, consisting of a denoising diffusion probability model, a multi-scale dynamic condition module and a Gaussian cross-fusion attention module. The denoising diffusion probability model simulates image degradation in the forward diffusion stage, and extracts features from standard normal distribution sampling in the reverse diffusion stage; the multi-scale dynamic condition module captures long-distance dependencies at different scales; the Gaussian cross-fusion attention module integrates the features of the encoder and multi-scale module.

Benefits of technology

The denoising diffusion probability model gradually removes noise, captures image details, and enhances the model's adaptability to complex structures of ultrasonic images; the multi-scale dynamic condition module improves the integration ability of image contrast and context information; the Gaussian cross-fusion attention module solves the incompatibility problem during feature fusion and improves network performance and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180826B_ABST
    Figure CN119180826B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical ultrasound image segmentation method of a diffusion model, a multi-scale and an attention module. The present invention uses an improved denoising diffusion probability model as its diffusion model, and the multi-scale and attention modules are a self-designed multi-scale dynamic condition module and a Gaussian cross fusion attention module. The method uses a denoising diffusion probability model to remove image noise and capture important detail information; uses a multi-scale dynamic condition module to improve image contrast and the ability to integrate contextual information of different scales; and uses a Gaussian cross fusion attention module to overcome the incompatibility when directly merging encoder features and dynamic condition module features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a medical ultrasound image segmentation method based on a diffusion model, a multi-scale and an attention module. Background Art

[0002] Deep learning technology can automatically learn deep information from images without human intervention. Methods based on convolutional neural networks (CNNs) have achieved excellent results in the field of medical ultrasound image segmentation. Among them, the U-Net proposed by Ronneberger et al. has become one of the most popular network architectures due to its symmetrical network structure and jump connection. In order to further improve the segmentation accuracy, Zhou et al. proposed the U-Net++ network, redesigned the jump connection structure, and introduced a deep supervision mechanism to prune different segmentation tasks, reducing the feature differences between the encoder and decoder. In addition, the Attention U-net proposed by Oktay et al. adds an integrated attention gate between the encoder and decoder, highlighting the local region features and suppressing the information of irrelevant regions, thereby improving the segmentation accuracy. The UNeXt proposed by Valanarasu et al. introduces a tokenized multilayer perceptron (MLP) block to complete the labeling and projection convolution operations, and uses the tokenized MLP to model the feature representation, effectively capturing local dependencies.

[0003] Although convolutional neural networks have strong image detail extraction capabilities, they are usually limited in modeling explicit long-range relationships. To overcome this limitation, the Transformer architecture for sequence-to-sequence prediction has become an alternative, and many variants have been proposed based on Vision Transformer and Swin Transformer, such as TransUNet, UNETR, Swin-UNet, and TransUNet+, which have all outperformed traditional U-Net in medical image segmentation tasks. Compared with traditional convolutional neural networks, Transformer is better at handling global dependencies, which is very beneficial for the segmentation of complex images. In particular, the TransUNet network proposed by Chen et al. combines the Transformer with the U-Net structure, which not only retains the local feature extraction capability of U-Net, but also enhances the global feature capture capability. However, since this method uses traditional convolutional neural networks for feature extraction in the initial encoding stage, the receptive field of subsequent convolutions is too large, which fails to give full play to the advantages of Transformer, and multi-scale information is not effectively integrated and utilized in the downsampling stage. The pure Transformer method still has the problem of feature loss.

[0004] As a powerful generative model, the denoising diffusion probability model has received widespread attention and popularity in recent years. This method simulates the evolution of data, gradually refines the segmentation boundaries, and generates high-quality images, especially when dealing with complex backgrounds and images with rich details. For example, the MedSegDiff method has achieved remarkable success and surpassed the most advanced segmentation methods before, greatly improving the segmentation effect of medical images. However, this method only uses static ground truth images of lesions as conditional information at each step, making it difficult to learn richer lesion information. Therefore, the single use of the denoising diffusion probability model still has limitations, and related modules need to be added to enhance the model's ability to learn lesion information, so that the generated images are closer to the ground truth images. This will further improve the accuracy and robustness of the segmentation. Summary of the invention

[0005] The purpose of the present invention is to solve the above-mentioned problems and to provide a medical ultrasound image segmentation method based on a diffusion model, a multi-scale and an attention module.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] It includes an encoder-decoder architecture, which consists of three core parts: a denoising diffusion probability model, a multi-scale dynamic condition module, and a Gaussian cross-fusion attention module. The characteristics are: the denoising diffusion probability model gradually adds Gaussian noise to the initial image during the forward diffusion stage. In the back diffusion stage, the image is sampled from the standard normal distribution. , and input into the encoder to extract features;

[0008] In the encoder part, first use the Stem module to For shallow feature extraction, the encoder consists of 3×3 convolution, batch normalization and linear rectification units. to In the second stage, four convolution modules are used to further extract shallow features and downsample them;

[0009] The multi-scale dynamic condition module captures the long-range dependencies of different scales from ultrasound images and effectively fuses the two parts of features through the Gaussian cross-fusion attention module;

[0010] In the decoder part, first In the first stage, the features of the Gaussian cross attention module and the skip connection features of the multi-scale dynamic condition module are concatenated and 1×1 convolved. The transformed image maintains the same resolution as the input. After that, it is upsampled by 3×3 transposed convolution and input to the next stage. arrive In the second stage, the input features and the jump connection features of the corresponding layer of the dynamic condition module are gradually restored to the original resolution through channel splicing, 1×1 convolution and transposed convolution, forming an optimized end-to-end network architecture.

[0011] Preferably, the specific steps are as follows:

[0012] Step 1: Use the denoising diffusion probability model to remove image noise and capture important details

[0013] The denoising diffusion probability model consists of a forward diffusion stage and a reverse diffusion stage;

[0014] In the forward diffusion process, Gaussian noise Gradually add to the input image ,implement After time steps, the original image Transformed into an image with almost all Gaussian noise , the forward diffusion process can be defined as:

[0015] (1)

[0016] (2)

[0017] Where: and Represents the step length and Noise image at the time ; is the coefficient of the noise level, which controls the proportion of noise added and varies with the step size decreases with the increase of is the identity matrix; represents Gaussian distribution;

[0018] In the back diffusion process, the noise image The noise is gradually removed through a series of iterative steps and restored to the original image. , assuming there is a reverse process distribution , which is a distribution with learnable parameters The neural network is parameterized to accurately fit the mean of the image at each time step and variance ;

[0019] (3)

[0020] (4)

[0021] Back diffusion process It can be expressed as:

[0022] (5)

[0023] In the training process of the denoising diffusion probability model, the key step is to minimize the estimated distribution and the true posterior distribution The KL divergence between the two probability distributions is measured by adjusting the neural network parameters to make the predicted distribution Close to minimizing the estimated distribution , the training objective of DDPM is expressed as:

[0024] (6)

[0025] The loss function of the DMA-USeg model can be expressed as:

[0026] (7)

[0027] in, Segment the image for the dynamic condition module; represents the fitting function of the neural network; E x0 , ε Indicates the mathematical expectation of the calculated value in brackets;

[0028] Step 2: Use a multi-scale dynamic condition module to improve image contrast and the ability to integrate contextual information at different scales

[0029] A multi-scale dynamic conditional module is used to capture image detail information, highlight the target area features, and serve as auxiliary information to guide the image generation process in the decoder. First, a Stem layer is used, and then four feature extraction modules are used. to Perform step-by-step feature extraction and downsampling, and finally, input the extracted features into the TransFuse module;

[0030] Then, through the multi-scale fusion module, the information between different scales interacts, captures the global context dependency, makes full use of deep features, and solves the problem of low contrast of ultrasound images. First, the output features of the Transformer module are resized by 1×1 convolution operation and expressed as , in the CNN module The output characteristics of the stage are expressed as . Then the soft gating that controls the integration degree of features at different scales It can be expressed as:

[0031] (13)

[0032] in, Represented as a linear map; Expressed as Activation function; It is represented as channel splicing;

[0033] Next, we use element-wise multiplication to transform the soft gate and Merge and then Perform element addition to enhance features and preserve details, and fuse images It can be expressed as:

[0034] (14)

[0035] in, It is the element dot product operation; Addition for elements;

[0036] Step 3: Use the Gaussian cross-fusion attention module to solve the incompatibility problem when fusing encoder features and dynamic condition module features

[0037] The Gaussian cross-fusion attention module is used to solve the incompatibility problem that may occur when directly merging the two features. First, the two parts of the input features are constrained to be Gaussian distributions. and , then and After normalization, it is used as the query variable and , then add the elements of the two feature maps to get the key K and value V, and feed them into the multi-head attention mechanism to get the output features after interaction. The calculation result is expressed as:

[0038] (15)

[0039] (16)

[0040] Finally, the features and After layer normalization and multi-layer perceptron processing, the network performance and stability are improved. Then, through channel splicing and 3×3 convolution processing, the size of the features is adjusted and input into the decoder to complete the final decoding output. The calculation results are as follows:

[0041] (17)

[0042] in, is the convolution operation; is a multi-layer perceptron; Normalize the layer;

[0043] Preferably, the TransFuse module in step 2 includes CNN and Swin Transformer modules, and the specific process is as follows:

[0044] First, the output features of the CNN module are divided into non-overlapping image blocks of size 4×4. These image blocks are projected to arbitrary dimensions through a linear embedding layer. Then, the dimensionally projected features are input into the SwinTransformer module to further extract the features in these image blocks and retain the spatial information. Finally, in the multi-scale fusion module, the detail information from different scales and levels is fused with the global information to form a more complete and diverse feature representation.

[0045] Swin Transformer is mainly responsible for feature representation learning. The core module consists of layer normalization, multi-head attention mechanism, residual connection and multi-layer perceptron MLP. The calculation process is as follows:

[0046] (8)

[0047] (9)

[0048] (10)

[0049] (11)

[0050] Where: Indicates Output features of each stage; and They represent the features output after the window-based multi-head attention mechanism (W-MSA) and the moving window multi-head attention mechanism (SW-MSA) and residual connection, respectively. and Respectively and The features are output after passing through a multi-layer perceptron and residual connections.

[0051] After the multi-head attention mechanism, the calculation formula of attention weight is:

[0052] (12)

[0053] Where: Matrices representing queries, keys, and values, respectively. is the number of feature dimensions; is the relative position encoding matrix, The function is used to normalize the attention weights;

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1) Combination of denoising diffusion probability model with multi-scale module and attention module: The denoising diffusion probability model is used to gradually remove noise, accurately capture the details in the image, and enhance the model's adaptability to the complex structure of ultrasound images;

[0056] 2) Multi-scale dynamic condition module: This module extracts multi-scale features from ultrasound images as auxiliary information to guide the image generation process, thereby improving the fusion of contextual information across different scales;

[0057] 3) Gaussian Cross-Fusion Attention Module: This module enhances the correlation between features, realizes the fusion and interaction of semantic features, and effectively reduces the information incompatibility problem that may occur when directly fusing the encoder features with the features of the multi-scale dynamic condition module. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flow chart of the method of the present invention;

[0059] Figure 2 It is a schematic diagram of the TransFuse module of the present invention;

[0060] Figure 3 This is a schematic diagram of the core modules of the Swin Transformer of the present invention;

[0061] Figure 4It is a schematic diagram of a multi-scale fusion module of the present invention;

[0062] Figure 5 It is a schematic diagram of the segmentation result of the BUSI data set of the present invention;

[0063] Figure 6 It is a schematic diagram of the DDTI dataset segmentation results of the present invention. DETAILED DESCRIPTION

[0064] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.

[0065] The present invention is specifically implemented through the following technical solutions:

[0066] Figure 1 This is a flow chart of the technical solution of the present invention (DMA-USeg).

[0067] The overall design of the technical solution of the present invention is: adopting an encoder-decoder architecture, which consists of three core parts: a denoising diffusion probability model, a multi-scale dynamic condition module, and a Gaussian cross-fusion attention module. In the forward diffusion stage, the denoising diffusion probability model gradually adds Gaussian noise to the initial image. In the back diffusion stage, the image is sampled from the standard normal distribution. , and input into the encoder to extract features.

[0068] In the encoder part, first use the Stem module to Perform shallow feature extraction. This module consists of 3×3 convolution, batch normalization (Batch Normalization, BN) and Rectified Linear Unit (Rectified Linear Unit, ReLU). to In the first stage, four convolution modules are used to further extract shallow features and downsample them. The multi-scale dynamic condition module captures the long-range dependencies of different scales from ultrasound images, and the two parts of features are effectively fused through the Gaussian cross-fusion attention module.

[0069] In the decoder part, first In the first stage, the features of the Gaussian cross attention module are concatenated with the skip connection features of the multi-scale dynamic condition module through channel concatenation and 1×1 convolution, and the transformed image maintains the same resolution as the input. After that, it is upsampled by 3×3 transposed convolution and input to the next stage. arrive In the second stage, the input features and the jump connection features of the corresponding layer of the dynamic condition module are gradually restored to the original resolution through channel splicing, 1×1 convolution and transposed convolution, forming an optimized end-to-end network architecture.

[0070] The implementation steps and key algorithms of the technical solution of the present invention are as follows:

[0071] Step 1: Use the denoising diffusion probability model to remove image noise and capture important detail information.

[0072] The denoising diffusion probability model is a type of generative model based on the Markov chain concept. The model consists of a forward diffusion stage and a backward diffusion stage. Aiming at the specific task of medical ultrasound image segmentation, the present invention designs a method to convert Gaussian noise into Gradually add to the input image ,implement After time steps, the original image Transformed into an image with almost all Gaussian noise , the forward diffusion process can be defined as:

[0073] (1)

[0074] (2)

[0075] Where: and Represents the step length and Noise image at the time ; is the coefficient of the noise level, which controls the proportion of noise added and varies with the step size decreases with the increase of is the identity matrix; represents a Gaussian distribution.

[0076] In the back diffusion process, the noise image The noise is gradually removed through a series of iterative steps and restored to the original image. . Assume that there is a reverse process distribution , which is a distribution with learnable parameters The neural network is parameterized to accurately fit the mean of the image at each time step and variance .

[0077] (3)

[0078] (4)

[0079] Back diffusion process It can be expressed as:

[0080] (5)

[0081] In the training process of the denoising diffusion probability model, the key step is to minimize the estimated distribution and the true posterior distribution The KL divergence between the two probability distributions is measured by adjusting the neural network parameters to make the predicted distribution As close as possible to minimizing the estimated distribution The training objective of DDPM can be expressed as:

[0082] (6)

[0083] The loss function of the DMA-USeg model in this paper can be expressed as:

[0084] (7)

[0085] in, Segment the image for the dynamic conditional module; represents the fitting function of the neural network; E x0 , ε Indicates the mathematical expectation of the calculated value in parentheses.

[0086] Step 2: Use a multi-scale dynamic condition module to improve image contrast and the ability to integrate contextual information of different scales.

[0087] In traditional image segmentation based on denoising diffusion probability model, noisy images are usually used as the only input information. Although this method can enhance the target area to a certain extent, its accuracy is still limited. In addition, although medical ultrasound images contain segmentation target area information, it is still very difficult to effectively separate the background from the target area. In order to solve this problem, the present invention uses a multi-scale dynamic condition module to capture image detail information, highlight the target area features, and use it as auxiliary information to guide the image generation process in the decoder. The multi-scale dynamic condition module is as follows: Figure 1 As shown in (a), a hierarchical architecture is adopted. The module first passes through a Stem layer, and then adopts four feature extraction modules to Perform step-by-step feature extraction and downsampling. Finally, the extracted features are input into the TransFuse module. The Stem layer adjusts the input image through a 7×7 convolution operation. The resolution is . The feature extraction module consists of multiple bottleneck layers, and the number of stacking times of bottleneck layers in each module is 3, 4, 6 and 3 respectively. Each bottleneck layer includes two 3×3 convolutions and one 1×1 convolution. This structure not only increases the number of channels, but also enriches the features through different convolution operations. As the network deepens, the resolution of the feature map gradually decreases, and the number of channels and receptive field gradually increase. Each feature extraction module not only passes the features to the next module, but also inputs them into the TransFuse module to achieve an effective combination of local information and global information. This module significantly improves the model's ability to capture the feature information of the target area in ultrasound images.

[0088] The above TransFuse module is as follows Figure 2 As shown in the figure, this module improves the expressiveness of features by combining the advantages of CNN and Swin Transformer. At the same time, it fuses multi-scale features to generate a more comprehensive feature representation, enhancing the hierarchy of features and the richness of information.

[0089] The present invention first divides the output features of the CNN module into non-overlapping image blocks of size 4×4. These image blocks are projected into arbitrary dimensions through a linear embedding layer. Then, the dimensionally projected features are input into the SwinTransformer module, which uses its powerful context understanding ability to further extract features from these image blocks and retain spatial information. Finally, in the multi-scale fusion module, detail information from different scales and levels is fused with global information to form a more complete and diverse feature representation, allowing the model to more effectively capture important detail information and global information in the image.

[0090] The core modules of Swin Transformer are as follows: Figure 3 As shown, it is mainly responsible for feature representation learning. The core module consists of layer normalization, multi-head attention mechanism, residual connection and multi-layer perceptron MLP. The calculation process is as follows:

[0091] (8)

[0092] (9)

[0093] (10)

[0094] (10)

[0095] Where: Indicates Output features of each stage; and They respectively represent the output features after the window-based multi-head attention mechanism (W-MSA) and the moving window multi-head attention mechanism (SW-MSA) and residual connection. and Respectively and The features are output after passing through a multi-layer perceptron and residual connections.

[0096] After the multi-head attention mechanism, the calculation formula of attention weight is:

[0097] (12)

[0098] Where: Represent the query, key, and value matrices respectively. is the number of feature dimensions; is the relative position encoding matrix, The function is used to normalize the attention weights;

[0099] Then, through the multi-scale fusion module, the information between different scales interacts and captures the global context dependency. Make full use of deep features to solve the problem of low contrast of ultrasound images. First, the output features of the Transformer module are resized after 1×1 convolution operation and expressed as , in the CNN module The output characteristics of the stage are expressed as . Then the soft gating that controls the integration degree of features at different scales It can be expressed as:

[0100] (13)

[0101] in, Represented as a linear map; Expressed as Activation function; Represented as channel concatenation.

[0102] Next, we use element-wise multiplication to transform the soft gate and Merge and then Element addition is performed to enhance features and preserve details. It can be expressed as:

[0103] (14)

[0104] in, It is the element dot product operation; is element addition, such as Figure 4 shown.

[0105] Step 3: Use the Gaussian cross-fusion attention module to solve the incompatibility problem when fusing encoder features and dynamic condition module features.

[0106] To ensure that the encoder output characteristics And the multi-scale dynamic condition module output features To keep the same in space and frequency, the present invention uses Gaussian cross fusion attention module to solve the incompatibility problem that may occur when directly merging the two features. First, the two parts of the input features are constrained to be Gaussian distributions. and , then and After normalization, it is used as the query variable and Then add the elements of the two feature maps to get the key K and value V, and feed them into the multi-head attention mechanism to get the output features after interaction. The calculation result is expressed as:

[0107] (15)

[0108] (16)

[0109] Finally, the features and After layer normalization and multi-layer perceptron processing, the network performance and stability are improved. Then, through channel splicing and 3×3 convolution processing, the size of the feature is adjusted and input into the decoder to complete the final decoding output. The calculation results are as follows:

[0110] (17)

[0111] in, is the convolution operation; is a multi-layer perceptron; Normalize the layer;

[0112] The datasets used in this paper are the BUSI ultrasound breast dataset and the DDTI ultrasound thyroid dataset. The BUSI dataset includes 780 breast ultrasound images of 600 female patients aged between 25 and 75 years old collected in 2018. The average size of each image is The DDTI dataset contains 637 pixel-level labeled ultrasound thyroid images from a single device provided by Pedraza et al., including 487 malignant lesions, 210 benign lesions, and 133 normal ultrasound images.

[0113] In order to comprehensively and accurately evaluate the performance of the model, the experiment of this invention adopts six commonly used image segmentation evaluation indicators, including mean intersection over union (MIoU), F1 score (F1 Score), precision (Precision), recall rate (Recall) and accuracy (Accuracy). The meanings of the evaluation indicators are as follows: MIoU measures the overlap between the predicted segmentation and the actual segmentation area; Precision calculates the proportion of correctly predicted samples in the total samples; F1 Score is the harmonic mean of precision and recall rate, which comprehensively reflects the balance between precision and recall rate; Recall measures the proportion of correctly predicted positive samples in all actually positive samples; Accuracy predicts the proportion of correctly predicted samples in the total samples.

[0114] (18)

[0115] (19)

[0116] (20)

[0117] (twenty one)

[0118] (twenty two)

[0119] Where: TP is the number of pixels correctly identified as the target area; TN is the number of pixels correctly identified as the background; FP is the number of pixels incorrectly identified as the target area; TN is the number of pixels incorrectly identified as the background.

[0120] Table 1 Evaluation results of segmentation algorithms of different network models under BUSI dataset. ↑ indicates that the larger the value, the better the corresponding segmentation effect; ↓ indicates that the larger the value, the better the corresponding segmentation effect. and Δ represent the percentage of decrease and increase of the evaluation index value compared with the proposed model.

[0121] Table 1 Evaluation results of different network model segmentation algorithms on the BUSI dataset

[0122]

[0123] Table 2 Evaluation results of different network model segmentation algorithms under DDTI dataset

[0124]

[0125] It can be seen from Tables 1 and 2 that the precision and accuracy of the networks using only CNN or Transformer architecture (Attention U-Net, Swin U-Net) are lower than that of the hybrid CNN and Transformer structure (UConvTrans). In addition, MedSegDiff based on the DDPM network has significantly better segmentation performance than other methods, achieving 96.87% and 97.92% accuracy on the two datasets respectively.

[0126] The method DMA-USeg of the present invention not only adopts a DDPM-based architecture, but also introduces a multi-scale dynamic condition module that mixes CNN and Transformer. Its performance is the best among all the compared algorithms. On the BUSI dataset, the accuracy value of the DMA-USeg method is 5.78% higher than that of Attention U-Net and 0.66% higher than that of MedSegDiff. On the DDTI dataset, the accuracy of DMA-USeg is 3.77% higher than that of Trans-UNet, 1.74% higher than that of TransAttUnet, and 0.44% higher than that of MedSegDiff. These results demonstrate the advantages of the DMA-USeg method in medical ultrasound image segmentation.

[0127] Figure 5 and Figure 6 The visualization results of different segmentation methods on the BUSI dataset and the DDTI dataset are shown. The red part in the figure represents the comparison between the lesion in the local area and the ground truth image. It can be seen from the figure that compared with other methods, AttentionUNet, TransUNet, Swin U-Net and TransAttUnet all have the problem of not being able to accurately segment the lesion area. Although UConvTrans can perform segmentation more accurately, it is still insufficient in detail extraction. MedSegDiff uses a diffusion model for image segmentation, which retains good boundary information, but its detail extraction ability is still weaker than this model. This shows that the DMA-USeg method can pay attention to small boundary information and prevent the loss of features, and obtain segmentation results that are closer to the ground truth image.

[0128] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0129] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A medical ultrasound image segmentation system based on a diffusion model, multi-scale and attention module, including an encoder-decoder architecture, consisting of three core parts: a denoising diffusion probability model, a multi-scale dynamic condition module, and a Gaussian cross-fusion attention module, characterized by: In the forward diffusion phase, the denoising diffusion probability model gradually adds Gaussian noise to the initial image. In, the gradual degradation process of the image is simulated; In the back diffusion stage, samples are drawn from the standard normal distribution. , and input into the encoder to extract features; In the encoder part, first use the Stem module to For shallow feature extraction, the Stem module consists of 3×3 convolution, batch normalization and linear correction units. Subsequently, the features extracted by the Stem module are sequentially to Four feature extraction modules, in to In the second stage, four convolution modules are used to further extract shallow features and downsample them; The multi-scale dynamic condition module captures the long-range dependencies of different scales from ultrasound images, and effectively fuses the features of the multi-scale dynamic condition module and the encoder through the Gaussian cross-fusion attention module; The multi-scale dynamic condition module adopts a hierarchical architecture; the module first passes through a Stem layer, and then uses four feature extraction modules to Perform step-by-step feature extraction and downsampling; finally, input the extracted features into the TransFuse module The above-mentioned TransFuse module improves the expressiveness of features by combining the advantages of CNN and Swin Transformer; In the decoder part, first In the first stage, the features of the Gaussian cross attention module and the skip connection features of the multi-scale dynamic condition module are concatenated and 1×1 convolved. The transformed image maintains the same resolution as the input. After that, it is upsampled by 3×3 transposed convolution and input to the next stage. arrive stage, where De1 to The stages are connected sequentially, and the input features and the jump connection features of the corresponding layer of the dynamic condition module are gradually restored to the original resolution through channel splicing, 1×1 convolution and transposed convolution, forming an optimized end-to-end network architecture.

2. A medical ultrasound image segmentation method, applied to the medical ultrasound image segmentation system based on diffusion model, multi-scale and attention module as claimed in claim 1, characterized in that: The specific steps are as follows: Step 1: Use the denoising diffusion probability model to remove image noise and capture important details The denoising diffusion probability model consists of a forward diffusion stage and a reverse diffusion stage; In the forward diffusion process, Gaussian noise Gradually add to the input image ,implement After time steps, the original image Transformed into an image with almost all Gaussian noise , the forward diffusion process can be defined as: (1) (2) Where: and Represents the step length and Noise image at the time ; is the coefficient of the noise level, which controls the proportion of noise added and varies with the step size decreases with the increase of is the identity matrix; represents Gaussian distribution; In the back diffusion process, the noise image The noise is gradually removed through a series of iterative steps and restored to the original image. , assuming there is a reverse process distribution , which is a distribution with learnable parameters The neural network is parameterized to accurately fit the mean of the image at each time step and variance ; (3) (4) Back diffusion process It can be expressed as: (5) In the training process of the denoising diffusion probability model, the key step is to minimize the estimated distribution and the true posterior distribution The KL divergence between the two probability distributions is measured by adjusting the neural network parameters to make the predicted distribution Close to minimizing the estimated distribution , the training objective of the denoising diffusion probability model is expressed as: (6) The loss function of the DMA-USeg model composed of the entire medical ultrasound image segmentation system can be expressed as: (7) in, Segment the image for the dynamic conditional module; represents the fitting function of the neural network; E x0 , ε Indicates the mathematical expectation of the calculated value in brackets; Step 2: Use a multi-scale dynamic condition module to improve image contrast and the ability to integrate contextual information at different scales A multi-scale dynamic conditional module is used to capture image detail information, highlight the target area features, and serve as auxiliary information to guide the image generation process in the decoder. First, a Stem layer is used, and then four feature extraction modules are used. to Perform step-by-step feature extraction and downsampling, and finally, input the extracted features into the TransFuse module; Then, through the multi-scale fusion module, the information between different scales interacts, captures the global context dependency, makes full use of deep features, and solves the problem of low contrast of ultrasound images. First, the output features of the Transformer module are resized by 1×1 convolution operation and expressed as , in the CNN module The output characteristics of the stage are expressed as , then the soft gating that controls the integration degree of features at different scales It can be expressed as: (13) in, Represented as a linear map; Expressed as Activation function; It is represented as channel splicing; Next, we use element-wise multiplication to transform the soft gate and Merge and then Perform element addition to enhance features and preserve details, and fuse images It can be expressed as: (14) in, It is the element dot product operation; Addition for elements; Step 3: Use the Gaussian cross-fusion attention module to solve the incompatibility problem when fusing encoder features and dynamic condition module features The Gaussian cross-fusion attention module is used to solve the incompatibility problem that may occur when directly merging the two features. First, the two parts of the input features are constrained to be Gaussian distributions. and , then and After normalization, it is used as the query variable and , then add the elements of the two feature maps to get the key K and value V, and feed them into the multi-head attention mechanism to get the output features after interaction. The calculation result is expressed as: (15) (16) in, Represent the query, key, and value matrices respectively. is the number of feature dimensions; is the relative position encoding matrix, The function is used to normalize the attention weights; Finally, the features and After layer normalization and multi-layer perceptron processing, the network performance and stability are improved. Then, through channel splicing and 3×3 convolution processing, the size of the features is adjusted and input into the decoder to complete the final decoding output. The calculation results are as follows: (17) in, is the convolution operation; is a multi-layer perceptron; Normalize the layer.

3. A medical ultrasound image segmentation method according to claim 2, characterized in that: The TransFuse module in step 2 contains CNN and Swin Transformer modules. The specific process is as follows: First, the output features of the CNN module are divided into non-overlapping image blocks of size 4×4. These image blocks are projected to arbitrary dimensions through a linear embedding layer. Then, the dimensionally projected features are input into the Swin Transformer module to further extract the features in these image blocks and retain the spatial information. Finally, in the multi-scale fusion module, the detail information from different scales and levels is fused with the global information to form a more complete and diverse feature representation. Swin Transformer is mainly responsible for feature representation learning. The core module consists of layer normalization, multi-head attention mechanism, residual connection and multi-layer perceptron MLP. The calculation process is as follows: (8) (9) (10) (11) Where: Indicates Output features of each stage; and They represent the features output after the window-based multi-head attention mechanism (W-MSA) and the moving window multi-head attention mechanism (SW-MSA) and residual connection, respectively. and Respectively and The features are output after passing through the multi-layer perceptron and residual connection; After the multi-head attention mechanism, the calculation formula of attention weight is: (12) Where: Represent the query, key, and value matrices respectively. is the number of feature dimensions; is the relative position encoding matrix, Function is used to normalize the attention weights.

Citation Information

Patent Citations

  • Image denoising method based on channel attention mechanism and feature pyramid

    CN110766632A

  • Image generation method and apparatus, computer readable storage medium, electronic device and computer program product

    WO2024131597A1