Medical image segmentation method based on region-boundary mutual learning diffusion model

By introducing a region-boundary mutual learning diffusion model in medical image segmentation, combining the region-perception module and the boundary-perception denoising network, the characteristics diversity and boundary ambiguity problems in the prior art are solved, and the segmentation accuracy is significantly improved.

CN120163832AActive Publication Date: 2025-06-17SHAANXI UNIV OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510154474.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-17
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing diffusion model ignores feature diversity when extracting conditional features, is unable to integrate features at different levels, and is difficult to deal with boundary blur in medical image segmentation, resulting in low segmentation accuracy.

Method used

A medical image segmentation method based on the region-boundary mutual learning diffusion model is proposed. By constructing the network model RBML-Diff, combining the region perception module and the boundary perception denoising network, multi-scale fusion of regional features and boundary features is achieved.

Benefits of technology

It significantly improves the recognition ability of different sizes and morphological lesions, reduces the impact of feature diversity on segmentation performance, generates more accurate segmentation image results, and solves the problems of large-scale changes and boundary ambiguity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163832A_ABST
    Figure CN120163832A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation method based on a region-boundary mutual learning diffusion model. The method comprises the following steps: 1, constructing a network model RBML-Diff and initializing hyper-parameters; 2, preprocessing the unmarked medical image, the marked medical image and the marked medical image truth value; 3, in the forward diffusion process of the network model, noise is added to the truth value of the marked medical image, and a noise mask graph is obtained; 4, inputting the noise mask image and the unmarked medical image into a conditional encoder, and outputting a feature map; 5, an intra-layer feature interaction module and an inter-layer feature cooperation module in the region sensing module carry out region sensing and feature refinement on a feature map output by a condition encoder; 6, extracting the boundary of the unlabeled medical image by using a Laplace operator, and repeating the region-boundary mutual learning module and the up-sampling operation to obtain a prediction result; and 7, gradually carrying out reverse diffusion denoising on a prediction result to obtain a segmented mask graph. The method effectively solves the problems of large-scale change and boundary fuzziness in medical image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and specifically relates to a medical image segmentation method based on a region-boundary mutual learning diffusion model. Background Art

[0002] Medical image segmentation technology plays a crucial role in the field of clinical medicine. It can extract the target regions or tissues in the image, thereby providing key basis for doctors to analyze pathological features and helping doctors diagnose diseases more accurately.

[0003] Traditional medical image analysis relies on doctors' experience and manual segmentation, which is not only time-consuming and laborious, but also easily affected by subjective factors. With the progress of medical image segmentation technology, automatic segmentation has become a research hotspot. However, due to the low contrast, the diversity of the shape and size of the target region, and the ambiguity of the boundary between the target region and the surrounding normal tissues in medical images, automatic medical image segmentation remains a highly challenging task.

[0004] With the development of deep learning technology, researchers have proposed a variety of automatic medical image segmentation models, which are mainly divided into convolutional neural network (CNN)-based and Transformer-based segmentation models. CNN-based segmentation models often fail to fully capture the relationships and interdependencies between different regions in medical images and lack a deep understanding of the global image. Transformer-based segmentation models are good at capturing global information and enhancing the network's understanding of global context, but they tend to ignore local details and are insufficient in identifying tiny lesion regions and dealing with fuzzy boundaries.

[0005] Diffusion models have excellent feature learning capabilities in the reverse denoising process and have received extensive attention. Different from CNN and Transformer, diffusion models model the evolution process of the target region as a parameterized Markov chain, which helps to learn the statistical distribution characteristics of the target region, thereby segmenting images more precisely. Some scholars have applied diffusion models to image segmentation tasks and achieved good segmentation results. For example: Amit et al. first applied the diffusion model to the image segmentation problem and introduced the concept of multiple generations, effectively improving the diffusion model and enhancing the performance of image segmentation; Wu et al. first integrated Transformer into the diffusion model and used it for medical image segmentation, and proposed SSFormer to simulate the interaction between noise and semantic features; Chen et al. proposed a conditional Bernoulli diffusion model. Different from the diffusion model using Gaussian noise, it uses Bernoulli noise as the diffusion kernel, which can effectively enhance the model's ability to handle binary segmentation tasks and has a faster sampling speed.

[0006] Although diffusion models have shown great potential and application value in image segmentation tasks, however, when applying diffusion models to medical image segmentation tasks, there are still four problems: (1) When extracting conditional features, the complexity of lesions of different sizes cannot be fully considered, resulting in poor segmentation performance for lesions of different sizes and relatively hidden areas; (2) Extracting conditional features from the original image or noise mask ignores the impact of feature diversity on segmentation performance; (3) Unable to integrate features at different levels according to the diversity of semantic features, resulting in the inability to fully capture the key semantic information of the lesion area boundary and low segmentation accuracy; (4) For the difficult-to-distinguish boundary problem, existing diffusion models cannot perform effective feature fusion, resulting in poor segmentation effects for lesions with blurred boundaries. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a medical image segmentation method based on a region-boundary mutual learning diffusion model, which can accurately identify lesions of various shapes, sizes, and blurred boundaries, and improve the segmentation accuracy of medical images.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions to implement:

[0009] A medical image segmentation method based on a region-boundary mutual learning diffusion model includes the following steps:

[0010] Step 1, construct a network model RBML-Diff and initialize hyperparameters;

[0011] The network model RBML-Diff includes a conditional encoder, a region perception module, and a boundary perception denoising network, where: the region perception module includes an intra-layer feature interaction module and an inter-layer feature collaboration module; the boundary perception denoising network includes a segmentation encoder and a region-boundary mutual learning module;

[0012] Step 2, preprocess the unlabeled medical image X to be segmented, the labeled medical image x, and the labeled medical image ground truth x0;

[0013] Step 3, input the preprocessed labeled medical image x and the labeled medical image ground truth x0 into the network model RBML-Diff, and use the forward diffusion process of the network model RBML-Diff to add noise to x0. After T-step iteration, obtain the noisy mask image x T ;

[0014] Step 4, input the noisy mask image x T and the unlabeled medical image X to be segmented into the conditional encoder for feature extraction, and obtain the feature maps and

[0015] Step 5: Input and into different intra-layer feature interaction modules in the region perception module respectively. The intra-layer feature interaction module performs feature mining on different directions and scales for the feature maps and to obtain feature maps and Input the feature maps and into two inter-layer feature collaboration modules of the region perception module respectively. The inter-layer feature collaboration module performs cross-layer feature refinement on the feature maps and to obtain refined multi-scale region feature maps and

[0016] Step 6: Extract the boundary of the unannotated medical image X to be segmented through the Laplace operator to obtain the boundary feature map and use the segmentation encoder in the boundary-aware denoising network to extract the feature of the noise mask map x T . Concatenate the output feature map of the segmentation encoder and the feature map , and after upsampling, obtain the predicted feature map

[0017] Set the predicted feature map output by the region-boundary mutual learning module in the boundary-aware denoising network as First, perform inversion on the predicted feature map containing respectively and extract the boundary through the Laplace operator to obtain the reverse attention map and the boundary attention map Then, through the region-boundary mutual learning module, the region features and are respectively cross-learned with the boundary feature map , the next-level reverse attention map and the boundary attention map . After refinement in the channel and space and then upsampling, obtain Repeat the operations of passing through the region-boundary mutual learning module and upsampling, and finally obtain the prediction result

[0018] Step 7: Gradually perform reverse diffusion denoising on the prediction result to obtain the final segmentation mask map

[0019] Further, the initialization hyperparameters in step 1 include: setting the training batch size to 32, setting the number of training iterations to 150, and during the training process, adopting a learning rate decay strategy and setting the initial learning rate to 0.0001.

[0020] Further, the specific process of step 2 is as follows:

[0021] First, perform data augmentation on the unlabeled medical image X to be segmented, and then perform normalization processing on the augmented unlabeled medical image X to be segmented;

[0022] Adjust the sizes of the labeled medical image x and the labeled medical image ground truth x0 to 256×256.

[0023] Further, the data augmentation includes random scaling, horizontal and vertical flipping, randomly selecting an angle for rotation, randomly filling with noise, and random cropping.

[0024] Further, the specific process of step 3 is as follows:

[0025] During the forward diffusion process of the network model RBML-Diff, gradually add Gaussian noise to the original labeled medical image ground truth x0 through continuous T-step iterations. The forward diffusion process is expressed as:

[0026]

[0027] where T represents the number of diffusion steps, x t and x t-1 represent the noise mask maps at time t and t - 1 respectively during the forward diffusion process, and x 1:t represents the noise mask map at t = 1, 2,..., T during the forward diffusion process;

[0028] During each iteration, the formula for adding Gaussian noise is expressed as:

[0029]

[0030] where β t is the parameter controlling Gaussian noise during the forward diffusion process, I is an n×n identity matrix, is the Gaussian distribution;

[0031] The process of sampling at any time t is expressed as:

[0032]

[0033] α t = 1 - β t

[0034] where α t is related to βt The relevant attenuation coefficient represents the product of all α during the forward diffusion process t .

[0035] Furthermore, the specific process of step 4 is as follows:

[0036] Step 4.1: Use the noise mask map x T and the unlabeled medical image X to be segmented as the input of the conditional encoder. The conditional encoder includes four stages, and each stage includes a fusion embedding layer and a Transformer encoder. The Transformer encoder of the first stage outputs a feature map

[0037] Step 4.2: Use the feature map as the second stage of the conditional encoder. The Transformer encoder of the second stage outputs a feature map

[0038] Step 4.3: Input the feature map into the third stage of the conditional encoder. The Transformer encoder of the third stage outputs a feature map

[0039] Step 4.4: Input the feature map into the fourth stage of the conditional encoder. The Transformer encoder of the fourth stage outputs a feature map

[0040] Furthermore, the specific process of step 5 is as follows:

[0041] Step 5.1: Input the feature maps and into the four intra-layer feature interaction modules of the region perception module respectively to achieve feature interaction in different directions and scales. The four intra-layer feature interaction modules respectively output feature maps and denoted as which is expressed as:

[0042]

[0043] wherein represents the i-th feature map generated by the backbone network, k represents the number of branches, represents the output of the k-th branch, represents element-wise addition, Conv s (·) represents asymmetric convolution, Conv1(·) represents 1×1 convolution, Conv3(·) represents 3×3 convolution, Represents the splicing of three branches, and ReLU is the ReLU activation function;

[0044] Step 5.2: Input the feature map and into the two inter-layer feature collaboration modules of the region perception module respectively to strengthen the shared feature expression, and obtain the output feature maps and which is expressed as:

[0045]

[0046] In the formula, κ(·) represents two ordinary convolutions, tanh(·) represents the hyperbolic tangent function, ||·||2 represents the L2 norm, represents element-wise multiplication, (·) T represents matrix transpose, represents the feature map after being processed by two 1×1 convolutions;

[0047] Furthermore, the specific process of step 6 is as follows:

[0048] Step 6.1: Extract the boundary features of the unlabeled medical image X to be segmented through the Laplacian operator, and obtain the Laplacian boundary feature map which is expressed as:

[0049]

[0050] f l = L1(X)

[0051] L k = X k - u(X k+1 )

[0052]

[0053] In the formula, X represents the input unlabeled medical image to be segmented, g(·) represents the convolution operator with a Gaussian filter, d(·) represents downsampling, u(·) represents upsampling, L k represents the Laplacian image of the k-th layer of the Laplacian pyramid, and X k represents the feature map of the unlabeled medical image X to be segmented after Gaussian convolution and downsampling;

[0054] Step 6.2: Input the noise mask map x T into the segmentation encoder of the boundary-aware denoising network, and then splice the output feature map of the segmentation encoder and the feature map After upsampling, obtain the predicted feature map

[0055] Step 6.3: Set the predicted feature map output by the region-boundary mutual learning module in the boundary-aware denoising network as Take the negation of the predicted feature map containing to highlight the insignificant regions and obtain the reverse attention map Meanwhile, use the Laplacian operator to extract the boundary of the predicted feature map to obtain the boundary attention map The reverse attention map and the boundary attention map are expressed as:

[0056]

[0057] In the formula, L0() represents the Laplacian operator;

[0058] Step 6.4: Through the region-boundary mutual learning module, the region feature map is cross-learned with the Laplacian boundary feature map the next-level reverse attention map and the boundary attention map respectively, which is expressed as:

[0059]

[0060] In the formula, σ() represents the Sigmoid activation function;

[0061] Step 6.5: Refine the output feature map in the channel and space, which is expressed as:

[0062]

[0063] In the formula, CA represents the channel attention operation, and SA represents the spatial attention operation;

[0064] Step 6.6: Upsample to obtain the result Repeat passing through the region-boundary mutual learning module and the upsampling operation to obtain the prediction result

[0065] Furthermore, the specific process of reverse diffusion denoising in step 7 is expressed as:

[0066]

[0067] In the formula, x t and x t-1 respectively represent the noise mask maps at times t and t-1 in the reverse diffusion process, and μ θ () represents at time t and the current noise mask map xt The mean predicted by the network model under the condition of σ 2 represents the variance of the Gaussian distribution in the reverse diffusion process, β t is a parameter for controlling Gaussian noise in the reverse diffusion process, α t is related to β t is the decay coefficient represents the product of all α t in the reverse diffusion process is the segmentation mask image

[0068] Compared with the prior art, the present invention has the following technical effects:

[0069] The network model RBML-Diff proposed by the present invention includes a region perception module and a boundary perception denoising network, wherein: the region perception module includes an intra-layer feature interaction module and an inter-layer feature collaboration module. The intra-layer feature interaction module focuses on the diversity within a single layer's characteristics, while the inter-layer feature collaboration module aims to fully utilize the potential of multi-layer characteristics, thereby facilitating the region perception module to deeply mine and improve multi-scale region features, significantly enhancing the recognition ability of the morphological diversity of the segmentation target, reducing the impact of feature diversity on the segmentation performance, and being more suitable for segmenting complex lesion images of different sizes and shapes; the boundary perception denoising network uses the Laplacian morphological operator to accurately capture boundary information, and uses the region-boundary mutual learning module to fuse multi-scale region features and high-frequency boundary features. Through this mutual learning, RBML-Diff fully utilizes the multi-level features from regions and boundaries in the image to capture the key semantic information of the lesion region boundary, and then generates an accurate segmentation image result. In short, the present invention effectively solves the problems of large-scale changes and boundary ambiguity in medical image segmentation, can accurately extract the contour of the target to be segmented in medical images, provides technical support for early clinical diagnosis and treatment, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 is a schematic structural diagram of the network model RBML-Diff of the present invention;

[0071] Figure 2 is a schematic structural diagram of the intra-layer feature interaction module of the present invention;

[0072] Figure 3 is a schematic structural diagram of the inter-layer feature collaboration module of the present invention;

[0073] Figure 4 is a schematic structural diagram of the region-boundary mutual learning module of the present invention;

[0074] Figure 5It is the visualization results of the network model RBML-Diff of the present invention and the existing network on the polyp segmentation dataset. Detailed implementation manners

[0075] The following further elaborates on the specific content of the present invention in conjunction with embodiments.

[0076] A medical image segmentation method based on a region-boundary mutual learning diffusion model includes the following steps:

[0077] Step 1, construct the network model RBML-Diff as shown in Figure 1 and initialize the hyperparameters;

[0078] The network model RBML-Diff includes a conditional encoder, a region perception module, and a boundary perception denoising network, where: the conditional encoder includes four stages, and each stage includes a fusion embedding layer and a Transformer encoder; the region perception module includes four intra-layer feature interaction modules as shown in Figure 2 and two inter-layer feature collaboration modules as shown in Figure 3 ; the boundary perception denoising network includes a segmentation encoder and three region-boundary mutual learning modules as shown in Figure 4 ;

[0079] The initialized hyperparameters include: setting the training batch size to 32, setting the number of training iterations to 150, and during the training process, adopting a learning rate decay strategy and setting the initial learning rate to 0.0001;

[0080] Step 2, preprocess the unlabeled medical image X to be segmented, the labeled medical image x, and the labeled medical image ground truth x0. The specific process is as follows:

[0081] Perform data augmentation on the unlabeled medical image X to be segmented, including random scaling, horizontal and vertical flipping, randomly selecting an angle for rotation, randomly filling noise, and random cropping, and then perform normalization processing;

[0082] Adjust the sizes of the labeled image x and the ground truth x0 of the labeled image to 256×256;

[0083] Step 3, input the preprocessed labeled medical image x and the labeled medical image ground truth x0 into the network model RBML-Diff, and use the forward diffusion process of the network model RBML-Diff to perform noise addition processing on x0. After T-step iteration, obtain the noisy mask image x T , and the specific process is as follows:

[0084] During the forward diffusion process of the network model RBML-Diff, Gaussian noise is gradually added to the original labeled medical image ground truth x0 through continuous T-step iterations. The forward diffusion process is expressed as:

[0085]

[0086] where T represents the number of diffusion steps, x t and x t-1 represent the noise mask maps at time t and t - 1 respectively during the forward diffusion process, and x 1:t represents the noise mask map for t = 1, 2,..., T during the diffusion process;

[0087] During each iteration, the formula for adding Gaussian noise is expressed as:

[0088]

[0089] where β t is the parameter controlling Gaussian noise during the forward diffusion process, I is an n×n identity matrix, is the Gaussian distribution;

[0090] The process of sampling at any time t is expressed as:

[0091]

[0092] α t = 1 - β t

[0093] where α t is the decay coefficient related to β t , represents the product of all α t during the forward diffusion process;

[0094] Step 4: Input the noise mask map x T and the unlabeled medical image X to be segmented into the conditional encoder for feature extraction, obtaining the feature maps and The specific process is as follows:

[0095] Step 4.1: Take the noise mask map x T and the unlabeled medical image X to be segmented as the input of the conditional encoder. The Transformer encoder in the first stage outputs the feature map

[0096] Step 4.2: Input the feature map into the second stage of the conditional encoder. The Transformer encoder in the second stage outputs the feature map

[0097] Step 4.3. Input the feature map into the third stage of the conditional encoder. The Transformer encoder in the third stage outputs a feature map

[0098] Step 4.4. Input the feature map into the fourth stage of the conditional encoder. The Transformer encoder in the fourth stage outputs a feature map

[0099] Step 5. Input and into the four intra-layer feature interaction modules of the region perception module as shown in Figure 2 . The intra-layer feature interaction modules perform feature mining in different directions and scales on the features and to obtain the feature maps and Input the feature map into one of the inter-layer feature collaboration modules in the region perception module as shown in Figure 3 . Input into the other inter-layer feature collaboration module in the region perception module as shown in Figure 3 . The two inter-layer feature collaboration modules respectively perform cross-layer feature refinement on the feature maps and to obtain the refined multi-scale region feature maps and The specific process is as follows:

[0100] Step 5.1. Respectively input the feature maps and into the four intra-layer feature interaction modules of the region perception module to achieve feature interaction in different directions and scales. The four intra-layer feature interaction modules respectively output the feature maps and Denoted as Expressed as:

[0101]

[0102] In the formula, represents the i-th feature map generated by the backbone network, k represents the number of branches, represents the output of the k-th branch, represents element-wise addition, Conv s (·) represents asymmetric convolution, Conv1(·) represents 1×1 convolution, Conv3(·) represents 3×3 convolution, represents as shown in Figure 2Concatenation of the three branches of the intra-layer feature interaction module shown, where ReLU is the ReLU activation function;

[0103] Step 5.2: Input the feature map and into the two inter-layer feature collaboration modules of the region perception module respectively to enhance the shared feature representation, obtaining the output feature maps and which is expressed as:

[0104]

[0105]

[0106] In the formula, κ(·) represents two ordinary convolutions, tanh(·) represents the hyperbolic tangent function, ||·||2 represents the L2 norm, represents element-wise multiplication, (·) T represents matrix transpose, represents the feature map after being processed by two 1×1 convolutions;

[0107] Step 6: Extract the boundary of the unlabeled medical image X to be segmented through the Laplacian operator, obtaining the boundary feature map and denoted as Then use the region-boundary mutual learning module shown in Figure 4 to mutually learn and fuse the features of the region and the boundary. The specific process is as follows:

[0108] Step 6.1: Extract the boundary features of the unlabeled medical image X to be segmented through the Laplacian operator, obtaining the Laplacian boundary feature map which is expressed as:

[0109]

[0110] f l = L1(X)

[0111] L k = X k - u(X k+1 )

[0112]

[0113] In the formula, X represents the input unlabeled medical image to be segmented, g(·) represents the convolution operator with a Gaussian filter, d(·) represents downsampling, u(·) represents upsampling, L k represents the Laplacian image of the k-th layer of the Laplacian pyramid, X kIt represents the feature map of the unannotated medical image X to be segmented after Gaussian convolution and downsampling;

[0114] Step 6.2: Input the noise mask map x T into the segmentation encoder of the boundary-aware denoising network, and then concatenate the output feature map of the segmentation encoder with the feature map and perform upsampling to obtain the predicted feature map

[0115] Step 6.3: Set the predicted feature map output by the region-boundary mutual learning module in the boundary-aware denoising network as For the predicted feature map containing take the inverse to highlight the insignificant regions and obtain the reverse attention map At the same time, use the Laplacian operator to extract the boundary of the predicted feature map to obtain the boundary attention map The reverse attention map and the boundary attention map are expressed as:

[0116]

[0117] In the formula, L0() represents the Laplacian operator;

[0118] Step 6.4: Through the region-boundary mutual learning module, the region feature map is cross-learned with the Laplacian boundary feature map the next-level reverse attention map and the boundary attention map respectively, which is expressed as:

[0119]

[0120] In the formula, σ( ) represents the Sigmoid activation function;

[0121] Step 6.5: Refine the output feature map in the channel and space, which is expressed as:

[0122]

[0123] In the formula, CA represents the channel attention operation, and SA represents the spatial attention operation;

[0124] Step 6.6: Perform upsampling on to obtain the result Repeat passing through the region-boundary mutual learning module and the upsampling operation, and finally obtain the predicted result

[0125] Step 7. For the prediction result Gradually perform reverse diffusion denoising to obtain the final segmentation mask image The process of reverse diffusion denoising is expressed as:

[0126]

[0127] In the formula, \(x\) t and \(x\) t-1 respectively represent the noise mask images at times \(t\) and \(t - 1\) during the reverse diffusion process. \(\mu\) θ () represents the mean predicted by the model under the condition of time \(t\) and the current noise mask image \(x\) t . \(\sigma\) 2 represents the variance of the Gaussian distribution during the reverse diffusion process. \(\beta\) t is a parameter for controlling Gaussian noise during the reverse diffusion process. \(\alpha\) t is the decay coefficient related to \(\beta\) t . \(\prod_{i = 1}^{t}\alpha_{i}\) represents the product of all \(\alpha\) t during the reverse diffusion process, and \(\hat{x}_{0}\) is the segmentation mask image.

[0128] To verify the performance of the network model RBML - Diff proposed in this embodiment during medical image segmentation, polyp images in the existing publicly available datasets Kvasir, CVC - ClinicDB, CVC - 300, CVC - ColonDB, and ETIS - LaribPolypDB are used for network training and testing. The training dataset has a total of 1450 images, among which 900 images are from Kvasir and 550 images are from CVC - ClinicDB; 100 images are randomly selected from the training set as the validation set to verify the performance of the network model RBML - Diff during medical segmentation. Some images are selected for testing in each of the five datasets. Specifically: 100 images are selected in Kvasir, 62 images are selected in CVC - ClinicDB, 60 images are selected in CVC - 300, 380 images are selected in CVC - ColonDB, and 196 images are selected in ETIS - LaribPolypDB, and the test data is not used for training. During the training process, the AdamW optimizer is used to update the parameters of the network model RBML - Diff.

[0129] To test the accuracy and superiority of the image segmentation method proposed in this embodiment, it is further illustrated through the following experiments, where:

[0130] 1) The verification environment is as follows: the CPU is Intel(R) Xeon(R) Gold 6226R; the memory is 32 GB; the GPU is Nvidia Geforce RTX 3090 with a video memory of 24 GB, and it is carried out on the Ubuntu 16.04.10 operating system. The deep learning framework adopted by the network model RBML-Diff in this embodiment is PyTorch;

[0131] 2) The segmentation performance is evaluated through the following metric parameters, which are respectively:

[0132]

[0133] S α = α·S o +(1 - α)·S r

[0134]

[0135] In the formula, TP, FP, and FN respectively represent the number of true positive, true negative, and false negative samples. P represents precision, R represents recall, mDice refers to the set similarity metric function, which is usually used to calculate the similarity between two samples; mIoU refers to the ratio of the predicted target to the true annotation, indicating the similarity or overlap degree between two samples; is the weighted average of precision and recall, S α is used to evaluate the structural similarity between the predicted segmentation result and the true segmentation. mE ξ is the average structural similarity over the entire image, and maxE ξ is the maximum structural similarity over the entire image. MAE refers to the mean absolute error, which is used to measure the difference between the true value and the predicted value. mDice, mIoU, S α , mE ξ and maxE ξ range from [0, 1]. The closer the value is to 1, the better the segmentation effect. The closer the MAE value is to 0, the better the segmentation effect.

[0136] First, ablation experiments are conducted to verify the roles of the intra-layer feature interaction module, the inter-layer feature collaboration module, and the region-boundary mutual learning module in RBML-Diff proposed in this embodiment. On five datasets, ablation experiments are carried out on RBML-Diff and the Baseline based on Pyramid Transformer with the same architecture. Specifically, these three modules are respectively removed from RBML-Diff and compared with the Baseline for analysis. The effectiveness of these three modules is verified by statistically calculating mDice and mIoU. The results of the ablation experiments are shown in Table 1.

[0137] Ablation experiment results of each module of RBML-Diff in Table 1

[0138]

[0139] As can be seen from Table 1, compared with RBML-Diff, after removing the intra-layer feature interaction module, the mDice and mIoU scores on the datasets Kvasir, CVC-ClinicDB, CVC-ColonDB, and ETIS-LaribPolypDB all decreased. Especially in the ETIS-LaribPolypDB dataset, the mDice decreased by 2.3%. However, the performance on the CVC-300 dataset was the best. By observing the images in the CVC-300 dataset, it can be seen that the sizes of the polyp images are similar and very regular, and the number of datasets is small. This is exactly the reason why better performance can be obtained without the intra-layer feature interaction module. Similarly, after removing the inter-layer feature collaboration module, the mDice and mIoU scores decreased on all datasets. Among them, the mDice score decreased by 1.6% in CVC-ClinicDB, the mDice score decreased by 1.4% in ETIS-LaribPolypDB, and the decrease in other datasets was relatively small. In addition, when the region-boundary mutual learning module was not included, the decrease in the mDice and mIoU scores on the five datasets was more significant, indicating the importance of the region-boundary mutual learning module in improving the model performance. Compared with the intra-layer feature interaction module and the inter-layer feature collaboration module, the region-boundary mutual learning module has a more significant impact on the model performance.

[0140] Tables 2 to 6 show the segmentation results of RBML-Diff and the mainstream segmentation network models in recent years on the existing five publicly available polyp datasets respectively. The optimal results are shown in bold red font, and the sub-optimal results are shown in bold blue font. The specific results are as follows:

[0141] Quantitative results of different network models on the Kvasir dataset in Table 2

[0142]

[0143] As can be seen from Table 2, on the Kvasir dataset, RBML-Diff not only achieved excellent results of 0.925 and 0.879 in mDice and mIoU, but also in S α 、mE ξ 、maxE ξThe highest evaluation scores were obtained by mDice and MAE respectively. Specifically, compared with the existing models Polyp-PVT and CTNet, the mDice of the network model RBML-Diff proposed in this embodiment increased by 0.8%, and compared with the existing model Polyp-PVT, the mIoU of RBML-Diff increased by 1.5%.

[0144] Table 3 Quantitative Results of Different Network Models on the CVC-ClinicDB Dataset

[0145]

[0146] As can be seen from Table 3, on the CVC-ClinicDB dataset, the mIoU of RBML-Diff maxE ξ and MAE both obtained the highest evaluation scores, and the scores on other metrics were also higher than those of most existing network models;

[0147] Table 4 Quantitative Results of Different Network Models on the CVC-300 Dataset

[0148]

[0149] As can be seen from Table 4, on the CVC-300 dataset, for RBML-Diff, the scores of mDice, mIoU, S α mE ξ and MAE were 0.910, 0.847, 0.898, 0.945, 0.976 and 0.005 respectively, all higher than those of other network models, and maxE ξ also achieved a sub-optimal score of 0.982;

[0150] Table 5 Quantitative Results of Different Network Models on the CVC-ColonDB Dataset

[0151]

[0152] As can be seen from Table 5, on the CVC-ColonDB dataset, all evaluation metrics of RBML-Diff exceeded those of all existing network models, and compared with CTNet ranked second in terms of scores, there were significant improvements of 2.4% and 2.7% in mDice and mIoU respectively; meanwhile, S α mE ξ maxE ξ and MAE also increased by 1.8%, 1.3%, 1%, 1.1% and 0.2% respectively;

[0153] Table 6 Quantitative results of different network models on the ETIS-LaribPolypDB dataset

[0154]

[0155] As can be seen from Table 6, in the ETIS-LaribPolypDB dataset, although the evaluation index score of RBML-Diff is lower than that of the model CTNet, its performance is very close to that of CTNet. This gap is related to the multi-target polyp images contained in the ETIS-LaribPolypDB dataset;

[0156] By analyzing the results in Tables 2-6, it can be seen that compared with other existing models, RBML-Diff exhibits good learning ability and generalization ability on different datasets.

[0157] Figure 5 The visualization results of the network model RBML-Diff proposed in this embodiment and the existing network models on the polyp segmentation dataset are shown. It can be seen that RBML-Diff always has robust segmentation ability on polyps of different sizes and types, and its adaptability and accuracy are better than those of other network models. In addition, RBML-Diff has superior perception performance for the blurred edges of polyps and can more accurately segment the target boundaries of polyps.

Claims

1. A medical image segmentation method based on region-boundary mutual learning diffusion model, characterized in that: The steps include: Step 1: Build the network model RBML-Diff and initialize the hyperparameters; The network model RBML-Diff includes a conditional encoder, a region-aware module and a boundary-aware denoising network, wherein: the region-aware module includes an intra-layer feature interaction module and an inter-layer feature collaboration module; the boundary-aware denoising network includes a segmentation encoder and a region-boundary mutual learning module; Step 2, preprocessing the unlabeled medical image X to be segmented, the labeled medical image x and the true value x0 of the labeled medical image; Step 3: Input the preprocessed labeled medical image x and the true value of the labeled medical image x0 into the network model RBML-Diff, and use the forward diffusion process of the network model RBML-Diff to add noise to x0. After T steps of iteration, the noise mask image x0 is obtained. T ; Step 4: The noise mask image x T And the unlabeled medical image X to be segmented is input into the conditional encoder for feature extraction to obtain the feature map and Step 5: and The different intra-layer feature interaction modules in the region perception module are input respectively, and the intra-layer feature interaction modules perform and Perform feature mining in different directions and scales to obtain feature maps and The feature map and The two inter-layer feature coordination modules of the region perception module are input respectively, and the inter-layer feature coordination module performs feature mapping on the feature map. and Perform cross-layer feature refinement to obtain a refined multi-scale regional feature map and Step 6: Use the Laplace operator to extract the boundary of the unlabeled medical image X to be segmented and obtain the boundary feature map And use the segmentation encoder in the boundary-aware denoising network to extract the noise mask image x T The features of the encoder are split into output feature maps and feature maps Splicing, after upsampling, obtain the predicted feature map The predicted feature map output by the region-boundary mutual learning module in the boundary-aware denoising network is set to First include Prediction feature map Invert them respectively and extract the boundaries through the Laplace operator to obtain the reverse attention map and boundary attention map Then, the regional features are transformed into and Respectively with the boundary feature map Reverse attention map for the next level and boundary attention map Cross-learning is performed, and after refinement in channels and space, upsampling is performed to obtain Repeat the region-boundary mutual learning module and upsampling operation to finally get the prediction result Step 7: Prediction results Step by step reverse diffusion denoising to obtain the final segmentation mask map 2. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The initialization hyperparameters of step 1 include: setting the training batch size to 32, setting the number of training iterations to 150, and during the training process, adopting the learning rate decay strategy, setting the initial learning rate to 0.0001.

3. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The specific process of step 2 is: First, data enhancement is performed on the unlabeled medical image X to be segmented, and then the enhanced unlabeled medical image X to be segmented is normalized; The size of the labeled medical image x and the labeled medical image truth x0 is resized to 256×256.

4. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 3 is characterized in that: The data augmentation includes random scaling, horizontal and vertical flipping, randomly selected angles for rotation, random filling noise, and random cropping.

5. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The specific process of step 3 is as follows: In the forward diffusion process of the network model RBML-Diff, Gaussian noise is gradually added to the original annotated medical image truth value x0 through continuous T-step iterations. The forward diffusion process is expressed as: In the formula, T represents the number of diffusion steps, x t and x t-1 Represent the noise mask images at time t and t-1 in the forward diffusion process, x 1:t Represents the noise mask image for t=1,2,...,T in the forward diffusion process; In each iteration, the formula for adding Gaussian noise is expressed as: In the formula, β t is the parameter for controlling Gaussian noise in the forward diffusion process, I is an n×n unit matrix, is a Gaussian distribution; The process of sampling at any time t is expressed as: α t =1-β t In the formula, α t is related to β t The associated attenuation coefficient, represents all α in the forward diffusion process t The product of .

6. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The specific process of step 4 is as follows: Step 4.1: The noise mask image x T and the unlabeled medical image X to be segmented as the input of the conditional encoder. The conditional encoder consists of four stages, each of which includes a fusion embedding layer and a Transformer encoder. The Transformer encoder in the first stage outputs a feature map Step 4.2: Feature map As the second stage of the conditional encoder, the second stage Transformer encoder outputs the feature map Step 4.3: Feature map Input conditional encoder stage 3, Transformer encoder output feature map of stage 3 Step 4.4: Feature map Input conditional encoder stage 4, Transformer encoder output feature map of stage 4 7. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The specific process of step 5 is as follows: Step 5.1: Separate feature maps and The input is sent to the four intra-layer feature interaction modules of the region perception module to realize the interaction of features in different directions and scales. The four intra-layer feature interaction modules output feature maps respectively. and Recorded as It is expressed as: In the formula, represents the i-th feature map generated by the backbone network, k represents the number of branches, represents the output of the kth branch, represents element addition, Conv s (·) represents asymmetric convolution, Conv1(·) represents 1×1 convolution, Conv3(·) represents 3×3 convolution, represents the concatenation of three branches, ReLU is the ReLU activation function; Step 5.2: Feature map and The two inter-layer feature collaboration modules are input into the region perception module to strengthen the shared feature expression and obtain the output feature map and It is expressed as: In the formula, κ(·) represents two ordinary convolutions, tanh(·) represents the hyperbolic tangent function, and ||·||2 represents the L2 norm. represents element-wise multiplication, (·) T represents the matrix transpose, represent Feature map after two 1×1 convolutions.

8. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The specific process of step 6 is as follows: Step 6.1: Extract the boundary features of the unlabeled medical image X to be segmented by Laplacian operator to obtain the Laplacian boundary feature map It is expressed as: f l =L1(X) L k =X k -u(X k+1 ) Where X represents the input unlabeled medical image to be segmented, g(·) represents the convolution operator with Gaussian filter, d(·) represents downsampling, u(·) represents upsampling, and L k represents the Laplacian image of the kth level of the Laplacian pyramid, X k Represents the feature map of the unlabeled medical image X to be segmented after Gaussian convolution and downsampling; Step 6.2: The noise mask image x T Input the segmentation encoder of the boundary-aware denoising network, and then add the output feature map of the segmentation encoder to the feature map After splicing and upsampling, the predicted feature map is obtained Step 6.3: Set the predicted feature map output by the region-boundary mutual learning module in the boundary-aware denoising network to To include Prediction feature map Negate and highlight the insignificant areas to obtain the reverse attention map At the same time, the Laplacian operator is used to extract the prediction feature map The boundary of Reverse Attention Map and boundary attention map It is expressed as: Where, L0( ) represents the Laplace operator; Step 6.4: Transform the regional feature map into Respectively with the Laplace boundary feature map Reverse attention map for the next level and boundary attention map Perform cross learning, expressed as: Where, σ( ) represents the Sigmoid activation function; Step 6.5: Output feature map The refinement is performed in channels and spaces, expressed as: In the formula, CA represents the channel attention operation, and SA represents the spatial attention operation; Step 6.6: Upsampling to get the result Repeat the region-boundary mutual learning module and upsampling operation to get the prediction result 9. The medical image segmentation method based on region-boundary mutual learning diffusion model according to claim 1, characterized in that: The specific process of reverse diffusion denoising in step 7 is expressed as follows: In the formula, x t and x t-1 Represent the noise mask images at time t and t-1 in the reverse diffusion process, μ θ ( ) represents the noise mask image x at time t and the current noise mask image x t The mean value predicted by the network model under the condition of 2 represents the variance of the Gaussian distribution during the back diffusion process, β t is the parameter controlling Gaussian noise in the back diffusion process, α t is related to β t The associated attenuation coefficient, Indicates that all α in the reverse diffusion process t The product of is the segmentation mask image.

Citation Information

Patent Citations

  • Three-dimensional medical image automatic segmentation method based on deep learning

    CN111311592A

  • Medical image segmentation method based on conditional diffusion model

    CN116596949A

  • Polyp image segmentation method based on boundary clue depth fusion

    CN119151968A

  • Semantic Segmentation Based on a Hierarchy of Neural Networks

    US20200302239A1

  • Medical image segmentation method based on u-network

    US20220398737A1