Dual-mode guided interactive diffusion medical image segmentation method

By introducing a dual-mode-guided interactive diffusion method in medical image segmentation, the deep interaction between the prior model and the mutual diffusion model is used to solve the problem of undersegment or oversegment of image segmentation in traditional Chinese medicine in the prior art, and the segmentation effect and stability are improved.

CN120219735APending Publication Date: 2025-06-27LANZHOU JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510205462.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing medical image segmentation methods deal with medical images with strong noise, complex morphology and irregular boundaries, they are prone to undersegment or oversegment problems, and have weak generalization ability, making it difficult to adapt to image segmentation tasks of different modes and different resolutions.

Method used

A dual-mode guided interactive diffusion medical image segmentation method (DM-GID) is proposed. Through the built-in prior model and mutual diffusion model, the regional consistency module, feature alignment module and dynamic interaction module are used to realize the joint optimization of the model and improve the segmentation effect.

Benefits of technology

By generating more accurate guiding information through the prior model, the diffusion model is more clearly guided in the denoising process, improving the model's performance in detail recovery and boundary definition, enhancing the understanding and adaptability of complex features, and significantly improving the accuracy and stability of medical image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219735A_ABST
    Figure CN120219735A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-mode guided interactive diffusion medical image segmentation method, and relates to the technical field of medical image segmentation. The invention provides a dual-mode guided interactive diffusion medical image segmentation method (DM-GID) on the basis of a de-noising diffusion probability model (DDPM), a DM-GID network architecture is composed of a built-in prior model and an interdiffusion model, more accurate guide information is generated by the prior model, the diffusion model has a clearer generation direction, and the image segmentation efficiency is improved. According to the method, deep interaction is carried out on the prior model segmentation process and the diffusion model generation process through the mutual diffusion model, and a better medical image segmentation effect is obtained through joint optimization of the prior model and the mutual diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and particularly relates to a dual-mode guided interactive diffusion medical image segmentation method. Background Art

[0002] Medical image segmentation aims to accurately extract organ, lesion or tissue regions from complex medical image data, providing important basis for disease diagnosis, treatment planning and surgical navigation. Existing segmentation methods can be roughly divided into manual segmentation, traditional segmentation and deep learning-based segmentation methods.

[0003] Manual segmentation methods rely on experts' knowledge and experience. Although they have a certain degree of accuracy, they are time-consuming and laborious, and subjective factors are inevitable. Traditional segmentation methods include threshold-based, region-based, clustering-based segmentation methods, etc. They mainly rely on low-level features of images and manually designed prior knowledge, and perform well in segmenting simple-structured images. However, for medical images with strong noise, complex morphology and irregular boundaries, under-segmentation or over-segmentation problems often occur. In addition, the generalization ability of traditional methods is weak, and it is difficult to adapt to image segmentation tasks of different modalities and resolutions.

[0004] With the development of deep learning technology, the superiority of deep learning-based segmentation methods has become increasingly prominent. U-Net is the earliest proposed classic model. Through its symmetric encoder-decoder structure and skip connections, it effectively fuses multi-scale features and has excellent performance on small-scale datasets. The success of U-Net has spawned a large number of variants, such as AttentionUNet and UNet+, which improve the segmentation performance by introducing attention mechanisms and multi-resolution modules respectively. VisionTransformer (ViT) proposed in 2021 shows the potential of Transformer in visual tasks by directly modeling image patches. SETR first introduces Transformer into image segmentation, making up for the limitations of convolutional neural networks (CNNs) in long-range dependence modeling. In addition, Swin Transformer balances computational efficiency and global modeling ability through a hierarchical window mechanism and has become one of the mainstream architectures of current segmentation models. In the exploration of the combination of CNN and Transformer, TransUNet combines the advantages of CNN and Transformer, not only retaining the advantages of U-Net in local feature extraction, but also enhancing the global feature capture ability, significantly improving the segmentation performance. The encoder of UNeXt gradually downsamples the input image into multiple low-resolution feature maps, and the decoder uses a Transformer model to achieve better feature representation and context modeling ability. However, when facing noise in medical images, they are often prone to overfitting or being interfered by noise.

[0005] In recent years, the Denoising Diffusion Probabilistic Model (DDPM) has received extensive attention as a powerful generative model. It realizes data generation by gradually adding noise to the data and learning the denoising process, and can effectively remove the noise in medical images and restore the true structure of the images. However, the denoising network of the traditional diffusion model is U-Net, which performs poorly in dealing with long-range dependencies. In response to this, TransDDPM makes the diffusion model better capture the global information of the image by introducing the Transformer structure. To improve the expressive ability of the denoising process, Resdiff extracts features with a pre-trained CNN and then inputs these features into the diffusion model for fine-grained generation optimization. ControlNet is added on the basis of the pre-trained large model itself, and ControlNet is responsible for learning to map additional conditional information into the pre-trained large model with fixed parameters. When the current research combines classical models based on deep learning with diffusion models, it cannot dynamically capture the changes in intermediate features during the diffusion process and has not achieved end-to-end joint optimization, thus limiting the adaptability of the model in multi-scale and complex noise environments.

[0006] To achieve the effective integration of classical models based on deep learning and diffusion models, the present invention proposes a dual-mode guided interactive diffusion medical image segmentation method (DM-GID) based on the Denoising Diffusion Probabilistic Model (DDPM). The DM-GID network architecture consists of a built-in prior model and a mutual diffusion model. The prior model generates more accurate guiding information, enabling the diffusion model to have a clearer generation direction. The mutual diffusion model deeply interacts with the segmentation process of the prior model and the generation process of the diffusion model. The method of the present invention aims to obtain better medical image segmentation results through the joint optimization of the prior model and the mutual diffusion model. Summary of the Invention

[0007] The purpose of the present invention is to provide a dual-mode guided interactive diffusion medical image segmentation method to solve the above problems.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows: It includes the DM-GID built-in prior model and the mutual diffusion model, and the two types of models are jointly optimized through deep interaction. The specific steps are as follows: Step 1: Use the prior model to generate more accurate guiding information, and add a Region Consistency Module (RCM) to this model. Taking the normal tissue image as a reference, learn the normal tissue features and compare them with the lesion tissue features, so as to highlight and enhance the lesion-related features and further precisely guide the diffusion generation process.

[0009] Step 2: Use the mutual diffusion model to interact the denoising process with the process of generating guidance information. Add a Feature Alignment Module (FAM) to this model, and use deeper semantic information and structural features for guidance to enhance the model's understanding and adaptation ability to complex features.

[0010] Step 3: Add a Dynamic Interaction Module (DIM) to the mutual diffusion model, which is used to capture the global dependencies of the prior embedding and strengthen its own feature expression. At the same time, the prior embedding and the diffusion embedding are interacted synergistically for information integration, reducing the gap between the noise features and the semantic features and improving the quality of the generated images.

[0011] Furthermore, the prior model consists of a dual-encoding decoding model and a regional consistency module; The dual-encoding in the decoding model has a dual-encoder divided into a lesion encoder and a normal tissue encoder, and the specific content is as follows: I M and I N are used as the inputs of the lesion encoder and the normal tissue encoder respectively, and then the features of the lesion encoder x i M, i = 1, 2, 3, 4, 5 and the features of the normal tissue encoder x i N , i = 1, 2, 3, 4, 5 are input into the regional consistency module to learn the feature inconsistency between the lesion area and the normal tissue area. The output result x i F , i = 1, 2, 3, 4, 5 of the regional consistency module is used as the input of the next layer of the lesion encoder. The normal tissue encoder only I N extracts features layer by layer and then passes them to the next layer without any other additional operations. The input of the first layer of the decoder is the output result of the last regional consistency module, that is, x 5 F , and the input of other layers is the graph obtained by merging the decoding output of the previous layer and the features obtained by the corresponding regional consistency module; The regional consistency module enhances the feature difference of the lesion tissue by introducing a contrastive learning mechanism and using the feature representation of the normal tissue as a reference. It consists of two parts: feature contrast and feature refinement: The feature contrast part consists of the encoder e , the projection layer p and the predictorm Construct positive sample pairs consisting of normal tissue encoder features and lesion encoder features { x i M , x i N} and feed them into the feature comparison part, and maximize the cosine similarity between the two output vectors p 1 and z 2: (1) In the formula, , and are L2-norms, The symmetric loss is defined as: (2) The feature comparison part is used in the training stage and the projection layer p and the predictor m are removed in the inference stage. The stop-gradient operation is used to ensure that the lesion encoder can be trained normally. Therefore, formula (2) is modified to: (3) In the formula, stopgrad (·) represents the operation of stopping backpropagation of gradients.

[0012] Feature refinement part: After contrastive learning, two pairs of features { p i 1, z i 2} and { p i 2, z i 1} are obtained; First, multiply these two pairs of features respectively. This operation can reflect x i M and x i N between similarities. Then multiply the similarity result by -1 to obtain the similarity metric map S i 1 and S i 2: (4) To highlight the lesion tissue features in x i M , an activation operation is used to enhance the performance of the similarity to obtain Si m , and further refine the weighting of features using convolution operations to obtain a weight map A i , and then A i is applied to the features of the lesion area x i M to obtain x i F : (5) In the formula, and respectively represent convolutional layers with a convolutional kernel of 3×3 and 5×5, represents the channel concatenation operation, represents the activation function Sigmoid, represents the activation function GELU. The regional consistency module gradually strengthens the features of the lesion tissue through the double-encoding stage of the prior model to improve the segmentation effect.

[0013] Furthermore, the mutual diffusion model consists of a denoising diffusion probability model DDPM, a feature alignment module, and a dynamic interaction module; The operation of the DDPM is divided into two processes: forward diffusion and reverse denoising; In the forward diffusion process, given the data distribution x 0~ q ( x ), noise is gradually added to the data, and a total of T steps are added to generate a series of noisy samples x 1, x 2,... x T . Its mathematical definition is: (6) In the formula, x t is the data at the t th step in the diffusion process, determines the mean and variance of the added noise, which increases as the time step t increases, I represents the covariance matrix of the noise, indicating that the noise is independent in each dimension and has the same variance; In order to directly derive the distribution of any time step x 0 from the initial data t without iterationx 0 to q ( x t | x 0), and its mathematical definition is: (7) In the formula, , is the cumulative noise term. This recursive form indicates that x t at any time step can be obtained by direct sampling: (8) In the formula, represents the noise added to the data.

[0014] In the reverse diffusion process, denoising starts from the Gaussian noise and finally restores to the original data distribution q ( x 0); each step of the reverse process is modeled as a conditional probability distribution: (9) In the formula, represents the mean of this normal distribution, represents the variance of this normal distribution.

[0015] As extended by Ho et al., by solving μ θ in x backward to obtain the new mean: (10) Finally, the parametric model is trained to predict to the intermediate cumulative noise , and its training objective is: (11) The feature alignment module structure has the functions of information compression, information calibration, and feature mapping; Information compression process: Using the significant features provided by the prior model as the input to extract more accurate and effective features: (12) In the formula, IC ( ) represents the information compression process; Information calibration process: Inputting the distilled information into the integrated focusing module to obtain the weight map An Apply A n to the first-layer encoded features of DDPM f d , which further enhances the attention to key features. This not only ensures the prediction range of the generation result but also provides further optimization: (13) In the formula, IF ( ) represents the integrated focusing sub-module, Sigmoid ( ) represents the activation function Sigmoid .

[0016] Feature mapping process: Effectively capture the complex non-linear relationships between features, and then generate richer feature representations: (14) In the formula, DWConv ( ) represents the depth convolution, GELU ( ) represents the activation function GELU , BN ( ) represents batch normalization; The dynamic interaction module consists of N blocks with the same architecture, and each block consists of a Stem block, a learnable attention LA, and an enhanced gated linear unit EGLU. The specific operation process is as follows: First, input the prior embedding x 5 F and the diffusion embedding x e D into the dynamic interaction module, where x e D is regarded as the focusing area, that is, input 1, x 5 F is regarded as the peripheral area, that is, input 2. Then, adopt a recursive mechanism, take the result of the enhanced gated linear unit as input 1, x e D as input 2, and input them into a group of learnable attention and enhanced gated linear units again. The output result is used as the input 1 of the next block, and repeat this process N times. Let N = 4, and the input 2 for each block is always x5 F , the purpose is to always provide an invariant reference benchmark throughout the multi-round process to prevent "excessive forgetting" of prior features; x e D Provide an invariant reference benchmark to prevent "excessive forgetting" of prior features; The Stem block contains convolutional, flattening, transposing, and layer normalization operations to convert the features into a format suitable for the attention input, defining Q A , K A and V A as the query, key, and value from Input 1, K D and V D as the key and value from Input 2; a learnable query embedding is added to facilitate x e D and x 5 F information integration between them and merge it into the original attention: (15) In the formula, represents the matrix multiplication operation, represents the transpose of represents the transpose of

[0017] Since the relationships and dependencies between features are highly variable, a learnable key embedding is added to capture the potential relative position relationships in the input data: (16) In the formula, represents the matrix multiplication operation, represents the transpose of

[0018] In the standard attention mechanism, by introducing a position bias focusing bias ( b A ), a dispersion bias ( b D ), and a local bias ( b L ), the attention mechanism can adjust the strength of the attention area through relative position weights to obtain the weight map A . Focus area weight mapA A and the peripheral region weight map A D is obtained by splitting A The learnable attention can be described as: (17) In the formula represents the channel splicing operation, is the channel type Softmax , is the splitting operation of the channel type, is the dimension of the query or key, is the scaling factor, represents the matrix multiplication operation.

[0019] Furthermore, the integrated focusing sub-module IF in the formula (13) is composed of specific channel attention and multi-path spatial attention, adopting the design order of channel first and then space, focusing on important feature channels and regions with high expression ability, and improving the calculation efficiency; In the sub-module of specific channel attention, the distribution of information between channels is further optimized. First, the input features are enhanced channel by channel through depth convolution and the activation function GELU to highlight the characteristics of each channel; then, weight distribution is performed on the importance of each channel; the final output highlights the key channel features, making the feature distribution more targeted; The multi-path spatial attention captures features from multiple angles by introducing a dual-branch mechanism of max pooling and average pooling. Max pooling focuses on local saliency and captures the extreme responses of features; while average pooling extracts global smooth information to enhance the perception of feature distribution.

[0020] According to one aspect of the present invention, a dual-mode guided interactive diffusion medical image segmentation method is provided. Compared with the prior art, the present invention has the following beneficial effects: (1) By introducing normal tissue pictures as references, the present invention better learns the data distribution characteristics of the lesion area, thus more accurately focusing on the lesion. It provides more explicit guidance for the diffusion process and improves the performance of the model in detail restoration and boundary definition.

[0021] (2) The present invention adds a feature alignment module to the mutual diffusion model. This module takes the salient features of the prior model as input and uses a joint attention mechanism to make the diffusion model pay more attention to the lesion area during training. This helps to avoid unreasonable regional segmentation by the diffusion model and makes the generated segmentation results more stable.

[0022] (3) The present invention adds a dynamic interaction module, which mimics the focus perception mode of biological vision, adopts a self-attention mechanism similar to the prior embedding, adopts a cross-attention mechanism similar to the prior embedding and the diffusion embedding, and introduces learnable queries, keys and position biases in the attention. This can make the prior features and the features in the diffusion process more closely combined, and solve the problem of the gap between the noise features and the semantic features. Description of the Drawings

[0023] Figure 1 It is a schematic flowchart of the method of the present invention; Figure 2 It is a schematic structural diagram of the regional consistency module of the present invention; Figure 3 It is the structure of the feature alignment module of the present invention; Figure 4 It is the structure of the dynamic interaction module of the present invention; Figure 5 It is a schematic diagram comparing the test results of different segmentation methods under the BUSI dataset of the present invention; Figure 6 It is a schematic diagram comparing the test results of different segmentation methods under the DDTI dataset of the present invention; Figure 7 It is a schematic diagram comparing the test results of different segmentation methods under the LIDC-IDRI dataset of the present invention. Detailed Embodiments

[0024] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0025] As follows Figures 1 - 7 As shown, the present invention provides a dual-mode guided interactive diffusion medical image segmentation method, Figure 1 shows the technical scheme flowchart of the method (DM-GID) of the present invention. DM-GID incorporates a prior model and a mutual diffusion model, and the two models are jointly optimized through deep interaction.

[0026] Step 1: Use the prior model to generate more accurate guidance information. A regional consistency module (RCM) is added to this model. Taking the normal tissue image as a reference, the normal tissue features are learned and compared with the lesion tissue features, so as to highlight and enhance the lesion-related features and further accurately guide the diffusion generation process.

[0027] Step 2: Use the mutual diffusion model to interact the denoising process with the process of generating guidance information. Add a Feature Alignment Module (FAM) to this model, and use deeper semantic information and structural features for guidance to enhance the model's understanding and adaptation ability of complex features.

[0028] Step 3: Add a Dynamic Interaction Module (DIM) to the mutual diffusion model, which is used to capture the global dependencies of the prior embedding and strengthen its own feature expression. At the same time, the prior embedding and the diffusion embedding are interacted synergistically for information integration, reducing the gap between the noise features and the semantic features and improving the quality of the generated images.

[0029] Model Composition, Function and Key Algorithms

[0030] ① Prior Model

[0031] The prior model consists of a double-encoding decoding model and a regional consistency module.

[0032] Double Encoding The double encoder in the decoding model is divided into a lesion encoder and a normal tissue encoder, as shown in Figure 1 (a). I M and I N are used as the inputs of the lesion encoder and the normal tissue encoder respectively, and then the features of the lesion encoder x i M ( i = 1, 2, 3, 4, 5) and the features of the normal tissue encoder x i N ( i = 1, 2, 3, 4, 5) are input into the regional consistency module to learn the feature inconsistency between the lesion area and the normal tissue area. The output result x i F ( i = 1, 2, 3, 4, 5) of the regional consistency module is used as the input of the next layer of the lesion encoder. The normal tissue encoder only I N extracts features layer by layer and then passes them to the next layer without any other additional operations. Since the regional consistency module learns the region of interest. Therefore, the input of the first layer of the decoder is the output result of the last regional consistency module (i.e., x 5 F ), and the input of other layers is the graph obtained by merging the decoding output of the previous layer and the features obtained by the corresponding regional consistency module.

[0033] The structure of the regional consistency module is as follows Figure 2 shown. By introducing a contrastive learning mechanism, the regional consistency module uses the feature representation of normal tissues as a reference to enhance the feature differences of lesion tissues. It consists of two parts: feature contrast and feature refinement.

[0034] The feature contrast part is as shown in Figure 2 (a), and it consists of an encoder e , a projection layer p , and a predictor m . We feed the positive sample pairs { x i M , x i N} composed of the normal tissue encoder feature and the lesion encoder feature into the feature contrast part, and maximize the cosine similarity between the two output vectors p 1 and z 2: (1) In the formula, , , and are L2-norms. It is equivalent to the mean square error of l 2 normalized vectors, and the range is 2. The symmetric loss is defined as: (2) The feature contrast part is used in the training stage and the projection layer p and the predictor m are removed in the inference stage. Since optimizing the feature contrast part simultaneously will cause the model to collapse when two samples are mapped to the same point in the low-dimensional space, the stop gradient operation is used to ensure the normal training of the lesion encoder. Therefore, formula (2) is modified to: (3) In the formula, stopgrad (·) represents the operation of stopping the backpropagation of gradients.

[0035] The feature refinement part is as shown in Figure 2 (b). After contrastive learning, two pairs of features { p i 1, z i 2} and { p i 2, z i 1} are obtained. First, these two pairs of features are multiplied respectively, and this operation can reflect x i M andx i N The similarity between them. Then multiply the similarity result by -1 to obtain the similarity metric map S i 1 and S i 2: (4) This can weaken the module's attention to the regions with higher similarity and pay more attention to the parts with larger differences.

[0036] To highlight x i M the lesion tissue features in, we adopt an activation operation to enhance the performance of the similarity to obtain S i m and use a convolution operation to further refine the weighting of the similarity to obtain the weight map A i Then A i is applied to the lesion region features x i M to obtain x i F : (5) In the formula and respectively represent the convolutional layers with convolution kernels of 3×3 and 5×5, represents the channel concatenation operation, represents the activation function Sigmoid, represents the activation function GELU. The regional consistency module gradually strengthens the lesion tissue features through the double-encoding stage of the prior model to achieve the purpose of improving the segmentation effect.

[0037] ② Mutual Diffusion Model

[0038] The mutual diffusion model consists of DDPM, a feature alignment module, and a dynamic interaction module.

[0039] (a) Denoising Diffusion Probability Model (DDPM)

[0040] The operation of DDPM is divided into two processes: forward diffusion and reverse denoising. In the forward diffusion process, given the data distribution x 0~q ( x ), gradually add noise to the data, a total of T steps are added, thus generating a series of noisy samples x 1. x 2. …, x T . Its mathematical definition is: (6) In the formula, x t is the data at the t step in the diffusion process, determines the mean and variance of the added noise, which increases as the time step t increases, I represents the covariance matrix of the noise, indicating that the noise is independent in each dimension and has the same variance.

[0041] In order to directly derive the distribution of any time step x 0 from the initial data t without iteration x 0~ q ( x t | x 0), its mathematical definition is: (7) In the formula, , is the cumulative noise term. This recursive form indicates that x t at any time step can be obtained by direct sampling: (8) In the formula, represents the noise added to the data.

[0042] In the reverse diffusion process, start from the Gaussian noise and gradually denoise it until finally restoring to the original data distribution q ( x 0). Each step of the reverse process is modeled as a conditional probability distribution: (9) In the formula, represents the mean of this normal distribution, represents the variance of this normal distribution.

[0043] As extended by Ho et al., by solving μ θ in formula (8) x0 To obtain a new mean: (10) Finally, by training the parametric model to predict to the accumulated noise in the middle , and its training objective is: (11) (b) Feature alignment module The structure of the feature alignment module is as shown in Figure 3 the figure, and it has the functions of information compression, information calibration, and feature mapping.

[0044] The information compression process is as shown in Figure 3 (a). It takes the saliency features provided by the prior model as the input, aiming to extract more accurate and effective features: (12) In the formula, IC ( ) represents the information compression process. The information calibration process is as shown in Figure 3 (b). It inputs the distilled information into the integrated focusing module to obtain the weight map A n . Applying A n to the first-layer encoded features f d of the DDPM further enhances the attention to key features. It not only ensures the prediction range of the generation result but also provides further optimization: (13) In the formula, IF ( ) represents the integrated focusing sub-module, Sigmoid ( ) represents the activation function Sigmoid .

[0045] The feature mapping process is as shown in Figure 3 (c). It effectively captures the complex non-linear relationships between features, and then generates a richer feature representation: (14) In the formula, DWConv ( ) represents the depth convolution, GELU ( ) represents the activation function GELU , BN( ) represents batch normalization.

[0046] The integrated focusing sub-module in Equation (13) IF is composed of specific channel attention and multi-path spatial attention, adopting a design order of channel first and then space, focusing on important feature channels and regions with high expression ability, and improving the computational efficiency.

[0047] Such as Figure 3 (d) shows that the specific channel attention focuses on enhancing the uniqueness of each channel and reallocating channel weights. Different from traditional channel attention, the present invention further optimizes the distribution of inter-channel information in this sub-module. First, the input features are enhanced channel by channel through depth convolution and the activation function GELU to highlight the characteristics of each channel. Then, weight distribution is performed on the importance of each channel. The final output highlights the key channel features, making the feature distribution more targeted.

[0048] Such as Figure 3 (e) shows that inspired by the multi-level dispersion spatial attention (MDSA), the present invention designs a multi-path spatial attention. It captures features from multiple angles by introducing a dual-branch mechanism of max pooling and average pooling. Max pooling focuses on local saliency and captures the extreme responses of features; while average pooling extracts global smooth information and enhances the perception of feature distribution.

[0049] (c) Dynamic interaction module

[0050] In the biological visual system, the fovea is the area on the retina that focuses on collecting high-resolution detailed information. At the same time, through the advanced visual pathway, the fovea realizes information interaction with the peripheral area of the retina, integrating the low-resolution background information provided by the peripheral area, so as to realize the dynamic optimization of overall visual perception. The present invention designs an attention module (i.e., the dynamic interaction module) that not only focuses on the characteristics of the focused area (i.e., the fovea area) but also can dynamically interact with the peripheral area. It is composed of N blocks with the same architecture. Each block consists of a Stem block, learnable attention (LA), and an enhanced gated linear unit (EGLU). Its structure is as shown in the dotted box of Figure 1 (b).

[0051] The specific operation process is as follows: First, the prior embedding x 5 F and the diffusion embedding x e D are input into the dynamic interaction module, where xe D As the focus area (i.e., input 1), x 5 F As the peripheral area (i.e., input 2). Since the enhanced gated linear unit can selectively suppress or amplify specific input signals and avoid the propagation of redundant information, the enhanced gated linear unit is selected to replace the MLP. Then, in order to x e D gradually obtain more delicate global context and semantic information while retaining its inherent noise characteristics, we adopt a recursive mechanism, that is, taking the result of the enhanced gated linear unit as input 1, x e D as input 2, and inputting them into a group of learnable attention and enhanced gated linear units again. The output result is used as input 1 of the next block, and this process is repeated N times. In the present invention, N is set to 4. The input 2 for each block is always x 5 F , with the aim of always providing an invariant reference benchmark for x e D throughout the multi-round process to prevent "excessive forgetting" of prior features.

[0052] The Stem block contains convolution, flattening, transposing, and layer normalization operations, aiming to convert the features into a format suitable for attention input. The detailed process of learnable attention is as shown in Figure 4 (a). Define Q A , K A and V A as the query, key, and value from input 1, K D and V D as the key and value from input 2. To better facilitate the x e D and x 5 F information integration between them, we add a learnable query embedding , and merge it into the original attention: (15) In the formula, represents the matrix multiplication operation, denotes the transpose of, denotes the transpose of.

[0053] Since the relationships and dependencies between features are highly variable, a learnable key embedding is added to capture the potential relative position relationships in the input data. In visual tasks, it is often used together with the focus region: (16) wherein, represents the matrix multiplication operation, denotes the transpose of.

[0054] In the standard attention mechanism, the query-key similarity result only depends on the content similarity between features, while ignoring the relative position relationships in the image. In visual information processing, adjacent pixels are usually more relevant, while pixels far away have less influence. By introducing position bias focusing bias ( b A ), dispersion bias ( b D ), and local bias ( b L ), the attention mechanism can adjust the strength of the attention region through relative position weights to obtain the weight map A . The focus region weight map A A and the peripheral region weight map A D are obtained by splitting A . Therefore, the learnable attention can be described as: (17) wherein represents the channel concatenation operation, is the channel-wise Softmax , is the channel-wise splitting operation, is the dimension of the query or key, is the scaling factor, represents the matrix multiplication operation.

[0055] The Gated Linear Unit (GLU) is a channel mixer that divides the input into two parts by introducing a gating mechanism: one part is used for information processing, and the other part is used to gate the information. We found that adding a convolution to the gating branch can enhance the local perception ability of the gating branch. In addition, the gated features may introduce biases, while residual connections can balance the information flow between the input and the gated operation to some extent, making the output more robust. Therefore, the present invention specifically designs an enhanced gated linear unit, and its structure is as shown in Figure 4 (b).

[0056] Compare the experimental data with the evaluation results: The datasets used in the present invention are the BUSI breast ultrasound dataset, the DDTI thyroid ultrasound dataset, and the LIDC-IDRI computed tomography lung nodule dataset. The BUSI dataset collected 780 breast ultrasound images of 600 female patients, including 437 benign masses, 210 malignant masses, and 133 normal masses. The DDTI dataset contains 637 pixel-level labeled thyroid ultrasound images from a single device provided by Pedraza et al. The LIDC-IDRI dataset includes a total of 1018 study instances. For each image in the instance, two-stage diagnostic annotations were performed by 4 experienced chest radiologists.

[0057] The method of the present invention was compared with eight other advanced medical image segmentation methods on the BUSI dataset, the DDTI dataset, and the LIDC-IDRI dataset. These eight comparison methods are AttentionUNet, TransUNet, SwinUNet, FRBNet, TGDAUNet, CFATransUnet, TransGuider, and MedSegDiff. To evaluate the effectiveness of the method of the present invention, five commonly used image segmentation evaluation metrics were adopted, including the F1 score, recall, precision, mean intersection over union (MIoU), and accuracy. The F1 score is a statistical metric for measuring the similarity between two sample sets; recall measures the proportion of correctly detected positive class samples among all actual positive class samples; precision calculates the proportion of correctly predicted samples in the total samples; MIoU measures the degree of overlap between the predicted segmentation and the actual segmentation region; and accuracy is the proportion of correctly predicted samples in the total samples.

[0058] Table 1 Quantitative result analysis of different segmentation methods on the BUSI dataset

[0059] Note: ↑ indicates that the larger the value, the better the corresponding segmentation effect, and ↓ indicates that the smaller the value, the better the corresponding segmentation effect. And Δ represents the percentage decrease and increase in the evaluation index value compared with the method of the present invention.

[0060] The quantitative results of the method DM-GID of the present invention and eight comparison methods on the BUSI dataset are shown in Table 1. It can be seen from the table that the segmentation performance of DM-GID on the BUSI dataset is better than other methods. The F1 Score of DM-GID is 90.14%, the MIoU is 82.59%, and the Precision is 89.41%, which are 0.77%, 1.24%, and 1.32% higher than the current advanced method MedSegDiff respectively. This shows that the method of the present invention has significant advantages in the accuracy and regional consistency of segmentation. At the same time, DM-GID also reaches 90.91% and 97.89% in Recall and Accuracy respectively, showing the best performance among all methods.

[0061] Table 2 Analysis of Quantitative Results of Different Segmentation Methods under the DDTI Dataset

[0062] Note: ↑ indicates that the larger the value, the better the corresponding segmentation effect, and ↓ indicates that the smaller the value, the better the corresponding segmentation effect. And Δ represents the percentage decrease and increase in the evaluation index value compared with the method of the present invention.

[0063] The quantitative results of the method of the present invention and eight comparison methods on the DDTI dataset are shown in Table 2. It can be seen from the table that the segmentation performance of DM-GID on the DDTI dataset is better than other methods. The Recall of DM-GID is 93.09%, which is 1.13% higher than Medsegdiff, indicating that DM-GID more effectively reduces false positives. The Precision of DM-GID is 92.11%, which is 0.78% higher than Medsegdiff, indicating that the detection of the target area is more comprehensive. The MIoU of DM-GID is 85.82%, which is 1.01% higher than Medsegdiff, indicating a significant improvement of DM-GID in complex boundaries and geometric consistency. The F1Score of DM-GID is 92.60% and the Accuracy is 98.32% are both the highest values, and the experimental results show that the method of the present invention reaches the best in terms of the accuracy, coverage, and integrity of target segmentation.

[0064] Table 3 Analysis of Quantitative Results of Different Segmentation Methods under the LIDC-IDRI Dataset

[0065] Note: ↑ indicates that the larger the value, the better the corresponding segmentation effect, and ↓ indicates that the smaller the value, the better the corresponding segmentation effect. And Δ represents the percentage decrease and increase in the evaluation index value compared with the method of the present invention.

[0066] The quantitative results of the method of the present invention and eight comparison methods on the LIDC-IDRI dataset are shown in Table 3. It can be seen from the table that DM-GID leads in multiple indicators, highlighting its excellent performance in the pulmonary nodule segmentation task. The F1Score of DM-GID reaches 84.47%, Recall is 85.90%, Precision is 83.06%, MIoU is 83.42%, and Accuracy is 99.85%. Although the Precision of DM-GID is lower than that of some algorithms, its Recall is higher, indicating that DM-GID is more suitable for medical image segmentation. The difference between Recall and Precision of CFATransUnet is too large, indicating that CFATransUnet is more inclined to avoid misjudgment and gives up recall. Compared with other methods, DM-GID can more comprehensively cover the pulmonary nodule area while maintaining high segmentation accuracy when segmenting small targets, thus showing stronger segmentation robustness.

[0067] The method DM-GID of the present invention is compared with eight other advanced medical image segmentation methods on the BUSI dataset, DDTI dataset and LIDC-IDRI dataset, including AttentionUNet, TransUNet, SwinUNet, FRBNet, TGDAUNet, CFATransUnet, TransGuider and MedSegDiff. All comparison methods are implemented using default settings. The first column in the figure is the image with the lesion area, and the second column is the standard lesion image marked by professional doctors. The next eight columns are the segmentation images obtained by the methods of AttentionUNet, TransUNet, SwinUNet, FRBNet, TGDAUNet, CFATransUnet, TransGuider and MedSegDiff respectively, and the last column is the segmentation image obtained by the method of the present invention.

[0068] Figure 5 and Figure 6The visualization test results comparison of different segmentation methods on the BUSI dataset and the DDTI dataset is respectively shown. The red edges in the figure represent the ground truth boundaries of the lesions. Specifically, AttentionUNet, TransUNet, SwinUNet, FRBNet, TGDAUNet, CFATransUnet, TransGuider, MedSegDiff, and DM-GID can basically locate the approximate positions of the lesions. However, AttentionUNet has missed segmentation when dealing with targets with complex shapes. TransUNet combines the global modeling ability of the Transformer structure, and the coherence of the segmentation region is enhanced, but there is a slight over-segmentation problem when processing images. FRBNet has a certain enhancement effect on specific texture features, but the geometric consistency is poor, and the target shape will be distorted. TGDAUNet has improved the modeling of geometric features, and the coherence of the segmentation region is strong, but there are still deficiencies in detail processing. The lesions segmented by SwinUNet and CFATransUnet are slightly lacking in the accuracy of edge details. TransGuider is not robust enough to the noise in complex backgrounds and is prone to introducing pseudo-segmentation regions. MedSegDiff uses a diffusion model for image segmentation, but local details are lost, especially in the segmentation of target boundaries, which is not ideal. Compared with other methods, the method of the present invention has higher segmentation accuracy for the target region, clear boundaries, and strong ability to fuse texture and geometric features.

[0069] Figure 7 This is the visualization test results comparison of different image segmentation methods on the LIDC-IDRI dataset. The red edges in the figure represent the ground truth boundaries of the lesions. To highlight the lesion regions that the model focuses on, small lesions are enlarged and placed in the green box in the lower right corner. From Figure 7 the segmentation results, it can be seen that there are certain over-segmentation or under-segmentation problems in various methods. Specifically, the AttentionUNet and FRBNet algorithms are prone to blurring or breaking when capturing the boundaries of small target regions. When detecting small targets, the SwinUNet and MedSegDiff algorithms are prone to missing some key regions. TGDAUNet performs unstably in the geometric shape modeling of some targets, with a large shape deviation. In samples with complex backgrounds, the segmentation regions of TransUNet and CFATransUnet are easily interfered. TransGuider has good semantic consistency in the segmentation shape of the target, but there are still problems of detail loss. The method of the present invention can basically accurately identify the boundaries of small target regions and demonstrates better geometric consistency.

[0070] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0071] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A dual-mode guided interactive diffusion medical image segmentation method, characterized by: Including the DM-GID built-in prior model and the mutual diffusion model, the two types of models are jointly optimized through deep interaction. The specific steps are as follows: Step 1: Use the prior model to generate more accurate guidance information. Add the regional consistency module (RCM) to the model, use the normal tissue image as a reference, learn the normal tissue features and compare them with the lesion tissue features, so as to highlight and enhance the lesion-related features and further accurately guide the diffusion generation process; Step 2: Use the mutual diffusion model to interact the denoising process with the process of generating guidance information; add the feature alignment module FAM to the model to use deeper semantic information and structural features for guidance, thereby enhancing the model's understanding and adaptability to complex features; Step 3: Add a dynamic interaction module DIM to the mutual diffusion model to capture the global dependencies of the prior embedding and strengthen its own feature expression. At the same time, the prior embedding and the diffusion embedding interact synergistically to integrate information, reduce the gap between noise features and semantic features, and improve the quality of the generated image.

2. The method of claim 1, wherein: The prior model consists of a double encoding The decoding model and a regional consistency module; The double coding The dual encoder in the decoding model is divided into a lesion encoder and a normal tissue encoder. The specific contents are as follows: I M and I N As the input of the lesion encoder and the normal tissue encoder respectively, the features of the lesion encoder are then x i M, i =1, 2, 3, 4, 5 and characteristics of normal tissue encoders x i N , i =1, 2, 3, 4, 5 are input into the regional consistency module to learn the feature inconsistency between the lesion area and the normal tissue area. The output of the regional consistency module is x i F , i = 1, 2, 3, 4, 5 are used as the input of the next layer of the lesion encoder, and the normal tissue encoder is only I N Features are extracted layer by layer and then passed to the next layer without any additional operations. The input of the first layer of the decoder is the output of the last region consistency module, that is, x 5 F , while the input of other layers is the combined features of the decoding output of the previous layer and the features obtained by the corresponding regional consistency module; The regional consistency module introduces a contrast learning mechanism and uses the feature representation of normal tissue as a reference to enhance the feature difference of lesion tissue. It consists of two parts: feature comparison and feature refinement: The feature comparison part is done by the encoder e , Projection layer p and predictor m The positive sample pairs composed of normal tissue encoder features and lesion encoder features are { x i M , x i N } is fed into the feature comparison part and maximizes the two output vectors p 1 and z Cosine similarity between 2: (1) In the formula, , and is L2-norm, The symmetric loss is defined as: (2) The feature comparison part is used in the training phase, and the projection layer is removed in the inference phase. p and predictor m , the gradient operation is stopped to ensure that the lesion encoder can be trained normally, so formula (2) is modified as follows: (3) In the formula, stopgrad(·) Indicates stopping the reverse gradient propagation operation; Feature refinement part: After contrastive learning, two pairs of features are obtained { p i 1, z i 2} and { p i 2, z i 1}; First, multiply the two pairs of features separately. This operation can reflect x i M and x i N The similarity between Then multiply the similarity result by -1 to get the similarity metric graph S i 1 and S i 2: (4) To highlight x i M The lesion tissue features in the image are obtained by using activation operation to enhance the similarity. S i m , and use the convolution operation to further refine the weighting of the features to obtain the weight map A i , then A i Applied to lesion area features x i M Go up, get x i F : (5) In the formula, and They represent convolution layers with convolution kernels of 3×3 and 5×5 respectively. Indicates the channel splicing operation, represents the activation function Sigmoid, represents the activation function GELU. The regional consistency module gradually strengthens the lesion tissue features through the dual encoding stage of the prior model to improve the segmentation effect.

3. The method of claim 1, wherein: The mutual diffusion model consists of the denoising diffusion probability model DDPM, the feature alignment module and the dynamic interaction module; The DDPM operation is divided into two processes: forward diffusion and reverse denoising; In the forward diffusion process, given the data distribution x 0~ q ( x ), gradually adding noise to the data, adding a total of T step, thus generating a series of noisy samples x 1. x 2. … x T ; Its mathematical definition is: (6) In the formula, x t The diffusion process t Step data, Determine the mean and variance of the noise added, as the time step t Increase and increase, I Represents the covariance matrix of the noise, indicating that the noise is independent in each dimension and has the same variance; In order to directly start from the initial data without iteration x 0Derivation of any time step t Distribution x 0~ q ( x t | x 0), which is mathematically defined as: (7) In the formula, , is the cumulative noise term; this recursive form shows that at any time step x t can be obtained by direct sampling: (8) In the formula, represents the noise added to the data; In the back diffusion process, from Gaussian noise Start to gradually denoise and finally restore to the original data distribution q ( x 0); Each step of the reverse process is modeled as a conditional probability distribution: (9) In the formula, represents the mean of the normal distribution, represents the variance of the normal distribution; According to formula (8), we can solve μ θ In x 0 to get the new mean: (10) Finally, the parameterized model is trained Go to predict arrive The accumulated noise in the middle , and its training objectives are: (11) The feature alignment module structure has the functions of information compression, information calibration and feature mapping; Information compression process: The salient features provided by the prior model As input, more accurate and effective features are extracted: (12) In the formula, IC ( ) represents the information compression process; Information calibration process: The distilled information is input into the integrated focusing module to obtain the weight map A n ;Will A n The first layer encoding features applied in DDPM f d In terms of the above, the focus on key features is further strengthened; it not only ensures the prediction range of the generated results, but also provides further optimization: (13) In the formula, IF ( ) represents the integrated focusing submodule, Sigmoid ( ) represents the activation function Sigmoid; Feature mapping process: effectively captures the complex nonlinear relationship between features, thereby generating richer feature representations: (14) In the formula, DWConv ( ) represents the depth convolution, GELU ( ) represents the activation function GELU , BN ( ) represents batch normalization; The dynamic interaction module consists of N Each block consists of a Stem block, a learnable attention LA, and an enhanced gated linear unit EGLU. The specific content is: First, embed the prior x 5 F and diffuse embedding x e D Enter the dynamic interaction module, where x e D As the focus area, input 1, x 5 F As the peripheral area, that is, input 2, then a recursive mechanism is used to take the result of the enhanced gated linear unit as input 1. x e D It is used as input 2 and fed into a set of learnable attention and enhanced gated linear units again. The output is used as input 1 of the next block and the process is repeated. N All over, set N = 4, input 2 of each block is always x 5 F The purpose is to always x e D Provide an unchanging reference benchmark to prevent "excessive forgetting" of prior features; The Stem block contains convolution, flattening, transposition, and layer normalization operations to transform the features into a format suitable for attention input. Q A , K A and V A are the query, key and value from input 1, K D and V D are the keys and values ​​from input 2; a learnable query embedding is added Promote x e D and x 5 F The information between the two is integrated and merged into the original attention: (15) In the formula, represents the matrix multiplication operation, express The transpose of express The transpose of Since the relationships and dependencies between features are highly variable, a learnable key embedding is added , to capture the potential relative position relationship in the input data: (16) In the formula, represents the matrix multiplication operation, express The transpose of In the standard attention mechanism, by introducing position bias, focus bias ( b A ), dispersion deviation( b D ) and local deviation ( b L ), the attention mechanism adjusts the strength of the attention area through the relative position weight, and obtains the weight map A , focus area weight map A A And the peripheral area weight map A D By A Split, the learnable attention can be described as: (17) In the formula Indicates the channel splicing operation, It is channel type Softmax , It is a channel split operation. is the dimension of the query or key, is the scaling factor, Represents a matrix multiplication operation.

4. The method of claim 1, wherein: The integrated focusing submodule in equation (13) IF It consists of specific channel attention and multi-path spatial attention, adopting the design order of channel first and space second, focusing on important feature channels and areas with high expressiveness, thus improving computational efficiency; The distribution of information between channels is further optimized in the submodule of specific channel attention. First, the input features are enhanced channel by channel through deep convolution and activation function GELU to highlight the characteristics of each channel. Then, the importance of each channel is weighted. The final output highlights the key channel features, making the feature distribution more targeted. Multi-path spatial attention captures features from multiple angles by introducing a dual-branch mechanism of maximum pooling and average pooling. Maximum pooling focuses on local saliency and captures the extreme response of features, while average pooling extracts global smooth information and enhances the perception of feature distribution.

Citation Information

Cited By

  • Heterogeneous double-flow fusion method and system for grading diabetic retinopathy

    CN121033041A

  • General image restoration model of Mama diffusion based on contour prior guidance

    CN122391018A