Medical image segmentation system based on multi-architecture fusion and diffusion model
By employing dual-channel encoding, cross-scale feature fusion, and diffusion model post-processing, the limitations of local and global information fusion in medical image segmentation are overcome, achieving high-precision and robust medical image segmentation applicable to various medical image segmentation tasks.
Patent Information
- Application Number
- CN202511206373.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing medical image segmentation methods have limitations in capturing local texture and global context, especially when dealing with large targets or complex scenes that require cross-regional information association. They also lack effective post-processing mechanisms to refine the segmentation mask.
Employing a dual-channel encoding module, a cross-scale feature fusion module, and a diffusion model post-processing module, this approach achieves efficient fusion of local and global information through dynamic deformable convolution and frequency domain information interaction, combined with CNN and Transformer architectures. It also performs iterative denoising optimization through a diffusion model.
It improves the precision and robustness of medical image segmentation, significantly enhances the accuracy of segmentation boundaries and the quality of segmentation results, and demonstrates good generalization ability.
Smart Images

Figure CN121053147A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision and artificial intelligence technology, and specifically relates to a medical image segmentation system based on a multi-architecture fusion and diffusion model. Background Technology
[0002] Medical image segmentation is a key technology in precision medicine. Its goal is to accurately identify and delineate the boundaries of pathological structures (such as tumors and polyps) or organs from medical images, providing a basis for quantitative analysis of diseases, surgical planning, and efficacy evaluation.
[0003] Existing deep learning segmentation methods primarily rely on a single network architecture. For example, convolutional neural network (CNN)-based methods, such as U-Net, excel at capturing local texture and detail features through encoder-decoder structures and skip connections. However, due to the inherent locality of convolutional operations, pure CNN models have limitations in capturing the global context and long-range dependencies of images, especially when dealing with large targets or complex scenes requiring cross-regional information association (such as multi-organ segmentation).
[0004] On the other hand, the Visual Transformer (ViT) architecture utilizes a self-attention mechanism, which can effectively model global contextual information and has shown great potential in natural image segmentation tasks. However, when applied to high-resolution medical images, it suffers from high computational complexity and low efficiency, and its sensitivity to local detail features (such as texture and edges) is not as good as that of CNNs.
[0005] To combine the advantages of both, some hybrid CNN-Transformer architectures have emerged. However, these hybrid models still face challenges in the feature fusion stage. For example, the simple skip connections (such as concatenation or addition) in U-Net that directly fuse shallow detail features and deep semantic features of the encoder may lead to semantic conflicts and noise interference. Especially in medical images with blurred target boundaries, such direct fusion can reduce segmentation accuracy.
[0006] In addition, the output of existing segmentation models often suffers from problems such as blurred boundaries, internal holes, or small artifacts, and lacks an effective post-processing mechanism to refine the segmentation mask.
[0007] Therefore, designing a medical image segmentation method that can efficiently integrate local and global information, dynamically adapt to targets of different scales, and possess refined post-processing capabilities is a pressing problem in the current technological field. Summary of the Invention
[0008] Purpose of the invention: The purpose of this invention is to overcome the shortcomings of the prior art and provide a medical image segmentation system based on a multi-architecture fusion and diffusion model. It aims to solve the limitations of a single network architecture in long-range dependency modeling and local detail capture, optimize the multi-scale feature fusion mechanism, and introduce a post-processing module to improve the precision and robustness of the segmentation results.
[0009] Technical Solution: The present invention discloses a medical image segmentation system based on a multi-architecture fusion and diffusion model, comprising a dual-channel encoding module, a cross-scale feature fusion module, a decoding module, and a diffusion model post-processing module. The dual-channel encoding module uses a dual-channel encoder to extract multi-scale features from the input medical image. The dual-channel encoder includes multiple scale levels, each level including a parallel first channel and a second channel. The cross-scale feature fusion module employs a frequency domain-based cross-scale feature fusion module, replacing the traditional skip connections, to enhance and fuse the multi-scale features output by the encoder. The decoding module receives the enhanced features through a decoder and generates a preliminary segmentation mask. The diffusion model post-processing module uses the preliminary segmentation mask as a noisy input and performs iterative denoising and optimization using a pre-trained denoising diffusion probability model.
[0010] Furthermore, the first channel employs standard convolutional layers and downsampling operations to extract local textures and basic features of the image.
[0011] Furthermore, the second channel uses a deformable convolutional module DCDA based on dynamic attention guidance for feature extraction; the DCDA module dynamically generates the offset of the sampling points according to the input features, so that the receptive field adaptively focuses on the key region and captures non-rigid deformation and long-range dependencies.
[0012] Furthermore, the specific implementation process of the cross-scale feature fusion module is as follows:
[0013] For the shallow feature map output by the encoder, its channels are divided into high-frequency branches and low-frequency branches; the high-frequency branches use dynamic filtering pooling and depthwise separable convolution to enhance local details; the low-frequency branches use a multi-head dilated window attention mechanism to model long-range semantic information in a sparse sampling manner.
[0014] The shallow features processed by high and low frequencies are integrated with the original deep feature map. A global multi-head self-attention mechanism is used to dynamically establish associations and assign weights between features at different scales, generating enhanced features with sufficient information interaction, which serve as the input to the decoder.
[0015] Furthermore, the decoder restores the feature map to its original resolution through progressive upsampling and convolution operations, generating a preliminary segmentation mask.
[0016] Furthermore, the diffusion model post-processing module is a pre-trained denoising neural network conditioned on the original medical image, denoted as ε. θ It is a process of adding Gaussian noise to a clean target segmentation mask So in a positive, progressive manner; this process lasts for T steps, and at each step t, noise is added according to a predefined variance sequence βt. Its single-step transformation process is described by the following probability distribution:
[0017]
[0018] Among them, S t It is the noisy mask at step t, where N represents the normal distribution and I is the identity matrix;
[0019] Training neural network ε θ Learn the inverse process of the above noise-adding process, that is, from a purely noisy image S T The process begins by gradually removing noise, eventually recovering a clean segmentation mask So. The single-step transformation of this reverse process is represented as follows:
[0020] p θ (S t-1 |S t )=N(S t-1 μ θ (S t ,t),∑ θ (S t ,t))
[0021] Where, μ θ and ∑ θ It is made by neural network ε θ The predicted mean and variance.
[0022] Furthermore, the denoising training neural network ε θ The training is conditional, i.e., the network ε θ Not only receive the current noisy mask S t The network ε takes time step t as input and also receives the original medical image I as conditional information; θ The noise ε is trained to predict the noise added to the mask at step t; the objective function of the training is to minimize the difference between the predicted noise and the real noise, as follows:
[0023]
[0024] in, E represents the feature embedding, and ε is the real noise sampled from a standard normal distribution; through training, the network ε θ It can utilize the contextual information of the original image to guide noise prediction and removal.
[0025] Furthermore, the iterative denoising and optimization process using a pre-trained denoising diffusion probability model is implemented as follows:
[0026] S1: Initial segmentation mask S c This is taken as the initial input to the module, initiating the reverse denoising process for S. c By applying T steps of noise, the noisy feature S is obtained. T ;
[0027] S2: Starting from time step t = T, iterate backwards to t = 1; at each time step t, perform the following operations:
[0028] The current noisy mask S t The original medical image I is input into the trained denoising neural network ε. θ In the process, the noise ε of the current step is predicted. θ (S t ,I,t);
[0029] Based on the predicted noise, the mask S from the previous step is calculated using the following formula. t-1 :
[0030]
[0031] in, It is α t The cumulative product;
[0032] S3: Repeat step S2 until t=1 to obtain the final denoising result, which is the final segmentation mask So after fine processing.
[0033] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] 1. Enhanced feature representation capability: Through a dual-channel encoder, especially the dynamically deformable convolutional channel, the model can simultaneously and efficiently capture local fine textures and global long-range dependencies, effectively handling complex structures and non-rigid deformations in medical images, and improving the comprehensiveness and adaptability of feature extraction.
[0035] 2. Superior feature fusion mechanism: The cross-scale fusion module based on frequency domain information decouples and selectively processes high and low frequency information, and uses global attention for cross-scale interaction, avoiding the semantic conflict problem of traditional skip connections, realizing effective alignment and complementarity of features at different levels, and significantly improving the accuracy of segmentation boundaries.
[0036] 3. Higher segmentation precision: The innovative introduction of a diffusion model as a post-processing module iteratively optimizes the initial segmentation results, effectively removing noise, filling holes, and smoothing boundaries, generating higher quality final segmentation results that are more consistent with anatomical structures;
[0037] 4. Excellent performance and generalization: It combines the advantages of CNN, Transformer and diffusion model. Experimental results on multiple public datasets show that the method of this invention outperforms the existing mainstream methods in key indicators such as Dice coefficient and MIoU. It also shows good robustness and generalization ability for medical image segmentation tasks with different modalities and different targets. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the overall architecture of the present invention;
[0039] Figure 2 This is a schematic diagram illustrating the working principle of Dynamic Attention-Guided Deformable Convolution (DCDA) in this invention.
[0040] Figure 3 This is a schematic diagram of the High and Low Frequency Domain Sensing Module (HLFPM) in this invention;
[0041] Figure 4 This is a schematic diagram illustrating the working principle of the multi-head attention mechanism (MHSA) in this invention;
[0042] Figure 5 This is a schematic diagram illustrating the working principle of the Multi-Head Inflated Window Attention (MDWA) in this invention. Detailed Implementation
[0043] The present invention will now be described in further detail with reference to the accompanying drawings:
[0044] This invention proposes a medical image segmentation system based on a multi-architecture fusion and diffusion model, the overall architecture of which is as follows: Figure 1 As shown, it mainly includes a dual-channel feature encoding module, a cross-scale feature fusion module, a decoding and preliminary segmentation module, and a diffusion model refinement post-processing module.
[0045] like Figure 1 As shown, the left-hand area illustrates the system's dual-channel feature encoding module. The system receives a medical image and performs multi-scale feature extraction on the image using a dual-channel encoder, generating a multi-scale feature map. The dual-channel encoder contains multiple scale levels, each including:
[0046] The first channel 101 uses a standard 3x3 convolutional layer for feature extraction, and then downsampling is performed using 2x2 max pooling to reduce spatial resolution and increase the receptive field.
[0047] The second channel 102 employs the Dynamically Deformable Convolution (DCDA) module designed in this invention. For example... Figure 2 As shown, for an input feature map, the DCDA module first uses a lightweight offset prediction network 201 to dynamically predict a two-dimensional offset (Δp) for each convolutional kernel sampling point. Based on the output offset, dynamic attention 202 samples features from the input feature map at the deformation points to obtain keys and values, making the feature map more focused on important regions in the image. Subsequently, convolutional operations sample and aggregate at these sparse locations with offsets, thereby achieving adaptation to the target shape. This design is more flexible than standard convolution and has a lower computational cost than global self-attention.
[0048] At the end of the last scale, the feature maps extracted from the first and second channels are concatenated or added along the channel dimension to form the fused feature output 103 for that scale, which is then passed to the decoder module and the cross-scale feature fusion module.
[0049] Cross-scale feature fusion employs a cross-scale feature fusion module based on frequency domain information to replace the traditional skip connections, thereby enhancing and fusing the multi-scale features output by the encoder.
[0050] The four feature maps (X1, X2, X3, X4) output by the encoder are fed into the cross-scale feature fusion module 104. The core of this module is to refine the shallow features X1 and X2 to better fuse them with the deep features X3 and X4.
[0051] For feature maps X1 and X2, first enter as follows: Figure 3 The high-frequency domain sensing module (HLFPM) 301 is shown. In this module, the feature channel is divided into a high-frequency branch 302 and a low-frequency branch 303.
[0052] In the high-frequency branch 302, operations such as depthwise separable convolution are used to focus on extracting and preserving high-frequency detail information such as edges and textures of the image.
[0053] In the low-frequency branch 303, a multi-head inflated window attention (MDWA) module is used. For example... Figure 5As shown, MDWA performs sparse sampling within a fixed-size window around the query patch at different inflation rates (r = 1, 2, 3...), and then computes self-attention. This approach captures long-range dependencies while avoiding the high computational cost of global attention. The processed high- and low-frequency features are then re-fused to obtain the enhanced shallow features X1' and X2'.
[0054] Finally, all features (X1', X2', X3, X4) at all four scales are flattened into a sequence and fed into the input shown in the figure. Figure 4 This is a standard multi-head self-attention (MHSA) layer. This layer computes dependencies between features at all scales, enabling global information interaction. The final output features are passed to the decoder as dynamic, content-aware skip connections.
[0055] like Figure 1 As shown, the right-hand side of the diagram illustrates the system's decoding module. This module takes the enhanced features, processed by the cross-scale feature fusion module, and inputs them into a decoder network. The decoder then restores the feature map to its original resolution through progressive upsampling and convolution operations, generating a preliminary segmentation mask.
[0056] The diffusion model refinement post-processing module 105 introduces a denoising diffusion probability model (DDPM)-based post-processing module after the decoder generates the initial segmentation result. This module is used to refine and correct the initial segmentation mask. The core of this post-processing module is a pre-trained denoising neural network conditioned on the original medical image, denoted as ε. θ Its working principle is based on a symmetrical noise addition and denoising process.
[0057] The denoising diffusion probability model is based on a positive, progressive process of adding Gaussian noise to a clean target segmentation mask So. This process lasts for T steps, and at each step t, noise is added according to a predefined variance sequence βt. The single-step transformation process can be described by the following probability distribution:
[0058]
[0059] Among them, S t It is the noisy mask at step t, where N represents the normal distribution and I is the identity matrix.
[0060] The goal of the post-processing module in this invention is to train a neural network ε. θ To learn the inverse process of the above noise-adding process, that is, from a purely noisy image S T The process begins by gradually removing noise, ultimately recovering a clean segmentation mask So. This single-step transformation of the reverse process can be represented as:
[0061] pθ (S t-1 |S t )=N(S t-1 μ θ (S t ,t),∑ θ (S t ,t))
[0062] Where, μ θ and ∑ θ It is made by neural network ε θ The predicted mean and variance.
[0063] To achieve accurate segmentation result correction, a denoising neural network ε θ The training of this network is conditional. Specifically, the network ε θ Not only receive the current noisy mask S t The network takes time step t as input and the original medical image I as conditional information. It is trained to predict the noise ε added to the mask at step t. The objective function of the training is to minimize the difference between the predicted noise and the actual noise, which can be expressed as:
[0064]
[0065] in, E represents the feature embedding, and ε is the real noise sampled from a standard normal distribution. Through this training, the network ε θ We learned to use the contextual information of the original image to guide noise prediction and removal.
[0066] During the inference phase, when the decoder generates the initial segmentation mask (denoted as the expected clean result S) c After that, the post-processing module performs the following iterative optimization steps:
[0067] a) Apply the initial segmentation mask S c This is considered the initial input to the module. To initiate the reverse denoising process, S is first... c By applying T steps of noise, the noisy feature S is obtained. T In a preferred embodiment, T is a preset total number of iterations.
[0068] b) Starting from time step t = T, iterate backwards to t = 1. At each time step t, perform the following operations:
[0069] i. Change the current noisy mask S t The original medical image I is input into the trained denoising neural network ε. θ In the process, the noise ε of the current step is predicted. θ (S t ,I,t).
[0070] ii. Based on the predicted noise, calculate the cleaner mask S from the previous step using the following formula. t-1 :
[0071]
[0072] in, It is α t The cumulative product.
[0073] c) Repeat step b) until t=1 to obtain the final denoising result So, which is the final segmentation mask after fine processing.
[0074] Through the above iterative denoising process, the present invention can effectively utilize the rich details and structural information of the original image to correct the boundary blurring, internal holes and small artifacts in the preliminary segmentation results, thereby significantly improving the accuracy and robustness of the final segmentation results.
[0075] To verify the effectiveness of this invention, its method was compared with various existing technology models on the publicly available Kvasir-SEG and Synapse datasets. Evaluation metrics included Dice similarity coefficient (Dice), precision, recall, and mean intersection-union ratio (MIoU); the comparison results are shown in Tables 1 and 2.
[0076] Table 1 shows the performance comparison results on the Kvasir-SEG dataset.
[0077]
[0078]
[0079] Table 2 shows the performance comparison results on the Synapse dataset.
[0080]
[0081] As shown in Table 1, this invention achieves the best performance in four key metrics—Dice similarity coefficient, precision, recall, and mean intersection-over-union ratio (MIoU)—on the Kvasir-SEG polyp segmentation task. Specifically, its Dice coefficient (94.67%) is 5.6 percentage points higher than the classic UNet (89.07%), and its MIoU (89.86%) is significantly improved by 9.55 percentage points. Even compared to the second-best performing advanced hybrid model UCTransNet (Dice 93.26%), it still maintains a significant advantage of 1.41 percentage points.
[0082] As shown in Table 2, in the more complex Synapse multi-organ segmentation task, the present invention again achieves the best performance in Dice coefficient (90.46%) and MIoU (83.45%). Compared with the suboptimal models in this task, such as Deeplabv3-plus (Dice 88.67%) and FusionUnet (MIoU 79.81%), the present invention achieves significant performance improvements of 1.79 and 3.64 percentage points, respectively.
[0083] The experimental results across the aforementioned datasets objectively demonstrate that the multi-architecture fusion scheme proposed in this invention, by organically combining dual-channel encoding, frequency-domain-based feature fusion, and diffusion model post-processing, exhibits superior accuracy and strong generalization ability, whether for the segmentation of polyps with relatively simple morphology or for the segmentation of multi-organs with complex structures and diverse categories. This fully reflects the inventiveness, advancement, and significant application value of this invention in various medical image segmentation scenarios.
[0084] It should be understood that the embodiments and descriptions above are only the principles, main features and advantages of the present invention. Various changes and modifications can be made to the present invention without departing from the spirit and scope of the invention, and all such changes and modifications fall within the protection scope of the present invention.
Claims
1. A medical image segmentation system based on a multi-architecture fusion and diffusion model, characterized in that, The system includes a dual-channel encoding module, a cross-scale feature fusion module, a decoding module, and a diffusion model post-processing module. The dual-channel encoding module uses a dual-channel encoder to extract multi-scale features from the input medical image. The dual-channel encoder includes multiple scale levels, each level including a first channel and a second channel in parallel. The cross-scale feature fusion module adopts a frequency domain-based cross-scale feature fusion module to replace the traditional skip connections, enhancing and fusing the multi-scale features output by the encoder. The decoding module receives the enhanced features through a decoder and generates a preliminary segmentation mask. The diffusion model post-processing module uses the preliminary segmentation mask as a noisy input and performs iterative denoising and optimization using a pre-trained denoising diffusion probability model.
2. The medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 1, characterized in that, The first channel uses standard convolutional layers and downsampling operations to extract local textures and basic features of the image.
3. The medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 1, characterized in that, The second channel uses a deformable convolutional module DCDA based on dynamic attention guidance for feature extraction; the DCDA module dynamically generates the offset of the sampling points according to the input features, so that the receptive field adaptively focuses on the key region and captures non-rigid deformation and long-range dependencies.
4. The medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 1, characterized in that, The specific implementation process of the cross-scale feature fusion module is as follows: For the shallow feature map output by the encoder, its channels are divided into high-frequency branches and low-frequency branches; the high-frequency branches use dynamic filtering pooling and depthwise separable convolution to enhance local details; the low-frequency branches use a multi-head dilated window attention mechanism to model long-range semantic information in a sparse sampling manner. The shallow features processed by high and low frequencies are integrated with the original deep feature map. A global multi-head self-attention mechanism is used to dynamically establish associations and assign weights between features at different scales, generating enhanced features with sufficient information interaction, which serve as the input to the decoder.
5. A medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 1, characterized in that, The decoder restores the feature map to its original resolution through progressive upsampling and convolution operations, generating a preliminary segmentation mask.
6. The medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 1, characterized in that, The diffusion model post-processing module is a pre-trained denoising neural network conditioned on the original medical image, denoted as ε. θ It is a process of adding Gaussian noise to a clean target segmentation mask S0 in a positive, progressive manner; this process lasts for T steps, and at each step t, noise is added according to a predefined variance sequence βt. Its single-step transformation process is described by the following probability distribution: Among them, S t It is the noisy mask at step t, where N represents the normal distribution and I is the identity matrix; Training neural network ε θ Learn the inverse process of the above noise-adding process, that is, from a purely noisy image S T The process begins by gradually removing noise, eventually recovering a clean segmentation mask S0. The single-step transformation of this reverse process is represented as follows: p θ (S t-1 |S t )=N(S t-1 ;μ θ (S t ,t),∑ θ (S t ,t)) Where, μ θ and ∑ θ It is made by neural network ε θ The predicted mean and variance.
7. A medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 6, characterized in that, The denoising training neural network ε θ The training is conditional, i.e., the network ε θ Not only receive the current noisy mask S t The network ε takes time step t as input and also receives the original medical image I as conditional information; θ The noise ε to be added to the mask at step t is trained; the objective function of the training is to minimize the difference between the predicted noise and the real noise, as follows: in, E represents the feature embedding, and ε is the real noise sampled from a standard normal distribution; through training, the network ε θ It can utilize the contextual information of the original image to guide noise prediction and removal.
8. A medical image segmentation system based on a multi-architecture fusion and diffusion model according to claim 1, characterized in that, The iterative denoising and optimization process using a pre-trained denoising diffusion probability model is as follows: S1: Initial segmentation mask S c This is taken as the initial input to the module, initiating the reverse denoising process for S. c By applying T steps of noise, we obtain the noisy feature S. T ; S2: Starting from time step t = T, iterate backwards to t = 1; at each time step t, perform the following operations: The current noisy mask S t The original medical image I is input into the trained denoising neural network ε. θ In the process, the noise ε of the current step is predicted. θ (S t ,I,t); Based on the predicted noise, the mask S from the previous step is calculated using the following formula. t-1 : in, It is α t The cumulative product; S3: Repeat step S2 until t=1 to obtain the final denoising result, which is the final segmentation mask S0 after fine processing.
Citation Information
Patent Citations
Medical image segmentation method based on conditional diffusion model
CN116596949A
Image anomaly detection method based on self-supervised learning and diffusion generation model
CN118037711A
Medical image segmentation method based on dynamic multi-scale conditional diffusion model
CN118247509A
Medical image segmentation method and system based on diffusion model and domain adaptation
CN118799337A
Multilevel feature fusion medical image segmentation method and system based on diffusion model
CN119151969A
Cited By
Multi-modal medical image segmentation method based on sparse hypergraph diffusion network
CN121305094A
Multimodal medical image segmentation method based on sparse hypergraph diffusion network
CN121305094B
Casting body sheet image segmentation method and system based on dynamic noise filtering and detail enhancement
CN121564352A
Diffusion model-based SAR (Synthetic Aperture Radar) marine oil spill intelligent identification method and system
CN121640298A