Mama-based spectrum dynamic fusion and double attention enhancement medical image segmentation method

By using spectral domain fusion and dual attention enhancement of the SDAU-Mamba model, the problems of spatial correlation loss and lesion focusing in Mamba architecture in medical image segmentation are solved, and higher accuracy segmentation of complex edge textures and small lesions is achieved.

CN120876849APending Publication Date: 2025-10-31SHAANXI UNIV OF SCI & TECH
View PDF 0 Cites 27 Cited by

Patent Information

Application Number
CN202510897718.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing medical image segmentation methods based on the Mamba architecture suffer from problems such as lack of spatial correlation and insufficient focus on key pathological features during block processing, resulting in low segmentation accuracy for complex edge textures and small lesions.

Method used

The SDAU-Mamba model is adopted, which combines spectral domain fusion and dual attention enhancement. Multi-scale feature modeling is performed through the MISAP module and pathological region focusing is performed through the BRA mechanism. Spatial-frequency domain feature fusion is achieved by using Fourier transform and self-attention pooling pyramid, and the expressive ability of lesion region is improved through bipolar routing attention mechanism.

Benefits of technology

It significantly improves the segmentation accuracy of complex edge textures and small lesions in medical images, with an average increase of 6.5% in the Dice coefficient. It solves the problem of fragmentation in frequency-spatial feature co-modeling and insufficient dynamic focusing of lesion areas in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876849A_ABST
    Figure CN120876849A_ABST
Patent Text Reader

Abstract

The invention discloses a Mama-based spectrum dynamic fusion and double-attention enhancement medical image segmentation method, which comprises the following steps of: firstly, constructing a Mama integrated spectrum domain and attention pyramid module, fusing spectrum dynamic characteristics and a self-attention pooling mechanism, and performing frequency domain information compensation and local characteristic enhancement to obtain a spectrum dynamic fusion image; the spatial correlation loss caused by image blocking processing is relieved; secondly, designing a layered enhanced U-shaped architecture, deploying an MISAP module in a shallow layer of an encoder to capture multi-scale global context features, introducing a bipolar routing attention mechanism in a deep layer, and dynamically allocating sparse attention weights to focus a key pathological region; according to the method, the segmentation precision of complex edge textures and tiny lesions in medical images can be remarkably improved, and the Dice coefficient in breast tumor, polyp and abdominal organ segmentation tasks is averagely improved by 6.5%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of image processing and computer vision, specifically relating to a medical image segmentation method based on Mamba-based spectral dynamic fusion and dual attention enhancement. Background Technology

[0002] Medical image segmentation, as a core technological component of intelligent diagnosis and treatment systems, plays a crucial role in the era of precision medicine by analyzing anatomical structures and lesion features from massive amounts of image data. Its technological evolution has consistently revolved around two core contradictions: on the one hand, the complex topological structure of human tissues requires models to possess multi-scale spatial continuity modeling capabilities to capture the morphological relationships between organs and lesions; on the other hand, the real-time processing requirements of high-resolution images in clinical scenarios impose stringent constraints on computational resource efficiency. Traditional convolutional neural network-based methods are limited by the progressive feature aggregation mechanism of local receptive fields, exhibiting significant shortcomings in modeling long-range spatial dependencies; while the visual Transformer architecture, although overcoming the global interaction bottleneck through self-attention mechanisms, suffers from computational complexity that is quadratic with image resolution. Therefore, constructing a novel medical image segmentation framework that combines high computational efficiency with strong robustness is not only crucial for providing pixel-level quantitative analysis tools for early lesion screening but also holds significant strategic importance for promoting the leapfrog development of precision medicine from theoretical innovation to clinical application.

[0003] In the field of medical image segmentation, traditional techniques have evolved into two representative technical routes: deep optimization methods based on Convolutional Neural Networks (CNNs) and hybrid modeling paradigms incorporating Transformers. Regarding CNN architecture innovation, improved models, represented by the U-Net family (U-Net++, UNet3+, etc.), establish cross-layer feature transfer mechanisms to optimize the semantic gap problem by constructing multi-scale feature pyramids (such as dense skip connections and full-scale gradient flow architectures). While these methods can construct hierarchical feature representations through layer-by-layer convolutional kernel stacking, they suffer from systemic defects in modeling long-range spatial relationships across organs due to the inherent characteristics of local receptive fields. On the other hand, in the direction of hybrid Transformer models, representative works such as TransUNet and Swin-UNet, although overcoming the global semantic modeling bottleneck by introducing self-attention mechanisms, require mandatory image block operations in their sequential processing paradigm, leading to pixel-level spatial continuity breaks. It is noteworthy that this disruption of topological structure easily causes morphological distortions in segmentation tasks of medical images with anatomical continuity (such as vascular networks and nerve fiber bundles).

[0004] In the current evolution of medical image segmentation technology, the Mamba architecture based on State Space Models (SSM) is leading a new generation of efficient modeling paradigms. The core breakthrough of this approach lies in its dual advantages of linear computational complexity and long-range spatial dependency modeling capabilities, providing an innovative solution for high-resolution medical image processing. Representative technological advancements include: U-Mamba embeds differentiable state space modules into the bottleneck layer of a U-shaped network, constructing a dynamic global context representation through a time-varying parameter system; VM-Unet innovatively adopts a pure visual state space block to construct an encoder-decoder architecture, inheriting the topology-preserving properties of U-Net while achieving cross-dimensional global feature interaction through gated recurrent units; CM-UNet proposes a hierarchical hybrid modeling paradigm, preserving the local inductive bias of convolutional neural networks in shallow networks to capture anatomical structural details, while deep networks establish global semantic associations through state space differential equations; and Pan-Mamba innovatively integrates local morphological features and global pathological semantic information by constructing a multi-resolution state space pyramid. This series of technological breakthroughs marks a paradigm shift in medical image segmentation from static feature extraction to dynamic differential equation modeling, providing a new theoretical framework for balancing long-range dependency modeling and computational efficiency in medical data.

[0005] While current medical image segmentation methods based on the Mamba architecture have made significant progress, they still face two key technical bottlenecks: First, the spatial-frequency domain collaborative representation mechanism is flawed. Existing methods employ image block-based sequential processing paradigms (such as unfolding 2D slices into 1D sequences), which not only disrupt the spatial continuity of anatomical structures through forced dimensionality reduction but also limit feature extraction efficiency in two dimensions: in the spatial domain, the grid effect caused by block operations leads to the loss of detailed features of local microstructures; in the frequency domain, the lack of a feature fusion mechanism in conjunction with Fourier transform results in the ineffective capture of high-frequency texture information. This insufficient feature extraction in both spatial and frequency domains is particularly pronounced in image segmentation tasks with complex anatomical topologies, easily leading to morphological distortion of key anatomical landmarks and pixel-level edge localization deviations. Second, the dynamic focusing ability of lesion regions is insufficient. Existing global state transition equations employ a homogenized feature modeling strategy, lacking a multi-scale dynamic attention allocation mechanism for key lesion regions. Therefore, achieving cross-domain feature fusion and dynamic decoupling of anatomical structures is a key breakthrough direction for improving model robustness and generalization ability. Summary of the Invention

[0006] The purpose of this invention is to provide a medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement, which solves the problems of spatial correlation loss and insufficient focus on key pathological features caused by block processing in existing state space models in medical image segmentation, and can significantly improve the segmentation accuracy of complex edge textures and small lesions in medical images.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A medical image segmentation method based on Mamba dynamic spectral fusion and dual attention enhancement includes the following steps:

[0009] Step 1: Construct the SDA U-Mamba model;

[0010] The SDAU-Mamba model adopts a U-shaped encoder-decoder structure. The encoder path extracts multi-scale features step by step through four downsampling operations. In the shallow stage, the Mamba Integrated Spectrum and Attention Pyramid (MISAP) module is deployed. The MISAP module includes a visual state space module, Spectral Dynamic Feature Fusion (SDFF), and Self-Attention Pyramid Pooling (SAPP). In the deep stage, a Bipolar Routing Attention (BRA) mechanism is introduced.

[0011] Step 2: Preprocess the images in the multi-source medical image dataset. According to the official predefined partitioning scheme of each dataset, divide the dataset into training set, validation set and test set.

[0012] Step 3: Train the SDAU-Mamba model;

[0013] Step 3.1: The input image is divided into fixed-size 16×16 image blocks and embedded. The resulting embedded features are used as encoder input and passed layer by layer to the shallow encoding stages Stage 1 to Stage 3. In the shallow encoding stage, the image block features are first modeled for multi-scale features by the MISAP module. The Visual State Space Module (VSS-Block) captures long-range spatial dependencies through a four-way scanning mechanism. The Spectral Dynamic Feature Fusion (SDFF) module dynamically weights and fuses the global features extracted from the frequency domain with the local detail features obtained by deformable convolution. It also uses the Self-Attention Pyramid Pooling (SAPP) strategy to achieve effective aggregation of multi-scale contextual information.

[0014] Step 3.2: In the deep coding stages Stage 4 to Stage 5, the SDAU-Mamba model uses a bipolar route attention mechanism (BRA) to refine and focus the features extracted in the previous stage. The coarse-grained attention routing mechanism selects the top 4 image block regions with the highest semantic relevance to the pathological region. Then, fine-grained attention calculation is performed in the selected regions to enhance the expression of key lesion features and suppress the interference of background regions.

[0015] Step 3.3: Calculate the total loss and update the parameters of the SDAU-Mamba model;

[0016] Step 4: Evaluate the performance of the SDAU-Mamba model using the validation set and optimize the parameters;

[0017] Step 5: Input the medical image to be segmented into the trained SDAU-Mamba model and output the segmentation result.

[0018] Furthermore, the spectral dynamic feature fusion module in step 3.1 includes a frequency domain filtering enhancement module, a deformable convolution modeling module, and a cross-domain dynamic fusion module. The frequency domain filtering enhancement module is used to extract global frequency information from the input feature map and enhance key frequency components through a dynamic filtering mechanism. The deformable convolution modeling module is used to capture the geometric deformation features of the lesion region in the spatial domain and enhance the model's ability to perceive structural changes. The cross-domain dynamic fusion module is used to fuse features from the frequency domain and the spatial domain to take into account both global perception and local structural modeling, thereby improving the detail restoration capability of the segmentation results.

[0019] Furthermore, the specific steps of the frequency domain filtering enhancement module include:

[0020] (a) Frequency domain transformation: transforming the input feature map X∈r H×W×C By mapping the two-dimensional Fast Fourier Transform (FFT) to the frequency domain, the frequency domain feature representation is obtained:

[0021] F(X) = FFT2D(X)

[0022] Where H, W, and C are the height, width, and number of channels of the image, respectively, and F(X) is the frequency domain output;

[0023] (b) Channel attention weighting: Channel compression of frequency domain features is performed using 1×1 convolution, and a dynamic filtering weight matrix W is generated by the Softmax function. θ :

[0024] W θ =Softmax(Conv 1×1 (F(X))

[0025] (c) Frequency Domain Reconstruction: Perform a two-dimensional inverse Fourier transform (IFFT) on the weighted frequency domain features to obtain the enhanced spatial domain features F. freq :

[0026] F freq =IFFT2D(W θ ⊙F(X))

[0027] Here, ⊙ represents element-wise multiplication.

[0028] Furthermore, the processing flow of the deformable convolution modeling module includes:

[0029] (a) Offset generation: Apply a 3×3 convolution to the input image X to generate the corresponding spatial offset field Δp. n ∈R H ×W×C :

[0030] △p n =Conv 3×3 (X)

[0031] (b) Deformable Sampling: The sampling position of the standard convolution kernel is adjusted according to the offset to achieve adaptive modeling of local deformable structures and output feature F. 6ef (p0) is defined as follows:

[0032]

[0033] Where p0 represents the current convolution center position, R is the receptive field region, and p n At the location of each sampling point within the receptive field R, w n These are the convolution weights.

[0034] Furthermore, the specific fusion strategy of the cross-domain dynamic fusion module is as follows:

[0035] F out =αF freq +(1-α)·F def

[0036] Wherein, the fusion coefficient α∈[0,1] can be obtained through learning or fixed settings, reflecting the relative contributions of frequency domain and spatial domain features in the fusion process, F out This represents the fusion features of the final output.

[0037] Furthermore, the self-attention pyramid pooling (SAPP) strategy in step 3.1 includes multi-scale feature aggregation, context fusion representation construction, spatial attention weight generation, and upsampling / downsampling fusion output, as detailed below:

[0038] (1) Multi-scale feature aggregation: For the intermediate feature map X∈R output by the encoderH×W×C Spatial pyramid pooling is performed, and different scale parameters k∈{1,3,6,8} are selected to control the size of the pooling grid, generating a multi-scale context representation.

[0039] Two-dimensional adaptive average pooling (AdaptivePool2d) is used, combined with the image aspect ratio. The pooling region is dynamically adjusted, and the pooling process is expressed as follows:

[0040] Y k =AdaptivePool2d(X,(k,max(1,[ar×k])))

[0041] Where Y k This represents the output feature map after adaptive pooling.

[0042] (2) Context-fusion expression construction

[0043] The feature maps Y1, Y2, Y3, and Y4 after pooling at each scale are upsampled to the original resolution H×W using bilinear interpolation, and then stitched together along the channel dimension to form a unified pyramid fusion representation.

[0044] z=Concat(Flatten(Y1),Flatten(Y2),Flatten(Y3),Flatten(Y4))

[0045] Where z represents the final fused feature map, Flatten means flattening each two-dimensional feature map into a one-dimensional channel sequence, and Concat means concatenating the channels.

[0046] (3) Spatial attention weight generation

[0047] A lightweight spatial attention mechanism is applied to the fused feature map z, generating a dynamic weighting matrix A through local perception and nonlinear mapping. k :

[0048] A k =Sigmoid(Conv 1×1 (GeLU(Conv 3×3 (z)))

[0049] Among them, Conv_{3×3} is used to capture local spatial relationships, GeLU activation function improves nonlinear expression ability, Conv_{1×1} generates spatial weights after channel compression, and finally the attention weights are obtained by normalization through Sigmoid activation function.

[0050] (4) Feature reconstruction and output fusion

[0051] The generated spatial weight Ak The fused feature map z is element-wise weighted, then upsampled to restore the spatial resolution to match the input image, and this is used as the output Z for the current stage.

[0052] Z = Upsample(A k ⊙z)

[0053] Where ⊙ represents element-wise multiplication, and Upsample is the bilinear upsampling function.

[0054] Furthermore, the Bipolar Routing Attention (BRA) mechanism in step 3.2 includes region partitioning and projection coding, region affinity calculation, dynamic route index generation, and local token aggregation, with the specific process as follows:

[0055] (1) Region division and feature projection coding

[0056] First, input feature map X∈R H×W×C The area is uniformly divided into S×S non-overlapping regions, each region containing HW / S 2 Each token (i.e., a pixel block or embedding block) is used; subsequently, the input feature map is encoded in three ways using a learnable linear projection matrix to generate query Q, key K, and value V vectors respectively:

[0057] Q = XW q K = XW k V = XW v

[0058] Among them W q W k W v ∈R c×c The projection matrix;

[0059] (2) Regional-level semantic affinity modeling

[0060] Perform an average pooling operation on all tokens within each region to obtain the aggregated representation of the query and key for each region, denoted as follows: and

[0061]

[0062] Where Ω i Represents the set of token indices for the i-th region, |Ω i | represents a set| Ω i The number of elements in |Q j and K j Let represent the query vector and key vector of the j-th token, respectively;

[0063] Next, the semantic similarity matrix between all regions is calculated. Each element Indicate the semantic relevance between region i and region j:

[0064] A r =Q r (K r ) T

[0065] Q r and K r These represent the query matrix and the key matrix, respectively.

[0066] (3) Generation of regional dynamic routing index

[0067] Based on the regional similarity matrix A r For each region, the top k most relevant regions are selected as its attention scope to generate a routing index matrix. Where the i-th row Indicates the k region indices most relevant to region i:

[0068] I r =topkIndex(A r )

[0069] Among them, topkIndex(A r ) is a function call used to extract data from matrix A. r Select the indices of the k elements with the highest values;

[0070] (4) Token-level attention computation and interaction

[0071] According to route index I r Collect the corresponding keys and values ​​from the entire graph to construct a local key-value cluster representation:

[0072] K g =gather(K,I) r ),V g =gather(V,I) r )

[0073] Where K g , That is, each region only participates in token interactions within its relevant region, gather(K,I) r ) and gather(V,I r ) is a function call used to collect data from matrices K and V by I. r The element at the specified position;

[0074] Finally, for each original query token, attention weights and feature weighted summations are performed only on its corresponding set of aggregate keys to generate the final output:

[0075] Furthermore, in step 3.3, the loss function L of the total loss... total A linear combination strategy of binary cross-entropy and Dice loss is adopted, specifically in the form of:

[0076] L total =λL BCE +(1-λ)L Dice

[0077] Where λ = 0.5, L BCE For binary cross-entropy loss, L Dice This is a loss for Dice.

[0078] Compared with the prior art, the present invention has the following beneficial effects:

[0079] (1) This invention addresses the two major technical bottlenecks of spatial topological fragmentation and insufficient lesion focusing in existing Mamba architectures for medical image segmentation, proposing a spectral dynamic fusion combined with attention-enhanced U-Mamba network (SDAU-Mamba). By constructing a spectral domain fusion attention pyramid (MISAP) module, and utilizing the synergistic mechanism of Fourier transform and self-attention pooling pyramid, dynamic fusion of spatial and frequency domain features is achieved while retaining the linear computational complexity advantage of Mamba. The designed hierarchical enhanced U-shaped architecture uses a three-level spectral domain fusion attention pyramid module to construct a multi-scale feature pyramid in the early stage of encoding, and introduces a bipolar routing attention mechanism in the later stage of encoding to suppress redundant background through dynamic sparse masking. Experiments show that SDAU-Mamba improves the Dice coefficient by an average of 6.5% in breast tumor, polyp, and abdominal organ segmentation tasks.

[0080] (1) The Mamba integrated spectrum and attention pyramid module of this invention solves the problem of fragmentation in frequency-spatial feature collaborative modeling in traditional methods through a dual-path collaborative mechanism. The Spectral Dynamic Feature Fusion (SDFF) path combines Fast Fourier Transform (FFT) with deformable convolution. The former enhances the global continuity of lesion texture through dynamic weighting of frequency components, while the latter accurately captures the geometric deformation of organ edges through adaptive sampling point prediction. The two achieve deep fusion of spectral oscillation features and spatial deformation features through a cross-domain feature calibration matrix. The Self-Attention Pyramid Pooling (SAPP) path constructs a four-level progressive feature distillation system and uses a spatial attention mechanism to alleviate the problem of insufficient extraction of spatial structural information caused by Mamba.

[0081] (2) The U-shaped network proposed in this invention achieves a dynamic balance between global and local features through a phased optimization strategy. The shallow encoder embeds a MISAP module, uses a multi-resolution spectral pyramid to construct an anatomical prior knowledge base, suppresses noise interference through dynamic frequency domain filtering, and establishes a cross-modal feature alignment mechanism. The deep encoder introduces a bipolar routing attention mechanism, focuses on key pathological regions based on a sparse routing strategy, and enhances the semantic coherence of small-scale targets through a cross-layer interactive network. This hierarchical design strengthens the global topology modeling capability at the shallow layer and achieves refined representation of lesion regions at the deep layer, avoiding computational redundancy in traditional models and solving the feature fragmentation problem, forming a progressive learning framework from macroscopic anatomy to microscopic pathology. Attached Figure Description

[0082] Figure 1 This is a model framework diagram of the present invention;

[0083] Figure 2 This is a schematic diagram of the spectral dynamic feature fusion SDFF module of the present invention. Detailed Implementation

[0084] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0085] like Figure 1 As shown in the figure, the specific steps of the medical image segmentation method based on Mamba dynamic fusion and dual attention enhancement described in this embodiment are as follows:

[0086] Step 1: Construct the SDAU-Mamba model

[0087] The SDAU-Mamba model employs a U-shaped encoder-decoder structure. The encoder path progressively extracts multi-scale features through four downsampling operations: In the shallow stage, the Mamba Integrated Spectrum and Attention Pyramid (MISAP) module is deployed, consisting of a Visual State Space Block (VSS-Block), Spectral Dynamic Feature Fusion (SDFF), and Self-Attention Pyramid Pooling (SAPP). It uses a four-way orthogonal scanning strategy to unfold two-dimensional features into pseudo-sequences for state space modeling, and integrates the local deformation perception of deformable convolution with the global adjustment capability of the Fourier frequency domain. In the deep stage, a Bipolar Routing Attention (BRA) mechanism is introduced, utilizing dynamic sparse weight allocation to focus on key pathological regions. The decoder progressively restores resolution through patch embedding and fuseds the multi-scale features passed from the encoder via skip connections. Specifically, the Visual State Space Block (VSS-Block) uses four-way scanning unfolding to achieve cross-directional long-range dependency modeling; Spectral Dynamic Feature Fusion (SDFF) enhances local feature adaptability through deformable convolution offset learning; and Self-Attention Pyramid Pooling (SAPP) establishes global spatial associations using multi-scale pooling and spatial attention. This architecture effectively solves the problems of spatial continuity breakage and local-global feature imbalance in medical image segmentation through joint frequency-spatial domain modeling and hierarchical attention mechanism.

[0088] Step 2: Preprocess the images in the multi-source medical image dataset. Based on the officially defined partitioning scheme of each dataset, divide the dataset into training set, validation set, and test set.

[0089] This invention selects publicly available datasets covering three mainstream medical imaging modalities for experimental validation: the BUSI dataset contains 600 breast ultrasound images, mainly used for the task of identifying microcalcification features in benign and malignant lesions; the CVC-ClinicDB dataset contains 612 colonoscopy video keyframes, suitable for the task of polyp edge segmentation under dynamic lighting and mucosal high reflectivity conditions; and the CHAOS-Liver dataset contains 2,280 abdominal CT / MRI multimodal volumetric data, mainly used to validate the robustness and continuity of anatomical structure modeling under the environment of organ deformation and cross-modal texture differences.

[0090] During data preprocessing, the original medical images underwent standardized processing, including image resizing and intensity normalization. All input images were uniformly scaled to a resolution of 224×224 pixels. Basic geometric augmentation strategies were also introduced, including random rotation (±15°), horizontal flipping, and vertical flipping, to increase data diversity and mitigate the risk of model overfitting. To ensure consistency and objectivity across modal evaluations, no additional post-processing algorithms or domain adaptation techniques were employed. The training, validation, and test sets for each dataset were partitioned using the officially provided predefined partitioning criteria.

[0091] Step 3: Train the SDAU-Mamba model

[0092] The preprocessed medical images are input into the SDAU-Mamba model, which sequentially completes phased tasks such as feature extraction, model training, region focusing, and segmentation prediction. This SDAU-Mamba model comprehensively optimizes the expressive power of cross-scale features and the discrimination performance of key regions by fusing spectral dynamic features and a multiple attention mechanism.

[0093] Step 3.1: The input image is divided into fixed-size 16×16 image blocks and embedded. The resulting embedded features are used as encoder input and passed layer by layer to the shallow encoding stages Stage 1 to Stage 3. In the shallow encoding stage, the image block features are first modeled for multi-scale features by the MISAP module. The Visual State Space module (VSS-Block) captures long-range spatial dependencies through a four-way scanning mechanism. The Spectral Dynamic Feature Fusion (SDFF) module dynamically weights and fuses the global features extracted from the frequency domain with the local detail features obtained by deformable convolution, and uses the Self-Attention Pyramid Pooling (SAPP) strategy to effectively aggregate multi-scale contextual information.

[0094] (I) The training process of the spectral dynamic feature fusion module proposed in this invention:

[0095] This invention proposes a feature modeling method that combines a dynamic frequency-domain enhancement path with a spatial-domain geometric adaptive path to improve the recognition of blurred boundaries and heterogeneous lesions in medical images. This method, through an end-to-end training framework, fully exploits global information in the frequency domain and local geometric changes in the spatial domain while maintaining spatial semantic continuity, thereby achieving efficient modeling and accurate segmentation of complex lesion features. Figure 2 As shown, the overall structure consists of three stages: a frequency domain filtering enhancement module, a deformable convolution modeling module, and a cross-domain dynamic fusion module. The specific implementation process is as follows:

[0096] (1) Frequency domain filtering enhancement module

[0097] This stage aims to extract global frequency information from the input feature map and enhance key frequency components through a dynamic filtering mechanism. The specific steps are as follows:

[0098] (a) Frequency domain transformation: transforming the input feature map X∈R H×W×C By mapping the two-dimensional Fast Fourier Transform (FFT) to the frequency domain, the frequency domain feature representation is obtained:

[0099] F(X) = FFT2D(X)

[0100] Where H, W, and C are the height, width, and number of channels of the image, respectively, and F(X) is the frequency domain output.

[0101] (b) Channel attention weighting: Channel compression of frequency domain features is performed using 1×1 convolution, and a dynamic filtering weight matrix W is generated by the Softmax function. θ :

[0102] W θ =Softmax(Conv 1×1 (F(X))

[0103] (c) Frequency Domain Reconstruction: Perform a two-dimensional inverse Fourier transform (IFFT) on the weighted frequency domain features to obtain the enhanced spatial domain features F. freq :

[0104] F freq =IFFT2D(W θ ⊙F(X))

[0105] Here, ⊙ represents element-wise multiplication.

[0106] (2) Deformable Convolution Modeling Module

[0107] This stage is used to capture the geometric deformation features of the lesion region in the spatial domain, enhancing the model's ability to perceive structural changes. The processing flow is as follows:

[0108] (a) Offset generation: Apply a 3×3 convolution to the input image X to generate the corresponding spatial offset field Δp. n ∈R H ×W×C :

[0109] △p n =Conv 3×3 (X)

[0110] (b) Deformable Sampling: The sampling position of the standard convolution kernel is adjusted according to the offset to achieve adaptive modeling of locally deformed structures. Output Feature F def (p0) is defined as follows:

[0111]

[0112] Where p0 represents the current convolution center position, R is the receptive field region, and p n At the location of each sampling point within the receptive field R, w n These are the convolution weights.

[0113] (3) Cross-domain dynamic fusion module

[0114] This stage fuses features from the frequency and spatial domains to balance global perception and local structure modeling, thereby improving the detail reproduction capability of the segmentation results. The specific fusion strategy is as follows:

[0115] F out =αF freq +(1-α)·F def

[0116] Wherein, the fusion coefficient α∈[0,1] can be obtained through learning or fixed settings, reflecting the relative contributions of frequency domain and spatial domain features in the fusion process, F out This represents the fusion features of the final output.

[0117] (II) The training process of pyramid spatial attention proposed in this invention:

[0118] To further enhance the model's ability to model cross-scale spatial dependencies, this invention proposes a context enhancement method based on the synergistic optimization of multi-scale pyramid pooling and spatial attention mechanisms, to supplement the shortcomings of state-space networks (such as Mamba) in spatial information modeling. By applying multi-granularity context awareness and spatially guided weighting to feature maps, the sensitivity to lesion structures at different scales is effectively improved. The overall process includes four stages: multi-scale feature aggregation, context fusion representation construction, spatial attention weight generation, and upsampling / downsampling fusion output, as detailed below:

[0119] (1) Multi-scale feature aggregation

[0120] First, the intermediate feature map X∈R output by the encoder... H×W×C Spatial pyramid pooling is performed, and different scale parameters k∈{1,3,6,8} are selected to control the size of the pooling grid, generating a multi-scale context representation. Specifically, two-dimensional adaptive average pooling (AdaptivePool2D) is used, combined with the image aspect ratio. The pooling region is dynamically adjusted, and the pooling process is expressed as follows:

[0121] Y k =AdaptivePool2d(X,(k,max(1,[ar×k])))

[0122] Where Y k This represents the output feature map after adaptive pooling; this operation can extract multi-granular spatial features from local details to global semantics, providing multi-level information support for subsequent attention mechanisms.

[0123] (2) Context-fusion expression construction

[0124] The feature maps Y1, Y2, Y3, and Y4 after pooling at each scale are upsampled to the original resolution H×W using bilinear interpolation, and then stitched together along the channel dimension to form a unified pyramid fusion representation.

[0125] z=Concat(Flatten(Y1),Flatten(Y2),Flatten(Y3),Flatten(Y4))

[0126] Here, z represents the final fused feature map output, Flatten indicates flattening each two-dimensional feature map into a one-dimensional channel sequence, and Concat indicates the concatenation operation along the channel dimension. This step achieves the joint representation of multi-scale semantic context, fusing information from different receptive fields to enhance spatial discrimination capabilities.

[0127] (3) Spatial attention weight generation

[0128] To assign different importance weights to different spatial locations, a lightweight spatial attention mechanism is applied to the fused feature map z, generating a dynamic weighting matrix A through local perception and nonlinear mapping. k :

[0129] A k =Sigmoid(Conv 1×1 (GeLU(Conv 3×3 (z)))

[0130] Among them, Conv_{3×3} is used to capture local spatial relationships, GeLU activation function enhances nonlinear expression ability, Conv_{1×1} generates spatial weights after channel compression, and finally attention weights are obtained by Sigmoid normalization.

[0131] (4) Feature reconstruction and output fusion

[0132] The generated spatial weight A k The fused feature map z is element-wise weighted, then upsampled to restore the spatial resolution to match the input image, and this is used as the output Z for the current stage.

[0133] Z = Upsample(A k ⊙z)

[0134] Here, ⊙ represents element-wise multiplication, and Upsample is the bilinear upsampling function. This output feature map integrates multi-scale contextual information and spatially guided attention, significantly enhancing the model's ability to perceive lesion structures of different sizes and blurred boundaries.

[0135] Step 3.2: In the deep encoding stages (Stages 4 to 5), the SDA U-Mamba model employs a bipolar routed attention mechanism (BRA) to refine and focus the features extracted in the previous stage. A coarse-grained attention routing mechanism selects the top-4 image regions with the highest semantic relevance to the pathological area. Subsequently, fine-grained attention computation is performed within the selected regions to enhance the expression of key lesion features and suppress interference from background areas.

[0136] The training process of the two-level routing attention mechanism in this invention:

[0137] To further improve the global modeling efficiency and content awareness of the model in large-size medical images, this invention introduces a two-level routing attention mechanism based on region-level coarse screening and token-level fine computation, effectively alleviating the computational and storage bottlenecks faced by traditional self-attention mechanisms under high-resolution input. This mechanism performs coarse-grained screening at the region level and performs token-level fine computation within a subset, combined with a hardware-friendly dense matrix implementation, significantly reducing computational complexity while maintaining semantic globality. It is particularly suitable for scenarios where key regions are sparsely distributed in medical images. It mainly includes four stages: region partitioning and projection encoding, region affinity calculation, dynamic routing index generation, and local token aggregation interaction. The specific implementation process is as follows:

[0138] (1) Region division and feature projection coding

[0139] First, input feature map X∈R H×W×C The area is uniformly divided into S×S non-overlapping regions, each region containing HW / S 2 Each token (i.e., a pixel block or embedding block) is then used to encode the input feature map in three ways using a learnable linear projection matrix, generating query Q, key K, and value V vectors respectively:

[0140] Q = XW q K = XW k V = XW v

[0141] Among them W q W k W v ∈R c×c This is the projection matrix.

[0142] (2) Regional-level semantic affinity modeling

[0143] Perform an average pooling operation on all tokens within each region to obtain the aggregated representation of the query and key for each region, denoted as follows: and

[0144]

[0145] Where Ω i Represents the set of token indices for the i-th region, |Ω i | represents a set| Ω i The number of elements in |Q j and K j Let represent the query vector and key vector of the j-th token, respectively.

[0146] Next, the semantic similarity matrix between all regions is calculated. Each element Indicate the semantic relevance between region i and region j:

[0147] A r =Q r (K r ) T

[0148] Q r and K r These represent the query matrix and the key matrix, respectively.

[0149] (3) Generation of regional dynamic routing index

[0150] Based on the regional similarity matrix A r For each region, select the k most relevant regions as its attention scope. Generate a routing index matrix. Where the i-th row Indicates the k region indices most relevant to region i:

[0151] I r =topkIndex(A r )

[0152] This index matrix is ​​used to limit the candidate range for subsequent attention calculations, thereby effectively controlling computational complexity.

[0153] (4) Token-level attention computation and interaction

[0154] According to route index I r Collect the corresponding keys and values ​​from the entire graph to construct a local key-value cluster representation:

[0155] K g =gather(K,I) r ),V g=gather(V,6) r )

[0156] Where K g , That is, each region only participates in token interactions within its relevant region.

[0157] Finally, for each original query token, attention weights and feature weighted summations are performed only on its corresponding set of aggregate keys to generate the final output:

[0158]

[0159] This localized attention mechanism effectively focuses on the context of key regions while avoiding dense attention computation across the entire graph, thus balancing efficiency and expressive power.

[0160] Step 3.3: Calculate the total loss and update the parameters of the SDAU-Mamba model.

[0161] The loss function L of this invention total A linear combination strategy of binary cross-entropy and Dice loss is employed to collaboratively optimize the class imbalance and gradient stability issues in medical image segmentation. The specific form is as follows:

[0162] L total =λL BCE +(1-λ)L Dice

[0163] Where λ = 0.5, L BCE For binary cross-entropy loss, pixel-level classification gradients are provided to enhance the detailed learning of lesion edges; L Dice To mitigate the Dice loss, the foreground-background pixel imbalance is alleviated by softening the IoU metric.

[0164] Step 4: Evaluate the performance of the SDAU-Mamba model using the validation set and optimize the parameters.

[0165] After model training, the proposed SDAU-Mamba network is evaluated using a pre-defined validation set to measure its generalization ability on unseen data. Evaluation metrics (such as Dice coefficient, IoU, HD, and HD95) on the validation set are calculated to quantitatively compare the model's segmentation results with the ground truth annotations, analyzing its performance in key structure recognition and boundary reconstruction. Simultaneously, model hyperparameters (such as learning rate, regularization coefficient, and loss weights) are dynamically adjusted based on the validation results to further optimize network performance, ensuring the final model achieves an optimal balance between accuracy and stability.

[0166] Step 5: Input the medical image to be segmented into the trained SDAU-Mamba model and output the segmentation result.

[0167] The medical image to be segmented is input into a pre-trained SDA U-Mamba model. The model performs feature extraction and semantic parsing through its end-to-end inference structure, ultimately outputting the corresponding segmentation result. This result can be used to annotate lesion regions, organ boundaries, or other medical structures, providing a high-precision image segmentation reference for subsequent clinical diagnosis, auxiliary analysis, or downstream tasks.

[0168] The effects of this invention can be further illustrated by the following experiments.

[0169] Analysis of Experimental Results for Medical Image Segmentation:

[0170] (a) Experimental Setup and Evaluation Metrics: This experiment was conducted on three benchmark medical imaging datasets (BUSI breast ultrasound, CVC-ClinicDB colonoscopy polyps, and CHAOS-Liver multimodal abdominal CT / MRI), with a uniform input resolution of 224×224. The PyTorch framework and NVIDIA 3090 GPU platform were used. The optimizer was Adam (learning rate 0.0001, batch size 2), with 400 training epochs, and only basic data augmentation (rotation / flipping) was applied. The evaluation system included four core metrics:

[0171]

[0172]

[0173] HD(A,B)=max(sub a∈A inf b∈B ||ab||,sub b∈B inf a∈A ||ba||)

[0174] Where Y represents the true labeled binary segmentation mask, The model represents the segmentation mask predicted by the model. A and B represent the surface point sets of the ground truth and the predicted result, respectively. sup represents the supremum (maximum value), inf represents the infremum (minimum value), and ||ab|| is the Euclidean distance between two points a and b. IoU measures the spatial overlap between the predicted region and the ground truth. This metric is sensitive to the integrity of the segmented target, especially in ultrasound images with blurred edges, and has important clinical significance. Dice evaluates the foreground pixel matching accuracy. It outperforms IoU in evaluating the segmentation performance of small targets (such as microcalcifications) and can effectively reflect the model's adaptability to imbalanced data. HD95 quantifies the maximum local deviation between the segmentation boundary and the ground truth contour. It is calculated using the 95th percentile of the distance distribution. This metric is directly related to the stringent requirements for the continuity of anatomical structures in clinical diagnosis, such as topological conformality of the liver vascular system. HD measures the maximum spatial mismatch between two point sets.

[0175] (b) Experimental Results Analysis: In the breast ultrasound segmentation task (BUSI), as shown in Table 1, SDA U-Mamba significantly outperformed mainstream methods (U-Net: 58.39% IoU / 72.56% Dice; Swin-Unet: 58.92% IoU / 74.72% Dice) with 65.73% IoU and 78.72% Dice, and HD95 was reduced by 47.7% compared to VM-Unet (7.20 vs 13.77). This indicates that the MISAP module effectively enhances the representation ability of blurred boundaries in ultrasound images through multi-scale spectral domain feature fusion, while the BRA mechanism suppresses the interference of speckle noise on the localization of pathological areas.

[0176] For colonoscopy polyp segmentation (CVC-ClinicDB), as shown in Table 2, the model achieved a new performance record with 86.46% IoU and 92.52% Dice (compared to Rolling-Unet: 86.17% IoU / 91.70% Dice), while maintaining an excellent HD95 score of 3.32. Experiments verified that the attention pyramid pooling strategy of the SAPP module can maintain the spatial continuity of polyp boundaries under dynamic lighting and mucosal reflection interference, solving the problem of local semantic breaks when processing endoscopic images using traditional Mamba blocks.

[0177] In the multimodal liver segmentation (CHAOS-Liver), as shown in Table 3, SDA U-Mamba achieved 96.67% IoU and 98.28% Dice (Swin-U-Mamba: 95.79% IoU / 97.32% Dice), with an HD95 of 0.043, representing a 4.78-fold improvement over VM-Unet (0.043 vs 0.009 → 0.056). Although this dataset already possesses high baseline performance due to its high-quality annotations, the model significantly optimizes the topological consistency of the liver parenchyma and vascular system through global spectral domain modeling in shallow MISAP and cross-layer feature routing in deep BRA, demonstrating its robustness in segmenting complex anatomical structures.

[0178] Cross-dataset comparisons show that the model improves the average IoU by 6.5% across the three modalities of ultrasound (BUSI), endoscopy (CVC), and CT / MRI (CHAOS), validating its core advantage of balancing global dependence and local discriminative features.

[0179] Table 1 compares this invention with other advanced segmentation methods on the BUSI dataset.

[0180]

[0181] Table 2 compares this invention with other advanced segmentation methods on the CVC dataset.

[0182]

[0183]

[0184] Table 3 compares this invention with other advanced segmentation methods on the CHAOS dataset.

[0185]

[0186] This invention constructs a Mamba integrated spectral domain and attention pyramid module, fusing spectral dynamic features with a self-attention pooling mechanism. Through frequency domain information compensation and local feature enhancement, it alleviates the spatial correlation loss caused by image block processing. Secondly, it designs a hierarchical enhanced U-shaped architecture, deploying a MISAP module in the shallow layer of the encoder to capture multi-scale global contextual features, and introducing a bipolar routing attention mechanism in the deep layer to dynamically allocate sparse attention weights to focus on key pathological regions. This invention can not only significantly improve the segmentation accuracy of complex edge textures and small lesions in medical images, but also maintain efficient inference speed with linear computational complexity, providing a highly reliable analysis tool for clinical pathological diagnosis.

Claims

1. A medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement, characterized in that, Includes the following steps: Step 1: Construct the SDA U-Mamba model; The SDAU-Mamba model adopts a U-shaped encoder-decoder structure. The encoder path extracts multi-scale features step by step through four downsampling operations. In the shallow stage, the Mamba Integrated Spectrum and Attention Pyramid (MISAP) module is deployed. The MISAP module includes a visual state space module, Spectral Dynamic Feature Fusion (SDFF), and Self-Attention Pyramid Pooling (SAPP). In the deep stage, a Bipolar Routing Attention (BRA) mechanism is introduced. Step 2: Preprocess the images in the multi-source medical image dataset. According to the official predefined partitioning scheme of each dataset, divide the dataset into training set, validation set and test set. Step 3: Train the SDAU-Mamba model; Step 3.1: The input image is divided into fixed-size 16×16 image blocks and embedded. The resulting embedded features are used as encoder input and passed layer by layer to the shallow encoding stages Stage 1 to Stage 3. In the shallow encoding stage, the image block features are first modeled for multi-scale features by the MISAP module. The Visual State Space Module (VSS-Block) captures long-range spatial dependencies through a four-way scanning mechanism. The Spectral Dynamic Feature Fusion (SDFF) module dynamically weights and fuses the global features extracted from the frequency domain with the local detail features obtained by deformable convolution. It also uses the Self-Attention Pyramid Pooling (SAPP) strategy to achieve effective aggregation of multi-scale contextual information. Step 3.2: In the deep coding stages Stage 4 to Stage 5, the SDAU-Mamba model uses a bipolar route attention mechanism (BRA) to refine and focus the features extracted in the previous stage. The coarse-grained attention routing mechanism selects the top 4 image block regions with the highest semantic relevance to the pathological region. Then, fine-grained attention calculation is performed in the selected regions to enhance the expression of key lesion features and suppress the interference of background regions. Step 3.3: Calculate the total loss and update the parameters of the SDAU-Mamba model; Step 4: Evaluate the performance of the SDAU-Mamba model using the validation set and optimize the parameters; Step 5: Input the medical image to be segmented into the trained SDAU-Mamba model and output the segmentation result.

2. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 1, characterized in that, The spectral dynamic feature fusion module in step 3.1 includes a frequency domain filtering enhancement module, a deformable convolution modeling module, and a cross-domain dynamic fusion module. The frequency domain filtering enhancement module is used to extract global frequency information from the input feature map and enhance key frequency components through a dynamic filtering mechanism. The deformable convolution modeling module is used to capture the geometric deformation features of the lesion region in the spatial domain and enhance the model's ability to perceive structural changes. The cross-domain dynamic fusion module is used to fuse features from the frequency domain and the spatial domain to take into account both global perception and local structural modeling, thereby improving the detail restoration ability of the segmentation results.

3. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 2, characterized in that, The specific steps of the frequency domain filtering enhancement module include: (a) Frequency domain transformation: transforming the input feature map X∈R H×W×C By mapping the two-dimensional Fast Fourier Transform (FFT) to the frequency domain, the frequency domain feature representation is obtained: F(X) = FFT2D(X) Where H, W, and C are the height, width, and number of channels of the image, respectively, and F(x) is the frequency domain output; (b) Channel attention weighting: Channel compression of frequency domain features is performed using 1×1 convolution, and a dynamic filtering weight matrix W is generated by the Softmax function. θ : W θ =Softmax(Conv 1×1 (F(X)) (c) Frequency Domain Reconstruction: Perform a two-dimensional inverse Fourier transform (IFFT) on the weighted frequency domain features to obtain the enhanced spatial domain features F. freq : F freq =IFFT2D(W θ ⊙F(X)) Here, ⊙ represents element-wise multiplication.

4. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 3, characterized in that, The processing flow of the deformable convolution modeling module includes: (a) Offset generation: Apply a 3×3 convolution to the input image X to generate the corresponding spatial offset field Δp. n ∈R H×W×C : △p n =Conv 3×3 (X) (b) Deformable Sampling: The sampling position of the standard convolution kernel is adjusted according to the offset to achieve adaptive modeling of local deformable structures and output feature F. def (p0) is defined as follows: Where p0 represents the current convolution center position, R is the receptive field region, and p n At the location of each sampling point within the receptive field R, w n These are the convolution weights.

5. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 4, characterized in that, The specific fusion strategy of the cross-domain dynamic fusion module is as follows: F out =αF freq +(1-α)·F def Wherein, the fusion coefficient α∈[0,1] can be obtained through learning or fixed settings, reflecting the relative contributions of frequency domain and spatial domain features in the fusion process, F out This represents the fusion features of the final output.

6. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 1, characterized in that, The self-attention pyramid pooling (SAPP) strategy in step 3.1 includes multi-scale feature aggregation, context fusion representation construction, spatial attention weight generation, and upsampling / downsampling fusion output, as detailed below: (1) Multi-scale feature aggregation: For the intermediate feature map X∈R output by the encoder H×W×C Spatial pyramid pooling is performed, and different scale parameters k∈{1,3,6,8} are selected to control the size of the pooling grid, generating a multi-scale context representation. Two-dimensional adaptive average pooling (AdaptivePool2d) is used, combined with the image aspect ratio. The pooling region is dynamically adjusted, and the pooling process is expressed as follows: Y k =AdaptivePool2d(X,(k,max(1,[ar×k]))) Where Y k This represents the output feature map after adaptive pooling. (2) Context-fusion expression construction The feature maps Y1, Y2, Y3, and Y4 after pooling at each scale are upsampled to the original resolution H×W using bilinear interpolation, and then stitched together along the channel dimension to form a unified pyramid fusion representation. z=Concat(Flatten(Y1),Flatten(Y2),Flatten(Y3),Flatten(Y4)) Where z represents the final fused feature map, Flatten means flattening each two-dimensional feature map into a one-dimensional channel sequence, and Concat means concatenating the channels. (3) Spatial attention weight generation A lightweight spatial attention mechanism is applied to the fused feature map z, generating a dynamic weighting matrix A through local perception and nonlinear mapping. k : A k =Sigmoid(Conv 1×1 (GeLU(Conv 3×3 (z))) Among them, Conv_{3×3} is used to capture local spatial relationships, GeLU activation function improves nonlinear expression ability, Conv_{1×1} generates spatial weights after channel compression, and finally the attention weights are obtained by normalization through Sigmoid activation function. (4) Feature reconstruction and output fusion The generated spatial weight A k The fused feature map z is element-wise weighted, then upsampled to restore the spatial resolution to match the input image, and this is used as the output Z for the current stage. Z=Upsample(A k ⊙z) Where ⊙ represents element-wise multiplication, and Upsample is the bilinear upsampling function.

7. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 1, characterized in that, The Bipolar Routing Attention (BRA) mechanism in step 3.2 includes region partitioning and projection coding, region affinity calculation, dynamic route index generation, and local token aggregation. The specific process is as follows: (1) Region division and feature projection coding First, input feature map X∈R H×W×C The area is uniformly divided into S×S non-overlapping regions, each region containing HW / S 2 Each token (i.e., a pixel block or embedding block) is used; subsequently, the input feature map is encoded in three ways using a learnable linear projection matrix to generate query Q, key K, and value V vectors respectively: Q=XW q ,K=XW k ,V=XW v Among them W q W k W v ∈R c×c The projection matrix; (2) Regional-level semantic affinity modeling Perform an average pooling operation on all tokens within each region to obtain the aggregated representation of the query and key for each region, denoted as follows: and Where Ω i Represents the set of token indices for the i-th region, |Ω i | represents a set| Ω i The number of elements in |Q j and K j Let represent the query vector and key vector of the j-th token, respectively; Next, the semantic similarity matrix between all regions is calculated. Each element Indicate the semantic relevance between region i and region j: A r =Q r (K r ) T Q r and K r These represent the query matrix and the key matrix, respectively. (3) Generation of regional dynamic routing index Based on the regional similarity matrix A r For each region, the top k most relevant regions are selected as its attention scope to generate a routing index matrix. Where the i-th row Indicates the k most relevant region indices for region i: I r =topkIndex(A r ) Among them, topkIndex(A r ) is a function call used to extract data from matrix A. r Select the indices of the k elements with the highest values; (4) Token-level attention computation and interaction According to route index I r Collect the corresponding keys and values ​​from the entire graph to construct a local key-value cluster representation: K g =gather(K,I r ),Vc g =gather(V,I r ) in That is, each region only participates in token interactions within its relevant region, gather(K,I) r ) and gather(V,I r ) is a function call used to collect data from matrices K and V by I. r The element at the specified position; Finally, for each original query token, attention weights and feature weighted summations are performed only on its corresponding set of aggregate keys to generate the final output:

8. The medical image segmentation method based on Mamba spectral dynamic fusion and dual attention enhancement according to claim 1, characterized in that, The loss function L of the total loss in step 3.3 total A linear combination strategy of binary cross-entropy and Dice loss is adopted, specifically in the form of: THE total =λL BCE +(1-λ)L 1ice Where λ = 0.5, L bCE For binary cross-entropy loss, L Dice This is a loss for Dice.

Citation Information

Cited By

  • Medical image segmentation method and device based on self-supervised reconstruction assistance, and medium

    CN121169962A

  • Breast ultrasonic image segmentation method and system based on decoupling learning

    CN121190505A

  • RGB-T image multi-modal semantic segmentation method based on state space

    CN121305091A

  • Image region-of-interest extraction method and system based on Mama architecture

    CN121353650A

  • A method and system for extracting regions of interest in images based on the Mamba architecture

    CN121353650B