A medical image segmentation method based on a hybrid state space model

CN122820741APending Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610876840.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明提供一种基于混合状态空间模型的医学图像分割方法,旨在解决现有三维医学图像分割方法中,基于卷积神经网络的模型难以建模长程依赖导致大尺度解剖结构分割精度受限,而基于Transformer的模型计算复杂度和内存开销巨大难以处理高分辨率三维数据,从而无法兼顾模型高效性与分割高精度的技术问题

Benefits of technology

1、提升了分割精度与计算效率的平衡性:通过构建深度融合三维卷积局部建模优势与视觉Mamba架构全局建模能力的混合编码器,从根本上解决了传统卷积网络感受野有限和Transformer类模型计算开销大的矛盾。实验表明,在公开数据集上,本方法在分割精度上显著优于nnU-Net、Swin-UNETR等先进方法,同时参数量和计算量(GFLOPs)大幅降低,实现了高精度与高效率的协同优化,更利于临床实际部署。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820741A_ABST
    Figure CN122820741A_ABST
Patent Text Reader

Abstract

The application discloses a medical image segmentation method based on a hybrid state space model, which comprises the following steps: performing initial convolution feature extraction on an original three-dimensional medical image to obtain initial shallow layer features; inputting the initial shallow layer features into an encoder path to perform multi-level down-sampling and feature abstraction, and constructing a multi-scale feature pyramid; inputting the deepest layer features output by the encoder path into a bottleneck layer, performing deep feature transformation and refinement through stacked enhanced residual visual state space blocks; performing step-by-step up-sampling and feature fusion on the features output by the bottleneck layer through a decoder path to restore the spatial resolution; and performing convolution and normalization processing on the features finally output by the decoder path to generate a final three-dimensional medical image segmentation probability map. The application effectively breaks through the limitations of traditional methods by systematically integrating the local perception advantages of three-dimensional convolution and the global linear complexity sequence modeling advantages of a state space model in a neural network architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and medical image processing, specifically to a three-dimensional medical image segmentation method based on a hybrid state space model, which is particularly suitable for automated and high-precision segmentation of organs and lesions in three-dimensional medical images such as computed tomography (CT) and magnetic resonance imaging (MRI). Background Technology

[0002] Revolutionary advancements in medical imaging technology have provided unprecedented visualization capabilities for clinical diagnosis. Three-dimensional imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) can generate three-dimensional volumetric data containing hundreds or even thousands of slices at sub-millimeter resolution. This high-dimensional data plays an irreplaceable role in critical medical processes such as disease screening and surgical planning. Medical image segmentation, as a core task of medical image analysis, aims to assign each voxel in an image to a specific anatomical structure or lesion region, and is a prerequisite for subsequent quantitative analysis.

[0003] Three-dimensional medical image data is massive; a typical CT scan may contain 512×512×300 voxels. Processing such large-scale data requires efficient algorithm design. While traditional 3D segmentation models based on convolutional neural networks (CNNs) can effectively extract local features, their limited receptive field makes it difficult to model long-range dependencies, limiting their accuracy when dealing with large-scale anatomical structures or scenarios requiring global contextual understanding. Furthermore, the computational cost of 3D convolution increases cubically with the data size, resulting in a huge number of parameters and high computational resource requirements.

[0004] Existing technologies struggle to achieve a good balance between model efficiency and segmentation accuracy when handling 3D medical image segmentation tasks, or their performance is limited under conditions of scarce labeled data. There is a lack of a unified solution that can systematically balance high accuracy, high efficiency, and low label dependency. Therefore, an innovative approach is urgently needed to address these challenges.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] This invention provides a medical image segmentation method based on a hybrid state-space model, aiming to solve the technical problem in existing 3D medical image segmentation methods that, due to the difficulty of modeling long-range dependencies by convolutional neural network-based models, the segmentation accuracy of large-scale anatomical structures is limited, while Transformer-based models have huge computational complexity and memory overhead, making it difficult to handle high-resolution 3D data, thus failing to balance model efficiency and high segmentation accuracy.

[0007] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0008] According to a first aspect of the present invention, a medical image segmentation method based on a hybrid state-space model is provided, the method comprising: Initial convolutional feature extraction is performed on the original 3D medical image to obtain initial shallow features; The initial shallow features are input into the encoder path for multi-level downsampling and feature abstraction to construct a multi-scale feature pyramid. The encoder path contains multiple cascaded downsampling stages, and the core of each downsampling stage is a residual visual state space block. The residual visual state space block is deeply fused with a deep separable convolutional operation and a selective state space model to extract local spatial features in parallel and model long program sequence dependencies. The deepest features output from the encoder path are input to the bottleneck layer, and the deep features are transformed and refined through stacked enhanced residual visual state space blocks; wherein, the enhanced residual visual state space block includes grouped learnable nonlinear enhancement modules, which are used to apply adaptive nonlinear transformations to the input features; The features output by the bottleneck layer are progressively upsampled and fused through the decoder path to restore spatial resolution. The decoder path includes multiple upsampling stages corresponding to the encoder stages, and the core of each upsampling stage is a lightweight and efficient decoding module. The lightweight and efficient decoding module optimizes the upsampled features through depthwise separable convolution and fuses detailed features from the same level of the encoder through adaptive gated skip connections. The features output by the decoder path are convolved and normalized to generate the final three-dimensional medical image segmentation probability map.

[0009] In some exemplary embodiments, the residual visual state space block includes: The local feature extraction branch is used to extract local spatial patterns and details of the input features through depthwise separable convolution operations; The global sequence modeling branch is used to expand the input features into a sequence through a multi-directional scanning strategy, and to model the long-range dependencies in the sequence with linear complexity using a selective state-space model, and then reassemble them into a feature map; The feature fusion unit is used to adaptively weight and fuse the local features output by the local feature extraction branch with the global features output by the global sequence modeling branch to output enhanced features.

[0010] In some exemplary embodiments, the lightweight and efficient decoding module includes: The upsampling unit is used to perform trilinear interpolation upsampling on the input low-resolution features to restore them to the target spatial size; A lightweight convolutional optimization unit is used to perform nonlinear optimization on the upsampled features through cascaded depthwise separable convolutional layers to obtain optimized decoder features; An adaptive gating fusion unit is used to automatically learn a spatial dimension gating weight map based on the features of the optimized decoder and the detailed features from the same level of the encoder, and to use the weight map to perform weighted modulation on the encoder detailed features. The feature addition unit is used to add the weighted modulated encoder detail features to the optimized decoder features element by element to achieve initial fusion; The residual enhancement unit is used to apply a residual connection consisting of depthwise separable convolutions to the features after the initial fusion, to perform final feature refinement and enhancement, and to output the result of this decoding stage.

[0011] In some exemplary embodiments, the group-learnable nonlinear enhancement module includes: The channel normalization and projection submodule is used to perform layer normalization on the input features and map the number of channels to the hidden dimension through linear projection. The grouping and intra-group transformation submodule is used to uniformly divide the projected features into multiple groups along the channel dimension, and independently apply an intra-group learnable nonlinear transformation layer to each group; wherein, each transformation layer performs nonlinear mapping on the intra-group features through multiple learnable univariate spline functions; The inter-group merging submodule is used to merge the transformed features of each group along the channel dimension. The linear projection and residual submodule is used to activate and linearly project the merged features to restore the original number of output channels, and then residually connect the projection results with the original input of the module to output the final enhanced features.

[0012] In some exemplary embodiments, the multi-directional scanning strategy in the global sequence modeling branch includes: horizontal scanning, vertical scanning, main diagonal scanning, and sub-diagonal scanning; wherein, each scanning method unfolds the three-dimensional feature map into a one-dimensional sequence along a specific direction.

[0013] In some exemplary embodiments, the feature fusion unit achieves adaptive weighted fusion through the following steps: concatenating the local features output by the local feature extraction branch with the global features output by the global sequence modeling branch in the channel dimension; processing the concatenated features through a light quantum network consisting of two 1×1×1 convolutional layers to generate a gated weight map in the spatial dimension; and using the gated weight map to dynamically weight and fuse the local features and the global features to obtain the output features.

[0014] In some exemplary embodiments, initial convolutional feature extraction of the original three-dimensional medical image involves using a three-dimensional convolutional kernel to perform spatial filtering and feature upscaling on the original three-dimensional medical image to generate an initial shallow feature map with a fixed number of basic channels.

[0015] According to a second aspect of the present invention, a medical image segmentation apparatus based on a hybrid state-space model is provided, comprising: The initial feature extraction module is used to perform initial convolutional feature extraction on the original 3D medical image to obtain initial shallow features; An encoder is used to perform multi-level downsampling and feature abstraction on the initial shallow features to construct a multi-scale feature pyramid; wherein, the encoder contains multiple cascaded residual visual state space blocks; The bottleneck layer is used to receive the deepest features output by the encoder and perform deep feature transformation and refinement through stacked enhanced residual visual state space blocks; wherein, the enhanced residual visual state space blocks include grouped learnable nonlinear enhancement modules. The decoder is used to progressively upsample and fuse multi-scale features output by the bottleneck layer to restore spatial resolution; wherein, the decoder includes multiple lightweight and efficient decoding modules; the segmentation head is used to perform convolution and normalization processing on the features finally output by the decoder to generate the final three-dimensional medical image segmentation probability map.

[0016] According to a third aspect of the present invention, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the medical image segmentation method based on a hybrid state-space model as described in the first aspect.

[0017] According to a fourth aspect of the present invention, a computer program product is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the medical image segmentation method based on a hybrid state-space model as described in the first aspect above.

[0018] According to a fifth aspect of the present invention, an electronic device is provided, comprising: Processor; and Memory for storing the executable instructions of the processor; The processor is configured to implement the medical image segmentation method based on a hybrid state-space model as described in the first aspect by executing the executable instructions.

[0019] Compared with the prior art, the beneficial effects of this application are as follows: 1. Improved balance between segmentation accuracy and computational efficiency: By constructing a hybrid encoder that deeply integrates the advantages of 3D convolutional local modeling with the global modeling capabilities of the visual Mamba architecture, the contradiction between the limited receptive field of traditional convolutional networks and the high computational cost of Transformer-like models is fundamentally resolved. Experiments show that on public datasets, this method significantly outperforms advanced methods such as nnU-Net and Swin-UNETR in segmentation accuracy, while greatly reducing the number of parameters and computational cost (GFLOPs), achieving a synergistic optimization of high accuracy and high efficiency, which is more conducive to practical clinical deployment.

[0020] 2. Enhanced modeling capabilities for multi-scale anatomical structures and complex boundaries: By introducing a local-global dual-path context-aware module, local details and global semantics are processed explicitly and in parallel, and dynamic fusion is achieved using adaptive gating. This enables the model to accurately depict multi-scale information from fine textures to the topology of the entire organ, as well as complex anatomical structures with blurred boundaries and varied morphologies. As a result, it demonstrates excellent robustness and accuracy in challenging tasks such as liver tumors and thin-walled myocardium.

[0021] 3. Achieved lighter and smarter feature decoding and reconstruction: Through the design of a lightweight and efficient decoding module, the parameter-dense transposed convolution is abandoned. Instead, a depthwise separable convolution and adaptive gating fusion mechanism are adopted. While significantly reducing the computational complexity of the decoding path, it can intelligently filter and fuse multi-scale features from the encoder, effectively ensuring the quality of detail reconstruction and avoiding the decoder becoming a bottleneck for model efficiency.

[0022] 4. Provides powerful adaptive nonlinear fitting capabilities: By introducing a grouped learnable nonlinear enhancement module, the traditional fixed activation function is replaced by a grouped computation paradigm and a learnable spline function based on the Kolmogorov-Arnold representation theorem. This gives the model the ability to automatically learn the best nonlinear transformation based on the data, which greatly improves the network's fitting accuracy for nonlinearities such as complex gray-level distributions and irregular tissue boundaries in medical images, and further ensures the accuracy of the segmentation results.

[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0025] Figure 1 This is a schematic diagram of the overall network architecture and information flow of the method described in this invention; Figure 2 This is a flowchart illustrating the specific architecture and data processing of the lightweight and efficient decoding module in this invention. Figure 3 This is a diagram illustrating the parallel processing architecture of the local-global dual-path context awareness module in this invention. Figure 4 This is a diagram showing the internal structure of the grouped learnable nonlinear enhancement module in this invention. Figure 5 This is a comparison of the visualization results of the method of the present invention on the LiTS liver tumor segmentation dataset; Figure 6 This is a comparison chart showing the visualization results of the method of the present invention on the ACDC heart segmentation dataset. Detailed Implementation

[0026] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the invention will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0027] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0028] This application provides a medical image segmentation method based on a hybrid state-space model. The method includes: extracting initial convolutional features from the original 3D medical image to obtain initial shallow features; performing multi-level downsampling and feature abstraction on the initial shallow features through an encoder path to construct a multi-scale feature pyramid, wherein each downsampling stage of the encoder includes a residual visual state-space block, which achieves synergistic enhancement of local perception and global semantics by deeply fusing the local feature extraction capability of depthwise separable convolution with the selective scanning long-range dependency modeling capability of the visual state-space module; inputting the deepest features output by the encoder to a bottleneck layer, and performing depth feature transformation and refinement through stacked enhanced residual visual state-space blocks, wherein the feedforward network in the enhanced residual visual state-space blocks is replaced by... The module is replaced with a grouped learnable nonlinear enhancement module. This module provides adaptive, highly expressive nonlinear transformations through grouping strategies and learnable spline functions to accurately fit complex tissue boundaries. The refined features output from the bottleneck layer are progressively upsampled and multi-scale feature fused through the decoder path to restore spatial resolution. Each decoding stage corresponds to an encoder stage, with a lightweight and efficient decoding module at its core. This module first upsamples the low-resolution input features, then optimizes the features through depthwise separable convolutional blocks, and finally dynamically fuses spatially detailed features from the same encoder level through adaptive gated skip connections to achieve intelligent detail reconstruction. The high-resolution features output from the decoder path are convolved and normalized to generate the final 3D medical image segmentation probability map.

[0029] The residual visual state space block includes: a local feature extraction branch, used to extract local spatial patterns and details of the input features through depthwise separable convolution operations; a global sequence modeling branch, used to unfold the input features into a sequence through a multi-directional scanning strategy, and to model the long-range dependencies in the sequence using a selective state space model with linear complexity, and then reassemble them into a feature map; and a feature fusion unit, used to adaptively weight and fuse the local features output by the local feature extraction branch and the global features output by the global sequence modeling branch, and output an enhanced feature that has both local accuracy and global consistency.

[0030] The lightweight and efficient decoding module includes: an upsampling unit for performing trilinear interpolation upsampling on the input low-resolution features to restore them to the target spatial size; a lightweight convolution optimization unit for performing nonlinear optimization on the upsampled features through cascaded depthwise separable convolutional layers to eliminate aliasing and enhance feature consistency, resulting in optimized decoder features; an adaptive gated fusion unit for automatically learning a spatial dimension gate weight map based on the optimized decoder features and detail features from the same level of the encoder, and using this weight map to perform weighted modulation on the encoder detail features; a feature addition unit for element-wise addition of the weighted modulated encoder detail features and the optimized decoder features to achieve initial fusion; and a residual enhancement unit for applying a residual connection consisting of a depthwise separable convolution and a learnable scale parameter to the features after initial fusion, performing final feature refinement and enhancement, and outputting the result of this decoding stage.

[0031] The grouped learnable nonlinear enhancement module includes: a channel normalization and projection submodule, used to perform layer normalization on the input features and map the number of channels to the hidden dimension through linear projection; a grouping and intra-group transformation submodule, used to uniformly divide the projected features into multiple groups along the channel dimension and independently apply an intra-group learnable nonlinear transformation layer to each group, wherein each transformation layer performs highly expressive nonlinear mapping on the intra-group features through multiple learnable univariate spline functions; an inter-group merging submodule, used to merge the transformed features of each group along the channel dimension; and a linear projection and residual submodule, used to perform linear projection on the merged features to restore the original number of output channels and perform residual connection between the projection result and the original input of the module to output the final enhanced features.

[0032] The first embodiment of the present invention provides a medical image segmentation method based on a hybrid state-space model, the overall network architecture and information flow of which are as follows: Figure 1 As shown. Specifically, the method includes the following steps: Step S1 involves performing initial convolutional feature extraction on the original 3D medical image to obtain initial shallow features. Specifically, given an original 3D medical image (such as a CT or MRI scan). (Taking a single channel as an example), through a three-dimensional convolutional kernel (e.g., with a size of...) Initial feature extraction is performed. This convolutional operation performs preliminary spatial filtering and feature dimensionality upscaling on the original voxel data, mapping the input to a shallow feature space with a fixed base number of channels (e.g., 32), generating an initial shallow feature map. This step aims to efficiently capture the most basic low-level information in the image, such as edges and textures, laying the foundation for subsequent deep feature abstraction.

[0033] Step S2 involves multi-level downsampling and feature abstraction of the initial shallow features through the encoder path to construct a multi-scale feature pyramid. Each downsampling stage of the encoder contains a residual visual state space block. The encoder path consists of four cascaded downsampling stages. Each stage contains two cascaded residual visual state space blocks and performs a downsampling operation (e.g., through convolution or pooling with a stride of 2) at the end of the stage. As the network depth increases, the spatial size of the feature map (depth) increases. ,high ,width The number of channels is halved at each level, while the number of channels is doubled accordingly (e.g., from 32 to 64, 128, 256), thus constructing a hierarchical multi-scale feature pyramid. The residual visual state space block is the core component of the encoder of this invention, and its specific architecture is as follows: Figure 3 As shown, it is used to deeply integrate the local convolutional feature extraction capability with the global state space sequence modeling capability.

[0034] The residual visual state space block specifically includes the following sub-steps: Step S21, Local Feature Extraction Branch: Used to extract local spatial patterns and details of the input features through depthwise separable convolution operations. Given input features... First, it uses 3D depthwise separable convolution for processing. Depthwise separable convolution decomposes standard 3D convolution into depthwise convolution and pointwise convolution. Depthwise convolution uses... Three-dimensional convolutional kernels (e.g., size 1) Perform independent spatial convolutions on each channel of the input to generate intermediate features. :

[0035] Then pointwise convolution uses a convolution kernel right Perform channel mixing to output local features as follows

[0036] Depthwise separable convolutions significantly reduce the number of parameters and computational complexity while maintaining the ability to extract local features.

[0037] Step S22, Global Sequence Modeling Branch: Used to model long-range dependencies of input features. First, a multi-directional scanning strategy is used to model the three-dimensional feature map. Expanded into a sequence. Specifically, this includes: horizontal scanning. (fixed depth) and height Along the width Traversal), Vertical Scan (fixed depth) and width Along height Traversal), main diagonal scan and sub-diagonal scan Each scan transforms the feature map into a one-dimensional sequence. ( Then, each sequence is input into a Selective StateSpace Model (SSM), i.e., the Mamba module, for processing. The core of the Mamba module is the selective scan mechanism, whose discrete-time statespace equation is:

[0038] Where the coefficient parameter and implicit time step The current input The function implements input-dependent, content-aware sequence modeling. This operation has linear computational complexity. The processed sequence... Through reverse scanning operation Reconstructing into a 3D feature map Finally, the feature maps obtained from the four directions are concatenated along the channel dimension and then processed through a... Convolutional processes are used to fuse the data and obtain global features. :

[0039] Step S23, Feature Fusion Unit: Used to process the local features output by the local feature extraction branch. The global features output by the global sequence modeling branch Perform adaptive weighted fusion. First, [the following steps are taken]. and Concatenate along the channel dimension, then pass through a pair of... The lightweight quantum network, consisting of convolutional layers (containing ReLU activations in between), finally generates a spatially dimensional gate weight using a sigmoid function. Final output features Generated by dynamic weighted fusion of gated weights:

[0040] in This indicates element-wise multiplication. This allows the model to autonomously decide on the signal-to-noise ratio in regions with clear boundaries and to rely on global semantics in structurally complex regions.

[0041] Step S3: Input the deepest features output by the encoder into the bottleneck layer, and perform deep feature transformation and refinement through stacked enhanced residual visual state space blocks. In this step, the feedforward network in the enhanced residual visual state space block is replaced by a grouped learnable nonlinear enhancement module.

[0042] The bottleneck layer receives features (e.g., size) output from the fourth stage of the encoder. This layer consists of multiple (e.g., four) stacked enhanced residual visual state space blocks. Compared to the residual visual state space blocks in step S2, the enhancement lies in replacing the standard linear feedforward network (FFN) with a Grouped Learnable Nonlinearity Enhancement Module (GNT) to provide a more powerful adaptive nonlinear transformation unit. The internal structure of this module is as follows... Figure 4 As shown, the specific steps include: Step S31, Channel Normalization and Projection Submodule: For input features Perform layer normalization (LayerNorm): Then through a The linear projection layer will increase the number of channels from Mapped to a larger hidden dimension : ,

[0043] Step S32, Grouping and Component Transformation Submodule: Transform the projected features Divided evenly along the channel dimension into A number of non-overlapping groups, i.e. ,in , For each set of features Each unit independently applies a learnable nonlinear transformation unit within the group. This transformation unit is based on the Kolmogorov-Arnold representation theorem and uses a learnable family of fringe functions. As an activation unit. For output channels within the group. Its value is determined by all input channels within the group. The sum of the contributions yields:

[0044] Each spline function A group spline basis functions and their corresponding learnable coefficients Parameterization: This allows the network to adaptively learn the optimal nonlinear mapping.

[0045] Step S33, Combining Sub-module: Combining the transformed features of each group Re-merge along the channel dimension to obtain .

[0046] Step S34, Linear Projection and Residual Submodule: For the merged features Applying an activation function (such as GELU) yields And then through a The linear projection layer maps the number of channels back to the original number of output channels. ; Finally, the module output is connected via residual connections. With the original input Add them together to get the final output: This module significantly improves the network's fitting accuracy for complex grayscale distributions and irregular boundaries in medical images.

[0047] Step S4: The refined features output by the bottleneck layer are progressively upsampled and fused with multi-scale features through the decoder path to restore spatial resolution. Each decoding stage corresponds to an encoder stage, the core of which is a lightweight and efficient decoding module.

[0048] The decoder path comprises four upsampling stages corresponding to the encoder. The core of each stage is a lightweight and efficient decoding module, the specific architecture of which is shown in the figure. For a specific decoding stage... (corresponding encoder number) This module receives two inputs: 1) from a deeper layer (the first layer). Low-resolution features output by the layer decoding module ;2) Detailed features from the same-level encoder passed through skip connections. The specific steps for this module are as follows: Step S41, Upsampling unit: for low-resolution features of the input Perform trilinear interpolation upsampling to restore its spatial dimensions to the same level as... same scale ,get No parameters need to be learned for this operation.

[0049] Step S42, Lightweight Convolutional Optimization Unit: Use a lightweight depthwise separable convolutional block pair Feature optimization is performed. This convolutional block consists of two cascaded depthwise separable convolutions, each followed by a ReLU activation function. Depthwise separable convolutions have a much lower computational complexity than standard convolutions, eliminating upsampling aliasing and enhancing feature consistency with minimal parameter cost, outputting optimized decoder features. .

[0050] Step S43, Adaptive Gating Fusion Unit: First, the encoder detailed features are... With optimized decoder features Perform channel-dimensional splicing. Then, through a light quantum network (containing two...) The convolutional layer (with ReLU activation in the middle) processes the concatenated features, and finally a spatially gated weight is generated by applying a Sigmoid threshold. This weight represents the importance of the encoder's detailed features in the current decoding context. Utilizing... Dynamically modulate the encoder's detailed features: .

[0051] Step S44, Feature Addition Unit: Add the hardened encoder features With optimized decoder features The initial fusion is achieved by adding elements one by one. .

[0052] Step S45, Residual Enhancement Unit: For the initial fused features A residual connection is applied for final refinement. This residual connection consists of a depthwise separable convolutional layer and a learnable scaling factor. Composed of ReLU activation. Output characteristics The calculation can be expressed as: This module significantly reduces the computational complexity of the decoding path while achieving intelligent detailed reconstruction through adaptive gating.

[0053] Step S5 involves convolution and normalization of the high-resolution features output from the decoder path to generate the final 3D medical image segmentation probability map. After progressive upsampling and feature fusion through four decoding stages, a high-dimensional feature map with the same spatial resolution as the input image is obtained. Finally, a... The convolutional layer maps the number of channels to the number of target categories. (For example, in liver tumor segmentation) (Corresponding to background, liver and gallbladder, tumor), then normalization is performed using the Softmax function to generate a probability map of each voxel belonging to each category. This allows for the segmentation of three-dimensional medical images.

[0054] In this embodiment, a hybrid encoder (residual visual state space block) that deeply integrates the advantages of local modeling in 3D convolution with the global modeling capabilities of the visual state space model (Mamba) fundamentally solves the contradiction between the limited receptive field of traditional convolutional networks and the high computational cost of Transformer-like models. Simultaneously, the designed lightweight and efficient decoder abandons parameter-dense transposed convolutions and adopts a depthwise separable convolution and adaptive gating fusion mechanism. This significantly reduces the computational complexity of the decoding path while intelligently selecting and fusing multi-scale features from the encoder, ensuring the quality of detailed reconstruction. Furthermore, the introduced grouped learnable nonlinear enhancement module provides powerful adaptive nonlinear fitting capabilities, further guaranteeing the accuracy of segmenting complex tissue boundaries.

[0055] Example 1: Network Training and Testing To verify the effectiveness of the method described in this invention, we conducted detailed network training and testing experiments in this embodiment. This embodiment selected two internationally recognized, publicly available 3D medical image segmentation datasets with different levels of challenge for experimental verification.

[0056] 1. LITS (Liver and Liver Tumor Segmentation) Dataset: This dataset originates from the MICCAI 2017 challenge and contains 201 enhanced abdominal CT scans. Each image has between 42 and 1026 slices, with most images having an in-plane size of 512×512 pixels. The dataset features anisotropic resolution and includes manually drawn annotations of the liver and liver tumors by experts. The challenge of this dataset lies in simultaneously and accurately segmenting the liver (a large organ) with vastly different scales and potentially small tumors within it, given the diverse tumor morphology, contrast, and complex backgrounds.

[0057] 2. ACDC (Automated Cardiac Diagnosis Challenge) Dataset: This dataset contains short-axis Cine-MRI sequences from 100 patients. Each patient is provided with images at two time points: end-diastolic (ED) and end-systolic (ES), each typically containing 6 to 10 slices. Target structures for segmentation include the left ventricular (LV), right ventricular (RV), and myocardium of both ventricles. The challenge of this dataset lies in segmenting thin-walled, dynamic, and contrast-varying cardiac structures, particularly the irregular shape of the right ventricle and the thinness of the myocardium. (2) Experimental environment Table 1 Experimental Environment Configuration Information

[0058] All experiments were conducted on a unified hardware and software platform to ensure the fairness and reproducibility of the results. Open-source libraries such as NumPy, NiBabel, SimpleITK, and scikit-learn were used. Lesions in each dataset were randomly divided into training, validation, and test sets in a 7:2:1 ratio. Intensity normalization (window truncation followed by normalization for CT, and z-score normalization for MRI) and isotropic resampling were performed on the input 3D datasets. During training, online data augmentation was applied using random rotation, random scaling, random elastic deformation, and random gamma correction. The network used the AdamW optimizer with an initial learning rate of 1e-3 and a cosine annealing learning rate scheduling strategy. The batch size was set to 2. The loss function was a composite loss of Dice loss and weighted cross-entropy loss with a weight ratio of 0.5:0.5. The total number of training epochs was set to 300, and an early stopping strategy (patience value of 50 epochs) was used on the validation set. The weight decay coefficient was set to 1e-5.

[0059] (3) Experimental results The Dice similarity coefficient (DSC) and 95% Hausdorff distance (HD95) were used as the main metrics for evaluating segmentation accuracy. The number of parameters and the number of floating-point operations (GFLOPs) were also reported to measure the model complexity.

[0060] On the LiTS dataset (Table 2): the method of this invention achieves the best results in liver segmentation DSC (96.25%), tumor segmentation DSC (72.52%), and average DSC (83.99%), while having the lowest average HD95 (10.87 mm). In terms of efficiency, the number of parameters (71.5M) and computational cost (1150.40 GFLOPs) of this invention are significantly lower than those of nnU-Net (88.62M, 7240.26 GFLOPs) and U-Mamba (173.53M, 4057.20 GFLOPs), achieving an excellent balance between accuracy and efficiency.

[0061] On the ACDC dataset (Table 3): the method of this invention achieves an average DSC of 92.85% and an HD95 of 1.05mm, which is also better than all the comparison methods, verifying its ability to finely segment thin-walled, dynamic heart structures, while maintaining leading computational efficiency.

[0062] Figure 5 and Figure 6 The presentation shows a visual comparison of the segmentation results of the method of this invention on the LiTS and ACDC test sets with those of nnU-Net, Swin-UNETR, and U-Mamba.

[0063] LiTS results ( Figure 5 For the liver, the method of this invention provides the closest and most continuous alignment between the predicted boundary and the ground truth (GT), accurately distinguishing the liver from adjacent organs. For small intrahepatic tumors, contrastive methods suffer from undersegmentation or coarse boundaries, while the method of this invention fully delineates the tumor outline and better preserves its internal fine structures.

[0064] ACDC results ( Figure 6 For thin-layered myocardium, the contours segmented by the method of this invention are the smoothest, most continuous, and have the most uniform thickness. When distinguishing between the left ventricular cavity, myocardium, and right ventricular cavity, the results of this invention show clear boundaries and accurate morphology, significantly outperforming the comparative methods. Visualization results intuitively demonstrate the comprehensive advantages of the method of this invention in addressing challenges such as multi-scale, complex boundaries, and thin-walled structures.

[0065] Table 2 Performance comparison of different methods on the LiTS dataset

[0066] Table 3 Performance comparison of different methods on the ACDC dataset

[0067] This application discloses a medical image segmentation method based on a hybrid state-space model. It achieves efficient and high-precision automatic segmentation of 3D medical images by constructing a hybrid encoder-decoder network architecture that deeply integrates the advantages of local convolutional inductive bias and global state-space sequence modeling. Specifically, the method includes the following steps: First, initial 3D convolutional feature extraction is performed on the original 3D medical image (such as CT or MRI volume data) to obtain an initial shallow feature map containing basic edge and texture information. Second, the initial shallow features are input into the encoder path for multi-level downsampling and feature abstraction, gradually constructing a multi-scale feature pyramid. The core of the encoder path consists of multiple cascaded residual visual state-space blocks. Within each block, the local spatial feature extraction capability of deep separable convolution and the selective scanning long-range dependency modeling capability of the visual state-space model (such as the Mamba architecture) are deeply fused through parallel branches to achieve synergistic enhancement of local details and global semantics. Next, the deepest and most semantically abstract features output by the encoder are fed into the bottleneck layer. Deep feature transformation and refinement are performed through stacked enhanced residual visual state space blocks. Innovatively, these enhanced blocks employ a grouped learnable nonlinear enhancement module instead of a traditional feedforward network. This module, based on grouped computation and learnable spline functions, provides adaptive, highly expressive nonlinear transformations to accurately fit the complex tissue boundaries and grayscale distributions in medical images. Then, the refined features output from the bottleneck layer are progressively upsampled and multi-scale feature fusioned through the decoder path to restore high spatial resolution. The core of this decoder path is a lightweight and efficient decoding module. This module abandons parameter-dense transposed convolutions and adopts lightweight optimization based on depthwise separable convolutions and an intelligent feature fusion mechanism based on adaptive gated skip connections. This significantly reduces the computational complexity of the decoding path while ensuring dynamic and selective reconstruction of spatial details from features at the same level as the encoder. Finally, the high-resolution features output by the decoder are convolved and normalized to generate a probability map of each voxel belonging to different anatomical structures or lesion categories, which is the final 3D medical image segmentation result. This method effectively overcomes the limitations of traditional methods by systematically integrating the local perceptual advantages of 3D convolution with the global linear complexity sequence modeling advantages of state-space models within a neural network architecture. Experiments on publicly available authoritative datasets demonstrate that, compared to the classic method nnU-Net, the advanced Transformer method Swin-UNETR, and U-Mamba, this method achieves significantly better segmentation accuracy on liver tumor segmentation (LiTS dataset) and heart structure segmentation (ACDC dataset) tasks, in key metrics such as Dice similarity coefficient and 95% Hausdorff distance. Furthermore, thanks to the lightweight decoder design and efficient state-space modeling, the number of parameters and computational cost (GFLOPs) are effectively controlled, achieving an excellent balance between high accuracy and high efficiency.This method is particularly suitable for automated and high-precision segmentation of organs and lesions with complex morphology, multi-scale features and blurred boundaries in three-dimensional medical images such as CT and MRI, providing reliable technical support for clinical diagnosis, surgical planning and other applications.

[0068] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.

[0069] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is defined only by the appended claims.

Claims

1. A medical image segmentation method based on a hybrid state-space model, characterized in that, The method includes: Initial convolutional feature extraction is performed on the original 3D medical image to obtain initial shallow features; The initial shallow features are input into the encoder path for multi-level downsampling and feature abstraction to construct a multi-scale feature pyramid. The encoder path contains multiple cascaded downsampling stages, and the core of each downsampling stage is a residual visual state space block. The residual visual state space block is deeply fused with a deep separable convolutional operation and a selective state space model to extract local spatial features in parallel and model long program sequence dependencies. The deepest features output from the encoder path are input to the bottleneck layer, and the deep features are transformed and refined through stacked enhanced residual visual state space blocks; wherein, the enhanced residual visual state space block includes grouped learnable nonlinear enhancement modules, which are used to apply adaptive nonlinear transformations to the input features; The features output by the bottleneck layer are progressively upsampled and fused through the decoder path to restore spatial resolution. The decoder path includes multiple upsampling stages corresponding to the encoder stages, and the core of each upsampling stage is a lightweight and efficient decoding module. The lightweight and efficient decoding module optimizes the upsampled features through depthwise separable convolution and fuses detailed features from the same level of the encoder through adaptive gated skip connections. The features output by the decoder path are convolved and normalized to generate the final three-dimensional medical image segmentation probability map.

2. The medical image segmentation method based on a hybrid state-space model according to claim 1, characterized in that, The residual visual state space block includes: The local feature extraction branch is used to extract local spatial patterns and details of the input features through depthwise separable convolution operations; The global sequence modeling branch is used to expand the input features into a sequence through a multi-directional scanning strategy, and to model the long-range dependencies in the sequence with linear complexity using a selective state-space model, and then reassemble them into a feature map; The feature fusion unit is used to adaptively weight and fuse the local features output by the local feature extraction branch with the global features output by the global sequence modeling branch to output enhanced features.

3. The medical image segmentation method based on a hybrid state-space model according to claim 1, characterized in that, The lightweight and efficient decoding module includes: The upsampling unit is used to perform trilinear interpolation upsampling on the input low-resolution features to restore them to the target spatial size; A lightweight convolutional optimization unit is used to perform nonlinear optimization on the upsampled features through cascaded depthwise separable convolutional layers to obtain optimized decoder features; An adaptive gating fusion unit is used to automatically learn a spatial dimension gating weight map based on the features of the optimized decoder and the detailed features from the same level of the encoder, and to use the weight map to perform weighted modulation on the encoder detailed features. The feature addition unit is used to add the weighted modulated encoder detail features to the optimized decoder features element by element to achieve initial fusion; The residual enhancement unit is used to apply a residual connection consisting of depthwise separable convolutions to the features after the initial fusion, to perform final feature refinement and enhancement, and to output the result of this decoding stage.

4. The medical image segmentation method based on a hybrid state-space model according to claim 1, characterized in that, The group-learnable nonlinear enhancement module includes: The channel normalization and projection submodule is used to perform layer normalization on the input features and map the number of channels to the hidden dimension through linear projection. The grouping and intra-group transformation submodule is used to uniformly divide the projected features into multiple groups along the channel dimension, and independently apply an intra-group learnable nonlinear transformation layer to each group; wherein, each transformation layer performs nonlinear mapping on the intra-group features through multiple learnable univariate spline functions; The inter-group merging submodule is used to merge the transformed features of each group along the channel dimension. The linear projection and residual submodule is used to activate and linearly project the merged features to restore the original number of output channels, and then residually connect the projection results with the original input of the module to output the final enhanced features.

5. The medical image segmentation method based on a hybrid state-space model according to claim 2, characterized in that, The multi-directional scanning strategy in the global sequence modeling branch includes: horizontal scanning, vertical scanning, main diagonal scanning, and sub-diagonal scanning; wherein, each scanning method unfolds the three-dimensional feature map into a one-dimensional sequence along a specific direction.

6. The medical image segmentation method based on a hybrid state-space model according to claim 2, characterized in that, The feature fusion unit achieves adaptive weighted fusion through the following steps: concatenating the local features output by the local feature extraction branch with the global features output by the global sequence modeling branch in the channel dimension; processing the concatenated features through a light quantum network consisting of two 1×1×1 convolutional layers to generate a gated weight map in the spatial dimension; and using the gated weight map to dynamically weight and fuse the local features and the global features to obtain the output features.

7. The medical image segmentation method based on a hybrid state-space model according to claim 1, characterized in that, Initial convolutional feature extraction of the original 3D medical image involves using a 3D convolution kernel to perform spatial filtering and feature dimensionality upscaling on the original 3D medical image, generating an initial shallow feature map with a fixed number of basic channels.

8. A medical image segmentation device based on a hybrid state-space model, characterized in that, include: The initial feature extraction module is used to perform initial convolutional feature extraction on the original 3D medical image to obtain initial shallow features; An encoder is used to perform multi-level downsampling and feature abstraction on the initial shallow features to construct a multi-scale feature pyramid; wherein, the encoder contains multiple cascaded residual visual state space blocks; The bottleneck layer is used to receive the deepest features output by the encoder and perform deep feature transformation and refinement through stacked enhanced residual visual state space blocks; wherein, the enhanced residual visual state space blocks include grouped learnable nonlinear enhancement modules. The decoder is used to progressively upsample and fuse multi-scale features output by the bottleneck layer to restore spatial resolution; wherein, the decoder includes multiple lightweight and efficient decoding modules; the segmentation head is used to perform convolution and normalization processing on the features finally output by the decoder to generate the final three-dimensional medical image segmentation probability map.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the medical image segmentation method based on the hybrid state-space model as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the medical image segmentation method based on the hybrid state-space model as described in any one of claims 1 to 7.