Method for improving medical image segmentation based on state space model network

Through the U-shaped network algorithm combined with the attention mechanism, the practicality of the existing medical image segmentation method in low contrast and boundary blur processing is solved, the segmentation accuracy and computing efficiency are improved, and it is suitable for real-time processing of high-resolution medical images.

CN120431118APending Publication Date: 2025-08-05GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510510419.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have poor practicality when dealing with characteristics such as low contrast, boundary blur and individual differences. Especially when high-resolution medical image processing, the calculation complexity is high, which is difficult to meet the real-time and hardware deployment requirements of clinical scenarios.

Method used

The U-shaped network algorithm using a state space model network combined with attention mechanism is used, and the CMamba module serves as a bridge between the encoder and the decoder. The Rise_trends module is used to train in the upsampling stage of the decoder, and the public multi-organ segmentation data set is used to optimize the network structure to improve segmentation efficiency.

Benefits of technology

It realizes the enhancement of texture continuity while maintaining sharp edges of the image, improves the accuracy and computing efficiency of medical image segmentation, and meets the real-time needs of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431118A_ABST
    Figure CN120431118A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning and medical image segmentation, in particular to a method for improving medical image segmentation based on a state space model network, and the method comprises the steps: enabling an innovation module CAMba to serve as a middle bridge between a network frame encoder and a decoder; the innovation module Risetends is used for an up-sampling stage of a decoder of a network framework; a network is established through the new module, and a public multi-organ segmentation data set is used to realize 8 segmentation of an abdominal organ; and testing and accepting the trained data set by using a loss function, and finding out an optimal experimental parameter to obtain an experimental result, thereby solving the problem that the existing medical image segmentation accuracy is not matched with the calculated amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and in particular relates to a method for improving medical image segmentation based on a state-space model network. Background Art

[0002] In recent years, the rapid development of deep learning, which uses multi-layer neural network architectures to adaptively learn complex data features, has shown strong potential for application in medical image analysis. Medical image segmentation, a key technology in precision medicine, aims to accurately delineate the tissue structure of each region in images such as CT and MRI. However, the technical difficulty lies in effectively handling the inherent low contrast, blurred boundaries, and individual differences of medical images.

[0003] While traditional deep learning fully convolutional U-Net networks hold a significant position, their limited long-range modeling capabilities make them difficult to handle complex medical anatomical structures, where organs often have similar morphologies and fuzzy boundaries. While attention mechanisms can compensate for the lack of global information in CNNs, their quadratic computational complexity poses significant challenges when processing high-resolution medical images, especially given the stringent real-time and hardware deployment requirements of clinical scenarios.

[0004] This paper uses the state space model and combines the attention mechanism U-type network algorithm to present a new image segmentation network, which effectively improves the efficiency of image segmentation. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for improving medical image segmentation based on a state-space model network, aiming to solve the problem of poor practicality of existing medical image segmentation methods.

[0006] To achieve the above object, the present invention provides a method for improving medical image segmentation based on a state-space model network, comprising the following steps: Step S1: The innovative CMamba module serves as the intermediate bridge between the encoder and decoder of the network framework. CMamba integrates the Mamba block and the channel attention block. While retaining the global modeling advantages of Mamba, it effectively balances the relationship between capturing long-range dependencies and local pixel associations through the local receptive field constraints of convolution operations and the feature calibration of the channel attention mechanism. This makes it particularly suitable for low-level vision tasks such as super-resolution that require processing texture details and structural consistency simultaneously. Step S2: The innovative module Rise_trends is used in the decoder upsampling phase of the network framework, achieving a breakthrough balance between computational efficiency and reconstruction quality. By integrating a dual-style mechanism, the "lp" mode excels at local detail enhancement and the "pl" mode is stronger at global context fusion, enabling the network to maintain sharp edges while enhancing texture continuity. Step S3: Build a network using the new module and use a public multi-organ segmentation dataset to segment eight abdominal organs. The public dataset, Synapse Multi-Organ Segmentation Dataset, is primarily used for multi-organ CT image segmentation tasks in the abdomen, covering several key organs such as the liver, kidneys, and pancreas. It serves as the basis for testing the network. The dataset uses 2212 medical images as training samples and 12 medical images as test samples. The 2212 images are used to train on the device to obtain weight parameters, and the 12 test images are used to obtain the training results. Step S4: Use the loss function to test and accept the trained data set to find the optimal experimental parameters and obtain the experimental results. The loss function includes Dice, HD95%, parameter quantity, running time, etc., among which the larger the Dice, the better (0-1) and the smaller the HD95, the better (0-1). A series of evaluation indicators are used to verify the final level of the network. The analysis results finally obtain the network weight, thereby improving the accuracy of medical segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0008] Figure 1 A flowchart of a method for improving medical image segmentation based on a state-space model network; Figure 2 This is a schematic diagram of the CMamba module of the present invention; Figure 3 This is a schematic diagram of the Rise_trends module of the present invention; Figure 4 This is a diagram of the network framework structure of the present invention; Figure 5 This is an example diagram of a multi-organ segmentation dataset; Figure 6 This is the effect diagram of real multi-organ segmentation; Figure 7 This is the effect diagram obtained after training using the network framework used in the present invention. Specific implementation

[0009] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and do not constitute a limitation of the present invention. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0010] See also Figures 1 to 7 The present invention provides a method for improving medical image segmentation based on a state space model network. The specific implementation process is as follows: Figure 1 As shown, the following steps are included.

[0011] Step S1: Innovate the CMamba module as the intermediate bridge between the network framework encoder and decoder:

[0012] The state space model (SSM) originates from modern control theory systems and is used to describe the state representation of a sequence at each time step and predict the next state based on the input. Therefore, the original theory deals with continuous time functions, similar to RNN networks, but the parameters ABC are fixed, and the formula is:

[0013] h'(t)=Ah(t)+Bx(t); y(t)=Ch(t)

[0014] Where x(t) is the input sequence, h(t) is the hidden state representation, y(t) is the predicted output sequence, and A∈R N * N , B∈R N * 1 , C∈R 1 * N However, neural network transmission is a nonlinear discrete model, so the continuous time function SSM needs to be converted into a discrete state. The conversion results are as follows:

[0015]

[0016] According to the above formula, we can deduct it from the reverse thinking and substitute it into the prediction equation to get:

[0017]

[0018] To overcome the possible loss of local details caused by pure sequence modeling, the module introduces a channel attention convolution block (CAB). This component enhances local features by cascading 3×3 convolutional layers and channel attention gating. The channel attention weight is calculated as:

[0019] CAB(x)=x·σ(W4·ReLU(W3·GAP(W2·GeLU(W1·x))))

[0020] Where x is the input, σ is the activation function, W4 and W3 are 1×1 convolutions, and W2 and W1 are 3×3 convolutions.

[0021] Therefore, while retaining the advantages of Mamba's global modeling, the entire processing flow effectively balances the relationship between capturing long-range dependencies and local pixel associations through the local receptive field constraints of convolution operations and the feature calibration of the channel attention mechanism. It is particularly suitable for low-level visual tasks such as super-resolution that require simultaneous processing of texture details and structural consistency. Its dynamic weight mechanism enables the network to adaptively adjust the contribution of different features. The overall formula is as follows:

[0022] X out =(x·r1+y(LN1(x)))·r2+CAB(LN2(x3))

[0023] Where r1 and r2 are skip_scale parameters, x3 = (x·r1+y(LN1(x)))·r2, so the overall network Cmamba module structure is as follows Figure 2 shown.

[0024] Step S2, the innovative module Rise_trends is used in the decoder upsampling stage of the network framework:

[0025] While traditional interpolation methods (such as bilinear and nearest neighbor interpolation) are computationally efficient, they suffer from insufficient feature adaptability and limited detail recovery. Existing dynamic upsampling methods (such as CARAFE and FADE) also incur significant computational overhead when improving performance. To address this, this paper proposes the Rise_trends dynamic upsampling framework, achieving a breakthrough balance between computational efficiency and reconstruction quality. It innovatively proposes a dual-style fusion mechanism, leveraging the "lp" mode's expertise in local detail enhancement and the "pl" mode's superiority in global context fusion. Its formula is as follows:

[0026] X offset =(Conv 1×1 ·X)·0.25+init_pos

[0027] X up =grid_sample(X,pixelshuffle(coords+X offset ))

[0028] X out =Conv 1×1 concat(X up_lp (X),X up_pl (pixelshuffle(X)))

[0029] Where init_pos is the initial grid position of input X, coords is coordinate normalization, pixelshuffle is pixel rearrangement upsampling, and grid_sample is bilinear interpolation. Its design is based on the feature fusion module of point-by-point convolution, which integrates the output features of the two upsampling paths through a dynamic weighting strategy, so that the network can maintain the sharp edges of the image while enhancing the texture continuity. Figure 3 shown.

[0030] Step S3: Build a network using the new module and use the public multi-organ segmentation dataset to segment the 8 abdominal organs:

[0031] The complete network framework is obtained by analyzing steps one and two, such as Figure 4 As shown in the figure. In the network framework, the image is first compressed to reduce the size of the image tensor, and this information is transferred through residual connections. In the encoder, the downsampling operation effectively highlights important features by gradually reducing the image size and increasing the depth of the convolutional layer. At the connection between the encoder and decoder, the CMamba module is introduced to balance the relationship between long-range dependency capture and local pixel association, serving as a bridge between the encoder and decoder to promote the effective transfer of information. In the decoder, Rise_trends dynamic upsampling is used to establish a balance between computational efficiency and reconstruction quality. Texture continuity is enhanced through a dual-style fusion mechanism. At the same time, the information retained by the encoder through residual connections and the upper and lower layer relationships obtained through cascade connections are combined to make the information completion more complete and accurate.

[0032] The dataset used in the experiment is the Synapse Multi-Organ Segmentation Dataset, which was compiled and made public by a research team supported by the National Institutes of Health (NIH). It was first released around 2015 and is a public dataset widely used in medical image segmentation research. It is mainly used for automatic segmentation research of abdominal organs, covering multiple key organs (aorta, gallbladder, left kidney, right kidney, liver, pancreas, spleen, stomach), such as Figure 5 Part of the data is shown. 12 of them are test sets and 2212 are training sets. The role of the designed network module is verified through training. The real result segmentation example is shown in Figure 6 shown.

[0033] Step S4: Use the loss function to test the trained data set and find the optimal experimental parameters to obtain the experimental results:

[0034] The evaluation metrics used in this study primarily include training accuracy, cross-validation loss, total loss, Dice, HD95%, and the number of network framework parameters. The Dice coefficient and HD95 (95% Hausdorff distance) are two widely used evaluation metrics in medical image segmentation tasks. The Dice coefficient measures segmentation accuracy by calculating the overlap between the predicted result and the true label; higher values indicate better overlap. HD95 evaluates segmentation quality from the perspective of spatial boundary deviation; smaller HD95 values indicate a closer fit between the segmentation boundary and the true contour.

[0035] Through the collected experimental data, after designing the network framework, we run the Synapse dataset to test whether the network is successfully built and to correct various errors in the network. We judge whether the network is feasible based on the evaluation indicators, and then use medical data to train the code to obtain a training parameter list. We use the parameter list to test the expected network code, analyze the test results and compare them with other networks, as shown in Table 1 Experimental Results (all indicators are expressed in %, and bold black indicates the best result). The experimental results are shown in Figure 1. Figure 7 shown.

[0036] Table 1 Experimental results

[0037]

Claims

1. A method for improving medical image segmentation based on a state-space model network, characterized in that: The following steps are involved: The innovative module CMamba serves as the intermediate bridge between the network framework encoder and decoder; The innovative module Rise_trends is used in the decoder upsampling stage of the network framework; By building a network with new modules, we can segment 8 abdominal organs using a public multi-organ segmentation dataset. Use the loss function to test the trained data set and find the optimal experimental parameters to obtain the experimental results.

2. The method for improving medical image segmentation based on a state-space model network according to claim 1, wherein: The method uses CMamba as a bridge to connect the encoder and decoder, where CMamba integrates Mamba blocks and channel attention blocks. While retaining the global modeling advantages of Mamba, it effectively balances the relationship between long-range dependency capture and local pixel association through the local receptive field constraint of the convolution operation and the feature calibration of the channel attention mechanism. It is particularly suitable for low-level vision tasks such as super-resolution that require simultaneous processing of texture details and structural consistency to prevent the loss of local details.

3. The method for improving medical image segmentation based on a state-space model network according to claim 1, wherein: The innovative module Rise_trends is used in the decoder stage of the network framework to achieve a breakthrough balance between computational efficiency and reconstruction quality. By fusing a dual-style mechanism, the network enhances texture continuity while maintaining sharp edges in the image, thereby overcoming the global and local data loss.

4. The method for improving medical image segmentation based on a state-space model network according to claim 1, wherein: Building a network framework using new modules means using the CMamba module as a network bridge and Rise_trends as an upsampling module. The public Synapse Multi-Organ Segmentation Dataset, primarily used for abdominal multi-organ CT image segmentation tasks, covers several key organs such as the liver, kidneys, and pancreas, and serves as the foundation for testing the network.

5. The method for improving medical image segmentation based on a state-space model network according to claim 1, wherein: The loss functions include Dice, HD95%, etc. The network is tested on the data set, and then the loss function is used to verify the final level of the network. The analysis results finally obtain the optimal network weight, thereby improving the accuracy of medical segmentation.