Tumor segmentation method based on Mama-guided multi-encoder fusion

By adopting the Mamba-guided multi-encoder fusion method in tumor segmentation, combined with the advantages of convolutional layer and SSM, the time-consuming and labor-intensive and automatic segmentation methods in complex morphological processing is solved, and efficient and accurate tumor segmentation and calculation cost reduction are achieved.

CN120070891APending Publication Date: 2025-05-30QIQIHAR UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510135098.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional tumor segmentation methods are time-consuming and labor-intensive and subjective, and the existing automatic segmentation methods are insufficient in dealing with complex tumor morphology.

Method used

The multi-encoder fusion method based on Mamba guidance is adopted to integrate the three-branch encoder, the Mamba guidance fusion attention module, the pixel attention feature fusion module and the expanded multi-scale fusion module, combining the local feature extraction capability of the convolutional layer and the global long-distance dependency capture capability of the SSM.

Benefits of technology

It realizes efficient and accurate tumor segmentation, improves feature expression and decoding capabilities, reduces computational costs, and performs better than U-Net, Transformer and Mamba related models on multiple benchmark data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070891A_ABST
    Figure CN120070891A_ABST
Patent Text Reader

Abstract

The invention discloses a tumor segmentation method based on Mama-guided multi-encoder fusion. The method comprises the following steps: 1, constructing a multi-encoder segmentation model fusing Mama and a convolutional neural network (CNN); 2, training the model on a liver and lung cancer CT image data set and optimizing model parameters; and 3, carrying out rapid positioning and accurate segmentation on an input CT image by using the trained model so as to obtain a binary image of a tumor segmentation result. According to the method, local and global information is fully utilized, multi-scale features are modeled, and the tumor segmentation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a tumor segmentation method based on the fusion of multiple encoders guided by Mamba. Background Art

[0002] Medical image analysis plays a crucial role in cancer diagnosis, treatment planning, and surgical intervention. Among them, accurate segmentation of tumors can provide important information such as tumor size, shape, and its spatial relationship with surrounding tissues, helping clinicians formulate effective treatment strategies, such as radiotherapy, chemotherapy, and surgical resection, thereby improving the treatment effect and prognosis of patients. However, traditional tumor segmentation methods rely on radiologists to manually delineate the tumor boundaries layer by layer in CT images, which is time-consuming and laborious, and prone to errors due to subjective factors and differences between observers. With the growing demand for efficient and automated tumor segmentation methods, researchers have begun to explore more advanced segmentation techniques.

[0003] Early tumor segmentation methods were mainly based on traditional image processing techniques and manual feature extraction, such as methods like fuzzy C-means (FCM) clustering, thresholding, region growing, and graph cuts. These methods perform poorly in dealing with tumors with complex morphologies and are sensitive to noise. In recent years, convolutional neural networks (CNNs) have become the mainstream in tumor segmentation research due to their powerful automatic feature extraction capabilities. Among them, U-Net has become the most widely used model in medical image segmentation due to its efficient and accurate segmentation performance. However, the convolutional operation of CNNs inherently limits its ability to capture long-range global dependencies, making it difficult to handle tumor regions with complex boundaries and structures. Although methods such as dilated convolution, self-attention mechanism, and feature pyramid have tried to solve this problem, they still fail to fully model long-term dependencies.

[0004] Transformer-based models have gradually attracted attention in the field of medical image segmentation due to their excellent global context capture capabilities. However, when dealing with large medical datasets, Transformer models have high computational costs and it is difficult to balance efficiency and accuracy. In addition, single-encoder architectures face significant challenges in simultaneously capturing local and global features, especially when dealing with tumors with irregular shapes, unclear boundaries, or fragmented regions. Summary of the Invention

[0005] The present invention aims to provide an efficient and accurate tumor segmentation method to solve the problems of time-consuming and laborious traditional manual segmentation, strong subjectivity, and the deficiencies of existing automatic segmentation methods in dealing with complex tumor morphologies.

[0006] To achieve the above object, the present invention proposes a tumor segmentation method based on Mamba-guided multi-encoder fusion, which integrates multiple parallel encoders to extract complementary features, combines the local feature extraction ability of the convolutional layer with the advantage of SSM in capturing global long-distance dependencies, and effectively overcomes the limitations of the single-encoder architecture. The method includes the following steps:

[0007] Construct a multi-encoder segmentation model that fuses Mamba and a convolutional neural network (CNN);

[0008] Train the model on the liver and lung cancer CT image datasets and optimize the model parameters;

[0009] Use the trained model to perform rapid localization and accurate segmentation based on CT images to obtain a binary map of the tumor segmentation result.

[0010] Further, it specifically includes the following steps:

[0011] Obtain a dataset, and preprocess the obtained dataset; divide the labeled data in the preprocessed dataset into a training set, a validation set, and a test set for the deep learning network; finally, build a multi-encoder segmentation model that fuses Mamba and a convolutional neural network (CNN), where building the multi-encoder segmentation model that fuses Mamba and a convolutional neural network (CNN) includes a three-branch encoder, a Mamba-guided fusion attention module (MGFA), a pixel attention feature fusion module (PAFF), and a dilated multi-scale fusion module (DMSF).

[0012] Further, building the multi-encoder segmentation model that fuses Mamba and a convolutional neural network (CNN) specifically includes the following steps: The model is based on an encoder-decoder structure and includes a three-branch encoder, a bottleneck module, three decoder blocks, and skip connections. In stage 0, each input image undergoes shallow feature extraction through two parallel convolutional modules: a series of 3×3 standard convolutional layers and a depthwise convolution (DWConv) layer. The three-branch encoder extracts detail features, context features, and guiding features respectively, and these features are fused through the Mamba-guided fusion attention (MGFA) module in stage 4. In stage 5, the dilated multi-scale fusion (DMSF) module further enhances the feature representation by capturing multi-scale context details. The decoder gradually restores the feature map to the original image size, and connects the downsampling blocks at the corresponding scales with the help of skip connections to ensure the restoration of fine-grained details, so that the network can generate accurate and reliable tumor segmentation predictions.

[0013] Further, training and optimizing the parameters of the built model specifically includes initializing the deep learning network using the CUDNN convolutional layer in Nvidia and the standard normal initialization method;

[0014] Set the total number of training rounds to 100 and save the final training weights;

[0015] In the test phase, output the target probability images of the validation set and the test set;

[0016] Finally, use the stochastic gradient descent algorithm to train the deep learning network.

[0017] Furthermore, randomly divide the labeled data in the preprocessed dataset into the training set and the test set of the deep learning network; specifically, perform dataset preprocessing on the arterial phase enhanced CT image datasets of liver tumors and lung cancers, and then divide them into the training set, the validation set and the test set according to the ratio of 7:1:2, and train the network segmentation model based on the nnunet framework.

[0018] Furthermore, the loss function is designed as a combination of cross-entropy loss and Dice loss. The objective function is defined as follows:

[0019] L = αL CE + βL Dice

[0020] L CE = -∑glog(p)

[0021]

[0022] where α and β are weight coefficients, p and g are the predicted probability and the corresponding groundtruth respectively, σ ∈ [0,1] is a smoothing factor used to avoid division by zero. In the experiments of this paper, both α and β are set to 1, and σ is set to 0.00001.

[0023] This objective function combines the pixel-level classification ability of cross-entropy and the contour matching ability of Dice loss, enabling the segmentation network to improve the accuracy of regional overlap while ensuring the global accuracy.

[0024] Furthermore, finally use the stochastic gradient descent algorithm to train the deep learning network, and its parameter formula is as follows:

[0025]

[0026] where, t represents the model parameters at the t-th iteration, and η is the learning rate (stepsize), and is the gradient of the loss function L at θ.

[0027] Furthermore, when using the trained model to quickly locate and accurately segment the CT image to obtain the binary segmentation map of the liver tumor or lung cancer lesion area, the specific implementation steps are as follows:

[0028] Image preprocessing and enhancement: The original CT images are sequentially preprocessed and enhanced to generate processed images.

[0029] Feature extraction: The images after online enhancement are input into a feature extraction network composed of a convolutional layer, a pooling layer, a normalization layer, and an SSM module, and a reconstructed feature map is output.

[0030] Sliding window pixel-by-pixel prediction: The reconstructed feature map is input into a classifier, and a sliding window strategy is used to predict each pixel in the feature map one by one, thereby generating two pixel-level label prediction score maps with the same size as the original image.

[0031] Probability conversion processing: The ReLU function is used to non-negativize the prediction scores. At the same time, through normalization operations, the possibility of each pixel belonging to different classes is accurately reflected.

[0032] Generating a binary segmentation map: For each pixel, the subscript corresponding to the class with the highest probability is selected as the label of the pixel, thereby achieving rapid localization of liver tumors or lung lesions while obtaining the final binary segmentation map.

[0033] The beneficial effects of the present invention are as follows:

[0034] A tumor segmentation method based on Mamba-guided multi-encoder fusion provided by the present invention uses a three-branch encoder structure, and simultaneously integrates a detail branch and a context branch based on CNN and a guided branch encoder based on Mamba, making up for the limitations of a single encoder and effectively capturing global dependencies and low-level spatial details. Through key modules such as a pixel attention feature fusion module, a Mamba-guided fusion attention module, and a dilated multi-scale fusion module, the feature expression and decoding capabilities are comprehensively improved. We conducted comparative experiments on four public and private benchmark datasets (covering liver tumors and lung cancers). The results show that MMEFU-Net outperforms U-Net, Transformer, and Mamba-related models in terms of metrics such as Dice similarity coefficient (DSC) and intersection over union (IoU), and significantly reduces the computational cost. Description of the Drawings

[0035] Figure 1 is a schematic diagram of the overall model architecture;

[0036] Figure 2 is a schematic diagram of the structure of the residual depthwise separable module (RDS);

[0037] Figure 3 is a schematic diagram of the structure of the improved residual vision Mamba layer (RVM) of the present invention;

[0038] Figure 4It is a schematic diagram of the Mamba-guided fusion attention module (MGFA) structure;

[0039] Figure 5 It is a schematic diagram of the dilated multi-scale fusion module (DMSF) structure;

[0040] Figure 6 It is a schematic diagram of the pixel attention feature fusion module (PAFF) structure;

[0041] Figure 7 It is a schematic diagram of the Dysample dynamic upsampler module structure;

[0042] Figure 8 It is a schematic diagram of some cases in the liver tumor dataset;

[0043] Figure 9 It is a schematic diagram of some cases in the lung cancer dataset;

[0044] Figure 10 It is a schematic diagram of the comparison between the original data and the preprocessed data;

[0045] Figure 11 It is a schematic diagram of the comparison of the segmentation effects of different models on the liver tumor dataset;

[0046] Figure 12 It is a schematic diagram of the comparison of the segmentation effects of different models on the lung cancer dataset. Detailed implementation manners

[0047] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] In the following description, specific details such as specific internal programs and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0049] As Figure 1 shown, a tumor segmentation method based on Mamba-guided multi-encoder fusion provided by the present invention includes the following steps:

[0050] Step 100, build a multi-encoder segmentation model that fuses Mamba and a convolutional neural network (CNN);

[0051] Step 200: Train and optimize the parameters of the established model on the liver and lung cancer CT image datasets;

[0052] Step 300: Use the trained model to perform rapid positioning and accurate segmentation based on CT images to obtain the binary maps of liver tumor and lung cancer segmentation results.

[0053] Among them, in Step 100, a multi-encoder segmentation model integrating Mamba and convolutional neural network (CNN) is established;

[0054] Among them, establishing a multi-encoder segmentation model integrating Mamba and convolutional neural network (CNN) includes a three-branch encoder, a Mamba-guided fusion attention module (MGFA), a pixel attention feature fusion module (PAFF), a dilated multi-scale fusion module (DMSF), and a decoder.

[0055] Establishing a multi-encoder segmentation model integrating Mamba and convolutional neural network (CNN) includes a three-branch encoder, specifically including the following steps:

[0056] A multi-branch encoder with a hierarchical structure is adopted, and the resolution of its feature maps gradually decreases as the network depth increases. To make up for the information loss caused by the reduction of spatial resolution, the number of channels of the feature maps gradually increases during the increase of the network depth. The three branches include a detail branch (D) and a context branch (C) based on CNN and a guidance branch (G) based on mamba. The encoder is divided into four stages, named stage 0-3 respectively, where the feature map resolution and channel number of each stage are different. In stage 0, the detail branch and context branch based on CNN first extract the shallow features of the input CT image through a cascaded residual backbone network; in stages 1-3, the detail branch consists of two 3×3 convolutional layers and a bottleneck residual block, and extracts details with 64, 64, and 128 feature channels at a resolution of 64×64 respectively.

[0057] The context branch uses three cascaded residual depthwise separable (RDS) modules to capture long-range dependencies and generate rich feature representations, such as Figure 2As shown. It extracts features at gradually decreasing resolutions (32×32, 16×16, and 8×8), with corresponding feature channels 128, 256, and 512. The Mamba-based guiding branch first uses a depthwise convolution (DWConv) layer to extract shallow features, generating a feature map with 32 fixed filters. Next, this branch employs three consecutive encoder blocks to extract deeper features, with each block doubling the number of channels of the feature map while halving the resolution. Thus, the guiding branch generates three multi-resolution feature maps at resolutions of 256×256, 128×128, and 64×64, with corresponding channel sizes 64, 128, and 256.

[0058]

[0059] where x i refers to the input feature, Conv 3 refers to a standard 3×3 convolutional layer with a stride of 2, Conv 1,G refers to pointwise convolution followed by the GeLU activation function, DWConv BN refers to depthwise convolution followed by batch normalization, refers to the output feature map.

[0060] To further enhance long-range spatial modeling, the Mamba-based guiding branch adopts an improved Residual Vision Mamba (RVM) layer to optimize semantic feature extraction, as Figure 3 (a) shown. By using advanced residual connections and scaling factors, the RVM layer achieves performance improvement with minimal parameter and computational complexity increase.

[0061]

[0062] After that, the RVM layer applies another LayerNorm for normalization and projects the enhanced features into a deeper representation space through a projection layer: the Visual State Space (VSS) module, as Figure 3 (b) shown, which efficiently models long-range spatial dependencies. This design enables the G branch to enhance feature fusion in the network, facilitating the integration of strong spatial and semantic information, thereby improving the segmentation performance.

[0063] W 1 = LayerNorm(SSM(SiLU(DWConv(Linear(Win)))))

[0064] W 2 = SiLU(Linear(W in ))

[0065] where ⊙ represents the Hadamard product.​​

[0066] Furthermore, the Mamba Guided Fusion Attention (MGFA) module is used to balance local details and global context features as Figure 4 shown. The MGFA module utilizes the long-range global dependency information extracted by the G branch to dynamically guide the fusion of the D branch and the C branch. Specifically, the D branch retains rich spatial and geometric details by maintaining the feature map resolution, while the C branch extracts semantically rich global context features. However, the C branch tends to lose high-frequency spatial details, especially in small tumor and boundary regions. Through the attention mechanism, it highlights the key regions, enabling the model to rely more on the local features of the detail branch in the boundary regions while using the context features to fill other regions. This effectively achieves efficient feature fusion and detail retention. This design ensures the dynamic fusion of multi-branch features, improves the representation ability of the network, while retaining detailed spatial information and deep semantic information. It shows significant advantages especially in segmentation tasks involving small tumors and complex boundaries.

[0067] σ = Sigmoid(x g )

[0068]

[0069] where f refers to the combination of convolution, batch normalization, and ReLU. When σ > 0.5, the model relies more on detail features, otherwise it biases towards context information.

[0070] Furthermore, the Dilated Multi-Scale Fusion (DMSF) module is used to enhance multi-scale context information modeling and feature fusion within the bottleneck structure, as Figure 5 shown. It employs parallel dilated convolutions with different dilation rates to capture features with different receptive fields. The outputs of these parallel paths are then spatially aligned and fused, effectively integrating multi-scale features into a unified representation. This design improves the model's representation ability for complex tumor structures, ensuring a more comprehensive and robust feature extraction process.

[0071]

[0072] where f r (·) represents a dilated convolution operation with a dilation rate of r. Then, multi-scale features are fused through channel-wise concatenation, followed by using a 1×1 convolution to compress and reorganize the features. Concat(·) represents the channel concatenation operation.

[0073] Furthermore, the Pixel Attention Feature Fusion (PAFF) module dynamically fuses the feature representations of the D branch encoder and the C branch encoder, as Figure 6As shown. Since the D branch contains fewer layers and channels, the C branch serves as a supplementary source through lateral connections, selectively learning supplementary information from the context branch. This ensures the accuracy and reliability of detail parsing while enhancing the discriminative ability of the encoder features. The formula is as follows;

[0074]

[0075] where σ i represents the likelihood that these two pixels belong to the same object. If σ i is higher, we rely more on x d because the C branch is semantically rich and accurate; otherwise, we rely more on x c .

[0076] Furthermore, the decoder block is responsible for decoding the feature map and gradually restoring the image resolution. Specifically, the decoder module receives two inputs: a feature map from the skip connection and a feature map from the output of the previous module. These two feature maps are first fused through an addition operation. Subsequently, the module applies depthwise convolution (DWConv), residual connections, and the ReLU activation function to further decode the feature map, while introducing an adjustment factor to enhance the decoding ability of the residual connection. Its computational process is expressed as follows:

[0077]

[0078] Furthermore, to overcome the challenge of accurately reconstructing the tiny structures in the target region, we introduce the DySample dynamic upsampler in the decoding stage (as Figure 7 shown). Different from traditional upsampling methods, DySample can adaptively adjust the positions and weights of the sampling points, thus significantly improving the accuracy and efficiency of upsampling. DySample performs particularly well in the tumor segmentation task, especially in cases of complex backgrounds and overlapping organ tissues, and can more precisely restore and reconstruct image details.

[0079] Step 200, train the built model and optimize the parameters. Among them,

[0080] Step 201, obtain the dataset, preprocess the obtained dataset, and divide the labeled data in the preprocessed dataset into the training set and test set of the deep learning network;

[0081] Step 202, finally use the training set as the input of the network to train and optimize the network.

[0082] Specifically, further, the built model is trained and optimized for parameters, specifically including initializing the deep learning network using the CUDNN convolutional layer and the standard normal initialization method in Nvidia; setting the batchsize to 4, the maximum number of iterations to 100, the initial learning rate to 0.0001, and saving the final training weights;

[0083] In the test phase, the target probability images of the validation set and the test set are output;

[0084] Finally, the deep learning network is trained using the stochastic gradient descent algorithm.

[0085] Among them, step 202 includes step 2021, randomly dividing the labeled data in the preprocessed dataset into the training set and the test set of the deep learning network; specifically, performing dataset preprocessing on the arterial phase enhanced CT image datasets of liver tumors and lung cancers, and then dividing them into the training set, the validation set, and the test set according to the ratio of 7:1:2, and training the network segmentation model based on the nnunet framework.

[0086] Step 2022, the loss function is designed as a combination of cross-entropy loss and Dice loss. The objective function is defined as follows:

[0087] L = αL CE + βL Dice

[0088] L CE = -∑glog(p)

[0089]

[0090] Among them, α and β are weight coefficients, p and g are the predicted probability and the corresponding ground truth respectively, σ ∈ [0,1] is a smoothing factor used to avoid division by zero. In the experiments in this paper, both α and β are set to 1, and σ is set to 0.00001.

[0091] Step 2023 optimization function: Finally, the deep learning network is trained using the stochastic gradient descent algorithm, and its parameter formula is as follows:

[0092]

[0093] Among them, represents the model parameters at the t-th iteration, is the learning rate (step size), is the gradient of the loss function at ;

[0094] Step 300, using the trained model to quickly locate and accurately segment the CT image to obtain a binary segmentation map of the liver tumor or lung cancer lesion area. The specific implementation steps are as follows:

[0095] Image preprocessing and enhancement: The original CT images are successively subjected to data preprocessing and enhancement to generate processed images.

[0096] Feature extraction and automatic positioning: The images after online enhancement processing are input into a feature extraction network composed of a convolutional layer, a pooling layer, a normalization layer, and an SSM module, and a reconstructed feature map is output.

[0097] Sliding window pixel-by-pixel prediction: The reconstructed feature map is input into a classifier, and a sliding window strategy is used to predict each pixel in the feature map one by one, thereby generating two pixel-level label prediction score maps with the same size as the original image.

[0098] Probability conversion processing: The ReLU function is used to perform non-negativity processing on the prediction scores to ensure that the output values are always positive, thereby avoiding negative value interference. At the same time, through normalization operations, the output better conforms to the probability distribution and more accurately reflects the possibility of each pixel belonging to different classes.

[0099] Generating a binary segmentation map: For each pixel, the subscript corresponding to the class with the highest probability is selected as the label of the pixel, thereby achieving rapid localization of liver tumors or lung lesions while obtaining the final binary segmentation map.

[0100] The following will give a specific example to illustrate a method for segmenting liver tumor and lung cancer lesion regions in CT images based on deep learning according to an embodiment of the present invention, which specifically includes the following steps:

[0101] Step 1) Build a multi-encoder segmentation model that integrates Mamba and a convolutional neural network (CNN). Among them, building a multi-encoder segmentation model that integrates Mamba and a convolutional neural network (CNN) includes a three-branch encoder, a Mamba-guided fusion attention module (MGFA), a pixel attention feature fusion module (PAFF), a dilated multi-scale fusion module (DMSF), and a decoder.

[0102] Step 2) Divide the labeled data in the preprocessed dataset into a training set and a test set for the deep learning network.

[0103] Step 21) In this embodiment, we evaluated the performance of the model on public datasets and private datasets, including LiTS2017 and three private datasets: QMLiTS, QMLUAD, and QMSCC. Figure 8 and Figure 9 Shows schematic diagrams of some cases of the liver tumor and lung cancer segmentation datasets.

[0104] LiTS2017: This dataset contains 131 enhanced abdominal CT scan volumes and their corresponding liver tumor annotation information. We selected 131 abdominal CT images from the LiTS dataset and sliced them into 7,182 axial slices containing tumors. Among them, 5,028 slices were used for the training set, 718 slices for the validation set, and 1,436 slices for the test set.

[0105] QMLiTS: This dataset contains 57 pathologically confirmed hepatocellular carcinoma cases, with the data collection time ranging from June 2017 to September 2024, sourced from the Affiliated Hospital of Qiqihar Medical College and the Fourth Hospital of Harbin Medical University. The dataset contains arterial phase enhanced abdominal CT images, randomly divided into 40 cases for training, 6 cases for validation, and 11 cases for testing.

[0106] QMLUAD: This dataset contains 57 pathologically confirmed pulmonary adenocarcinoma cases, with the data collection time ranging from January 2023 to June 2024, sourced from the Affiliated Hospital of Qiqihar Medical College. The selected images are arterial phase enhanced CT scans, and the dataset is randomly divided into 40 cases for training, 6 cases for validation, and 11 cases for testing.

[0107] QMSCC: This dataset contains 43 pathologically confirmed pulmonary squamous cell carcinoma cases, with the data collection time ranging from January 2023 to June 2024, sourced from the Affiliated Hospital of Qiqihar Medical College. The selected images are arterial phase enhanced CT scans, and the dataset is randomly divided into 30 cases for training, 4 cases for validation, and 9 cases for testing.

[0108] To eliminate the interference of other organ tissues and highlight the tumor area, we truncated and normalized the image intensity values of all axial slices to the range of [-200, 250] HU. The preprocessed input images were resized to 512×512 pixels and converted to RGB format (as Figure 10 shown).

[0109] Step 3) Model training and parameter optimization

[0110] Step 31) Initialize the deep learning network using the CUDNN convolutional layer in Nvidia and the standard normal initialization method; set the batch size to 4, the maximum number of iterations to 100, the initial learning rate to 0.0001, and save the final training weights;

[0111] Step 32), the loss function is designed as a combination of cross-entropy loss and Dice loss. The objective function is defined as follows:

[0112] L = αL CE + βL Dice

[0113] L CE = -∑g log(p)

[0114]

[0115] Where α and β are weight coefficients, p and g are the predicted probability and the corresponding ground truth respectively, and σ ∈ [0, 1] is a smoothing factor used to avoid division by zero. In the experiments of this paper, both α and β are set to 1, and σ is set to 0.00001.

[0116] Step 33) Optimization function: Finally, use the stochastic gradient descent algorithm to train the deep learning network, and its parameter formula is as follows:

[0117]

[0118] Where represents the model parameters at the t-th iteration, is the learning rate (step size), is the gradient of the loss function at ;

[0119] Step 34) Use the trained model to quickly locate and accurately segment the CT image to obtain the binary segmentation result of the liver tumor or lung cancer lesion area

[0120] Step 35) Compare the trained model with other segmentation models, and use multiple key indicators such as the Dice similarity coefficient (DSC, %) and the intersection over union (IoU, %) to evaluate the performance of the proposed network model, and obtain accurate tumor segmentation results. The experimental results are shown in Table 1 and Table 2, and their visualization results are as Figure 11 , Figure 12 shown.

[0121] Table 1

[0122]

[0123] Table 2

[0124]

[0125]

[0126] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of this application based on the concept of the present invention through logical analysis, reasoning or limited experiments should be within the protection scope determined by the claims.

Claims

1. A tumor segmentation method based on Mamba-guided multi-encoder fusion, characterized in that: The following steps are involved: Construct a multi-encoder segmentation model that integrates Mamba and convolutional neural network (CNN); Training the model and optimizing model parameters on liver and lung cancer CT image datasets; The trained model is used to quickly locate and accurately segment the input CT image to obtain a binary image of the tumor segmentation result.

2. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 1, characterized in that: The specific steps include: Build a multi-encoder segmentation model that integrates Mamba and convolutional neural network (CNN). The multi-encoder segmentation model that integrates Mamba and convolutional neural network (CNN) includes a three-branch encoder, Mamba guided fusion attention module (MGFA), pixel attention feature fusion module (PAFF), dilated multi-scale fusion module (DMSF) and decoder; Acquire a data set and preprocess the acquired data set; The labeled data in the preprocessed data set is divided into a training set and a test set for the deep learning network, and the training set is used as the input of the network to train and optimize the network; Finally, the optimized segmentation network is used to segment the CT image to obtain the segmentation results of the liver tumor or lung cancer lesion area.

3. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 2, characterized in that: Building a multi-encoder segmentation model that combines Mamba and a convolutional neural network (CNN) includes the following steps: The model is based on an encoder-decoder structure, which includes a three-branch encoder, a bottleneck module, three decoder blocks, and skip connections; In stage 0, each input image undergoes shallow feature extraction through two parallel convolutional modules: a series of 3×3 standard convolutional layers and a depthwise convolutional (DWConv) layer. In stages 1 to 3, the three-branch encoder extracts detail features, context features, and guidance features respectively, which are fused by the Mamba Guided Fusion Attention (MGFA) module in stage 4; In stage 5, the dilated multi-scale fusion (DMSF) module further enhances the feature representation by capturing multi-scale context details; The decoder restores the feature map to the original image size by upsampling layer by layer, adopts the Dysample dynamic upsampler, and uses jump connections to fuse the downsampled feature blocks of corresponding scales, so as to fully restore the fine-grained details and ensure that the network generates accurate and reliable tumor segmentation predictions.

4. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 2, characterized in that: The three-branch encoder combines the local feature extraction of CNN with the global dependency capture capability of SSM, and includes a detail branch (D), a context branch (C), and a Mamba guided branch (G), which overcomes the limitations of a single encoder model. The resolution of its feature map decreases with depth, while the number of channels gradually increases to compensate for information loss, effectively capturing global dependencies and shallow spatial details.

5. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 2, characterized in that: The Mamba Guided Fusion Attention (MGFA) module is used to cope with the semantic differences of encoders in different branches and balance local details and global context features. The MGFA module uses the global long-distance dependency information extracted by the G branch to dynamically guide the fusion of the D branch and the C branch, thereby improving the representation ability of the network while retaining detailed spatial information and deep semantic information. σ=Sigmoid(x g ) Here, f refers to the combination of convolution, batch normalization and ReLU. When σ>0.5, the model relies more on detail features, otherwise it is biased towards contextual information.

6. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 2, characterized in that: The dilated multi-scale fusion (DMSF) module is used to enhance the multi-scale context information modeling and feature fusion within the bottleneck structure; parallel dilated convolutions with different dilation rates are used to capture features of different receptive fields, thereby improving the model's ability to represent complex tumor structures. Among them, f r (·) represents an expansion convolution operation with a dilation rate of r. Then, multi-scale features are fused through channel-level connections, and then 1×1 convolution is used to compress and reorganize the features. Concat(·) represents a channel concatenation operation.

7. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 2, characterized in that: The pixel attention feature fusion (PAFF) module is used to dynamically fuse the feature representations of the D-branch encoder and the C-branch encoder; the detail branch is enabled to selectively learn supplementary information from the context branch, ensuring the accuracy and reliability of detail parsing while enhancing the discriminative ability of the encoder features. The formula is as follows: Among them, σ i represents the possibility that these two pixels belong to the same object. If σ i Higher, we rely more on x d , because the C branch is semantically rich and accurate; otherwise, we rely more on x c .

8. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 2, characterized in that: Train and optimize parameters of the built model, including: Initialize the deep learning network using the CUDNN convolutional layer and the standard Shota initialization method from Nvidia; Set the total number of training rounds to 100 and save the final training weights; In the testing phase, target probability images of the validation set and the test set are output; Finally, the stochastic gradient descent algorithm is used to train the deep learning network.

9. The tumor segmentation method based on Mamba-guided multi-encoder fusion as claimed in claim 1, characterized in that: When using the trained model to quickly locate and accurately segment CT images to obtain a binary segmentation map of the liver tumor or lung cancer lesion area, the specific implementation steps are as follows: Image preprocessing and enhancement: preprocess and enhance the original CT image to generate the processed image; Feature extraction: Input the enhanced image into the feature extraction network and output the reconstructed feature map; Sliding window prediction: Through the sliding window strategy, the feature map is predicted pixel by pixel to generate a prediction score map with the same size as the original image; Probability conversion: Use the ReLU function for non-negativity and normalize to reflect the possibility of pixels belonging to different categories; Generate segmentation map: Select the category corresponding to the maximum probability of each pixel as the label to obtain a binary segmentation map to locate liver tumors or lung lesions.

Citation Information

Cited By

  • 2D medical image segmentation method and system based on Mama and UNet

    CN120997233A

  • A 2D medical image segmentation method and system based on Mamba and UNet

    CN120997233B

  • Early-stage amyloid nephropathy recognition and diagnosis method based on deep learning

    CN121686056A