Super-pixel segmentation method and system based on Mamba state space model
Through the two-stage training strategy of the Mamba state space model, combined with the global network and local network, the limitations of global representation in the existing technology are solved, and high-precision superpixel segmentation is achieved, which improves the robustness and generalization of the model.
Patent Information
- Application Number
- CN202510419182.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
The existing deep learning superpixel segmentation algorithm has limitations in capturing global representations, resulting in low area capture capabilities and difficulty in generating high-precision superpixel results.
A two-stage training strategy based on Mamba state space model is adopted to capture the coarse and fine-grained features of the image through the global network and the local network respectively, and feature fusion is realized through feature fusion units to generate superpixel results.
The model's perception of boundary details and area attributes is improved, the robustness and generalization of superpixel segmentation are improved, and high-quality superpixel results are generated.
Smart Images

Figure CN120279271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing technology and pattern recognition, and more particularly, to a superpixel segmentation method and system based on a Mamba state space model. Background Art
[0002] As a powerful image simplification technology, the core idea of the superpixel algorithm is to divide an image into regions with semantic coherence through features such as color, texture, and spatial similarity. Compared with traditional pixel-level image representation methods, superpixels are more in line with human visual perception characteristics, can effectively reduce data redundancy, and provide key spatial support for multiple computer vision tasks such as saliency detection, object tracking, and semantic segmentation. Due to its strong boundary preservation ability and good generalization, in recent years, how to develop efficient and reliable superpixel segmentation algorithms has become a research hotspot. The superpixel segmentation methods in the existing literature can be mainly divided into two categories: traditional algorithms and deep learning algorithms.
[0003] Traditional superpixel algorithms first initialize grid cells or seed points, and then use methods such as k-means, graph cuts, and watershed transformation to iteratively optimize pixel allocation. Finally, the superpixel results are constructed. However, these methods are limited by hand-crafted feature design, making it difficult for them to be embedded in a trainable deep learning framework. In recent years, researchers have used deep neural networks for superpixel segmentation, providing more efficient alternatives to traditional algorithms, such as SEAL (Segmentation-Aware Loss) and SSN (Superpixel Sampling Networks). Both of them use fully convolutional networks to extract image features and use traditional clustering to generate superpixels. However, the backpropagation process of SEAL and SSN involves two independent stages, resulting in a high algorithm complexity. To solve the above problems, the SCN (Superpixel Segmentation with Fully Convolutional Networks) algorithm based on the UNet architecture has been proposed in the prior art. This method realizes feature modeling by directly predicting the membership relationship between image pixels and preset regular grid cells. Inspired by this, subsequent studies have successively proposed the AINet (Association Implantation for Superpixel) algorithm containing an association implantation module and the ESNet (Efficient Superpixel Network) algorithm integrating a pyramid gradient structure. However, existing deep learning superpixel algorithms generally use CNN as the basic architecture. Although this architecture performs excellently in local feature extraction, it has limitations in capturing global representations, resulting in a low regional capture ability. Therefore, it is urgent to design an efficient deep network model to obtain more accurate superpixel results. Summary of the Invention
[0004] To solve the above problems, the purpose of the present invention is to provide a superpixel segmentation technology based on the Mamba state space model, aiming to synchronously capture global context information and local detail features by adopting a two-stage training strategy, thereby realizing accurate superpixel segmentation.
[0005] To achieve the above technical purpose, the present application provides a superpixel segmentation method based on the Mamba state space model, including the following steps:
[0006] Use a two-stage training strategy to capture the coarse-grained features and fine-grained features of the image. Among them, in the first stage, a global network is adopted, and with the help of image cropping operation, the global features are accurately captured by the Mamba module; in the second stage, a local network is adopted, and with the help of non-overlapping sliding window sampling operation, the Mamba module is promoted to efficiently extract local features.
[0007] Obtain the pixel segmentation result by fusing global features and local features.
[0008] Preferably, in the first stage, the original image is processed by a CNN, and the cropped image sequence is processed using the Mamba module.
[0009] Preferably, in the first stage, the CNN encoder consists of a convolutional block, an activation function, and a pooling layer. Among them, the convolutional block uses a 3×3 convolution, the activation function uses LELU, and the pooling layer is max pooling. The Mamba encoder uses the VMamba framework.
[0010] Preferably, in the first stage, the CNN decoder consists of two convolutional blocks. Each convolutional block uses a 3×3 convolution and includes the activation function LELU and a batch normalization layer.
[0011] Preferably, in the second stage, the quarter-sized image is processed by a CNN, and the image sequence sampled by non-overlapping sliding windows is processed using the Mamba module.
[0012] Preferably, when performing feature fusion, through the feature fusion unit FFU, fuse the coarse-grained global features and fine-grained local features to enhance the global perception ability of local cues and at the same time strengthen the local understanding of global attributes.
[0013] Preferably, when fusing the coarse-grained global features and fine-grained local features, obtain the feature information of the global network and the local network through the PAB module for feature fusion. Among them, the PAB module consists of two layers of 1×1 convolutions and one layer of activation function.
[0014] The present invention discloses a superpixel segmentation system based on the Mamba state space model, including:
[0015] A feature capture module for capturing the coarse-grained features and fine-grained features of an image using a two-stage training strategy. Among them, in the first stage, a global network is used, and with the help of image cropping operations, the global features are accurately captured through the Mamba module; in the second stage, a local network is used, and with the help of non-overlapping sliding window sampling operations, the Mamba module is promoted to efficiently extract local features.
[0016] A feature fusion module for obtaining the pixel segmentation result by fusing global features and local features.
[0017] The present invention discloses the following technical effects:
[0018] The present invention has strong robustness, and the model's perception ability of boundary details and regional attributes is improved through a two-stage architecture.
[0019] The present invention has strong generalization ability and can obtain high-quality superpixel results in different modal datasets. Brief Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 is the superpixel segmentation model based on Mamba according to the present invention;
[0022] Figure 2 is the convolutional block in the decoding process according to the present invention;
[0023] Figure 3 is a comparison schematic diagram of the original curves and their first derivatives of the log-sum-exp linear unit and other common activation functions according to the present invention;
[0024] Figure 4 is the feature fusion unit according to the present invention;
[0025] Figure 5 are the quantization metrics of the method according to the present invention and the existing methods on four datasets of BSDS500, NYU, KITTI, and DRIVE;
[0026] Figure 6 is a visualized comparison experimental diagram of the method according to the present invention and the existing methods on four datasets of BSDS500, NYU, KITTI, and DRIVE. Detailed Embodiments
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present application. Generally, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0028] Such as Figure 1-6As shown in the figure, the present invention provides a superpixel segmentation model based on the Mamba architecture (Superpixel Segmentation with Mamba, SSMamba), which adopts a two-stage training strategy to synchronously capture global context information and local detail features, thereby achieving accurate superpixel segmentation. Due to the need for long-distance modeling, the computational complexity of the Mamba model increases linearly with the feature resolution. Therefore, the present invention uses a CNN to extract high-resolution features and Mamba to capture low-resolution attributes, forming a complementary mechanism of advantages. The specific content is as follows:
[0029] Use a two-stage training strategy to capture the coarse-grained and fine-grained features of the image. In the first stage of training, a global network is adopted. The original image is processed by a CNN, while Mamba processes the image sequence blocks obtained by the cropping operation to capture global context information. In the second stage of training, a local network is adopted. The CNN processes the image with a quarter size, while Mamba processes the non-overlapping sliding window sampling operation to obtain local image sequence blocks and extract local inherent attributes. Finally, the two-stage features are fused through a Feature Fusion Unit (FFU) and fed into the superpixel segmentation head to generate the superpixel result. The specific implementation steps are as follows:
[0030] Step 1: Construct a global network, configure the operating parameters of the global network, and preprocess the input data;
[0031] Step 2: Train the global network, input the image data in Step 1 into the global network for the first stage of training;
[0032] Step 3: Construct a local network, configure the operating parameters of the local network, and load the training weights of the global network.
[0033] Step 4: Train the local network, input the image data in Step 1 into the local network for the second stage of training;
[0034] Step 5: Global-local feature fusion, fuse the features output in Step 2 and Step 4 through a feature fusion unit;
[0035] Step 6: Generate superpixels. After preprocessing the test image, input it into the trained SSMamba network model to generate superpixels and output the segmentation result.
[0036] Embodiment: Aiming at the limitations and insufficient accuracy of CNN in the superpixel segmentation task, the present invention proposes a superpixel segmentation method based on the Mamba state space model, which is specifically described as follows:
[0037] (1) Construct the global network: The learning rate η of the global network model is 4e-4, the weight w is 5e-5, the momentum factor β is 0.9, the mini-batch input m is 32, and the Mamba module uses the pre-trained weights of VMamba. In terms of the dataset, BSDS500 is selected as the training data. This dataset usually contains multiple label annotations, so the images are extended to match the corresponding number of labels, and a total of 1087 training samples are generated. Before inputting the network for training, the images are resized to 208×208 and randomly flipped and cropped for data augmentation. The same operation strategy is also applied to the images in the test set.
[0038] (2) Train the global network: As shown in Figure 1 , the global network uses an encoder-decoder architecture to extract features from the given images, obtaining image features at different scales; the CNN encoder C ge is used to obtain high-resolution features. As shown in Figure 1 , it consists of a convolutional block, an activation function, and a pooling layer. Among them, the convolutional block uses a 3×3 convolution, the activation function uses LELU, and the pooling layer is a max pooling. Since the Mamba encoder has high computational resource requirements for high-resolution features, the two-dimensional image is converted into a sequence of two-dimensional image blocks Z, and the Mamba encoder M ge is used to mine low-resolution features on the sequence block Z. The combined global encoder CM ge is calculated as follows:
[0039]
[0040] where represents the low-level features output by C ge , and represents the high-level features output by M ge . To maintain the accuracy of spatial information, the resolution of is consistent with the input image X. The resolution of decreases exponentially.
[0041] In the decoding stage, the CNN decoder C gd is used to reconstruct the feature map. The Conv Block in it is as shown in Figure 2 , which consists of two convolutional blocks. Each convolutional block uses a 3×3 convolution and includes an activation function and a batch normalization layer. The formula for global decoding is as follows:
[0042] F G = C gd {CM ge} = C gd {{C ge (X)}, {Mge (Z)}} (2)
[0043] Among them, F G represents the coarse-grained global feature obtained by integrating CNN and Mamba.
[0044] In addition, during the training of the global network, a new activation function designed by the present invention, called the logarithmic sum exponential linear unit (LELU), is used. Its distribution curve is as Figure 3 shown. LELU has smoothness, continuity, and non-monotonicity. The calculation formula is as follows:
[0045]
[0046] Among them, τ represents the neuron in the input layer. LELU adopts a self-gating mechanism, multiplying the original input by the output after non-linearly operating on this input. This activation function has a small and negative tail feature, which helps the efficient flow of information. At the same time, the function is continuously differentiable, eliminating singularities and ensuring the smooth propagation of gradients. The expression of LELU can be further modified as:
[0047]
[0048] Among them, softplus(τ) = ln(1 + e τ ), sigmoid(τ) = 1 / (1 + e -τ ). Therefore, LELU can be regarded as a variant of the existing activation function. Compared with the Leaky ReLU used in SCN, AINet, and ESNet, LELU provides greater flexibility in adapting to different scenarios.
[0049] (3) Construct the local network: The parameters and data processing of the local network model are similar to those of the global network. It should be noted that in this stage, only the parameters of the local network are optimized, and the weights of the global network are frozen to focus on the extraction and optimization of local features.
[0050] (4) Train the local network: As shown in Figure 1 , the local network adopts an encoder-decoder architecture to extract features from the given image, obtaining image features at different scales; similarly, using the CNN encoder C le to obtain high-resolution features, which focus on the processing of one-fourth of the original image Q ∈ {X1, X2, X3, X4}, extracting fine-grained features; using a non-overlapping sliding window to obtain local image sequence blocks P, thereby enabling the Mamba encoder M le to extract low-resolution features. The calculation formula of the combined local encoder CM le is as follows:
[0051]
[0052] Among them, represents the low-level features output by C le and represents the high-level features output by M le The decoding process is similar to that of the decoder C of the global network gd and the calculation formula is as follows:
[0053] F L = C ld {CM le}} = C ld {{C le (Q)}, {M le (P)}} (6)
[0054] Among them, F L represents the fine-grained local features obtained through the integration of CNN and Mamba.
[0055] In addition, during the training of the local network, the same as the global network, LELU is used as the activation function.
[0056] (5) Global-local feature fusion: The feature fusion unit (FFU) is as Figure 4 shown, which is used to fuse the coarse-grained global features and the fine-grained local features, so as to enhance the global perception ability of local clues and at the same time enhance the local understanding of global attributes. The calculation formula of FFU is as follows:
[0057] FFU = ω1(PAB(F G )·F L + PAB(F G )) + ω2(PAB(F L )·F G + PAB(F L )) (7)
[0058] Among them, ω1 and ω2 are learnable parameters; F G and F L are the feature information obtained by the global network and the local network; PAB (Pointwise Activation Block) is as Figure 4 shown in, which consists of two layers of 1×1 convolutions and one layer of activation function.
[0059] (6) Generating superpixels: Input the test set images into the trained SSMamba network model to generate superpixels, and finally output the pixel segmentation results.
[0060] The effects of the present invention can be further illustrated by the following experiments:
[0061] To test the effectiveness and superiority of the superpixel segmentation of the present invention, the hardware experimental platform is Intel Platinum 8255C, 2.5 GHz, 32 GB of memory, NVIDIA GTX 3090 GPU, and the software platform is PyTorch. We compared SSMamba with traditional superpixel algorithms (including ERS (Entropy Rate Superpixel Segmentation), SLIC (Simple Linear Iterative Clustering), and SNIC (Simple Non-Iterative Clustering)) and deep superpixel algorithms (including SEAL, SSN, SCN, AINet, ESNet, and CDS (Content Disentangle Superpixel)). All methods use their official open-source codes. To evaluate the quality of the superpixels, four metrics are adopted in all datasets in the present invention: segmentation accuracy (Achievable Segmentation Accuracy, ASA), boundary recall-precision (Boundary Recall-Precision, BR-BP), under-segmentation error (Under-Segmentation Error, UE), and compactness (Compactness, CO). Among these evaluation metrics, the higher the ASA, BR-BP, and CO values, and the lower the UE value, the better the superpixel segmentation performance.
[0062] Figure 5 Shows the quantitative performance comparison of different superpixel segmentation algorithms on four datasets. With the help of the Mamba structure and activation function, SSMamba proposed by the present invention is significantly superior to other algorithms in key metrics such as ASA, BR-BP, and UE, fully demonstrating its excellent superpixel segmentation quality. In addition, in terms of the CO metric, SSMamba is also highly competitive, second only to the CDS method, indicating that the superpixel regions generated by the present invention have good connectivity and better boundary integrity than most of the comparison algorithms. Generally speaking, SSMamba combines the advantages of global context modeling and local feature perception, significantly improving the accuracy and coherence of superpixel partitioning. Experimental results on multiple datasets show that compared with other advanced methods, SSMamba can maintain high accuracy in different scenarios and datasets, effectively suppress overfitting, and ensure reliability and consistency in real applications.
[0063] Figure 6Shows the qualitative comparison results of ten superpixel segmentation algorithms, where (a) and (b) are the test image and its corresponding ground truth label respectively; (c)-(l) correspond to the superpixel segmentation results of ERS, SLIC, SNIC, SEAL, SSN, SCN, AINet, ESNet, CDS, and SSMamba proposed in the present invention respectively. From the visual effect, compared with other methods, SSMamba performs excellently in maintaining the object contour, boundary clarity, and regional consistency, and can generate more accurate superpixel regions that conform to the object structure. Further verifies the effectiveness and superiority of SSMamba in the superpixel segmentation task. Especially on the DRIVE dataset, this method shows higher accuracy and stronger robustness, which makes it show a broader application prospect in complex scenarios such as medical images, highlighting the potential value of the present invention in practical applications.
[0064] Table 1
[0065] Activation function ASA BR BP UE CO ReLU 0.9649 0.8335 0.1352 0.0694 0.3846 Leaky ReLU 0.9649 0.8358 0.1354 0.0695 0.3790 SiLU 0.9662 0.8416 0.1358 0.0668 0.3698 Mish 0.9655 0.8350 0.1383 0.0683 0.3972 LELU 0.9668 0.8461 0.1364 0.0657 0.3702
[0066] In addition, Table 1 shows the average performance comparison of different activation functions on BSDS500, where the range of the number of superpixels is 50-1900. It can be seen from Table 1 that LELU obtains 3 optimal indicators and Mish obtains 2 optimal indicators, further illustrating the effectiveness of the activation function of the present invention for superpixel results.
[0067] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0068] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0069] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A superpixel segmentation method based on the Mamba state space model, characterized in that It includes the following steps: Use a two-stage training strategy to capture the coarse-grained and fine-grained features of the image. Specifically, in the first stage, a global network is adopted. With the help of image cropping operation, the global features are captured by the Mamba module. In the second stage, a local network is adopted. With the help of non-overlapping sliding window sampling operation, the Mamba module is prompted to extract local features; Obtain the pixel segmentation result by fusing the global feature and the local feature.
2. The superpixel segmentation method based on the Mamba state space model according to claim 1, characterized in that: In the first stage, the original image is processed by CNN, and the cropped image sequence is processed by the Mamba module.
3. The superpixel segmentation method based on the Mamba state space model according to claim 2, characterized in that: In the first stage, the CNN encoder consists of a convolutional block, an activation function, and a pooling layer. Among them, the convolutional block uses 3×3 convolution, the activation function uses LELU, and the pooling layer is max pooling; the encoder of the Mamba module adopts the VMamba framework.
4. The superpixel segmentation method based on the Mamba state space model according to claim 3, characterized in that: In the first stage, the CNN decoder consists of two convolutional blocks. Each convolutional block uses 3×3 convolution and includes the activation function LELU and a batch normalization layer.
5. The superpixel segmentation method based on the Mamba state space model according to claim 4, characterized in that: In the second stage, the image of one-fourth size is processed by CNN, and the image sequence of non-overlapping sliding window sampling is processed by the Mamba module.
6. The superpixel segmentation method based on the Mamba state space model according to claim 5, characterized in that: When performing feature fusion, through the feature fusion unit FFU, the coarse-grained global feature and the fine-grained local feature are fused to enhance the global perception ability of local clues and at the same time enhance the local understanding of global attributes.
7. The superpixel segmentation method based on the Mamba state space model according to claim 6, characterized in that: When fusing the coarse-grained global feature and the fine-grained local feature, the feature information of the global network and the local network is obtained through the PAB module for feature fusion. Among them, the PAB module consists of two layers of 1×1 convolution and one layer of activation function.
8. A superpixel segmentation system based on the Mamba state space model, characterized in that, It includes: A feature capture module for using a two-stage training strategy to capture the coarse-grained and fine-grained features of the image. Specifically, in the first stage, a global network is adopted. With the help of image cropping operation, the global features are captured by the Mamba module. In the second stage, a local network is adopted. Through non-overlapping sliding window sampling operation, the Mamba module is prompted to extract local features; A feature fusion module for obtaining the pixel segmentation result by fusing the global feature and the local feature.
Citation Information
Patent Citations
Federal learning-based frequency domain adaptive CT pulmonary nodule image segmentation method
CN119579612A
Remote sensing image semantic segmentation method based on double-branch multi-scale fusion network
CN119579891A
Cited By
Photoacoustic image enhancement method and device combining Mama and CNN (Convolutional Neural Network)
CN120707580A