Medical image segmentation system and method
Through the medical image segmentation system combining KAN and SEM modules, the interpretability and diagnostic reliability of U-Net variants are solved, and high-precision and efficient medical image segmentation are achieved, which is suitable for the precise segmentation of complex medical images.
Patent Information
- Application Number
- CN202510282972.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing U-Net variants face explanatory and diagnostic reliability issues in kernel design and black box properties, affecting the performance and diagnostic effects of the model.
Combining the Kolmogorov-Arnold network (KAN) and selective scanning high-efficiency multi-scale (SEM) attention module, it enhances feature extraction and multi-scale feature fusion through encoder, decoder and jump connection, and uses KAN's high-dimensional mapping and SEM's dynamic attention mechanism to capture complex nonlinear features.
It significantly improves the accuracy and robustness of medical image segmentation, especially the segmentation effect of small targets and complex boundary areas. It is suitable for a variety of medical scenarios and provides efficient segmentation performance and clinical auxiliary tools.
Smart Images

Figure CN120259322A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to, but is not limited to, the technical field of medical image segmentation, and particularly relates to a medical image segmentation system and method. Background Art
[0002] In the past decade, significant progress has been made in medical image segmentation methods to meet the needs of computer-aided diagnosis and image-guided surgery systems. U-Net is a milestone in this field, demonstrating for the first time the effectiveness of encoder-decoder convolutional networks with skip connections in medical image segmentation. Subsequently, a series of improvements to U-Net have emerged, such as U-Net++, 3D U-Net, V-Net, etc., further expanding its application scope and performance. In addition, hybrid architectures like U-NeXt combine convolutional operations with MLP to improve the efficiency of the segmentation network, especially in resource-constrained environments.
[0003] At the same time, Transformer-based networks have attracted attention due to their ability to model long-range dependencies and global context. For example, TransUNet and Swin-UNet utilize Vision Transformers (ViT) and Swin Transformers, demonstrating powerful global modeling capabilities. However, despite their strong performance, these models require a large amount of data and computational resources and face challenges in cases of limited data or real-time processing. In recent years, structured state-space models (SSMs) have provided effective solutions for modeling long sequences with linear computational complexity. Models such as U-Mamba and SegMamba have shown performance improvements by combining SSMs with CNNs for medical image segmentation.
[0004] In view of the above analysis, the technical problems urgently needed to be solved in the prior art are as follows:
[0005] Existing U-Net variants still face fundamental challenges in terms of kernel design and black-box properties, which affect the interpretability and diagnostic reliability of the model. Summary of the Invention
[0006] Aiming at the problems existing in the prior art, the present invention provides a medical image segmentation system.
[0007] The present invention is implemented as follows. A medical image segmentation system, characterized in that the medical image segmentation system integrates a Kolmogorov - Arnold network (KAN) and a Selective - Scan Efficient Multi - scale (SEM) attention module. The system consists of an encoder, a decoder, and skip connections. The encoder and decoder are respectively composed of a convolutional module, a Selective - Scan Efficient Multi - scale (SEM) attention module, and a Symbolic Kolmogorov - Arnold network module (Tok - KAN module).
[0008] Further, for the encoder, the input image first undergoes three convolutional operations combined with the SEM module, and then is further processed through two symbolic MLP blocks.
[0009] Further, the decoder is composed of two symbolic KAN modules and three convolutional modules. Each encoder module halves the feature resolution, while each decoder module doubles the feature resolution.
[0010] Further, in the system, the skip connections are integrated between the encoder and the decoder, and simple addition is used to highlight the segmentation performance of the pure SSM model.
[0011] Further, the KAN can be expressed as:
[0012]
[0013] Where each layer Φ i consists of n in ×n out learnable activation functions φ:
[0014] Φ = {φ q,p}, p = 1, 2,..., n in , q = 1, 2,..., n out
[0015] The output from the k - th layer to the k + 1 - th layer can be represented in matrix form:
[0016] Z k+1 = Φ k Z k
[0017] Integrating the KAN as a bottleneck layer into the U - Net architecture, given the input feature map Z ∈ R C×H×W from the encoder, the KAN layer processes this feature as follows:
[0018] Z′ = LN(Z + DwConv(Φ(Z)))
[0019] where LN represents the normalization layer and DwConv represents the depth convolution.
[0020] Furthermore, the SEM attention module consists of two main parts: a feature extraction part and an attention mechanism part. In the feature extraction part, the improved scanning method not only includes the standard directions from top - left to bottom - right and from bottom - right to top - left, but also introduces an adaptive scanning strategy. The expressions for each scanning direction are as follows:
[0021] X (dir) = Scan(X, direction)
[0022] The directions include top - left to bottom - right, top - right to bottom - left, bottom - right to top - left, and bottom - left to top - right. The scanning results in each direction are subjected to feature extraction through the improved S6 module, where the S6 module dynamically adjusts parameters according to the input, filters and retains the most important features. Subsequently, a re - weighting operation sums and combines the sequences in the four directions to restore the output image to the same size as the input. The final feature integration process combines features from different directions as follows:
[0023]
[0024] The multi - scale module captures multi - scale spatial information using parallel 1x1 and 3x3 convolutions. For a given input X, the module divides it into G channel groups:
[0025]
[0026] The 1x1 branch uses global average pooling (GAP) to encode cross - channel dependencies and generates a channel attention map:
[0027]
[0028] where σ represents the Sigmoid function. The 3x3 branch has a larger kernel and captures local spatial interactions:
[0029]
[0030] The final output X″ of the EMA module aggregates these attention maps:
[0031]
[0032] Furthermore, the KAN bottleneck and the SEM module are combined to capture complex spatial dependencies and multi - scale information, enhancing the interpretability of feature representations:
[0033]
[0034] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:
[0035] The KM-UNet of the present invention is the first to explore the combination of the KANSSM model for medical image segmentation.
[0036] The present invention introduces a selective scanning efficient multi-scale (SEM) attention module, which can learn across space and integrate the multi-scale attention of the S6 module.
[0037] The present invention establishes a benchmark for the integration of the KAN neural network and the SSM model in the medical image segmentation task, providing valuable insights for the development of more efficient and effective KAN-based segmentation methods.
[0038] Experiments conducted on five datasets demonstrate that the present invention is highly competitive.
[0039] Traditional medical image segmentation methods often fail to accurately identify the boundaries or details of the target region due to limited feature extraction capabilities when dealing with complex scenarios (such as liver tumors, stroke regions, etc.), resulting in inaccurate segmentation results. By combining the Kolmogorov-Arnold network (KAN) and the selective scanning efficient multi-scale (SEM) attention module, the present invention enhances the system's ability to extract multi-scale features and performs high-dimensional symbolic processing on complex non-linear features, significantly improving the segmentation accuracy, especially in the segmentation of small targets or regions with weak signals.
[0040] Existing segmentation techniques usually lack an effective multi-scale feature fusion mechanism when dealing with medical images, making it difficult to effectively retain key features at different resolutions. The present invention fuses the high-resolution features extracted by the encoder with the low-resolution features of the decoder through skip connections to ensure the transmission and retention of important information during the segmentation process. At the same time, the SEM module further enhances the feature expression ability of key regions through a dynamic attention mechanism, fundamentally solving the problems of multi-scale feature fusion and feature loss.
[0041] There are often a large number of non-linear features in complex medical images, and existing segmentation methods are difficult to fully capture these features, resulting in poor adaptability of the segmentation results to regions with complex shapes. The present invention uses the symbolic Kolmogorov-Arnold network module (Tok-KAN module) to perform high-dimensional mapping and symbolic processing on features, significantly enhancing the model's ability to capture non-linear features and enabling it to better adapt to the diverse requirements of medical image segmentation.
[0042] Through multi-module collaborative optimization, the present invention not only significantly improves the segmentation accuracy and robustness, but also makes breakthroughs in efficiency and applicability. Compared with the prior art, the system of the present invention is more accurate in segmenting small targets, complex boundaries and low-signal regions, and is applicable to a variety of medical scenarios (such as liver tumor segmentation, stroke area recognition). In addition, the efficient segmentation performance of the system can significantly shorten the segmentation time, providing a high-quality auxiliary tool for clinical diagnosis and treatment planning, and showing broad industrial application prospects in the field of medical image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is the architecture diagram of the medical image segmentation system provided by the embodiment of the present invention;
[0044] Figure 2 is the schematic diagram of the SEM module provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0046] As Figure 1 shown, the embodiment of the present invention provides a medical image segmentation system (KM-UNet), which integrates a Kolmogorov-Arnold network (KAN) and a Selective-Scan Efficient Multi-scale (SEM) attention module. KM-UNet consists of an encoder, a decoder and skip connections. The encoder and decoder are respectively composed of a convolutional module, a Selective-Scan Efficient Multi-scale (SEM) attention module, and a symbolic Kolmogorov-Arnold network module (Tok-KAN module).
[0047] For the encoder, the input image first undergoes three convolutional operations combined with the SEM module, and then is further processed through two symbolic MLP blocks.
[0048] For the decoder, it consists of two symbolic KAN modules and three convolutional modules. Each encoder module halves the feature resolution, while each decoder module doubles the feature resolution.
[0049] In the system, skip connections are integrated between the encoder and the decoder to facilitate feature fusion. Simple addition is used to highlight the segmentation performance of the pure SSM model.
[0050] The described Kolmogorov - Arnold networks (KANs) aim to address the limitations of multi - layer perceptrons (MLPs), especially in terms of interpretability in parameter efficiency. A KAN can be represented as:
[0051]
[0052] where each layer Φ i consists of n in × n out learnable activation functions φ:
[0053] Φ = {φ q,p}, p = 1, 2,..., n in , q = 1, 2,..., n out The output from the k - th layer to the (k + 1) - th layer can be represented in matrix form:
[0054] Z k+1 = Φ k Z k
[0055] Different from traditional MLPs, KANs do not rely on a linear transformation matrix, which enables them to achieve comparable or better performance with fewer parameters. This architecture enhances the interpretability of the model and makes it suitable for applications that require understanding.
[0056] In an embodiment of the present invention, KANs are integrated as a bottleneck layer into the U - Net architecture to improve feature modeling and enhance the interpretability of the network. Given an input feature map Z ∈ R C×H×W from the encoder, the KAN layer processes this feature as follows:
[0057] Z′ = LN(Z + DwConv(Φ(Z)))
[0058] where LN represents the normalization layer and DwConv represents depth convolution. This enables the network to capture complex non - linear dependencies between features while maintaining the efficiency of parameter usage.
[0059] The Selective Scan Efficient Multi - Scale (SEM) attention module is designed to enhance the feature extraction and attention mechanism in both the encoder and decoder stages of the U - Net network. The SEM attention module consists of two main parts: a feature extraction part and an attention mechanism part. As Figure 2As shown, the feature extraction part draws inspiration from the SS2D module and improves its scanning method to achieve more efficient feature capture and multi-directional feature extraction. Specifically, the input feature map is unfolded into sequences along multiple directions to capture rich spatial information. The improved scanning method not only includes the standard directions from top left to bottom right and from bottom right to top left, but also introduces an adaptive scanning strategy to flexibly adjust the scanning order to adapt to different input features. The expression for each scanning direction is as follows:
[0060] X (dir) = Scan(X, direction)
[0061] The directions include top left to bottom right, top right to bottom left, bottom right to top left, and bottom left to top right. The scanning results in each direction are subjected to feature extraction through the improved S6 module, where the S6 module dynamically adjusts parameters according to the input, filters and retains the most important features. Subsequently, as Figure 2 shown, the reweighting operation sums and combines the sequences in four directions, restoring the output image to the same size as the input. The S6 module is derived from Mamba and introduces a selective mechanism based on S4 to achieve this by adjusting the parameters of the SSM. This enables the model to distinguish and retain relevant information while filtering out irrelevant information. The pseudocode of the S6 module is shown in Algorithm 1:
[0062]
[0063] The final feature integration process combines features from different directions as follows:
[0064]
[0065] The multi-scale module uses parallel 1x1 and 3x3 convolutions to capture multi-scale spatial information without reducing the channel dimension. For a given input X, the module divides it into G channel groups:
[0066]
[0067] The 1x1 branch uses global average pooling (GAP) to encode cross-channel dependencies and generate a channel attention map:
[0068]
[0069] where σ represents the Sigmoid function. The 3x3 branch has a larger kernel and captures local spatial interactions:
[0070]
[0071] The final output X″ of the EMA module aggregates these attention maps:
[0072]
[0073] To enhance global and local feature interactions, cross-space learning aggregates multi-scale information by combining the outputs of two branches, providing rich pixel-level attention to better understand semantic segmentation. The combination of the KAN bottleneck and the SEM module enables KM-UNet to efficiently capture complex spatial dependencies and multi-scale information, enhancing the interpretability of feature representations.
[0074]
[0075] The following is the detailed working principle of a medical image segmentation system based on the Kolmogorov-Arnold network (KAN) and the selective scanning efficient multi-scale attention module:
[0076] The medical image segmentation system consists of an encoder, a decoder, and skip connections. By combining the Kolmogorov-Arnold network (KAN) and the selective scanning efficient multi-scale (SEM) attention module, it achieves efficient segmentation of medical images. The system is designed in a hierarchical structure. The encoder is responsible for extracting multi-scale features from the input image, the decoder gradually restores the low-resolution features to a high-resolution segmentation result, and the skip connections are used to transfer key feature information between the encoder and the decoder.
[0077] In the encoder part, the input image first undergoes three convolutional operations combined with the SEM module to extract multi-scale local features. The SEM module enhances the representation ability of key region features through a selective scanning mechanism and an efficient attention strategy, thereby improving the capture accuracy of details in complex medical images. After completing the preliminary feature extraction, the encoder is further processed by two symbolic MLP blocks (Tok-KAN modules), which use the high-dimensional mapping ability of KAN to perform symbolic encoding on the features, enhancing the model's ability to capture non-linear features.
[0078] The decoder part consists of two symbolic KAN modules and three convolutional modules. The decoder restores the low-resolution features extracted by the encoder to high-resolution features through step-by-step decoding. Each decoder module performs high-dimensional feature decoding through the symbolic operation of the KAN module and combines convolutional modules to gradually double the feature resolution, thereby restoring a segmentation result with the same resolution as the input image.
[0079] The skip connection is integrated between the encoder and the decoder and is used for feature fusion between features of different resolutions. Through a simple addition operation, the high-resolution features extracted from the encoder are directly fused with the low-resolution features of the decoder, thereby highlighting the system's ability to retain detailed features. The design of the skip connection ensures that the segmentation model can still maintain high performance with a simple structure, while enhancing the model's adaptability to complex medical image segmentation tasks.
[0080] The Tok-KAN module is one of the core components of the system and is used for the feature symbolization processing of the encoder and the decoder. Based on the Kolmogorov-Arnold representation theory, this module maps the input features to a high-dimensional space through a symbolized MLP block, strengthening the model's expressive ability. The symbolization processing can effectively capture the non-linear relationships in medical images and improve the accuracy and robustness of the segmentation results.
[0081] The system enhances the attention to the details of medical images through the SEM module, improves the processing ability of non-linear features through the Tok-KAN module, and realizes efficient feature fusion through the skip connection, optimizing the segmentation performance as a whole. While maintaining high computational efficiency, the system can accurately segment small targets and subtle features in complex medical images, providing a high-quality auxiliary tool for medical diagnosis and treatment, and demonstrating significant advantages in the field of medical image processing.
[0082] In summary, this medical image segmentation system combines advanced attention mechanisms and symbolized feature representations to construct an efficient and robust image segmentation solution suitable for a variety of complex medical scenarios.
[0083] The embodiments of the present invention were extensively evaluated using three different and heterogeneous datasets, each with unique attributes, different data volumes, and different image reconstruction solutions. These datasets are commonly used in tasks such as image segmentation and generation, providing a comprehensive test platform for evaluating the effectiveness and adaptability of our method.
[0084] BUSI dataset: The BUSI dataset contains ultrasound images covering normal, benign, and malignant breast cancer cases, as well as corresponding segmentation maps. In this study, a subset of 647 ultrasound images was used, all of which represent benign and malignant breast tumors and were uniformly adjusted to 256×256 pixels. This dataset provides a comprehensive collection of images, which helps to detect and distinguish various types of breast tumors and provides valuable insights for medical professionals and researchers.
[0085] GlaS dataset: The GlaS dataset consists of 612 standard definition (SD) frames derived from 31 sequences, with each frame having a resolution of 384×288 pixels. These images are from 23 patients and were collected at the Hospital Clínic in Barcelona, Spain. The data was recorded using devices such as the Olympus Q 160AL and Q 165L, along with an ExtraII video processor. According to the established protocol, we selected 165 images from this dataset and uniformly adjusted them to 512×512 pixels for evaluation.
[0086] CVC-ClinicDB dataset: A resource commonly referred to as "CVC", is a publicly accessible resource widely used for colonoscopy video polyp detection. It contains 612 images with a resolution of 384×288 pixels, which are from 31 different colonoscopy sequences. This dataset provides a diverse representation of polyp instances and is particularly valuable for the development and evaluation of polyp detection algorithms. To ensure the consistency of all datasets in this study, we adjusted all images in the CVC-ClinicDB dataset to 256×256 pixels.
[0087] ISIC17 and ISIC18 datasets: The International Skin Imaging Collaboration 2017 and 2018 Challenge datasets (ISIC17 and ISIC18) are publicly available skin lesion segmentation datasets, containing 2,150 and 2,694 dermoscopic images and their corresponding segmentation masks respectively. We divided the datasets into training and test sets in a 7:3 ratio. Specifically, the ISIC 17 dataset was divided into a training set with 1,500 images and a test set with 650 images, while the ISIC 18 dataset contains a training set of 1,886 images and a test set of 808 images. For both datasets, we used metrics such as the mean intersection over union (mIoU) and Dice similarity coefficient (DSC) for a detailed evaluation to assess the performance of our method.
[0088] The present invention implements KM-UNet using PyTorch on NVIDIA RTX 4090 GPU. For Bushi, GlaS, CVCCliniCDB, ISIC17 and ISIC18 datasets, we set the batch size to 8 and use an initial learning rate of 1e-4, following a cosine annealing learning rate schedule with a minimum learning rate of 1e-5. The Adam optimizer is used for training, and the loss function combines binary cross entropy (BCE) and Dice loss to enhance performance. Each dataset is randomly divided into 80% for training and 20% for validation. The model was trained for a total of 300 epochs, and all results were averaged over three independent runs to ensure reliability. Only basic data augmentation was applied, including random rotation and flipping. We use metrics such as average intersection over union (IoU) and Dice similarity coefficient (F1 score) to qualitatively and quantitatively evaluate the segmentation results.
[0089] In comparative experiments, we evaluate the performance of KM-UNet and other existing methods on image segmentation tasks using two standard performance metrics: Intersection over Union (IoU) and F1 score. These metrics are used to evaluate the accuracy and robustness of segmentation models on different datasets.
[0090] We compare KM-UNet with several recent popular medical image segmentation methods, including classic convolutional neural network models such as U-Net and UNet++, and attention-based models such as AttUNet. In addition, we compare it with other efficient transformer variants such as U-Mamb a, and advanced MLP segmentation networks based on U-NeXt such as U-NeXt and Rolling-UNet.
[0091] Table 1: Comparison with state-of-the-art segmentation models on three heterogeneous medical scenarios.
[0092]
[0093] Table 2: Comparative experimental results on the ISIC17andISIC18dataset.(Boldindicates the best.)
[0094]
[0095] The results show that the embodiment of the present invention (KM-UNet) outperforms most of the comparison methods on all datasets, with significant improvements in both IoU and F1 score.
[0096] Compared with the traditional U-Net and UNet++ models, KMUNet performs excellently in capturing global context information, thus improving the segmentation accuracy and better retaining details. Compared with attention mechanism-based models such as AttUNet, KMUNet also performs outstandingly, especially when dealing with complex shapes and boundaries, reducing the problems of over-segmentation and under-segmentation.
[0097] Comparisons with efficient transformers and MLP-based models show that KM-UNet maintains strong segmentation capabilities, providing finer details in boundary segmentation, especially in challenging and diverse medical image datasets. This indicates that KM-UNet not only improves the segmentation performance but also significantly reduces false positives and false negatives in more complex situations.
[0098] Overall, KM-UNet shows superior performance in the image segmentation task, outperforming most existing models in both IoU and F1 score, highlighting its effectiveness and potential in medical image segmentation.
[0099] Example 1: Medical Image Segmentation System for Liver Tumors
[0100] The accurate segmentation of liver tumors is a key task in medical diagnosis and treatment planning. Due to the complex shape, blurred boundary and uneven distribution of liver tumors, traditional segmentation algorithms are difficult to meet the requirements of high-precision segmentation. In this example, the medical image segmentation system of the present invention is used to accurately segment the liver and tumor regions.
[0101] 1) Data preparation: The abdominal CT scan images of the patient are input into the system, and the encoder performs preliminary processing on the images.
[0102] 2) Multi-scale feature extraction: By combining the convolutional operations of the SEM module, multi-scale features of the liver and tumor regions are extracted to enhance the capture of tumor boundary details.
[0103] 3) Symbolic processing: The encoder performs high-dimensional mapping and symbolic processing on the features through the Tok-KAN module to improve the expression ability of complex non-linear features.
[0104] 4) Output of segmentation results: The decoder gradually restores the high-resolution features by combining skip connections, and compares the segmentation results with the original CT images to ensure that the segmentation accuracy meets medical requirements.
[0105] 5) Verification and application: After the segmentation results are verified by medical experts, they can be used for tumor volume measurement, lesion localization and surgical planning.
[0106] High precision: The system's ability to capture details of the liver and tumor regions is significantly enhanced, and the segmentation results are more in line with the actual morphology.
[0107] High efficiency: Multi-module collaborative optimization reduces the segmentation time and is suitable for the rapid clinical diagnosis requirements.
[0108] Example 2: Stroke Region Segmentation and Evaluation
[0109] The image analysis of stroke patients requires accurate segmentation of the damaged area to evaluate the degree of injury and formulate treatment plans. Stroke regions usually have characteristics such as irregular shapes and weak signals, and traditional segmentation methods are difficult to meet the requirements. In this example, the system of the present invention is used to segment and evaluate the stroke region.
[0110] 1) Data input: Input the brain MRI image of the patient, and the system automatically performs segmentation processing.
[0111] 2) Feature extraction and symbolization: The encoder combines with the SEM module to extract multi-scale features of the stroke region, and the Tok-KAN module symbolically encodes the features to enhance the recognition ability of weak signal regions.
[0112] 3) Skip connection fusion: The high-resolution features extracted by the encoder are fused with the features of the decoder through skip connections to further improve the segmentation accuracy.
[0113] 4) Segmentation output and visualization: The decoder gradually restores the features and generates the segmentation results, and the segmented image intuitively shows the location, size and shape of the stroke.
[0114] 5) Clinical application: Combine the segmentation results to calculate the stroke area and evaluate the treatment, such as formulating the dosage of thrombolytic drugs.
[0115] Strong robustness: The system shows excellent segmentation performance for stroke regions with complex shapes and weak signals.
[0116] High clinical value: The provided segmentation results provide high-quality auxiliary information for doctors' decision-making, improving the diagnostic efficiency and treatment effect.
[0117] These two examples respectively demonstrate the wide applicability and technical advantages of the present invention in complex medical image segmentation tasks, meeting the actual needs of precision medicine.
[0118] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or included in processor control code, such as provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.
[0119] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall all be covered by the protection scope of the present invention.
Claims
1. A medical image segmentation system, characterized in that, The system integrates the Kolmogorov - Arnold network (KAN) and the Selective Scanning Efficient Multi - scale (SEM) attention module. The system includes an encoder, a decoder, and skip connections. The encoder and decoder are respectively composed of a convolutional module, an SEM module, and a symbolized Kolmogorov - Arnold network module (Tok - KAN module), which are used to extract multi - scale features from the input image and perform step - by - step decoding to restore to the segmentation result. At the same time, feature fusion between the encoder and the decoder is achieved through skip connections.
2. The medical image segmentation system according to claim 1, wherein The encoder includes three convolutional modules combined with SEM modules, which are used to extract multi - scale local features, and two symbolized Tok - KAN modules, which are used to perform high - dimensional mapping and symbolic encoding processing on the extracted features.
3. The medical image segmentation system according to claim 1, wherein The decoder is composed of two symbolized Tok - KAN modules and three convolutional modules. Each decoder module doubles the feature resolution through symbolic operations and convolutional operations to gradually restore the high - resolution segmentation result.
4. The medical image segmentation system according to claim 1, wherein The skip connection uses a simple addition operation to fuse the high - resolution features output by the encoder with the low - resolution features of the decoder, so as to enhance the system's ability to retain detailed features.
5. The medical image segmentation system according to claim 1, wherein The SEM module adopts a selective scanning mechanism and an efficient multi - scale attention strategy, which is used to dynamically enhance the representation ability of key region features in the convolutional operation, thereby improving the capture accuracy of complex medical image details.
6. The medical image segmentation system according to claim 1, wherein The KAN can be expressed as: Among them, each layer Φ i consists of n in × n out learnable activation functions φ: Φ = {φ q,p}, p = 1, 2,..., n in , q = 1, 2,..., n out The output from the k - th layer to the k + 1 - th layer can be represented in matrix form: Z k+1 = Φ k Z k Integrate KAN as a bottleneck layer into the U-Net architecture. Given the input feature map $Z \in \mathbb{R}$ from the encoder, C×H×W the KAN layer processes this feature as follows: Z′ = LN(Z + DwConv(Φ(Z))) Where LN represents the normalization layer and DwConv represents the depth convolution.
7. The medical image segmentation system according to claim 1, characterized in that, The SEM attention module consists of two main parts: a feature extraction part and an attention mechanism part. In the feature extraction part, the improved scanning method not only includes the standard directions from top - left to bottom - right and from bottom - right to top - left, but also introduces an adaptive scanning strategy. The expression for each scanning direction is as follows: X (dir) = Scan(X, direction) The directions include top - left to bottom - right, top - right to bottom - left, bottom - right to top - left, and bottom - left to top - right. The scanning results in each direction are used for feature extraction through an improved S6 module, where the S6 module dynamically adjusts parameters according to the input, filters and retains the most important features. Subsequently, a re - weighting operation sums and combines the sequences in the four directions to restore the output image to the same size as the input. The final feature integration process combines features from different directions as follows: The multi - scale module uses parallel 1x1 and 3x3 convolutions to capture multi - scale spatial information. For a given input X, this module divides it into G channel groups: The 1x1 branch uses global average pooling (GAP) to encode cross - channel dependencies and generate a channel attention map: Where σ represents the Sigmoid function. The 3x3 branch has a larger kernel and captures local spatial interactions: The final output X″ of the EMA module aggregates these attention maps:
8. The medical image segmentation system according to claim 1, wherein The KAN bottleneck and the SEM module are combined to capture complex spatial dependencies and multi - scale information, enhancing the interpretability of feature representation:
9. A processing method based on medical image segmentation, characterized in that, Including the following steps: Preprocess the input medical image and input it into a medical image segmentation system integrated with a selective scanning efficient multi-scale attention module and a symbolic Kolmogorov-Amold network (KAN) module; In the encoder stage, extract features by combining the convolutional operations of the SEM module, and the convolutional operations include three convolutional modules and two symbolic multi-layer perceptron (MLP) blocks; In the decoder stage, use two symbolic KAN modules and three convolutional modules to gradually decode the extracted features, and each decoder module doubles the feature resolution; Transfer features between the encoder and the decoder through skip connections, and highlight the segmentation performance through addition operations.
10. The medical image segmentation method according to claim 9, wherein Adopt an improved scanning strategy to enhance the attention mechanism, specifically including: Use four scanning directions, namely from top left to bottom right, from top right to bottom left, from bottom right to top left, and from bottom left to top right, to perform multi-directional scanning on the input features; Extract the most important features using the S6 module with dynamically adjusted parameters in each direction, and merge the scanning results in each direction through reweighting operations to restore a feature map of the same size as the input; The multi-scale module further uses parallel 1×1 convolutions and 3×3 convolutions to capture multi-scale spatial information. The 1×1 convolution branch generates a channel attention map through global average pooling, and the 3×3 convolution branch captures local spatial interaction information; Aggregate these attention maps and output the final image characteristics with enhanced feature representations.
Citation Information
Cited By
Multi-expert collaborative medical image segmentation method based on multi-scale information fusion
CN121482404A
Multi-expert collaborative medical image segmentation method based on multi-scale information fusion
CN121482404B