Brain tumor region segmentation method based on adaptive boundary guided aggregation SAM

ABF-SAM addresses the challenges of complex brain tumor segmentation by integrating multi-modal MRI fusion and boundary detection, enhancing boundary robustness and segmentation accuracy through adaptive augmentation and dynamic fusion.

CN120318246APending Publication Date: 2025-07-15CHINA UNIV OF MINING & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510359476.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate segmentation due to modal differences and complex boundaries in brain tumor MRI image segmentation, especially in the absence of segmentation accuracy in noise and occlusion.

Method used

Adaptive boundary guidance aggregation SAM method is adopted to mine brain tumor boundary characteristics through boundary mining networks, combine dynamic fusion modules and adaptive boundary enhancement models, dynamically adjust information fusion weights, and improve segmentation accuracy.

Benefits of technology

The segmentation accuracy is improved under complex background and blurred boundaries, the robustness of the model is enhanced, and the segmentation performance of brain tumor areas is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318246A_ABST
    Figure CN120318246A_ABST
Patent Text Reader

Abstract

The invention discloses a brain tumor region segmentation method based on adaptive boundary guided aggregation SAM, and relates to the field of medical image analysis. The fusion output of the multi-modal magnetic resonance image is input into an image segmentation model to obtain image embedding, the boundary embedding of the brain tumor and the real boundary of the tumor are obtained through a boundary mining network, the boundary features are enhanced through an adaptive boundary enhancement model, the image embedding and the boundary embedding are output as fusion embedding by applying dynamic fusion, and the real boundary of the tumor is obtained. Inputting the fused embedded and adaptive enhanced boundary features and the features of the prompt encoder into a mask decoder of a segmentation model to obtain a segmentation result, and obtaining a total loss function in combination with a boundary mining network and a loss function based on region prediction; according to the brain tumor region segmentation method based on adaptive boundary guided aggregation SAM, accurate segmentation of the brain tumor region is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image analysis, and particularly to a method for segmenting brain tumor regions based on adaptive boundary-guided aggregation SAM. Background Art

[0002] Brain tumors are a type of central nervous system tumor with highly invasive and malignant characteristics, and are one of the main diseases causing death and neurological dysfunction globally. In clinical practice, magnetic resonance imaging (MRI) serves as the main detection and evaluation method for brain tumors, laying the foundation for the segmentation of brain tumor regions. The accurate segmentation of these regions is crucial for diagnosis, treatment planning, and prognosis evaluation. Especially in radiotherapy, accurate region segmentation can ensure that the radiation dose is highly concentrated in the tumor tissue while minimizing damage to normal brain tissue.

[0003] However, due to the significant modality differences between medical images and natural images, directly applying SAM (Segment anything model) to medical image segmentation yields unsatisfactory results, and generally only achieves good results in some modalities of medical image data. Therefore, although large models have powerful feature extraction capabilities, for the unique complexity of medical images, especially the brain tumor segmentation task, specific domain adaptation designs and optimization strategies need to be combined to fully unleash the potential of large models. In addition, brain tumors have complex morphologies and irregular boundaries in MRI images, and the heterogeneity of tumor sub-regions increases the segmentation difficulty. Therefore, accurately capturing these complex boundaries is the key to improving segmentation performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for segmenting brain tumor regions based on adaptive boundary-guided aggregation SAM, construct a framework of adaptive boundary-guided fusion SAM (ABF-SAM), achieve accurate segmentation of brain tumor regions, and enhance the robustness of the model to variable boundaries through ABAM.

[0005] To achieve the above purpose, the present invention provides a method for segmenting brain tumor regions based on adaptive boundary-guided aggregation SAM, including the following steps:

[0006] S1. Input the fusion output I of multimodal magnetic resonance images I4 fusion into the image segmentation model SAM to obtain the image embedding E sam ;

[0007] S2. Input the boundary-sensitive modal image I3 into the boundary mining network BDNet to mine the boundary features of each region of the brain tumor, and obtain the boundary embedding E of the brain tumor boundary and the true tumor boundary P b ;

[0008] S3. Enhance the boundary features through the adaptive boundary enhancement model ABAM and then provide them to the SAM model;

[0009] S4. Apply dynamic fusion DFM to reorganize the channel information and spatial information to dynamically fuse the image embedding E in S2 sam and the boundary embedding E in S3 boundary into the fusion embedding E DFM ;

[0010] S5. Input the fusion embedding E obtained in S4 DFM , the adaptively enhanced boundary features obtained in S3, and the features of the prompt encoder in the SAM model into the mask decoder in the SAM model to generate the segmentation results of each region of the brain tumor;

[0011] S6. Combine the boundary mining network and the region prediction-based loss function to obtain the total loss function of the SAM model.

[0012] Preferably, the process of obtaining the fusion output I fusion in S1 is as follows:

[0013] I fusion = MHA(Conv(I4))

[0014] where Conv(·) represents the convolution operation, MHA(·) is the multi-head self-attention mechanism, I4 ∈ R 4×H×W , H and W respectively represent the height and width of the image, and R represents the set of multimodal magnetic resonance images.

[0015] Preferably, the specific process in S2 is as follows:

[0016] S21. The boundary mining network BDNet uses a residual structure to build the network as the backbone network, and uses the boundary-sensitive modal image I3 as the input, where I3 ∈ Q 3×H×W , Q represents the set of boundary-sensitive modal images, and successively passes through four residual blocks encapsulated by residual connections. The output boundary clue of each residual block is where i ∈ [1, 2, 3, 4], C i ∈ [64, 128, 256, 512], B is the set of boundary clues, and the output b4 is upsampled once to obtain the boundary embedding E with a size of 256×16×16 boundary ;

[0017] S22. Further reconstruct the true boundary of the tumor through the guidance of the SAM model, and use the boundary clues of each residual block as the input of the boundary decoder, and layer by layer implement feature splicing during the decoding process, and finally decode to obtain the true boundary P of the tumor b .

[0018] Preferably, the process of enhancing the boundary features in S3 is as follows:

[0019] S31. Downsample and splice the boundary clues to form a comprehensive feature map, and use the channel attention mechanism to adjust the weights of the feature map to obtain the feature representation F of multi-scale boundary clues init , and respectively consider the two aspects of the deformation and displacement of the tumor to adaptively enhance the boundary clue representation;

[0020] S32. In order to capture the deformation and size changes of the tumor, use deformable convolution to adaptively process F init , use max pooling to highlight the strongest response area, use average pooling to retain the global average information, and splice the results of the two poolings to obtain a more complete deformation representation F d :

[0021] F d = Concat[Max(Dc(F init ))), Mean(Dc(F init ))]

[0022] where Dc(·) represents deformable convolution, Concat(·) represents the splicing function, and Max(·) and Mean(·) are max pooling and average pooling respectively;

[0023] S33. Considering the variability of brain tumor localization, respectively perform independent processing on F init and its transpose , divide it into smaller patches and embed them, then perform displacements in multiple directions, weight the features through convolution with a 1×1 convolution kernel, and aggregate the information in multiple directions using matrix addition to obtain the displacement feature

[0024]

[0025] where C d (·) represents performing a convolution operation on the feature with a d-direction offset, D = {right, right-up, left, left-down}, representing the set of offset directions;

[0026] S34. Use the multi-layer perceptron mlp to increase the displacement feature Non - linear expression, and obtain the displacement offset representation F through splicing and aggregation s :

[0027]

[0028] S35. Adopt element - wise multiplication operation to spatially combine the deformation information F d and the displacement offset information F s :

[0029] F ABA =F d ⊙F s

[0030] where ⊙ represents element - wise multiplication, and F ABA is the boundary feature after adaptive enhancement.

[0031] Preferably, the dynamic fusion process in S4 is as follows:

[0032] S41. Use the sigmoid function to set the gating mechanism of the learnable parameter α to achieve the information complementarity of the adaptive weighting of the image embedding E sam and the boundary embedding E boundary :

[0033] E G1 =(1 - σ(α))·E boundary +σ(α)·E sam

[0034] where E G1 is the initial fusion embedding, and σ(α) is the sigmoid function with the learnable parameter α;

[0035] S42. For channel information recombination, input the initial fusion embedding E G1 into two consecutive depth - wise separable convolutions to obtain the channel - sensitive output, and add it to the initial fusion feature E G1 again to obtain the channel embedding after dynamic recombination

[0036]

[0037] where D represents the channel convolution operation, and P represents the point - wise convolution operation;

[0038] S43. For spatial information recombination, use E G1 and its transpose as inputs respectively, divide the inputs into smaller patches, use the channel and spatial attention module CBAM to dynamically adjust the feature weights, and use the feature map with the same dimension as the patch as the learnable mapping, and then update the parameters by the way of gradient descent. The specific calculation is as follows:

[0039]

[0040] Among them, pe(·) represents patch embedding, cbam(·) represents applying channel attention and spatial attention to the input feature map in sequence to generate a weighted feature map, and l1 and l1′ are learnable mappings;

[0041] S44. Set the learnable mappings l2 and l2′ to be dynamically adjusted, and obtain by combining two types of spatial information through feature concatenation The specific calculation is as follows:

[0042]

[0043] S45. Use the second gating mechanism with learnable parameter α′ to fuse the channel recombination information and the spatial recombination information to achieve the dynamic fusion of boundary embedding and image embedding:

[0044]

[0045] Among them, E DFM is the fusion embedding of DFM.

[0046] Preferably, the specific process of calculating the loss function in S6 is as follows:

[0047] S61. Use the Dice coefficient as the optimization function of BDNet, and calculate the loss function of BDNet as follows:

[0048]

[0049] Among them, P B represents the predicted tumor boundary, G B represents the ground truth label of the boundary mask, N is the total number of pixels in the image, and i is the i-th pixel in the image;

[0050] S62. Considering that there is a large gap in the proportion of brain tumor and normal brain tissue regions, a region prediction-based optimization function is formulated using the Dice loss function and the Focal loss function, and the loss function is as follows:

[0051]

[0052] Among them, y is the true value, is the prediction result, r is the set of different tumor sub-regions, and r ∈ {WT, TC, ET}, WT is the region mask of the whole tumor, TC is the region mask of the tumor core, and ET is the region mask of the enhanced tumor;

[0053] S63. The overall loss function based on adaptive boundary-guided aggregation SAM is:

[0054] L = (1 - γ)·L boundary + γ·L mask

[0055] Among them, the weight hyperparameter γ = 0.8.

[0056] Therefore, the brain tumor region segmentation method based on adaptive boundary-guided aggregation SAM of the present invention has the following beneficial effects compared with the prior art:

[0057] 1. The adaptive boundary-guided aggregation SAM model ABF-SAM in this application can more accurately identify the edges of the target through the adaptive boundary guidance mechanism, especially performing well in complex backgrounds or cases where the target boundaries are blurred. In the face of interference such as noise and occlusion, the model can still maintain a high segmentation accuracy, showing strong robustness;

[0058] 2. The boundary mining network applied in this application can more accurately locate the edges of the target, especially performing excellently in complex scenarios or cases where the boundaries are blurred. By mining boundary information, it can better retain the detailed features of the target, improve the recognition ability of fine structures in segmentation or detection tasks, and the boundary mining network can significantly improve the accuracy of segmentation, detection and other tasks, especially standing out in tasks that require high-precision boundary output;

[0059] 3. The dynamic fusion technical solution in this application can capture richer information through dynamic fusion of multi-source or multi-level features. The DFM module can significantly improve the accuracy of classification, detection or segmentation and other tasks, can automatically adjust the fusion weights according to task requirements and data distribution, adapt to different scenarios and tasks, and reduce the dependence on manually designed fusion rules.

[0060] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0061] Figure 1 is a schematic diagram of the SAM framework of the present invention;

[0062] Figure 2 is a schematic diagram of the adaptive boundary enhancement module of the present invention;

[0063] Figure 3 is a schematic diagram of the dynamic fusion module of the present invention;

[0064] Figure 4 is a visual comparison diagram of brain tumor segmentation results obtained by different segmentation methods of the present invention;

[0065] Figure 5 is a visual comparison diagram of brain tumor segmentation results obtained by different basic models of the present invention. Detailed implementation manners

[0066] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0067] Embodiment

[0068] As Figures 1 - 3 shown, the brain tumor region segmentation method based on adaptive boundary-guided aggregation SAM of the present invention includes the following steps:

[0069] S1. Input the fusion output I fusion of the multimodal magnetic resonance image I4 into the image segmentation model SAM to obtain the image embedding E sam ; The process of obtaining the fusion output I fusion is as follows:

[0070] I fusion = MHA(Conv(I4))

[0071] where Conv(·) represents the convolution operation, MHA(·) is the multi-head self-attention mechanism, I4 ∈ R 4×H×W , H and W respectively represent the height and width of the image, and R represents the set of multimodal magnetic resonance images;

[0072] S2. Input the boundary-sensitive modal image I3 into the boundary mining network BDNet to mine the boundary features of each region of the brain tumor, and obtain the boundary embedding E boundary of the brain tumor and the true tumor boundary P b ;

[0073] S21. The boundary mining network BDNet uses a residual structure to build the network as the backbone network, and uses the boundary-sensitive modal image I3 as the input, where I3 ∈ Q 3×H×W , Q represents the set of boundary-sensitive modal images, and passes through four residual blocks encapsulated by residual connections in sequence. The boundary clue output by each residual block is where i ∈ [1, 2, 3, 4], C i ∈ [64, 128, 256, 512], B is the set of boundary clues, and the output b4 is upsampled once to obtain the boundary embedding E with a size of 256×16×16 boundary ;

[0074] S22. Further reconstruct the true boundary of the tumor through the guidance of the SAM model, and use the boundary clues of each residual block as the input of the boundary decoder, and layer by layer implement feature splicing during the decoding process, and finally decode to obtain the true boundary P of the tumor b

[0075] S3. Enhance the boundary features through the Adaptive Boundary Augmentation Model (ABAM) and then provide them to the SAM model;

[0076] S31. Downsample and splice the boundary clues to form a comprehensive feature map, and use the channel attention mechanism to adjust the weights of the feature map to obtain the feature representation F of multi-scale boundary clues init , considering the deformation and displacement of the tumor respectively, to adaptively enhance the boundary clue representation;

[0077] S32. In order to capture the deformation and size changes of the tumor, use deformable convolution to adaptively process F init , use max pooling to highlight the strongest response area, use average pooling to retain the global average information, and splice the results of the two poolings to obtain a more complete deformation representation F d :

[0078] F d = Concat[Max(Dc(F init )),Mean(Dc(F init ))]

[0079] where Dc(·) represents deformable convolution, Concat(·) represents the splicing function, Max(·) and Mean(·) are max pooling and average pooling respectively;

[0080] S33. Considering the variability of brain tumor localization, independently process F init and its transpose , divide them into smaller patches for embedding and then perform displacements in multiple directions, weight the features through convolution with a 1×1 convolution kernel, and aggregate the information in multiple directions using matrix addition to obtain the displacement feature

[0081]

[0082] where C d (·) represents performing a convolution operation on the features offset in the d direction, D = {right, right-up, left, left-down}, representing the set of offset directions;

[0083] S34. Add displacement features by a multi-layer perceptron (MLP) to increase the non-linear expression, and obtain the displacement offset representation F through splicing and aggregation s :

[0084]

[0085] S35. Use an element-wise multiplication operation to spatially combine the deformation information F d and the displacement offset information F s :

[0086] F ABA = F d ⊙ F s

[0087] where ⊙ represents element-wise multiplication, and F ABA is the boundary feature after adaptive enhancement;

[0088] S4. Apply dynamic fusion module (DFM) to reorganize channel information and spatial information, and embed the image of S2 into E sam and the boundary embedding E boundary to dynamically fuse into the fused embedding E DFM ;

[0089] S41. Use the sigmoid function to set the gating mechanism of the learnable parameter α to achieve the adaptive weighting of the image embedding E sam and the boundary embedding E boundary for information complementarity:

[0090]

[0091] where E G1 is the initial fused embedding, and σ(α) is the sigmoid function with the learnable parameter α;

[0092] S42. For channel information reorganization, input the initial fused embedding E G1 into two consecutive depthwise separable convolutions to obtain the channel-sensitive output, and add it to the initial fused feature E G1 to obtain the channel embedding after dynamic reorganization

[0093]

[0094] where D represents the channel convolution operation, and P represents the pointwise convolution operation;

[0095] S43. For spatial information reorganization, use E G1 and its transpose Take them as inputs respectively, divide the inputs into smaller patches, dynamically adjust the feature weights using the channel and spatial attention module CBAM, use a feature map with the same dimension as the patches as a learnable mapping, and then update the parameters by means of gradient descent. The specific calculation is as follows:

[0096]

[0097] Among them, pe(·) represents patch embedding, cbam(·) represents applying channel attention and spatial attention to the input feature map in sequence to generate a weighted feature map, and l1 and l1′ are learnable mappings;

[0098] S44. Set the learnable mappings l2 and l2′ to be dynamically adjusted, and obtain by combining two types of spatial information through feature concatenation The specific calculation is as follows:

[0099]

[0100] S45. Use the second gating mechanism with learnable parameter α′ to fuse the channel recombination information and the spatial recombination information to achieve the dynamic fusion of boundary embedding and image embedding:

[0101]

[0102] Among them, E DFM is the fusion embedding of DFM;

[0103] S5. Input the fusion embedding E DFM obtained in S4, the adaptive enhanced boundary features obtained in S3, and the features of the prompt encoder in the SAM model into the mask decoder in the SAM model to generate the segmentation results of each region of the brain tumor;

[0104] S6. Combine the boundary mining network and the region-based prediction loss function and obtain the total loss function of the SAM model;

[0105] S61. Use the Dice coefficient as the optimization function of BDNet, and calculate the loss function of BDNet as follows:

[0106]

[0107] Among them, P B represents the predicted tumor boundary, G B represents the ground truth label of the boundary mask, N is the total number of pixels in the image, and i is the i-th pixel in the image;

[0108] S62. Considering the large gap in the proportion of brain tumor and normal brain tissue regions, a Dice loss function and a Focal loss function are used to formulate an optimization function based on regional prediction. The loss function is as follows:

[0109]

[0110] where y is the ground truth, is the prediction result, r is the set of different tumor sub-regions, and r ∈ {WT, TC, ET}, WT is the regional mask of the whole tumor, TC is the regional mask of the tumor core, and ET is the regional mask of the enhanced tumor;

[0111] S63. The overall loss function based on the adaptive boundary-guided aggregation SAM is:

[0112] L = (1 - γ)·L boundary + γ·L mask

[0113] where the weight hyperparameter γ = 0.8.

[0114] In the specific implementation process, first, BDNet mines the features of each tumor region, aiming to obtain the tumor boundary embedding E boundary and the true tumor boundary P b , and enhances the robustness of the model to variable boundaries through ABAM; secondly, DFM aims to dynamically fuse the image embedding E sam and the boundary embedding E boundary to obtain the fused embedding E DFM to provide more tumor-characteristic information for the mask decoder; then, for the original SAM structure, we freeze the parameters of the image encoder while retaining the trainability of the prompt encoder and the mask decoder; finally, under the guidance of point, box, and true tumor boundary prompts, the modified mask decoder is used to directly generate the segmentation results of each region of the brain tumor.

[0115] The comparative experiment process is as follows:

[0116] 1. Dataset selection:

[0117] The BraTS dataset, released by MICCAI official, is widely used in the task of brain tumor segmentation; in this application, the proposed framework is evaluated on two public datasets, BraTS2019 and BraTS2023, which cover a variety of brain tumor cases, including different tumor deformations, sizes, and positions; the evaluation of the segmentation task is based on three brain tumor sub-regions: the whole tumor (WT = NCR / NET + ED + ET), the tumor core (TC = NCR / NET + ET), and the enhanced tumor (ET).

[0118] 2. Evaluation index setting:

[0119] The indexes include Dice Similarity Coefficient (Dice) and 95% Hausdorff distance (Hausdorff distance(95%), HD95); Generally speaking, an excellent segmentation method will produce a higher Dice and a lower HD95.

[0120] 3. Design and evaluate comparative experiments of different segmentation methods:

[0121] To verify the effectiveness of ABF-SAM, we selected 10 state-of-the-art segmentation methods to set up comparative experiments, including: U-Net, Attention U-Net, UNeXt, U-Net 3+, TransUNet, Swin-Unet, Convolution-Transformer hybrid optimization method CTO, Boundary-preserving assembled transformer UNet (BPAT-UNet), BiTransformer U-Net (BiTr-Unet), Dilated Hierarchical Decoupled Convolution Network with Attention (ADHDC); among them, both CTO and BPAT-UNet consider boundary information and have exclusive designs; BiTr-Unet and ADHDC use 3D data as input and can support three-dimensional information processing; the dataset division used in the comparative experiments is kept consistent; the quantitative results of different segmentation methods on the BraTS2019 and BraTS2023 datasets are shown in Table 1.

[0122] For easy observation, we provide the average of the segmentation results of the three tumor regions and mark the best result in bold; from the results, it can be observed that the method we proposed achieved impressive results. In terms of the average segmentation results of the WT, TC, and ET regions, ABF-SAM achieved the best result; on the BraTS2019 dataset, the average Dice score of the proposed method reached 88.03%, and the HD95 was 4.583, both significantly better than other comparison methods; similarly, on the BraTS2023 dataset, the average Dice score of the proposed method reached 89.84%, and the HD95 was 4.714, also achieving the optimal performance.

[0123] Secondly, ABF-SAM outperformed the models that used the boundary as guidance information on both datasets; among them, the CTO method used a boundary detection operator to obtain boundary information and used it as explicit supervision to guide learning, while BPAT-UNet enhanced boundary features and generated ideal boundary points to improve the segmentation effect; in contrast, ABF-SAM combined multiple boundary guidance schemes, covering decision-level and feature-level representations, including boundary embedding, adaptively enhanced boundary features, and hint information based on the true tumor boundary, thus improving the segmentation accuracy.

[0124] Then, although ADHDC and BiTr-Unet are models that utilize three-dimensional spatial information, the performance they obtained is sub-optimal; specifically, in the BraTs2019 and BraTs2023 datasets, compared with the sub-optimal ADHDC, our method increased the Dice score by 3.12% and 1.46% respectively, and reduced the HD95 by 5.271 and 6.559, indicating that even on two-dimensional data, ABF-SAM can still effectively capture key features and achieve better segmentation performance.

[0125] It is worth noting that compared with other advanced models, the proposed method has more obvious advantages in the BraTS2019 dataset; this phenomenon may be attributed to the strong robustness of the basic model SAM encoder, enabling it to achieve excellent results even under the condition of less training data, highlighting the great potential of fine-tuning the basic model; however, in the case of less data, the HD95 index of the WT region is not ideal; as the data volume increases, in the BraTS2023 dataset, this index has been significantly improved and finally achieved the best result.

[0126] Table 1 Experimental evaluation results of different brain tumor segmentation methods on BraTS2019 and BraTS2023

[0127]

[0128] Visual contrast analysis such as Figure 4As shown, by comparing the segmentation results with the ground truth annotations, it can be found that the method proposed in this application has better segmentation performance, especially in the boundary recognition of tumor sub-regions; this result verifies the importance of boundary information for brain tumor segmentation; from Figure 4 the first row of Figure 4 , it can be observed that due to the compression of the brain tissue by the edema, the shape of the body of the lateral ventricle is slightly deformed; and since cerebrospinal fluid and edema have similar imaging characteristics in multi-modal MRI, most models have over-segmentation phenomena in this region; in contrast, the model proposed in this application successfully avoids this problem, which benefits from the effective combination of various boundary information by ABF-SAM; in addition, in the segmentation of each tumor sub-region, the assistance of boundary information enables ABF-SAM to accurately identify and segment complex tumor sub-regions (such as Figure 4 the third row and the last row of Figure 4 ); based on these visualization results, it is obvious that the proposed method has significant robustness in the brain tumor segmentation task dealing with complex boundary information.

[0129] 4. Design and evaluate a comparative experiment of different derivative basic models based on SAM

[0130] This application selected four derivative basic models based on SAM for comparison, including the medical-specific segmentation any model MedSAM, the medical two-dimensional segmentation any model SAM-Med2D, the medical-adapted segmentation any model SAMed, and the ultrasound segmentation any model SAMUS; MedSAM: Adopts a full-parameter training strategy, uses the bounding box as the prompt information, and is trained on a large-scale medical dataset, showing strong medical image segmentation capabilities; SAM-Med2D: Performs parameter-efficient fine-tuning through an adapter; SAMed: Adopts a strategy based on low-rank decomposition to optimize the model; SAMUS: Introduces a parallel CNN branch into the SAM architecture and realizes the interaction between SAM and CNN features through a cross-branch attention mechanism; to avoid data leakage, this application strictly follows the original paper design and re-executes the model training on the brain tumor dataset.

[0131] The quantitative comparison results are shown in Table 2. The method proposed in this application is significantly superior to other basic models in all evaluation metrics, fully demonstrating its advantages in the brain tumor segmentation task. Specifically, MedSAM performs poorly in the multi-object brain tumor segmentation task. MedSAM only segments the WT, which belongs to a single-object segmentation task and is difficult to meet the requirements of brain tumor sub-region segmentation. Both SAM-Med2D and SAMed adopt a parameter-efficient fine-tuning strategy, and the performance difference on the BraTS dataset is small. SAMUS realizes effective feature interaction by introducing a parallel CNN branch and a cross-branch attention mechanism, and shows high segmentation performance on the BraTS2023 dataset with a large sample size. The average Dice coefficient is better than that of BraTS2019 (BraTS2019 vs. BraTS2023: 61.06% vs. 71.59%). However, the improvement of SAMUS in the HD95 metric is limited, indicating that there are still deficiencies in its boundary processing. In contrast, the method proposed in this application combines boundary embedding with a dynamic fusion module, and at the same time combines ABAM for adaptive boundary enhancement, making the model more sensitive and robust in the recognition and segmentation of complex tumor boundaries. In summary, in the medical image segmentation task, basic models not only need to carefully design fine-tuning strategies, but also need to be specifically designed in combination with the particularity of the task to achieve better segmentation performance.

[0132] Table 2 Experimental evaluation results of different basic models on BraTS2019 and BraTS2023 datasets

[0133]

[0134] The visual comparison results are as Figure 5 shown. Brain tumor segmentation faces significant challenges, mainly due to the irregularity of its shape and the complexity of its boundary, making it extremely difficult to accurately segment tumor sub-regions. It can be observed from Figure 5 that even with point and box prompt information, most models are still unable to accurately depict the boundaries of tumor sub-regions. This indicates that in high-precision segmentation tasks, the true boundary information of tumors is crucial. The method proposed in this application effectively fuses boundary embedding at the decision-making level, reconstructs the true boundary of the tumor and uses it as one of the display prompt information, making the model more accurate in identifying the location and boundary of tumor sub-regions, thus significantly improving the segmentation performance.

[0135] Therefore, the present invention adopts the above-mentioned brain tumor region segmentation method based on self-adaptive boundary-guided aggregation SAM, constructs a framework of self-adaptive boundary-guided aggregation SAM to achieve accurate segmentation of brain tumor regions, and enhances the robustness of the model to variable boundaries through ABAM.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An adaptive boundary-guided aggregation SAM-based brain tumor region segmentation method, characterized in that: S1. Input the fusion output I of the multi-modal magnetic resonance image I4 fusion into the image segmentation model SAM to obtain the image embedding E sam ; S2. Input the modality image I3 sensitive to boundaries into the boundary mining network BDNet to mine the boundary features of each region of the brain tumor, and obtain the boundary embedding E of the brain tumor boundary and the true tumor boundary P b ; S3. Enhance the boundary features through the adaptive boundary enhancement model ABAM and then provide them to the SAM model; S4. Recombine the channel information and spatial information by applying dynamic fusion DFM, and embed the image of S2 into E sam and embed the boundary of S3 into E boundary to dynamically fuse into the fused embedding E DFM ; S5. Embed the fusion obtained in S4 into E DFM Input the adaptive enhanced boundary features obtained in S3 and the features of the prompt encoder in the SAM model into the mask decoder in the SAM model to generate the segmentation results of each region of the brain tumor; S6. Combine the boundary mining network and the region prediction-based loss function to obtain the total loss function of the SAM model.

2. The adaptive boundary-guided aggregation SAM-based brain tumor region segmentation method according to claim 1, characterized in that: The process of obtaining the fused output I in S1 fusion is as follows: I fusion = MHA(Conv(I4)) where Conv(·) represents the convolution operation, MHA(·) is the multi-head self-attention mechanism, I4 ∈ R 4×H×W , H and W represent the height and width of the image respectively, and R represents the set of multimodal magnetic resonance images.

3. The adaptive boundary-guided aggregation SAM-based brain tumor region segmentation method according to claim 2, characterized in that: The specific process in S2 is as follows: S21. The boundary mining network BDNet uses a residual structure to build the network as the backbone network, and uses the boundary-sensitive modal image I3 as the input, where I3 ∈ Q 3×H×W , Q represents the set of boundary-sensitive modal images, and passes through four residual blocks encapsulated by residual connections in sequence. The output of each residual block is the boundary clue as where i ∈ [1, 2, 3, 4], C i ∈ [64, 128, 256, 512], B is the set of boundary clues. The output b4 is upsampled once to obtain the boundary embedding E with a size of 256×16×16 boundary ; S22. Further reconstruct the true boundary of the tumor through the guidance of the SAM model, and use the boundary clues of each residual block as the input of the boundary decoder. During the decoding process, feature splicing is realized layer by layer, and finally the true boundary P of the tumor is obtained by decoding b .

4. The adaptive boundary-guided aggregation SAM-based brain tumor region segmentation method according to claim 3, characterized in that: The process of enhancing the boundary features in S3 is as follows: S31. Downsample and splice the boundary cues to form a comprehensive feature map, and use the channel attention mechanism to adjust the weights of the feature map to obtain the feature representation F of multi-scale boundary cues init , considering the deformation and displacement of the tumor respectively to adaptively enhance the boundary cue representation; S32. To capture tumor deformation and size changes, deformable convolutions are used to adaptively process F init , max pooling is used to highlight the strongest response region, average pooling is used to retain the global average information, and the results of the two poolings are concatenated to obtain a more complete deformation representation F d : F d = Concat[Max(Dc(F init )),Mean(Dc(F init ))] Wherein, Dc(·) represents deformable convolution, Concat(·) represents the concatenation function, Max(·) and Mean(·) are max pooling and average pooling respectively; S33. Considering the variability of brain tumor localization, separately for F init and its transpose perform independent processing, divide it into smaller patches, embed them, perform displacements in multiple directions, weight the features through convolution with a convolution kernel of 1×1, and aggregate the information in multiple directions using matrix addition to obtain displacement features Among them, C d (·) represents performing a convolution operation on the feature offset in the d direction, D = {right, right-up, left, left-down}, representing the set of offset directions; S34. Increase the displacement features by the multi-layer perceptron mlp to obtain the non-linear expression, and obtain the displacement offset representation F by splicing and aggregation s : S35. Adopt an element-wise multiplication operation to spatially combine the deformation information F d and the displacement offset information F s : F ABA = F d ⊙F s where ⊙ represents element-wise multiplication, and F ABA is the boundary feature after adaptive enhancement.

5. The adaptive boundary-guided aggregation SAM-based brain tumor region segmentation method according to claim 4, characterized in that: The dynamic fusion process in S4 is as follows: S41. Set a gating mechanism for the learnable parameter α using the sigmoid function to implement the image embedding E sam and the boundary embedding E boundary for information complementation through adaptive weighting: E G1 = (1 - σ(α))·E boundary + σ(α)·E sam Among them, E G1 is the initial fusion embedding, and σ(α) is the sigmoid function with learnable parameter α; S42. For channel information recombination, the initial fusion embedding E G1 is input into two consecutive depthwise separable convolutions to obtain a channel-sensitive output, which is then added to the initial fusion feature E G1 to obtain the channel embedding after dynamic recombination Wherein, D represents the channel convolution operation, and P is the pointwise convolution operation; S43. For spatial information recombination, use E G1 and its transpose as inputs respectively. Divide the inputs into smaller patches, use the channel and spatial attention module CBAM to dynamically adjust the feature weights, use a feature map with the same dimension as the patch as the learnable mapping, and then update the parameters by gradient descent. The specific calculation is as follows: Wherein, pe(·) represents patch embedding, cbam(·) represents applying channel attention and spatial attention to the input feature map in sequence to generate a weighted feature map, and l1 and l1′ are learnable mappings; S44. Set the learnable mappings l2 and l2' to be dynamically adjusted, and obtain by feature concatenation by combining the two spatial information The specific calculation is as follows: S45. Use a second gating mechanism with learnable parameter α′ to fuse the channel recombination information and the spatial recombination information to achieve dynamic fusion of boundary embedding and image embedding: Among them, E DFM is the fusion embedding of DFM.

6. The adaptive boundary-guided aggregation SAM-based brain tumor region segmentation method according to claim 5, characterized in that: The specific process of calculating the loss function in S6 is as follows: S61. Use the Dice coefficient as the optimization function of BDNet, and calculate the loss function of BDNet as follows: where P B represents the predicted tumor boundary, G B represents the ground truth label of the boundary mask, N is the total number of pixels in the image, and i is the i-th pixel in the image; S62. Considering the large gap in the proportion of brain tumor and normal brain tissue regions, a region prediction-based optimization function is formulated using the Dice loss function and the Focal loss function, and the loss function is as follows: where y is the true value, is the predicted result, r is the set of different tumor sub-regions, and r ∈ {WT, TC, ET}, where WT is the region mask of the whole tumor, TC is the region mask of the tumor core, and ET is the region mask of the enhanced tumor; S63. The overall loss function of the adaptive boundary-guided aggregation SAM is: L = (1 - γ)·L boundary + γ·L mask Wherein, the weight hyperparameter γ = 0.8.

Citation Information

Cited By

  • Tumor model generation method and system based on adaptive marginal lattice arrangement adjustment

    CN120599159A

  • Tumor model generation method and system based on adaptive edge lattice arrangement adjustment

    CN120599159B

  • Multi-task brain glioma automatic segmentation and IDH genotyping method based on SAM

    CN120635122A