Brain tumor MRI image segmentation method based on boundary enhancement and dual-encoder fusion

Through the methods of boundary enhancement and dual encoder fusion, the problems of resolution reduction and feature redundancy in the existing technology are solved, and higher-precision brain tumor MRI image segmentation is achieved, especially the boundary capture capability in complex scenes.

CN120707452APending Publication Date: 2025-09-26BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251375.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing brain tumor MRI image segmentation methods suffer from problems such as reduced resolution, blurred boundary detail information, feature redundancy, and failure to fully integrate channel information, boundary features, and global image context information.

Method used

A method based on boundary enhancement and dual encoder fusion is adopted. The image enhancement module is used to improve the distinction between the target area and the background. The dual encoder structure is used to process the original MRI image and the boundary-enhanced image respectively, enhancing the learning ability of image details and global information. The feature map splicing and upsampling in the decoder stage are combined to restore the spatial resolution.

Benefits of technology

The accuracy and boundary capture capability of brain tumor MRI image segmentation have been improved, and the performance of the segmentation model has been significantly improved, especially in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707452A_ABST
    Figure CN120707452A_ABST
Patent Text Reader

Abstract

The invention discloses a brain tumor MRI image segmentation method based on boundary enhancement and dual-encoder fusion, and belongs to the field of computer vision. The method is generally divided into two stages of image enhancement and image segmentation. The first stage is image enhancement, parallel branches are adopted to process brightness and boundary features respectively, the brightness and the boundary features are fused to generate an enhanced image, and the distinction degree of a target area and a background is remarkably improved. In an image segmentation stage, an original MRI image and an image after boundary enhancement are respectively processed based on a segmentation model of a double-encoder structure. Each encoder extracts features of different levels through convolution operation so as to enhance the learning ability of image details and global information and improve the segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention designs a brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion, which belongs to the field of computer vision. Background Art

[0002] Brain tumor image segmentation is a key computer vision task aimed at automatically identifying and extracting brain tumor regions. Magnetic resonance imaging (MRI) is a noninvasive imaging technique that requires no surgery or radiation and enables repeated long-term imaging monitoring, making it the preferred method for brain monitoring. However, the quality of MRI images can vary depending on factors such as the scanning equipment, scanning parameters, body position, and scanning angle. Some MRI images may contain noise and artifacts, further increasing the complexity and difficulty of automatic segmentation.

[0003] With the recent advancements in deep learning, convolutional neural networks (CNNs) have been widely used in image segmentation tasks due to their powerful feature extraction capabilities. Deep learning models such as U-Net and FCN have become mainstream approaches in medical image segmentation. These networks are capable of learning complex feature representations and achieving more accurate image segmentation through end-to-end training. Currently, most brain tumor MRI image segmentation models use deep convolutional neural networks as their underlying architecture to improve segmentation accuracy and handle complex tumor morphologies. Specifically, U-Net, a classic architecture in medical image segmentation, effectively captures image context through its encoder-decoder structure and preserves details through skip connections. To further improve segmentation performance, many researchers have proposed new improvements and modules based on U-Net to enhance its expressive power and segmentation results. For example, Mobeen UrRehman et al. proposed a residual space pyramid pooling module based on the U-Net network. This module uses parallel layers of dilated convolutions and an attention gate module to recover the segmentation output from the extracted feature maps. This module combines knowledge from different feature maps while preserving local information to obtain a rich feature representation. Compared to segmentation models with a single encoder-decoder structure, segmentation models with a dual encoder-decoder structure can better capture global and local contextual information, thereby improving segmentation results. For example, the paper "DE-UFormer: U-shaped dual encoder architectures for brain tumor segmentation" by Yan Dong et al. proposes a dual encoder network model, DE-Uformer, which uses a CNN encoder and a Transformer encoder to obtain local features and global representations.

[0004] Although the above methods can complete the task of brain tumor MRI image segmentation, the existing segmentation methods still have the following challenges: (1) In the existing segmentation model, the downsampling operation extracts deep features through convolution and pooling, but at the same time it will lead to reduced resolution, blurred or even lost boundary detail information. (2) Most existing dual encoder models use the same image as input. This design often leads to feature redundancy. The features extracted by the two encoders may contain a large amount of repeated information, which fails to fully utilize the advantages of the dual encoder architecture. (3) The existing model fails to fully integrate channel information, boundary features and global image context information during the feature extraction process, resulting in its segmentation performance being limited in complex scenarios. Based on the above limitations, a brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion is proposed. Summary of the Invention

[0005] Aiming at the shortcomings of the existing technology, the present invention proposes a brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion. Figure 1 As shown in the figure, the method is generally divided into two stages: image enhancement and image segmentation. In the first stage, image enhancement, parallel branches are used to process brightness and boundary features respectively, fusing the two to generate an enhanced image, significantly improving the distinction between the target area and the background. In the image segmentation stage, a segmentation model based on a dual-encoder structure processes the original MRI image and the boundary-enhanced image separately. Each encoder extracts features at different levels through convolution operations to enhance the learning ability of image details and global information, thereby improving segmentation accuracy.

[0006] In order to achieve the above object of the invention, the present invention adopts the following technical solutions:

[0007] A brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion includes: a data preprocessing module, an image enhancement module, and an image segmentation module.

[0008] Data Preprocessing: This module consists of four parts: data resampling, 2D slice extraction, image screening, and image set partitioning. First, the 3D MRI brain image data is resampled and slices are extracted along the axial plane. Subsequently, all images are resized to 256×256 for subsequent segmentation. Finally, the image set is partitioned into training and test sets.

[0009] Image Enhancement Module: This module consists of three main parts: image brightness channel extraction and enhancement, superpixel-based boundary extraction, and a dual-branch fusion enhancement module. By extracting the brightness channel (L channel) of the MRI image, it enhances its contrast and optimizes the overall visual effect. A superpixel segmentation algorithm is used to generate candidate boundary regions for the target area, highlighting the edges of the target area. A parallel branch network is used to process brightness and boundary features separately, and the two are fused to generate the enhanced image.

[0010] Image segmentation module: This module uses a brain tumor image segmentation model based on a dual encoder structure. The model architecture is as follows Figure 2 As shown in Figure 1, the proposed method mainly consists of two parts: (1) a dual encoder architecture, in which the model’s two independent encoders process the original MRI image and the enhanced image, respectively. Each encoder extracts features at different levels through convolution operations to enhance the learning ability of image details and global information. (2) In the decoder stage, the feature maps from the two encoders are combined through splicing and upsampling to gradually restore the spatial resolution of the image and enhance the fusion of information, thereby improving the model’s accuracy in segmenting the target region boundaries and morphology. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is the overall architecture diagram of the brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion

[0012] Figure 2 Image segmentation network structure diagram based on dual encoder

[0013] Figure 3 This is a visualization of the prediction results on the BT2017 dataset.

[0014] Figure 4 This is a visualization of the prediction results on the BT-BTH dataset DETAILED DESCRIPTION

[0015] The specific implementation and exemplary embodiments of each step of the present invention will be described in detail below

[0016] Step 1: Data Preprocessing

[0017] Data preprocessing mainly involves image resampling, slice extraction, image screening, and image set division.

[0018] Step 1.1: First, resample the input 3D data to 3D size. Resample the 3D image of size 240×240×155 to 256×256×44. The resampling factor of each dimension is:

[0019]

[0020] Among them, i represents the dimension index, referring to the different dimensions of the image (x, y, z dimensions), originSize i is the size of the original image in the i-th dimension, newSize i is the size of the target image in the i-th dimension. Next, calculate the mapping from the target image coordinates to the original image coordinates. For each target coordinate (x′, y′, z′), use the following formula to convert the target image coordinates to the original image coordinates:

[0021] x=x′×f x (2)

[0022] y=y′×f y (3)

[0023] z=z′×f z (4)

[0024] Among them, (x, y, z) is the original image coordinates calculated from the target image coordinates. Finally, the nearest neighbor interpolation method is used to obtain the pixel value of the corresponding coordinates:

[0025] I′(x′,y′,z′)=I(x ≈ ,y ≈ ,z ≈ ) (5)

[0026] Where I(x,y,z) is the pixel value at the coordinate (x,y,z), (x ≈ ,y ≈ ,z ≈ ) represents (x, y, z) rounded to the nearest integer, and I′(x′, y′, z′) represents the resampled pixel value. Resampling reduces image redundancy, standardizes data, improves computational efficiency, and ensures that the image is compatible with subsequent deep learning segmentation models, thereby improving segmentation accuracy and efficiency.

[0027] Step 1.2: Slice the input 3D image data and slice the 256×256×44 3D MRI image into 44 256×256 2D images along the axial direction.

[0028] Step 1.3: Filter the slices based on the label. For each slice, check whether the target region exists in the label image. If the pixel value of a slice in the label image is greater than 0, it means that the slice contains the target region; otherwise, remove the slice.

[0029] Step 1.4: Randomly divide the image collection into training and test sets with a ratio of 8:2.

[0030] Step 2: Image Enhancement

[0031] Step 2.1: MRI images are typically stored in RGB format. However, in RGB space, brightness and color information are coupled together, making it difficult to enhance brightness alone. Therefore, the input RGB image is converted to CIELAB color space, where the L channel represents brightness and the A and B channels represent color information. The conversion from RBG color space to LAB color space is a two-step process. First, the RGB color space is converted to XYZ space:

[0032]

[0033] This matrix multiplication converts the color values ​​in the RGB color space Convert color values ​​to XYZ color space

[0034] Then, convert from XYZ space to LAB space:

[0035]

[0036] where X n , Y n , Z n It is the reference white point standard value, the default values ​​are 95.047, 100.0, 108.883. * , a * , b * It is the value of the three channels of the final LAB color space.

[0037] Secondly, the L channel is separated from the LAB image for subsequent contrast enhancement. The CLAHE (Contrast Limited Adaptive Histogram Equalization) method is used for local enhancement. The image is first divided into small grids of 8×8, and the pixel values ​​in each grid are adjusted by local histogram equalization to adjust the contrast. Assuming L * (x,y) is the original image L * The brightness value of the pixel with coordinate position (x, y) in the channel, and (x', y') is the position of the pixel in the divided grid, then L * (x′, y′) represents the original image channel L * The brightness value of the pixel at the (x′, y′) position in the grid. Assume that all pixel values ​​in the grid are L grid , then the local histogram H local (k) can be expressed as:

[0038]

[0039] Where k is the discretization level of the brightness value, k∈[0,255], and δ is the indicator function:

[0040]

[0041] In order to avoid excessive enhancement of local contrast, CLAHE introduces a contrast limit value clipLimit to limit the maximum cumulative frequency of the local histogram. The specific approach is to limit the cumulative frequency of the local histogram to a maximum value C. The clipping value C of each grid is:

[0042]

[0043] where N grid L is the total number of pixels in each small grid area, representing the size of the local area. max Is the maximum brightness value of the image, representing the upper limit of brightness in the image. In this method, the contrast limit value clipLimit is set to 2.0. If the local histogram value H local (k) exceeds C, then truncation is performed, and the local histogram value H of each brightness level local (k) is adjusted to:

[0044]

[0045] Next, the local histogram is equalized. By normalizing the cumulative distribution function (CDF) of each local histogram, a local equalization map can be obtained. The brightness value after local equalization is:

[0046]

[0047] The accumulation process here is performed in each grid, and the output brightness value will be adjusted to [0,L max -1] range. Finally, the L of all grids after equalization is * The channel values ​​are reassembled into a whole image. Each pixel L of the original image * (x,y) is replaced by its new value in the equalized local region:

[0048] L clahe (x,y)=L enhanced (x′,y′) (16)

[0049] Where (x′, y′) is the coordinate of the pixel (x, y) in the grid. The enhanced L is obtained through the above steps. clahe image.

[0050] Step 2.2: Then, the SLIC (Simple Linear Iterative Clustering) algorithm is used for superpixel segmentation. This algorithm divides the image into several superpixel regions of similar size and shape through iterative optimization. The basic steps of the SLIC algorithm include:

[0051] (1) Initialize the superpixel center. Suppose the image has N pixels and we want to divide the image into K superpixel regions. Then the size of each superpixel is N / K, and the distance between adjacent centers is:

[0052]

[0053] (2) Calculate the distance between each pixel and the superpixel center. The algorithm selects a 2S×2S area within the area of ​​cluster center distance S to calculate the distance between the pixel and each cluster center. The total distance D is composed of the color distance d lab and spatial distance d xy It consists of two parts. Among them, m is the compactness factor (m∈[1,40]), which is usually set to 10 and is used to measure the proportion of color value and spatial value in the similarity metric.

[0054]

[0055] Among them, l, a, and b are the three channels of the Lab color space, p represents the current pixel, e represents the superpixel seed point, and l o Indicates the brightness value of point o, a o Indicates the green-red value of point o, b o Indicates the blue-yellow value of point o, x o Indicates the x coordinate of point o, y o Indicates the y coordinate of point o;

[0056] (3) Iteratively update the superpixel center position. The specific steps are to classify the pixels and mark each pixel as the category of the cluster center with the smallest distance. Then, recalculate the cluster center and calculate the average vector value of all pixels belonging to the same cluster to obtain the cluster center again. This process is iterated until the maximum number of iterations is reached, and the clustering is completed. The maximum number of iterations in this method is set to 10.

[0057] After obtaining the superpixel segmentation result through the SLIC algorithm, the segmentation result is further used to extract the boundary information of the target area. The target area usually has color characteristics different from the non-target area. Therefore, by analyzing the average color value of the superpixel area, the feature differences between different areas can be effectively distinguished. In this study, we mainly focus on the B channel because it can effectively reflect the color difference of the target area. The specific approach is: for each superpixel area obtained after the SLIC algorithm, traverse the B channel in each area, calculate the average value of the B channel in each area, and replace the value of all pixels in this superpixel area with the average value. The pixel value in this superpixel area is:

[0058]

[0059] Among them, r is the current super pixel area, N r is the number of pixels in the current superpixel area, and b(x′, y′) represents the B channel value of the position (x′, y′) in the original image.

[0060] In this way, the color information within each superpixel is compressed into a single mean value that represents the characteristics of that region. This process helps extract the primary color features of the target region and, to a certain extent, removes noise and unnecessary details from the image. For example, if the superpixel region contains the target region, calculating the mean value can help extract the characteristics of the target region and prevent individual pixel anomalies from affecting the results.

[0061] Step 2.3: To enhance the image's feature expression and improve the target region segmentation, a dual-branch information fusion method is used to weightedly fuse the image information from different branches, thereby achieving comprehensive enhancement of image features. The specific method is: using weighted fusion technology, the image enhanced by CLAHE is weightedly fused with the image obtained after boundary enhancement.

[0062] I fused (x,y)=α·B(x,y)+β·L clahe (x,y) (22)

[0063] Among them, α and β are weighting coefficients, which are set to 0.5 in this study to ensure that the influence of the two types of information on the final image can achieve the best balance. B(x,y) is the image after boundary enhancement in step 2.2, L clahe (x, y) is the image after brightness enhancement in step 2.1. Finally, the weighted fusion image I fused (x,y) is applied to the three RGB channels to generate the final composite image.

[0064] Step 3: Image Segmentation

[0065] In this stage, the semantic information in the image is extracted through a deep learning network, and then a segmentation mask containing semantic feature information and boundary information is generated.

[0066] Step 3.1: Dual Encoder Stage. To extract a robust and representative set of features, this stage employs two independent VGG16 networks as encoders within the segmentation architecture, each extracting distinct features from the target region. The VGG network's deep convolutional structure and small 3×3 convolution kernel design effectively extract boundary information, morphological features, and global information, thereby enhancing the model's expressive power. Building on the VGG16 network, this stage utilizes only the convolutional layer as a feature extractor, omitting fully connected layers. This ensures that every feature map can participate in the decoding process, improving pixel-level prediction accuracy.

[0067] The model includes two input data, where input O1 is the original MRI image, represented as a 3-channel tensor (size 256×256×3), and input I2 is the image obtained after boundary enhancement (the method described in step 2), represented by a tensor (size 256×256×3). The two input data are used for two independent encoders to obtain complementary feature information. Encoder Enc1 is responsible for extracting global features from the original MRI image, while encoder Enc2 processes the image enhanced by the method described in step 2, focusing on boundaries and local features. Each encoder contains five convolution blocks. The first two convolution blocks each have two 3×3 convolution layers, and the last three convolution blocks each have three 3×3 convolution layers. Each convolution layer is followed by a ReLU activation function and down-sampled using a 2×2 maximum pooling layer. The calculation process of encoder Enc1 is as follows:

[0068] The convolution and pooling process of the first layer is:

[0069]

[0070] Where * represents the convolution operation, W1 is the convolution kernel weight of the first layer, b1 is the bias term, ReLU(x)=max(0,x) is the activation function, and MaxPool(x) represents the maximum pooling operation.

[0071] The subsequent convolution and pooling process is:

[0072]

[0073] The calculation process of encoder Enc2 is the same as that of encoder Enc1. Ultimately, each encoder extracts five levels of features:

[0074]

[0075] These features come from network layers of different depths. Contains more edge information, while deep features Contains more high-level semantic information. Enc1 primarily extracts global features, namely, the morphology, structure, and brightness distribution of the target region in the MRI image. Enc2 primarily extracts boundary features, namely, the edge information enhanced by the method described in step 2. This allows the model to focus more on the boundaries of the target region, improving segmentation accuracy.

[0076] Step 3.2: Decoder stage. The task of the decoder is to gradually restore the spatial resolution of the image while preserving the features extracted in the encoder stage. In the decoder stage, upsampling and concatenation operations are used to fuse the features extracted by the two encoders to restore the detailed information of the image. This stage consists of four components, each of which receives the features of the corresponding layers of the two encoders through skip connections. The specific method is as follows: First, obtain the features obtained from the encoder, assuming is the feature extracted by encoder Enc1 at layer i. Similarly, is the feature extracted by encoder Enc2. The feature after channel splicing is:

[0077]

[0078] Concat(*) represents concatenation along the channel dimension. Subsequently, the decoder uses skip connections and step-by-step upsampling to reconstruct features. The specific process is as follows:

[0079] The feature from the last layer of the encoder is P5, which is directly passed to the decoder layer as the initial input D5. In the decoder layer, the feature corresponding to the encoder layer is P i (i=1,2,3,4), and then upsample the output of the previous decoder layer to transform it into i (i=1,2,3,4) with the same spatial size for subsequent skip connections. The output after upsampling is:

[0080] U i =Interpolate(D i+1 ,size=P i ),i∈{1,2,3,4} (30)

[0081] Among them, Interpolate is an upsampling operation, and bilinear interpolation method is used here. i With encoder feature P i Concatenate in the channel dimension:

[0082] C i =Concat(U i ,P i ),i∈{1,2,3,4} (31)

[0083] Then C i Perform two 3×3 convolutions and ReLU activations to obtain the decoding output D of the current layer i :

[0084] D i =(ReLU(Conv3×3(ReLU(Conv3×3(C i)))) (32)

[0085] Among them, Conv3×3 represents 3×3 convolution, and ReLU is a nonlinear activation function. Finally, the top layer D1 undergoes 1×1 convolution to generate the final output Y:

[0086] Y=Conv1×1(D1) (33)

[0087] The segmentation model generated by the above steps is used to minimize the cross entropy loss L CE To train a deep segmentation network f:

[0088]

[0089] Where C is the number of categories and N is the total number of pixels. i,c is the one-hot encoded true label, y i,c =1 means pixel i belongs to category C, is the predicted probability after Softmax normalization.

[0090] The segmentation model uses Adaptive Moment Estimation (Adam) as the optimizer, with a learning rate decay type of cosine decay and a momentum of 0.9. The learning rate starts at 1e-4 and the model is trained for 100 epochs with a batch size of 16.

[0091] To fully validate the effectiveness of the proposed image segmentation model, experiments were conducted using partial data from BraTS2017 (referred to as BT2017) and data provided by Beijing Tiantan Hospital (referred to as BT-BTH). In step 3, the Dice similarity coefficient (Dice), 95% Hausdorff distance (HD95), and intersection over union (IoU) were used to evaluate the performance of the proposed method. For each metric, the mean and 95% confidence interval were calculated. Confidence intervals were determined using bootstrap analysis, including 5000 resampling iterations, to ensure statistical robustness and accuracy of the results.

[0092] In this study, different model algorithms were compared, and the experimental results are shown in Table 1:

[0093] Table 1 Experimental results of different model algorithms on two datasets

[0094]

[0095] The table above quantitatively compares the performance of the proposed method with three other methods on the BT2017 and BT-BTH datasets using three evaluation metrics. The proposed method achieves the best performance in terms of the Dice coefficient, HD95, and Intersection over Union (IoU), demonstrating its superiority in medical image segmentation tasks. SegNet also performs well, particularly on the BT-BTH dataset, where it outperforms U-Net and DeepLabv3+ in terms of Dice and IoU. DeepLabv3+ performs poorly on both datasets. U-Net performs well on the BT2017 dataset but performs only moderately well on the BT-BTH dataset, indicating that its performance is significantly affected by the characteristics of the dataset.

[0096] In order to visually compare the segmentation effects of different model algorithms, the prediction results of UNet, Deeplabv3+, SegNet and the method proposed in this paper on the BT2017 and BT-BTH datasets are visualized, as shown in the figure. Figure 3 and Figure 4 As shown in the figure, it can be seen that Deeplabv3+ has the lowest segmentation accuracy on both datasets, with large areas of misprediction. UNet performs relatively well in the boundary segmentation task, but fails to fully capture the target area in the prediction of some images. SegNet's segmentation accuracy in boundary areas is insufficient, and its boundary segmentation performance on the BT-BTH dataset decreases significantly. In contrast, the method proposed in this paper shows significant advantages in both segmentation accuracy and capturing boundary details, and can more accurately identify the boundary features of the target area. This effective capture of boundary information further verifies the superiority and robustness of this method.

[0097] Compared with the prior art, the present invention has the following beneficial effects:

[0098] First, this paper proposes a brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion. Two parallel branches are designed to process contrast and boundary information, respectively, processing the luminance and color channels of the MRI image to improve the separability of contrast and boundary information. By integrating these two types of information through weighted fusion, the model's ability to perceive the boundaries of the target region is strengthened.

[0099] 2. The present invention adopts a brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion, and proposes an image segmentation model based on a dual encoder structure. The model includes two inputs and two encoders, and combines boundary information and the global context information of the original image. The model can learn more comprehensive image features, thereby more effectively guiding the segmentation task.

[0100] 3. The brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion adopted in the present invention optimizes the encoder in the dual encoder structure to a VGG encoder with higher feature extraction capability, further enhances the feature representation capability of the model, and realizes refined segmentation of the target area.

Claims

1. A brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion, characterized in that It includes the following three modules: data preprocessing module (1), image enhancement module (2), and image segmentation module (3); (1) Data preprocessing module This module resamples the data, extracts two-dimensional slices, filters, and divides the data into training and test sets; (2) Image enhancement module This module first converts the input image from RGB color space to LAB color space, then enhances the brightness channel and color channel respectively, extracts and optimizes the brightness feature and boundary feature respectively, and fuses the two to generate the enhanced image; (3) Image segmentation module This module uses a brain tumor image segmentation model based on a dual-encoder structure. The dual-encoder architecture consists of two independent encoders, one for processing the original MRI image and the other for processing the enhanced image, to enhance the learning ability of image details and global information. In the decoder stage, the feature maps from the two encoders are combined to restore the spatial resolution of the image.

2. The brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion according to claim 1, characterized in that: In step (2), the method for performing image enhancement on the brain tumor MRI image is specifically as follows: 2.

1. Convert the MRI image from RGB color space to LAB color space, separate the luminance channel (L) from the LAB image, and perform local enhancement using the contrast-limited adaptive histogram equalization (CLAHE) method; 2.

2. The MRI image is segmented into superpixels using the Simple Linear Iterative Clustering (SLIC) algorithm. The superpixel segmentation results are then used to extract the boundary information of the target region. The average value of each superpixel segmentation region under the color channel is calculated and used to replace the values ​​of all pixels in the superpixel region. 2.

3. Using weighted fusion technology, the image after brightness enhancement is fused with the image after boundary enhancement to generate an enhanced image.

3. The brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion according to claim 1, characterized in that: Step (3) is specifically as follows: 3.

1. The model adopts a dual-encoder architecture, consisting of two independent encoders, one for processing the original MRI image and the other for processing the edge-enhanced image. Each encoder extracts image semantic features through convolution operations to enhance the learning ability of image details and global information. The encoder of the image segmentation model includes a dual encoder structure, wherein the dual encoder is composed of two independent feature extractors, at least one of which is a VGG network; The image segmentation model encoder uses only the convolutional part as a feature extractor and removes the fully connected layer to ensure that each layer of feature map participates in the decoding process, thereby improving pixel-level prediction accuracy. The dual encoder of the image segmentation model includes two input data: input I1 is the original MRI brain tumor image, represented as a 3-channel tensor; input I2 is the image processed by the boundary enhancement method, represented as a 3-channel tensor; encoder Enc1 is used to extract global features from the original MRI image, and encoder Enc2 is used to extract boundary and local features from the boundary-enhanced image to obtain complementary feature information; The dual encoder extracts features from the original image and the boundary-enhanced image through independent feature extraction processes, and combines the two through feature splicing or fusion operations; 3.

2. The decoder stage of the image segmentation model is used to fuse the different features extracted by the two encoders in the encoder stage and gradually restore the spatial resolution of the image; The feature splicing module is used to receive the corresponding layer features from encoder Enc1 and encoder Enc2 and splice them. The splicing is performed in the channel dimension, wherein the spliced ​​feature P i By formula: in and are the features extracted by the first and second encoders at layer i, i = 1, 2, 3, 4, 5; Each layer of the decoder receives features from the corresponding layers of the two encoders by using skip connections and restores the spatial information of the image through step-by-step upsampling and convolution operations.

4. The method according to claim 1, wherein The brain tumor MRI image segmentation method based on boundary enhancement and dual encoder fusion includes at least one of brightness channel enhancement, color channel enhancement, fusion of brightness enhancement and color enhancement, construction of an image segmentation model based on a dual encoder, independent extraction of image features by the dual encoder, and fusion of dual encoder features in the decoding stage.

5. The method according to claim 1, wherein The following steps are involved: Step 1: Data preprocessing is as follows: Step 1.1: First, resample the input 3D data to 3D size. Resample the 3D image of size 240×240×155 to 256×256×44. The resampling factor of each dimension is: Among them, i represents the dimension index, referring to the different dimensions of the image (x, y, z dimensions), originSize i is the size of the original image in the i-th dimension, newSize i is the size of the target image in the i-th dimension; next, the mapping of the target image coordinates to the original image coordinates is calculated. For each target coordinate (x′, y′, z′), the target image coordinates are converted to the original image coordinates using the following formula: x=x′×fx (2) y=y′×fy (3) z=z′×fz (4) Where (x, y, z) is the coordinate of the original image calculated from the coordinate of the target image; finally, the nearest neighbor interpolation method is used to obtain the pixel value of the corresponding coordinate: I′(x′,y′,z′)=I(x ≈ ,y ≈ ,z ≈ ) (5) Where I(x,y,z) is the pixel value at the coordinate (x,y,z), (x ≈ ,y ≈ ,z ≈ ) indicates that (x, y, z) is rounded to the nearest integer, and I′(x′, y′, z′) is the pixel value after resampling; Step 1.2: Slice the input 3D image data and slice the 256×256×44 3D MRI image into 44 256×256 2D images along the axial plane. Step 1.3: Filter the slices according to the labels. For each slice, check whether the target region exists in the label image. If the pixel value of a slice in the label image is greater than 0, it means that the slice contains the target region. Otherwise, remove the slice. Step 1.4: Randomly divide the image collection into training set and test set with a ratio of 8:2; Step 2: Image Enhancement Step 2.1: MRI images are typically stored in RGB format. However, in RGB space, brightness and color information are coupled together, making it difficult to enhance brightness alone. Therefore, the input RGB image is converted to CIELAB color space, where the L channel represents brightness and the A and B channels represent color information. The conversion from RBG color space to LAB color space is a two-step process. First, the RGB color space is converted to XYZ space: This matrix multiplication converts the color values ​​in the RGB color space Convert color values ​​to XYZ color space Then, convert from XYZ space to LAB space: where X n , Y n , Z n It is the reference white point standard value, the default values ​​are 95.047, 100.0, 108.883; L * , a * , b * It is the value of the three channels of the final LAB color space; Secondly, the L channel is separated from the LAB image for subsequent contrast enhancement operations; CLAHE (Contrast Limited Adaptive Histogram Equalization) is used for local enhancement. The image is first divided into small grids of 8×8, and the pixel values ​​in each grid are adjusted by local histogram equalization to adjust the contrast; assuming that L * (x,y) is the original image L * The brightness value of the pixel with coordinate position (x, y) in the channel, and (x', y') is the position of the pixel in the divided grid, then L * (x′, y′) represents the original image channel L * The brightness value of the pixel at the (x′, y′) position in the divided grid; Assume that all pixel values ​​in the grid are L grid , then the local histogram H local (k) is expressed as: Where k is the discretization level of the brightness value, k∈[0,255], and δ is the indicator function: In order to avoid excessive enhancement of local contrast, CLAHE introduces a contrast limit value clipLimit to limit the maximum cumulative frequency of the local histogram; the cumulative frequency of the local histogram is limited to a maximum value C, and the clipping value C of each grid is: where N grid is the total number of pixels in each small grid area, representing the size of the local area; L max is the maximum brightness value of the image, representing the upper limit of brightness in the image; in this method, the contrast limit value clipLimit is set to 2.0; if the local histogram value H local (k) exceeds C, then truncation is performed, and the local histogram value H of each brightness level local (k) is adjusted to: Next, the local histogram is equalized by normalizing the cumulative distribution function (CDF) of each local histogram. The brightness value after local equalization is: The accumulation process here is performed in each grid, and the output brightness value will be adjusted to [0,L max -1] range; finally, the equalized L of all grids * The channel values ​​are reassembled into a whole image; each pixel L of the original image * (x,y) is replaced by its new value in the equalized local region: L clahe (x,y)=L enhanced (x′,y′) (16) Where (x′, y′) is the coordinate of the pixel (x, y) in the grid; the enhanced L is obtained through the above steps clahe image; Step 2.2: Then, the SLIC (Simple Linear Iterative Clustering) algorithm is used for superpixel segmentation. This algorithm divides the image into several superpixel regions of similar size and shape through iterative optimization. The steps of the SLIC algorithm include: (1) Initialize the superpixel center point; suppose the image has N pixels and we want to divide the image into K superpixel regions. Then the size of each superpixel is N / K, and the distance between adjacent centers is: (2) Calculate the distance between each pixel and the superpixel center; the algorithm selects a 2S×2S area within the area of ​​the cluster center distance S to calculate the distance between the pixel and each cluster center; the total distance D is calculated by the color distance d lab and spatial distance d xy It consists of two parts; m is the compactness factor, which is used to measure the proportion of color value and spatial value in the similarity metric; Among them, l, a, and b are the three channels of the Lab color space, p represents the current pixel, e represents the superpixel seed point, and l o Indicates the brightness value of point o, a o Indicates the green-red value of point o, b o Indicates the blue-yellow value of point o, x o Indicates the x coordinate of point o, y o Indicates the y coordinate of point o; (3) Iteratively update the superpixel center position; the specific steps are to classify the pixels and mark the category of each pixel as the category of the cluster center with the smallest distance; then recalculate the cluster center, calculate the average vector value of all pixels belonging to the same cluster, and obtain the cluster center again; continue to iterate until the set maximum number of iterations is reached, and the clustering ends; the maximum number of iterations is set to 10; Traverse each superpixel region obtained by the SLIC algorithm, calculate the average value of the B channel in each region, and use the average value to replace the values ​​of all pixels in this superpixel region; the pixel value in this superpixel region is: Among them, r is the current super pixel area, N r is the number of pixels in the current superpixel area, b(x′, y′) represents the B channel value of the position (x′, y′) in the original image; Step 2.3: A dual-branch information fusion method is used to perform weighted fusion of image information from different branches, thereby achieving comprehensive enhancement of image features. The specific method is: using weighted fusion technology, the image enhanced by CLAHE is weightedly fused with the image obtained after boundary enhancement. I fused (x,y)=α·B(x,y)+β·L clahe (x,y) (22) Among them, α and β are weighting coefficients, and their values ​​are 0.5; B(x, y) is the image after boundary enhancement in step 2.2, and L clahe (x, y) is the image after brightness enhancement in step 2.1; finally, the weighted fusion image I fused (x,y) is applied to the three RGB channels to generate the final composite image; Step 3: Image Segmentation Step 3.1: Dual Encoder Stage: Two independent VGG16 networks are used as encoders in the segmentation architecture to extract feature information of the target region from different aspects. Based on the VGG16 network, this stage only uses the convolutional part as the feature extractor, without the fully connected layer. The model includes two input data, where input I1 is the original MRI image, represented as a 3-channel tensor with a size of 256×256×3, and input I2 is the image obtained after boundary enhancement, represented by a tensor with a size of 256×256×3; the two input data are used for two independent encoders to obtain complementary feature information; encoder Enc1 is responsible for extracting global features from the original MRI image, and encoder Enc2 processes the image enhanced by step 2; each encoder contains five convolution blocks, the first two convolution blocks each have two 3×3 convolution layers, and the last three convolution blocks each have three 3×3 convolution layers. Each convolution layer has a ReLU activation function and is downsampled using a 2×2 maximum pooling layer; the calculation process of encoder Enc1 is as follows: The convolution and pooling process of the first layer is: Where * represents the convolution operation, W1 is the convolution kernel weight of the first layer, b1 is the bias term, ReLU(x)=max(0,x) is the activation function, and MaxPool(x) represents the maximum pooling operation; The subsequent convolution and pooling process is: The calculation process of encoder Enc2 is the same as that of encoder Enc1. Ultimately, each encoder extracts five levels of features: These features come from network layers of different depths. Contains more edge information, while deep features Contains more high-level semantic information; Step 3.2: Decoder stage; This stage contains four components, each of which receives the features of the corresponding layers of the two encoders through a jump connection; the specific method is: first obtain the features obtained from the encoder, assuming is the feature extracted by encoder Enc1 at layer i. Similarly, It is the feature extracted by encoder Enc2; the feature after channel splicing is: Concat(*) represents concatenation along the channel dimension. Subsequently, the decoder uses skip connections and step-by-step upsampling to reconstruct features. The specific process is as follows: The feature from the last layer of the encoder is P5, which is directly passed to the decoder layer as the initial input D5; in the decoder layer, the feature corresponding to the encoder layer is P i (i=1,2,3,4), and then upsample the output of the previous decoder layer to transform it into i (i=1,2,3,4) with the same spatial size; the output after upsampling is: U i =Interpolate(D i+1 ,size=P i ),i∈{1,2,3,4} (30) Among them, Interpolate is an upsampling operation, and the bilinear interpolation method is used here; U i With encoder feature P i Concatenate in the channel dimension: C i =Concat(U i ,P i ),i∈{1,2,3,4} (31) Then C i Perform two 3×3 convolutions and ReLU activations to obtain the decoding output D of the current layer i : D i =ReLU(Conv3×3(ReLU(Conv3×3(C i )))) (32) Among them, Conv3×3 represents 3×3 convolution, ReLU is a nonlinear activation function; finally, the top layer D1 undergoes 1×1 convolution to generate the final output Y: Y=Conv1×1(D1) (33) The segmentation model generated by the above steps is used to minimize the cross entropy loss L CE To train a deep segmentation network f: Where C is the number of categories, N is the total number of pixels; i,c is the one-hot encoded true label, y i,c =1 means pixel i belongs to category C, is the predicted probability after Softmax normalization; The segmentation model uses adaptive moment estimation as the optimizer, the learning rate decay type is cosine decay, and the momentum is 0.9; the learning rate starts from 1e-4, and the model is trained for 100 epochs with a batch size of 16.