Semi-supervised medical image tumor segmentation method based on boundary enhancement
By introducing boundary enhancement technology and multi-task learning in the semi-supervised medical image segmentation model, combining boundary area segmentation and symbol distance field prediction tasks, the problem of malignant tumor area recognition is solved, and segmentation accuracy and boundary feature extraction quality are improved.
Patent Information
- Application Number
- CN202411982213.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
Existing semi-supervised medical image segmentation models are difficult to accurately identify malignant tumor areas, especially in diffuse weak boundaries and low contrast conditions, and traditional teacher-student models rely heavily on advanced semantic labels, fail to fully utilize the data and difficult to transfer the knowledge of the teacher model.
The boundary enhancement semi-supervised tumor segmentation network based on teacher knowledge accumulation is adopted, and the two auxiliary tasks of boundary area segmentation and symbol distance field prediction are combined. By designing boundary feature enhancement modules and multi-task learning, the weak boundary information and boundary structure recognition ability of tumor images are enhanced.
It effectively alleviates the problems of blurred boundary features and low contrast in tumor images, improves the quality of boundary feature extraction and tumor segmentation accuracy, and achieves more accurate identification of malignant tumor areas.
Smart Images

Figure CN120013957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision technology and image processing, and in particular to medical image tumor segmentation using semi-supervised learning and boundary enhancement. Background Art
[0002] Tumors are masses formed due to the uncontrolled proliferation of cells. Malignant tumors tend to grow infiltratingly and form diffuse excessively weak boundaries with surrounding normal tissues, which is an important basis for determining benign and malignant tumors in clinical practice. Segmenting diffuse weak boundaries helps clinicians inhibit the overall growth and spread of tumors. Magnetic resonance imaging (MRI) can provide high-resolution and multi-contrast medical images of the internal structure of the human body, thereby providing clinicians with more accurate information about tumors. However, malignant tumors have different shapes, diffuse edges, and low-contrast differences between foreground and background, which makes it difficult for traditional semi-supervised medical image segmentation models to accurately identify tumor areas. In addition, the traditional teacher-student semi-supervised model relies heavily on the high-level semantic labels provided by the teacher model. This method cannot make full use of the data and it is difficult to transfer the knowledge accumulated by the teacher model itself through multiple iterations to the student model.
[0003] In order to solve the above problems, the present invention adopts a boundary-enhanced semi-supervised tumor segmentation network based on teacher knowledge accumulation to perform tumor segmentation. At the same time, the present invention adopts two auxiliary tasks, boundary region segmentation (BR) and signed distance field (SDF) prediction, to guide the module to pay attention to the tumor boundary position information and structural information at the same time. Many algorithms have been proposed in the field of tumor segmentation, which are mainly divided into: uncertainty-based, boundary structure-aware and sample-balanced methods. Zou et al. proposed an end-to-end trusted medical image segmentation model TBraTS for brain tumors to quantify voxel-level uncertainty. The uncertainty is explicitly modeled using subjective logic theory, and the prediction of the backbone neural network is regarded as subjective opinion by parameterizing the segmented category probability as a Dirichlet distribution. Liu et al. proposed a boundary-aware consistent hidden representation learning network, in which the two branches share the same encoder and each has an independent decoder. While achieving consistency by perturbing the high-level hidden feature representation, a boundary-aware map is introduced to capture the organ boundary. You et al. proposed two loss weighting strategies, namely distribution-aware debiasing weighting (DistDW) and difficulty-aware debiasing weighting (DiffDW), using pseudo-label dynamic guidance models to solve data distribution and learning difficulty deviations.
[0004] In addition, the present invention designs a boundary feature enhancement module based on the traditional boundary extraction operation, which allows the teacher-student model to interact with the underlying semantics and utilizes the accumulated knowledge of the teacher model to enhance the intermediate features of the student model. Existing semi-supervised learning work is mainly divided into: based on pseudo-labels, based on consistency constraints, and based on prior knowledge learning. The work of Lee et al. proposed that pseudo-labels with higher confidence are usually closer to the distribution of true labels. Therefore, many uncertain measurement methods are proposed to generate more stable pseudo-labels. Hu et al. proposed attention-guided consistency, which encourages the attention maps of student models and teacher models to be consistent. Each image contains the same class objects, so different images share similar semantics in the feature space. Huang et al. added a reconstruction pre-training strategy from the proxy task to extract meaningful information to promote the next stage of learning. Yang et al. introduced a self-supervised jigsaw task in the semi-supervised training process to obtain better feature representation. The above researchers all designed semi-supervised models based on the traditional teacher model, proving the potential of the teacher model in semi-supervised learning. Summary of the invention
[0005] The purpose of the present invention is to provide a semi-supervised medical image tumor segmentation method based on boundary enhancement. By designing a semi-supervised learning network model and a multi-scale boundary enhancement module, the weak boundary information of the tumor image is enhanced, and multi-task learning is introduced. By designing boundary area segmentation tasks and distance field prediction tasks, the recognition ability of boundary structure and fine-grained position is enhanced, which effectively solves the problem of diffuse boundaries and low contrast inside and outside the boundaries in existing algorithms.
[0006] In order to achieve the above purpose, the present invention adopts the following technical scheme: The method of this paper first designs a semi-supervised learning network based on the traditional teacher-student model, and at the same time designs a boundary enhancement module based on 8-direction boundary extraction. The foreground boundary feature learning of medical images is enhanced at different scales. We input the medical image into the semi-supervised network model, extract the tumor features after boundary enhancement, and constrain the boundary expression ability of the features through the boundary area segmentation task and the distance field prediction task. Finally, the enhanced tumor features are constrained by cross-task and cross-pseudo-supervision, and finally decoded to the original image size and calculated the segmentation map.
[0007] A semi-supervised medical image tumor segmentation method based on boundary enhancement comprises the following steps:
[0008] Step 1: Acquisition of foreground tumor features of training data.
[0009] Step 1.1, input training images.
[0010] In step 1.2, the image passes through the teacher-student semi-supervised backbone model to obtain multi-scale image features.
[0011] Step 2: Boundary information enhancement.
[0012] Step 2.1, obtain the teacher-student tumor image features at different scales in the encoding stage.
[0013] Step 2.2, fuse the teacher-student features at the same level.
[0014] In step 2.3, the 8-direction boundary feature extraction method is used to extract boundary information from the fused features and calculate the attention score.
[0015] In step 2.4, the calculated attention scores are used to guide the teacher boundary features to be superimposed on the intermediate features of the student model, thereby enhancing the boundary information of the student model features.
[0016] Step 3: Boundary task constraints.
[0017] In step 3.1, the enhanced multi-scale tumor features are sent to the boundary task module.
[0018] In step 3.2, the boundary region segmentation task is applied to segment the tumor boundary part, and the distance field prediction task is applied to calculate the minimum distance from each pixel to the foreground area.
[0019] Step 3.2, compare the predicted boundary area segmentation map and distance field calculation map with the real segmentation map and calculation map, as well as the segmentation map and calculation map inferred by the teacher model to constrain the model.
[0020] Step 4: image to be segmented.
[0021] Step 4.1: Input the image to be segmented into the trained network model to obtain the image features after tumor boundary enhancement.
[0022] Step 4.2, upsample the output image features to obtain a segmentation map of the original image size as the segmentation result.
[0023] Compared with the prior art, the present invention has the following obvious advantages:
[0024] The present invention uses the proposed feature enhancement module at the encoding end to enhance the boundary features of multiple levels, and introduces an eight-directional Sobel operator to extract the edge information of the intermediate features of the teacher model, effectively alleviating the problems of blurred boundary features and low contrast of tumor images. The present invention designs boundary tasks, including additional boundary area segmentation tasks and distance prediction tasks, to further enhance the boundary enhancement capability, effectively alleviating the problems of poor recognition accuracy caused by complex tumor structure and variable shapes. Through the effective combination of the two methods proposed above, the quality of boundary feature extraction and the accuracy of tumor segmentation by the model are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flow chart of the overall method of the present invention;
[0026] Figure 2 It is a flow chart of the edge enhancement module Edgeformer of the present invention; DETAILED DESCRIPTION
[0027] The present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0028] The flow chart of the method of the present invention is as follows Figure 1 As shown, the specific steps include:
[0029] Step 1: Acquisition of foreground tumor features of training data.
[0030] Step 1.1, input training images.
[0031] In step 1.2, the image passes through the teacher-student semi-supervised backbone model to obtain multi-scale image features.
[0032] The images to be trained are processed by digital image processing using histogram normalization to control the grayscale value of the image to between 0 and 1 to prevent gradient vanishing and gradient explosion during the training process. The images to be trained are sent to the teacher model and the student model respectively, and the input images are processed using 3×3 convolution to convert the image channel to the hidden channel h. The encoder can extract high-level features with rich boundary information from each modality. After multiple steps, intermediate feature maps at different scales are obtained in the process of downsampling.
[0033] Step 2: Boundary information enhancement.
[0034] Step 2.1, obtain the teacher-student tumor image features at different scales in the encoding stage.
[0035] Step 2.2, fuse the teacher-student features at the same level.
[0036] In step 2.3, the 8-direction boundary feature extraction method is used to extract boundary information from the fused features and calculate the attention score.
[0037] In step 2.4, the calculated attention scores are used to guide the teacher boundary features to be superimposed on the intermediate features of the student model, thereby enhancing the boundary information of the student model features.
[0038] In the encoding phase, we design the EdgeFormer module based on MetaFormer to enhance the boundary information of each layer of intermediate features. Specifically, we apply multiple EdgeFormer modules alternately in the encoding layers of the teacher and student models to extract frontier boundary information at different scales. The teacher and student models share the parameters and gradients of the modules. During the training of the student model, the EdgeFormer module incorporates the temporal global historical boundary information features f of the teacher model into the encoding layer of the student model. q The temporal local boundary information feature f of the current iteration with the student model k interact. Then the 8-way sobel operator is used to calculate the fused boundary contour information. Finally, the attention score is calculated by superimposing the original features of the teacher model on the boundary features to guide the teacher's historical information to enhance the boundary information of the student model features. For the forward process of the teacher model, we adopt the self-attention method. The input of Edgeformer is all the intermediate features of the teacher model, which self-guide the model for enhancement.
[0039] Key Features k and query feature f q After two parallel 1×1 convolutional layers, preliminary feature extraction is performed and then concatenated in the channel dimension. The concatenated feature map is further processed by the Atrous Spatial Pyramid Pooling (ASPP) algorithm and a 3×3 convolutional layer to enhance the expressiveness of the features. Subsequently, the Sobel Multi-Directional Boundary Extractor (SMDBE) extracts and integrates the overall contour information by refining the Sobel operator into 8 directions. The original query feature f q The and contour features are combined through skip connections to enhance the effectiveness of information transmission, and then feature integration is performed through multi-point convolutional layers and batch normalization (BN) layers. Finally, the sigmoid function is used to obtain the attention score map of BR. The attention map is combined with the key feature f k Multiply them together to obtain features with enhanced boundary information.
[0040] Step 3: Boundary task constraints.
[0041] In step 3.1, the enhanced multi-scale tumor features are sent to the boundary task module.
[0042] In step 3.2, the boundary region segmentation task is applied to segment the tumor boundary part, and the distance field prediction task is applied to calculate the minimum distance from each pixel to the foreground area.
[0043] Step 3.2, compare the predicted boundary area segmentation map and distance field calculation map with the real segmentation map and calculation map, as well as the segmentation map and calculation map inferred by the teacher model to constrain the model.
[0044] This paper introduces the distance field (SDF) prediction task and the boundary region (BR) segmentation task, where SDF measures the Euclidean distance of each pixel to the nearest contour, which can constrain the continuity of the pixel-level boundary shape. The BSR segmentation task emphasizes the focus on the boundary region and identifies the location of the boundary pixel level. For the labeled image x∈D L , the corresponding foreground segmentation label is Y, and the BR segmentation label is B. The calculation from Y to B is express. The operation is calculated by subtracting the minimum pooling result from the maximum pooling result, with a pooling radius of 3. Therefore, B contains the inner and outer boundary areas of the foreground contour as shown in the following formula:
[0045]
[0046] During the student model training, the output probability maps of the BR segmentation task and the SDF prediction task are b stu and stu The corresponding true values are B and S. We define the SDF operation as Right now Following recent semi-supervised methods, we define a supervised BR segmentation loss and SDF prediction loss The formula is as follows:
[0047]
[0048]
[0049] in, It is the cross entropy loss focal loss that balances difficult samples and few samples. is the boundary distance metric loss, and is the mean square error loss. Applies only to tag sets.
[0050] For the output s of the SDF prediction task stu , which is guided by the foreground segmentation task and the BR segmentation task. The two task decoders output probability maps y stu and b stu. The prediction results are obtained through the argmax function and and its SDF and As pseudo labels for the SDF prediction task. The MSE loss is used to calculate the SDF cross-task pseudo-supervision loss, and the calculation formula is as follows:
[0051]
[0052] Cross-pseudo-supervision guidance is used for BR and foreground segmentation tasks. We use a mask boundary extraction operation. ), we use the traditional maximum pooling and minimum pooling difference operations Focus on capturing mutation information, thereby improving the model's ability to smooth non-boundary areas. ), we use the difference between average pooling and max pooling To smooth these areas, we use the focal loss and boundary distance metric to obtain the cross pseudo-supervision loss as follows: The calculation formula is as follows
[0053]
[0054]
[0055]
[0056]
[0057] The formula for the overall cross-task consistency loss is:
[0058] In addition to the cross-task pseudo-supervision loss, we also introduce a cross-model pseudo-supervision loss. Specifically, the pseudo foreground label and pseudo boundary label of the teacher model are the learning objectives of the corresponding tasks of the student model. The output of the foreground segmentation task of the teacher model is y tea , SDF prediction task output s tea And BR segmentation task output b tea The corresponding prediction results are as well as The cross-model pseudo-supervision loss formula for the corresponding task is as follows:
[0059]
[0060] Step 4: image to be segmented.
[0061] Step 4.1: Input the image to be segmented into the trained network model to obtain the image features after tumor boundary enhancement.
[0062] Step 4.2, upsample the output image features to obtain a segmentation map of the original image size as the segmentation result.
[0063] Through the above steps, the present invention has completed the construction and training of the model. Next is the test part of the model. The original medical image to be segmented is input into the network model, and the probability map of the original image size is obtained through the prediction of the model, in which each pixel value is expressed as the probability size of the predicted tumor area. The area part that is most likely to be the tumor area is calculated as the predicted foreground part, and the other parts are used as the background part to obtain the final predicted segmentation map as the final result.
[0064] In order to verify the generalization of the model on different data, we randomly selected a part of the data from the data set as the test set, and the rest as the training set, to ensure that the distribution of the training set and the test set is consistent to the greatest extent. Through multiple trainings, we found the optimal hyperparameters of the model. When the model is optimal, we randomly split the test set and the training set for multiple times to train the model, and tested the converged network on the test set to obtain multiple groups of test results. The average is taken as the test result of the final model for model evaluation. The experimental results are shown below:
[0065] The present invention conducted experiments on the BraTS2020 brain tumor segmentation dataset and the Kvasir-SEG gastroscopy tumor dataset. The experimental results are shown in Table 1 and Table 2 respectively.
[0066]
[0067] Table 1
[0068]
[0069] Table 2
[0070] The present invention uses several semi-supervised medical image segmentation benchmark methods for comparative experiments, including the fully supervised U-Net method, the URPC method based on uncertainty pseudo-label correction, the SASSNet method based on prior knowledge and structure perception, etc. We used four commonly used segmentation indicators, namely, accuracy Dice, intersection-over-union ratio Jaccard, boundary distance 95HD and ASD. Among them, the accuracy of the present invention is the highest among multiple methods. Experiments have proved that the algorithm of the present invention has achieved great performance improvement in the field of semi-supervised medical image tumor segmentation.
[0071] So far, the specific implementation process of the present invention has been described.
Claims
1. A semi-supervised medical image tumor segmentation method based on boundary enhancement, characterized in that: The following steps are involved: Step 1, obtaining foreground tumor features of training data; Step 1.1, input training image; Step 1.2, the image passes through the teacher-student semi-supervised backbone model to obtain multi-scale image features; The images to be trained are processed by digital image processing using histogram normalization to control the grayscale value of the images to between 0 and 1 to prevent gradient vanishing and gradient explosion during the training process. The images to be trained are sent to the teacher model and the student model respectively, and the input images are processed using 3×3 convolution to convert the image channels into hidden channels h. The encoder can extract high-level features with rich boundary information from each modality. After going through the above steps multiple times, intermediate feature maps at different scales are obtained during downsampling. Step 2: boundary information enhancement; Step 2.1, obtaining teacher-student tumor image features at different scales in the encoding stage; Step 2.2, fuse the teacher-student features at the same level; Step 2.3, use the 8-direction boundary feature extraction method to extract boundary information from the fused features and calculate the attention score; Step 2.4, using the calculated attention score to guide the teacher boundary features to be superimposed on the intermediate features of the student model, thereby enhancing the boundary information of the student model features; On the teacher-student network, a boundary enhancement module is designed based on the teacher's historical knowledge accumulation; An 8-dimensional Sobel operator is used to capture more time-accumulated boundary information from the teacher model in real time, calculate the attention score, and guide the boundary features of the teacher model to be integrated into the student model to enhance its weak boundary features; the features of each scale are enhanced and then sent to the next scale convolution layer for encoding; Step 3, boundary task constraints; Step 3.1, sending the enhanced multi-scale tumor features to the boundary task module; Step 3.2, applying the boundary region segmentation task to segment the tumor boundary part, and applying the distance field prediction task to calculate the minimum distance from each pixel to the foreground area; Step 3.3, compare the predicted boundary area segmentation map and distance field calculation map with the real segmentation map and calculation map, and the segmentation map and calculation map inferred by the teacher model to constrain the model; For the enhanced multiple scale features, they are upsampled to the same scale through transposed convolution, and point convolution is used for channel integration. The integrated features are sent to the boundary task module, and the boundary features enhanced by the distance field prediction task and boundary area segmentation task constraints are used respectively. The boundary area segmentation task enables the model to perceive the specific location of the boundary pixels, while the distance field prediction task provides prior knowledge of the tumor shape structure, so that the model pays more attention to the boundary features of the tumor. Step 4, image to be segmented; Step 4.1, input the image to be segmented into the trained network model to obtain the image features after tumor boundary enhancement; Step 4.2, upsampling the output image features to obtain a segmentation map of the original image size as the segmentation result; For the image to be segmented, the normalized medical image is input into the network model, and the probability map of the original image size is obtained through the model's prediction, in which each pixel value represents the probability size of the predicted tumor area; the area most likely to be the tumor area is calculated as the predicted foreground part, and the other part is used as the background part to obtain the final predicted segmentation map; the supervised loss uses focal loss to encourage the model to perform additional weighted processing on few samples and difficult samples, thereby increasing the model's attention to the tumor foreground; the unsupervised loss uses cross-task loss and cross-model pseudo-supervised loss to encourage different tasks of the student model to share information and consistency constraints, and at the same time use the pseudo-label of the teacher model as the training signal of the student model's unlabeled data.
2. The method for semi-supervised medical image tumor segmentation based on boundary enhancement according to claim 1, characterized in that: In steps 2.1 to 2.4, the intermediate features after boundary information enhancement are obtained: images are randomly extracted from the BraTS2020 dataset and the Kvasir-SEG data as training sets. For an image to be segmented, intermediate features of different scales are obtained through a multi-layer encoder, and the features of each scale are sent to EdgeFormer respectively. The sobel operator is used to extract the contour information and calculate the attention score, so as to guide the intermediate feature boundary information of the teacher model to be fused with the intermediate features of the student model.
3. The method for semi-supervised medical image tumor segmentation based on boundary enhancement according to claim 1, characterized in that: In steps 3.2-3.3, the boundary task constrains boundary features: randomly extract images from the BraTS2020 dataset and the Kvasir-SEG data as training sets, obtain enhanced features at different scales, upsample to the same scale for fusion, and use the distance field prediction task and boundary area segmentation task to constrain the boundary area and structural information, so that the model pays more attention to the boundary area in the segmentation results.
Citation Information
Cited By
Brain tumor segmentation method based on boundary perception mechanism
CN120976222A