Medical image segmentation system and method based on regional proposal

Through the medical image segmentation system based on the region proposal, the high brightness calibration and feature fusion module are used to solve the problem of inaccurate lesion region segmentation in the traditional method, and higher segmentation accuracy and understanding of the target position morphology are achieved, which is suitable for medical image segmentation tasks.

CN120339185APending Publication Date: 2025-07-18ZHEJIANG SHUREN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510314202.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-07
Filing Date
2025-03-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional medical image segmentation methods have poor effect on segmentation of lesion areas, especially when the lesion is small and not obvious, it is difficult to accurately segment. In addition, the position information is lost when the image resolution is reduced, and the U-shaped network has weak global context processing capabilities.

Method used

A medical image segmentation system based on region proposal is adopted, including a preprocessing module, an object detection module, a positioning module and a target segmentation module. The object detection module is used to obtain highlights, and a combination of image encoder, prompt encoder, mask encoder and boundary perception module is used to segment, and the attention module is introduced to capture feature dependencies, and the segmentation accuracy is improved through the feature fusion module and boundary perception module.

Benefits of technology

It improves the accuracy of target segmentation, solves the problem of rough mask boundaries in traditional methods, enhances the understanding of target position and morphology, and achieves higher segmentation accuracy and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339185A_ABST
    Figure CN120339185A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation system and method based on regional proposal. The medical image segmentation system comprises a preprocessing module, a target detection module, a positioning module and a target segmentation module. Performing target area detection on the image processed by the preprocessing module through a target detection module; a positioning module is used to acquire a highlight point of a target in the detected image according to the detected target area, and a target segmentation module is used to complete segmentation of the detected image according to the highlight point; according to the method, the highlight points are calibrated firstly, and then the detected image is segmented according to the highlight points, so that the problem of rough mask boundary caused by neglect of accurate segmentation of the target edge in a traditional target segmentation method is solved, and the accuracy of target segmentation is improved; meanwhile, an attention module is introduced into a target detection model and a mask decoder, so that the global dependency relationship among different features is better captured, and the model can better understand the position and form of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to a medical image segmentation system and method based on region proposal. Background Art

[0002] With the development of technology, medical image segmentation methods based on deep learning have become the cornerstone of building an efficient computer-aided diagnosis system, attracting extensive attention in the field of medical imaging. These methods greatly assist doctors in accurate diagnosis and treatment plan formulation by providing pixel-level accurate information of the lesion area, thereby improving the efficiency and accuracy of disease diagnosis and treatment. Their main purpose is to depict the lesion area of interest from the original image by labeling each pixel into a certain category, such as multiple abdominal organs and breasts, etc. This is one of the representative research topics in the fields of computer vision and medical image analysis. Accurate segmentation can provide reliable volume and shape information of the target structure, providing further assistance for clinical applications. Relevant research shows that its accuracy has exceeded that of human experts in specific fields.

[0003] The traditional methods of medical image segmentation mainly include the threshold method and the region edge detection method. However, these two methods are very sensitive to the contrast of the image. If the contrast between the lesion area and the surrounding tissues is not obvious, it may not be possible to accurately segment the lesion area. At the same time, they have poor adaptability to complex lesions. For complex lesions with irregular shapes and blurred boundaries, it is often difficult to accurately segment them. Especially in the case where the lesion is small and not obvious, the segmentation effect of the traditional method is not ideal. Therefore, many scholars have introduced mainstream deep learning methods such as convolutional neural networks and U-shaped networks into the field of medical image segmentation. Convolutional neural networks can automatically learn and extract the features of images. In a multi-layer structure, they can also learn feature hierarchies from simple to complex, deepening the understanding of the image content. However, in a deep network, due to continuous pooling operations, the spatial resolution of the image will gradually decrease, resulting in the loss of position information; while the U-shaped network combines low-level and high-level features through skip connections, which can effectively restore the details and local information of the image. At the same time, the encoder-decoder structure and skip connections enable the U-shaped network to fuse features at multiple scales, but the U-shaped network has relatively weak processing ability for the global context of the image. Summary of the Invention

[0004] The purpose of the present invention is to provide a medical image segmentation system and method based on region proposal.

[0005] In the first aspect, the present invention provides a medical image segmentation system based on region proposal, which includes a preprocessing module, a target detection module, a positioning module, and a target segmentation module; the target detection module is used to detect the target area of the preprocessed image; the positioning module is used to obtain the highlight points of the target in the measured image according to the detected target area;

[0006] The target segmentation module described above includes an image encoder, a prompt encoder, a mask encoder, a feature fusion module, and a boundary awareness module; the image encoder extracts image features of the preprocessed image; the prompt encoder extracts information features of the highlight positions; the mask encoder includes two decoder layers, an attention module, a transposed convolution layer, and a multi-layer perceptron; the two sequentially connected decoder layers process the image features and the information features to obtain an image feature vector and an information feature vector; the image feature vector is used as the input of the transposed convolution layer; the fusion result of the image feature vector and the information feature vector is used as the input of the attention module; the output of the attention module is used as the input of the multi-layer perceptron; the output result of the mask encoder is the result of multiplying the outputs of the transposed convolution layer and the multi-layer perceptron.

[0007] Preferably, the highlight is the pixel point corresponding to the maximum luminance value in the target area.

[0008] Preferably, the target detection module includes a backbone network, a head network, and multiple attention modules; the results of processing multiple outputs of the backbone network by the attention modules respectively are used as the input of the head network.

[0009] In a second aspect, the present invention provides a medical image segmentation method based on region proposals, which includes the following steps:

[0010] Step 1: Construct a medical image dataset and perform labeling processing on the images in the dataset;

[0011] Step 2: Construct a target detection model and use the dataset to train the target detection model;

[0012] Step 3: Use the trained target detection model to obtain the highlights of the target area;

[0013] Step 4: Construct a target segmentation model, which includes an image encoder, a prompt encoder, a mask encoder, a feature fusion module, and a boundary awareness module; the image encoder processes the image input into the target segmentation model through a twenty-four-layer self-attention mechanism feature extraction module to obtain image features; the prompt encoder is used to process the highlights to obtain information features; the mask encoder includes two decoder layers, an attention module, a transposed convolution layer, and a multi-layer perceptron; the two sequentially connected decoder layers process the image features and the information features respectively to obtain an image feature vector and an information feature vector; the image feature vector is used as the input of the transposed convolution layer to output a mask; the fusion result of the image feature vector and the information feature vector is used as the input of the attention module to output a token, and the token is input into the multi-layer perceptron; the output of the multi-layer perceptron is multiplied by the mask to obtain the mask output by the mask encoder;

[0014] The feature fusion module includes two upsampling layers in parallel; the two upsampling layers respectively process the shallow features and deep features output by the two self-attention mechanism feature extraction modules in the image encoder to obtain features with the same size as the mask; the fusion result of the features output by the two upsampling layers and the mask is used as the output of the feature fusion module; the boundary awareness module is used to process the output of the feature fusion module to obtain the output of the target segmentation model.

[0015] Step Five: Use the dataset to train the target segmentation model.

[0016] Step Six: Use the trained target detection model to perform target detection on the measured image to obtain the target area, and extract the highlight points in the target area; use the trained target segmentation model to complete the segmentation of the measured image based on the highlight points.

[0017] Preferably, in the second step, the target detection model includes a backbone network, a head network, and multiple attention modules; the results of processing the multiple outputs of the backbone network by the attention modules respectively are used as the input of the head network.

[0018] Preferably, in the first step, the label of the image includes the contour of the lesion area and the detection box containing the lesion area.

[0019] Preferably, use the dataset with the detection box of the lesion area as the image label to train the target detection model; use the dataset with the contour of the lesion area as the image label to train the target segmentation model.

[0020] Preferably, in the second step, the loss function L CIOU has the following expression:

[0021]

[0022] where IOU is the intersection over union; ρ(,·,) is the Euclidean distance; b and b gt respectively represent the center points of the predicted bounding box and the ground truth bounding box; c is the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box; a is the balance coefficient; v is the aspect ratio consistency between the predicted bounding box and the ground truth bounding box.

[0023] Preferably, the highlight point is the pixel point corresponding to the maximum brightness value in the target area.

[0024] Preferably, in the first step, the image is preprocessed before constructing the data set. The preprocessing method is as follows: Gaussian filtering is used for noise removal, and the Otsu algorithm is applied for global threshold segmentation to separate the breast tissue from the background; the mosaic data augmentation method is used to enhance the image.

[0025] The beneficial effects of the present invention are as follows:

[0026] 1. The present invention uses the target detection module to propose regions in the measured image, calibrates the highlight points in the target region of the measured image, and uses the target segmentation module to segment the measured image based on the highlight points, solving the problem of rough mask boundaries caused by neglecting the accurate segmentation of the target edge in the traditional target segmentation method and improving the accuracy of target segmentation.

[0027] 2. The present invention constructs a target segmentation model to segment the target region according to the calibrated highlight points, and introduces an attention module into the target detection model and the target segmentation model, which can better capture the global dependence relationship between different features, enabling the model to better understand the target position and shape. Description of the Drawings

[0028] Figure 1 is a flowchart of the present invention.

[0029] Figure 2 is a schematic structural diagram of the target detection model in the present invention.

[0030] Figure 3 is a flowchart of the target detection model in the present invention for detecting highlight points.

[0031] Figure 4 is a schematic structural diagram of the mask encoder in the present invention. Detailed Embodiments

[0032] The present invention will be further described below with reference to the accompanying drawings.

[0033] A region proposal-based medical image segmentation system includes a preprocessing module, an object detection module, a localization module, and an object segmentation module; the preprocessing module is used to preprocess the measured image; the object detection module is used to detect the target region of the preprocessed image; the object detection module adopts the Yolov7 model, which includes a backbone network, a head network, and multiple attention modules (CBAM); the results of processing the multiple outputs of the backbone network by the attention modules respectively are used as the input of the head network; the localization module is used to obtain the highlight points of the target in the measured image according to the detected target region; the object segmentation module includes an image encoder, a prompt encoder, a mask encoder, a feature fusion module, and a boundary awareness module; the image encoder is used to extract the image features of the preprocessed image; the prompt encoder is used to obtain information features according to the highlight points of the target; the mask encoder includes two decoder layers, an attention module, a deconvolution layer, and a multi-layer perceptron; the two decoder layers connected in sequence process the image features and information features respectively to obtain an image feature vector and an information feature vector; the image feature vector is used as the input of the deconvolution layer; the fusion result of the image feature vector and the information feature vector is used as the input of the attention module; the output of the attention module is used as the input of the multi-layer perceptron; the output result of the mask encoder is the result of multiplying the outputs of the deconvolution layer and the multi-layer perceptron.

[0034] As Figure 1 shown, the medical image segmentation method performed by the medical image segmentation system includes the following steps:

[0035] Step 1: Construct a dataset

[0036] Obtain FFDM (Full-Field Digital Mammography) images; use Gaussian filtering to remove noise, and apply the Otsu threshold method (Otsu algorithm) for global threshold segmentation to separate the breast tissue from the background; use the Mosaic data augmentation method to enhance the FFDM images; use the manually labeled lesion regions as labels to construct a dataset of labeled FFDM images; the labeled labels include the lesion region contour and the detection box containing the lesion region.

[0037] Step 2: As Figure 2 shown, construct an object detection model; the object detection model is used to detect the target region of the input FFDM image; in this embodiment, the Yolov7 model is used as the object detection model; the Yolov7 model includes a backbone network, a head network, and multiple attention modules; the results of processing the multiple outputs of the backbone network by the attention modules respectively are used as the input of the head network; the attention module includes a channel attention module and a spatial attention module;

[0038] 2-1. Channel attention module

[0039] The input of the Channel Attention Module (CAM) is the feature map F∈R output by the backbone network C×H×W ; where C is the number of channels; H and W are the height and width respectively; the feature map F is processed through a global average pooling layer, a fully connected layer and an activation function in sequence to obtain the channel attention weight. The channel attention weight can be multiplied by each channel of the input feature map to strengthen or weaken the feature representation of each channel; the channel attention module is expressed by the formula as follows:

[0040]

[0041] F FC = ReLU(W2δ(W1F avg ))

[0042] W C = sigmoid(F FC )

[0043] where F avg is the feature map after global average pooling processing; F c (i,j) represents the elements of all channels of the feature map F at the i-th row and j-th column; F FC is the feature map after fully connected layer processing; ReLU represents the ReLU activation function; W1 and W2 are learnable weight matrices; δ represents the activation function; W C is the channel attention weight; sigmoid represents the sigmoid activation function.

[0044] 2-2. Spatial Attention Module

[0045] The input of the Spatial Attention Module (SAM) is the feature map F∈R output by the backbone network C×H×W ; the feature map F is processed by max pooling and average pooling respectively, and the processing results are concatenated by channel and then an activation function is used to generate the spatial attention weight; the spatial attention weight is multiplied by the feature map F to strengthen or weaken the spatial position of the feature map. The spatial attention module is expressed by the formula as follows:

[0046] F concat = concat(F max F avg )

[0047] W S = sigmoid(W3δ(W4F concat ))

[0048] X att = W S ·X

[0049] Among them, F max and F avg are the results of performing max pooling and average pooling on the feature map output by the backbone network respectively; concat is the concatenation operation; F concat is the feature map after concatenating the results of max pooling and average pooling; W S is the spatial attention weight; W3 and W4 are learnable weight matrices; X att is the feature map output by the spatial attention module; X is the feature map output by the backbone network.

[0050] Taking the result of multiplying the channel attention weight W C and the feature map X att as the output of the attention module.

[0051] Using the dataset with the detection boxes labeled with the lesion regions to train the object detection model, the expression of the loss function L CIOU during the training process is:

[0052]

[0053] Among them, IOU is the intersection over union; ρ(,·,) is the Euclidean distance; b and b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively; c is the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box; a is the balance coefficient; v is the aspect ratio consistency between the predicted bounding box and the ground truth bounding box.

[0054] Step three, as Figure 3 shown, use the trained object detection model to process the dataset, detect the lesion regions of the FFDM images in the dataset, and obtain the detection boxes containing the lesion regions; take the maximum brightness value of all pixels in the detection box as the highlight point, and take the coordinates of the highlight point as its position feature.

[0055] Step four, as Figure 4As shown in the figure, a target segmentation model is constructed. The target segmentation model includes an image encoder, a prompt encoder, a mask encoder, a feature fusion module, and a boundary awareness module. The image encoder processes the FFDM image input into the target segmentation model through twenty-four consecutive Transformer layers to obtain image features. The prompt encoder is used to process the position features of the high-brightness points to obtain information features. The mask encoder includes two decoder layers, an attention module, a deconvolution layer, and a multi-layer perceptron. The two consecutive decoder layers process the image features and information features respectively to obtain an image feature vector and an information feature vector. The image feature vector is used as the input of the deconvolution layer to output a mask. The fusion result of the image feature vector and the information feature vector is used as the input of the attention module to output a token, and the token is input into the multi-layer perceptron. The output of the multi-layer perceptron is multiplied by the mask to obtain the mask output by the mask encoder.

[0056] The feature fusion module includes two upsampling layers in parallel. The two upsampling layers process the features output by the sixth layer and the twenty-fourth layer Transformer in the image encoder respectively to obtain features with the same size as the mask. The fusion result of the features output by the two upsampling layers and the mask is used as the output of the feature fusion module. The boundary awareness module processes the output of the feature fusion module to obtain the output of the target segmentation model. The boundary awareness module focuses on boundary refinement in the image and precise segmentation of the target. Especially when dealing with small targets (such as microcalcification lesions), it can improve the segmentation accuracy.

[0057] Step 5: Use the dataset with the label of the lesion area contour to train the target segmentation model.

[0058] Step 6: Use the trained target detection model to perform target detection on the measured FFDM image to obtain the target area, and extract the high-brightness points in the target area. Use the trained target segmentation model to complete the segmentation of the measured FFDM image based on the high-brightness points.

[0059] Step 7: Model evaluation

[0060] The present invention and the existing image segmentation model are respectively used to segment the image, and the accuracy (Accuracy), recall (Recall), specificity (Specificity), precision (Precision), Dice coefficient (DiceCoefficient), and Matthews correlation coefficient (Matthews Correlation Coefficient, MCC) are used to evaluate the segmentation results.

[0061] The accuracy ACC represents the proportion of samples correctly predicted by the model in the total samples, and it is one of the most commonly used evaluation criteria. Its expression is:

[0062]

[0063] Among them, TP is the number of samples correctly judged to contain microcalcification clusters; TN is the number of samples correctly judged not to contain microcalcification clusters; FP is the number of samples judged to contain microcalcification clusters but actually do not; FN is the number of samples judged to be normal but actually contain microcalcification clusters.

[0064] Recall, also known as true positive rate, represents the proportion of positive class samples correctly identified by the model among all samples that are actually positive class, and its expression is:

[0065]

[0066] Specificity, also known as true negative rate, represents the proportion of negative class samples correctly identified by the model among all samples that are actually negative class, and its expression is:

[0067]

[0068] Precision Pr represents the proportion of actually positive class samples among all samples predicted as positive class by the model, and its expression is:

[0069]

[0070] Dice coefficient is used to measure the similarity between two sample sets and is usually used in medical image segmentation tasks. Its expression is:

[0071]

[0072] Matthews correlation coefficient MCC comprehensively considers the impacts of true positive class, false positive class, true negative class, and false negative class, and is a relatively comprehensive evaluation index. Its expression is:

[0073]

[0074] Through these indexes, the performance of the model in the classification task can be comprehensively evaluated from multiple dimensions. Especially in the case of data imbalance, it can avoid the misleading caused by a single index. The evaluation results of different image segmentation methods are shown in Table 1.

[0075] Table 1 Evaluation results of different image segmentation methods

[0076]

[0077] As can be seen from Table 1, the present invention performs excellently in the image segmentation task, outperforming other models in terms of recall, specificity, precision, Dice coefficient, and Matthews correlation coefficient. The recall reaches 90.84%, indicating that the model is very sensitive to the recognition of positive regions and can effectively reduce missed detections; the precision is 89.32%, indicating high accuracy when predicting positive regions and reducing missegmentation. The Dice coefficient is 90.08%, indicating a high degree of overlap between the segmentation result and the true label, ensuring segmentation accuracy. The MCC is 0.9143, indicating excellent comprehensive performance of the model and being able to comprehensively reflect the segmentation effect. Therefore, the present invention performs outstandingly in the image segmentation task and is suitable for application scenarios with high-precision requirements.

Claims

1. A region proposal-based medical image segmentation system, comprising a preprocessing module and an object segmentation module; characterized in that: It also includes a target detection module and a positioning module; the target detection module is used to detect the target area of the preprocessed image; the positioning module is used to obtain the highlight point of the target in the image to be measured according to the detected target area; The target segmentation module includes an image encoder, a prompt encoder, a mask encoder, a feature fusion module, and a boundary awareness module; the image encoder extracts the image features of the preprocessed image; the prompt encoder extracts the information features of the highlight position; the mask encoder includes two decoder layers, an attention module, a deconvolution layer, and a multi-layer perceptron; the two decoder layers connected in sequence process the image features and information features to obtain an image feature vector and an information feature vector; the image feature vector is used as the input of the deconvolution layer; the fusion result of the image feature vector and the information feature vector is used as the input of the attention module; the output of the attention module is used as the input of the multi-layer perceptron; the output result of the mask encoder is the result of multiplying the outputs of the deconvolution layer and the multi-layer perceptron.

2. The medical image segmentation system based on region proposal according to claim 1, wherein: The highlight point is the pixel point corresponding to the maximum brightness value in the target area.

3. The medical image segmentation system based on region proposal according to claim 1, wherein: The target detection module includes a backbone network, a head network, and multiple attention modules; The results of processing the multiple outputs of the backbone network by the attention modules respectively are used as the input of the head network.

4. A region proposal-based medical image segmentation method, characterized in that: It includes the following steps: Step 1, construct a medical image dataset and perform labeling processing on the images in the dataset; Step 2, construct a target detection model and use the dataset to train the target detection model; Step 3, use the trained target detection model to obtain the highlight point of the target area; Step 4, construct a target segmentation model, which includes an image encoder, a prompt encoder, a mask encoder, a feature fusion module, and a boundary awareness module; the image encoder processes the image input to the target segmentation model through a twenty-four-layer self-attention mechanism feature extraction module connected in sequence to obtain image features; the prompt encoder is used to process the highlight point to obtain information features; the mask encoder includes two decoder layers, an attention module, a deconvolution layer, and a multi-layer perceptron; the two decoder layers connected in sequence process the image features and information features respectively to obtain an image feature vector and an information feature vector; the image feature vector is used as the input of the deconvolution layer to output a mask; the fusion result of the image feature vector and the information feature vector is used as the input of the attention module to output a token, and the token is input into the multi-layer perceptron; the output of the multi-layer perceptron is multiplied by the mask to obtain the mask output by the mask encoder; The feature fusion module includes two upsampling layers connected in parallel; the two upsampling layers process the shallow features and deep features output by two layers of the self-attention mechanism feature extraction module in the image encoder respectively to obtain features with the same size as the mask; The fusion result of the features output by the two upsampling layers and the mask is used as the output of the feature fusion module; the boundary awareness module is used to process the output of the feature fusion module to obtain the output of the target segmentation model; Step 5, use the dataset to train the target segmentation model; Step 6: Use the trained object detection model to perform object detection on the image to be measured, obtain the target area, and extract the high-brightness points in the target area; use the trained object segmentation model to complete the segmentation of the image to be measured based on the high-brightness points.

5. The method for medical image segmentation based on region proposal according to claim 4, wherein: In step 2 described above, the object detection model includes a backbone network, a head network, and multiple attention modules; The results of processing the multiple outputs of the backbone network by the attention modules respectively are used as the input of the head network.

6. The method for medical image segmentation based on region proposal according to claim 5, wherein: In step 1 described above, the labels of the images include the contour of the lesion area and the detection box containing the lesion area.

7. The method for medical image segmentation based on region proposal according to claim 6, wherein: Use the dataset with the detection box of the lesion area as the image label to train the object detection model; use the dataset with the contour of the lesion area as the image label to train the object segmentation model.

8. A method for medical image segmentation based on region proposal according to claim 5, characterized in that: In the second step described above, the loss function L during the training process CIOU has the following expression: where, IOU is the intersection over union; ρ(,·,) is the Euclidean distance; b and b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively; c is the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box; a is the balance coefficient; v is the aspect ratio consistency between the predicted bounding box and the ground truth bounding box.

9. A method for medical image segmentation based on region proposal according to claim 4, wherein: The high-brightness point is the pixel point corresponding to the maximum brightness value in the target area.

10. A method for medical image segmentation based on region proposal according to claim 4, characterized in that: In step 1 described above, preprocess the image before constructing the dataset. The preprocessing method is: use Gaussian filtering to remove noise, and apply the Otsu algorithm for global threshold segmentation to separate the breast tissue from the background; use the mosaic data augmentation method to enhance the image.