Medical image segmentation post-processing method and system based on double-branch evidence fusion and graph convolutional neural network
Patent Information
- Application Number
- CN202410888284.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-07-04
AI Technical Summary
由于MedSAM的输出仅能提供确定性的预测值,而无法表达模型对于预测结果的不确定性
[0054]本发明的有益效果为:本发明利用预训练好的基础分割模型MedSAM,结合证据理论和主观逻辑,将MedSAM的输出表示为对分割目标不同类别的信任度以及一个总的不确定性,通过双分支证据融合的不确定感知的图卷积神经网络对MedSAM的分割结果进行后处理,解决了医学小目标分割中由于MedSAM模型的过度自信导致的分割错误,提升了MedSAM的分割结果中不确定高的区域的分割精度。
Smart Images

Figure CN118941786B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing, and in particular to a method and system for post-processing medical image segmentation based on bi-branch evidence fusion and graph convolutional neural networks. Background Technology
[0002] Image segmentation is a fundamental task in medical imaging analysis, involving the identification and delineation of regions of interest (ROIs) in various medical images, such as organs, lesions, and tissues. Accurate segmentation is crucial for many clinical applications, including disease diagnosis, treatment planning, and disease progression monitoring. Manual segmentation has long been the gold standard for delineating anatomical structures and pathological regions, but this process is time-consuming, labor-intensive, and often requires a high level of expertise. In recent years, convolutional neural networks (CNNs) and Transformers have made significant strides in medical image segmentation. However, a significant limitation of many current medical image segmentation models is their task specificity. These models are typically designed and trained for specific segmentation tasks, and their performance can degrade significantly when applied to new tasks or different types of imaging data. This lack of generality poses a substantial obstacle to the widespread application of segmentation models in clinical practice.
[0003] As a revolutionary general-purpose model in the field of artificial intelligence, the basic model has been fully pre-trained on large-scale datasets and has demonstrated strong zero-shot generalization ability. In the field of image segmentation, Segment Anything (SAM), proposed by Meta in 2023, aims to serve as a foundational model for unifying the entire image segmentation task. To better adapt to image segmentation tasks in the medical field, Jun Ma et al. proposed MedSAM based on SAM. Using a dataset of over one million medical images, MedSAM outperforms existing state-of-the-art segmentation models and demonstrates performance comparable to or better than that of professional physicians. However, MedSAM still shows low confidence and robustness when facing segmentation tasks involving small targets such as brain tumors. This is because MedSAM's output can only provide deterministic predictions and cannot express the model's uncertainty regarding the prediction results. This lack of uncertainty assessment mechanism often leads to the model being overconfident when making incorrect predictions, especially in the boundary regions of small targets, easily generating many false positive (FP) and false negative (FN) regions, limiting the model's application in clinical practice.
[0004] Uncertainty estimation is considered crucial for assessing model reliability, quantifying the probability of model errors in prediction. Common methods for evaluating the uncertainty of neural network models include Bayesian neural networks (BNNs), ensemble-based methods, and dropout-based methods. In 2018, Sensoy et al. proposed a novel uncertainty assessment method that treats the neural network output as evidence based on Dempster-Shafer Evidence Theory (DST) and subjective logic-parameterized Dirichlet distributions, providing each object with a confidence value belonging to a different category and an overall uncertainty value, demonstrating impressive performance in model uncertainty quantification. Furthermore, evidence fusion, as an interpretable technique, plays a vital role in information integration in deep learning classifiers. Evidence-based methods can better utilize confidence assignment and uncertainty estimation to acquire information from medical images; compared to processing methods relying solely on a single network, evidence fusion techniques can achieve more accurate uncertainty analysis. Therefore, by utilizing evidence theory and subjective logic, the uncertainty of the MedSAM model can be evaluated, and further improvements in the accuracy of MedSAM model segmentation results can be achieved through uncertainty-aware fusion post-processing methods. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention provides a medical image segmentation post-processing method and system based on bi-branch evidence fusion and graph convolutional neural networks. Utilizing a pre-trained basic segmentation model MedSAM, two independent output branches are reconstructed using Dropout technology. Pixel classification evidence from both branches is obtained and fused to obtain a more reliable uncertainty assessment and working region of interest (ROI). Pixels with low uncertainty in the ROI are then used as labeled nodes in the graph, while pixels with high uncertainty are used as unlabeled nodes. A semi-supervised graph convolutional neural network (GCN) is then trained to classify the unlabeled points, improving the segmentation accuracy of regions with high uncertainty in the original MedSAM segmentation result.
[0006] The technical solution adopted by this invention to solve its technical problem is:
[0007] A post-processing method for medical image segmentation based on bi-branch evidence fusion and graph convolutional neural networks includes the following steps:
[0008] Step 1: Medical image preprocessing. First, process the intensity value v of each pixel i in the input medical image x. i Normalize:
[0009]
[0010] Where v min v is the minimum pixel intensity value.max The adjusted pixel intensity value v is the maximum pixel intensity value. i ∈[0,255], and then the size of x is adjusted to H×W×3 using bicubic interpolation, where H and W are the height and width of the processed medical image x, respectively;
[0011] Step 2: Obtain the dual-branch output. Input the processed medical image x into the MedSAM encoder, and then obtain the lower branch output m through the MedSAM decoder mask decoder 1. 1 The lower branch output m is obtained through the decoder mask decoder 2. 2 Mask decoder 2 shares network parameters with mask decoder 1. Mask decoder 2 is obtained by dropout operation based on mask decoder 1. That is, the output after two transposed convolutions and upsampling in mask decoder 1 is spatially multiplied by the output of the fully connected MLP layer to obtain m. 1 The output of the mask decoder 2 after two transposed convolution upsampling operations is spatially multiplied by the output of the last MLP layer after dropout to obtain m. 2 ;
[0012] Step 3: Fuse the two branches of evidence to obtain evidence from the upper and lower branches. Calculate the trust value of the upper and lower branches. and uncertainty Where K is the total number of categories of the segmented target, n=1 represents the upper branch, n=2 represents the lower branch, and the uncertainty of the upper and lower branches is fused based on evidence theory to obtain the fused uncertainty.
[0013] Step 4: Mark the pixels according to their level of uncertainty after fusion, and determine the working area Ub. ROI and Ub ROI Pixel labels with low internal uncertainty;
[0014] Step 5: Using the work area Ub ROI Each pixel in the graph is treated as a node, and a weighted graph with semi-labeled nodes is constructed by connecting these nodes to obtain Ub. ROI The connection matrix A and characteristic matrix X of the corresponding graph;
[0015] Step Six: Connect the work area USB ROI The adjacency matrix A and feature matrix X of the corresponding graph are input into a two-layer graph convolutional neural network (GCN) to obtain the output predicted probability Z. The parameters W = {W} of the GCN are updated by minimizing the joint optimization loss function L. (0)W (1)}, until the maximum number of iterations is reached:
[0016]
[0017] L = L CE +L Dice
[0018]
[0019] Among them W (0) and W (1) This is the weight parameter matrix of the first and second layers of the GCN. Sigmoid and ReLU are both activation functions. That is, adding an identity matrix I to the adjacency matrix A. N L CE and L Dice These are cross-entropy loss and Dice loss, respectively. Let y be the label corresponding to pixel i. i =1, where N is the predicted probability of the working area Ub. ROI The total number of pixels with labels, F = {F i |i∈Ub ROI} is the segmentation result after the output predicted probability Z of GCN is transformed into binary value through the sign function sgn (threshold is 0.5), F i ∈{0,1}, where 1 represents the foreground and 0 represents the background;
[0020] Step 7: Connect the work area USB ROI The adjacency matrix A and feature matrix X of the corresponding graph are input into a trained two-layer graph convolutional neural network (GCN) to obtain a binarized segmentation result F' = {F i '|i∈Ub ROI}, in the work area Ub ROI The original labeled pixel categories are preserved, while the categories of unlabeled pixels are replaced with the GCN segmentation results, finally yielding the final segmentation result F. * ={F i * |i∈Ub ROI}:
[0021]
[0022] Furthermore, the process of step three is as follows:
[0023] 3.1 Convert the output m of the dual branches 1 and m 2 Two-branch evidence e is obtained using the Softplus activation function. n :
[0024] e n =Softplus(m n ), n=1,2 (2)
[0025] 3.2 Based on the two-branch evidence e n Taking a uniform distribution, the parameters of the bi-branch Dirichlet distribution are obtained. and Dirichlet distribution intensity in
[0026] α n =e n +1, n=1,2 (3)
[0027] 3.3 Calculate the classification confidence value of each pixel i in the dual-branch output. and uncertainty for:
[0028]
[0029] 3.4 The uncertainties of the upper and lower branches are fused to obtain the fused uncertainty.
[0030]
[0031] Furthermore, the process of step four is as follows:
[0032] 4.1 Based on the uncertainty after fusion Uncertainty level labeling Ub = {Ub i |i=1,2,...,H×W}:
[0033]
[0034] Where τ is the uncertain threshold, if Ub i =1 indicates that the i-th pixel is a pixel with high uncertainty, while Ub i =0 indicates that the i-th pixel is a pixel with low uncertainty;
[0035] 4.2 Determine the working area Ub ROI :
[0036] Ub ROI =Ub∪Eb (8)
[0037] P = Sigmoid(m 1 (9)
[0038] Where, P = {P} i,j|i=1,2,...,H×W,j=1,2,...,K} is the lower branch output m of MedSAM. 1 After applying the Sigmoid activation function, the predicted probability value for pixel classification is obtained, Eb={Eb i |i=1,2,...,H×W} is the segmentation result of P after being transformed into a binary value by the sign function sgn (with a threshold of 0.5), where Eb i ∈{0,1}, where 1 represents the foreground and 0 represents the background. ∪ indicates that the two are first added together, and then the sign function sgn (with a threshold of 1) is used to convert the result into a binary mask, where 0 represents the background with low uncertainty, 1 represents pixels with high uncertainty and the foreground with low uncertainty, and finally, the region of all 1s is used as Ub. ROI ;
[0039] 4.3 Determine Ub based on the level of uncertainty. ROI Labels for pixels with low uncertainty:
[0040]
[0041] The "unlabeled" option indicates that there is no label.
[0042] Furthermore, the process of step five is as follows:
[0043] 5.1 For any i-th pixel, create connections to its four nearest neighbors (up, down, left, and right). Additionally, in Ub... ROI Randomly select some nodes and connect them, using the pixel values v = {v} of the original medical image. i The predicted probability value P corresponding to the output of the MedSAM lower branch and |i=1,2,...,H×W} is {P}. i,j We use the values of |i=1,2,...,H×W,j=1,2,...,K} to calculate the weights of the edges, and obtain Ub. ROI The connection matrix A of the corresponding graph is {w} i,j |i∈Ub ROI ,j∈Ub ROI}:
[0044]
[0045] weight2(x i ,x j )=||v i -v j || 2 (13)
[0046] weight3(x i ,x j )=||x i -xj || 2 (14)
[0047] Where, x i ,x j They are Ub ROI The positions of the i-th and j-th pixels are given, λ is a balance factor, weight1(·) is determined by the diversity between the two pixels, and weight2(x) is determined by the diversity between the two pixels. i ,x j ) and weight3(x i ,x j ) are respectively x i ,x j The pixel weights and distance weights between two points, σ1 and σ2, are respectively weight2(x i ,x j ) and weight3(x i ,x j The variance of ) is given in equation (12). When the segmented target consists only of foreground and background, K = 2, where P i,1 Let P be the probability value of the i-th pixel predicted as foreground and the probability value of the i-th pixel predicted as background from the lower branch of MedSAM. i,2 =1-P i,1 In equations (13) and (14), |||| represents the Euclidean distance;
[0048] 5.2 For Ub ROI Each pixel in the model has a feature value including the output m of the lower branch of the MedSAM model. 1 The standardized weighted eigenvalues M = {M i |i∈Ub ROI}, the predicted probability value P = {P i,j |i∈Ub ROI j = 1, 2, ..., K} and the uncertainty after fusion Get Ub ROI The corresponding characteristic matrix X:
[0049]
[0050] Among them, v μ It is the lower branch output m of MedSAM 1 The average value, v α It is m 1 The variance.
[0051] A medical image segmentation post-processing system based on bi-branch evidence fusion and graph convolutional neural network includes an image preprocessing module, a bi-branch output module, an evidence fusion module, an ROI labeling module, a graph construction module, a GCN training module, and a GCN segmentation module. Each module sequentially implements the processing steps one through seven of the method.
[0052] The technical concept of this invention is as follows: The decoder of the basic segmentation model MedSAM is reconstructed through a Dropout operation to obtain two independent output branches. Pixel classification evidence for each branch is obtained according to evidence theory and subjective logic. Then, the uncertainty of each pixel classification is quantified through evidence. By fusing the evidence from the two branches, a more reliable uncertainty of pixel classification is obtained. Then, the working region is determined by combining the segmentation results of MedSAM, and a weighted semi-label map corresponding to the working region is constructed. Pixels with low uncertainty directly take the segmentation results of MedSAM as labels. Then, a graph convolutional neural network (GCN) is trained as a post-processing model for the MedSAM segmentation results to re-predict the classification of pixels with low uncertainty.
[0053] A post-processing method for medical image segmentation based on bi-branch evidence fusion and graph convolutional neural networks (GCNs) comprises three parts: a bi-branch evidence fusion unit, a ROI weighted graph construction unit, and a GCN training and segmentation unit. The bi-branch evidence fusion unit, based on a pre-trained base model, reconstructs two independent output branches through a dropout operation, acquiring pixel classification evidence from each branch and fusing them to obtain the fused uncertainty. The ROI weighted graph construction unit uses pixels within the working region (ROI) as nodes, labeling them according to their fused uncertainty. Nodes with low uncertainty are labeled as the segmentation result of the base model, while nodes with high uncertainty remain unlabeled. A weighted semi-labeled graph is constructed through graph connections. The GCN training and segmentation unit uses the labeled node outputs from the weighted semi-labeled graph to calculate a joint optimization loss including cross-entropy loss and Dice loss, trains the graph convolutional neural network (GCN), classifies unlabeled points, and fuses the labels of nodes with low original uncertainty to obtain the final segmentation result. This invention solves the segmentation error caused by the overconfidence of the base model MedSAM in medical image segmentation.
[0054] The beneficial effects of this invention are as follows: This invention utilizes a pre-trained basic segmentation model MedSAM, combined with evidence theory and subjective logic, to represent the output of MedSAM as the degree of trust in different categories of the segmented target and a total uncertainty. The segmentation results of MedSAM are post-processed through a graph convolutional neural network with uncertainty perception through dual-branch evidence fusion, which solves the segmentation error caused by the overconfidence of the MedSAM model in the segmentation of small medical targets and improves the segmentation accuracy of regions with high uncertainty in the segmentation results of MedSAM. Attached Figure Description
[0055] Figure 1 This is a flowchart of the algorithm of the present invention.
[0056] Figure 2 This is a flowchart of the present invention.
[0057] Figure 3 The images shown are comparisons of the segmentation results of the present invention, where (a) is a brain tumor image, (b) is the gold standard, (c) is the segmentation result of the base model MedSAM, and (d) is the segmentation result of the method of the present invention. Detailed Implementation
[0058] The invention will now be further described with reference to the accompanying drawings.
[0059] Reference Figures 1-3 A post-processing method for medical image segmentation based on bi-branch evidence fusion and graph convolutional neural networks, taking brain tumor image segmentation as an example, includes the following steps:
[0060] Step 1: Preprocess the 2D slice x of the brain tumor, and convert the intensity value v of each pixel i. i Normalize:
[0061]
[0062] Where v min v is the minimum pixel intensity value. max The adjusted pixel intensity value v is the maximum pixel intensity value. i ∈[0,255], and then the size of x is adjusted to H×W×3 using bicubic interpolation, where H and W are the height and width of the processed medical image x, respectively, and H and W are set to 1024 here;
[0063] Step 2: Input the processed medical image x into the encoder of the base segmentation model MedSAM, and then obtain the next branch output m through the decoder mask decoder 1 of MedSAM. 1 The lower branch output m is obtained through the decoder mask decoder 2. 2 Mask decoder 2 shares network parameters with mask decoder 1. Mask decoder 2 is obtained by dropout operation based on mask decoder 1. That is, the output after two transposed convolutions and upsampling in mask decoder 1 is spatially multiplied by the output of the fully connected MLP layer to obtain m. 1The output of the mask decoder 2 after two transposed convolution upsampling operations is spatially multiplied by the output of the last MLP layer after dropout to obtain m. 2 ;
[0064] Step 3: Obtain evidence from the upper and lower branches Calculate the trust value of the upper and lower branches. and uncertainty Where K is the total number of categories of the segmented target, n=1 represents the upper branch, n=2 represents the lower branch, and the uncertainty of the upper and lower branches is fused based on evidence theory to obtain the fused uncertainty. The process is as follows:
[0065] 3.1 Convert the output m of the dual branches 1 and m 2 Two-branch evidence e is obtained using the Softplus activation function. n :
[0066] e n =Softplus(m n ), n=1,2 (2)
[0067] 3.2 Based on the two-branch evidence e n Taking a uniform distribution, the parameters of the bi-branch Dirichlet distribution are obtained. and Dirichlet distribution intensity in
[0068] α n =e n +1, n=1,2 (3)
[0069] 3.3 Calculate the classification confidence value of each pixel i in the dual-branch output. and uncertainty for:
[0070]
[0071] 3.4 The uncertainties of the upper and lower branches are fused to obtain the fused uncertainty.
[0072]
[0073] Step 4: Mark the pixels according to their level of uncertainty after fusion, and determine the working area Ub. ROI and Ub ROI Pixel labels with low internal uncertainty; the process is as follows:
[0074] 4.1 Based on the uncertainty after fusion Uncertainty level labeling Ub = {Ub i |i=1,2,...,H×W}:
[0075]
[0076] Where τ is the uncertain threshold, if Ub i =1 indicates that the i-th pixel is a pixel with high uncertainty, while Ub i =0 indicates that the i-th pixel is a pixel with low uncertainty;
[0077] 4.2 Determine the working area Ub ROI :
[0078] Ub ROI =Ub∪Eb (8)
[0079] P = Sigmoid(m 1 (9)
[0080] Where, P = {P} i,j |i=1,2,...,H×W,j=1,2,...,K} is the lower branch output m of MedSAM. 1 After applying the Sigmoid activation function, the predicted probability value for pixel classification is obtained, Eb={Eb i |i=1,2,...,H×W} is the segmentation result of P after being transformed into a binary value by the sign function sgn (with a threshold of 0.5), where Eb i ∈{0,1}, where 1 represents the foreground and 0 represents the background. ∪ indicates that the two are first added together, and then the sign function sgn (with a threshold of 1) is used to convert the result into a binary mask, where 0 represents the background with low uncertainty, 1 represents pixels with high uncertainty and the foreground with low uncertainty, and finally, the region of all 1s is used as Ub. ROI ;
[0081] 4.3 Determine Ub based on the level of uncertainty. ROI Labels for pixels with low uncertainty:
[0082]
[0083] Where unlabeled indicates that there is no label;
[0084] Step 5: Using the work area Ub ROI Each pixel in the graph is treated as a node, and a weighted graph with semi-labeled nodes is constructed by connecting these nodes to obtain Ub. ROI The connection matrix A and characteristic matrix X of the corresponding graph; the process is as follows:
[0085] 5.1 For any i-th pixel, create connections to its four nearest neighbors (up, down, left, and right). Additionally, in Ub... ROI Randomly select some nodes and connect them, using the pixel values v = {v} of the original medical image. i The predicted probability value P corresponding to the output of the MedSAM lower branch and |i=1,2,...,H×W} is {P}. i,j We use the values of |i=1,2,...,H×W,j=1,2,...,K} to calculate the weights of the edges, and obtain Ub. ROI The connection matrix A of the corresponding graph is {w} i,j |i∈Ub ROI ,j∈Ub ROI}:
[0086]
[0087] weight2(x i ,x j )=||v i -v j || 2 (13)
[0088] weight3(x i ,x j )=||x i -x j || 2 (14)
[0089] Where, x i ,x j They are Ub ROI The positions of the i-th and j-th pixels are given, λ is a balance factor, weight1(·) is determined by the diversity between the two pixels, and weight2(x) is determined by the diversity between the two pixels. i ,x j ) and weight3(x i ,x j ) are respectively x i ,x j The pixel weights and distance weights between two points, σ1 and σ2, are respectively weight2(x i ,x j ) and weight3(x i ,x j The variance of ) is given in equation (12). When the segmented target consists only of foreground and background, K = 2, where P i,1 Let P be the probability value of the i-th pixel predicted as foreground and the probability value of the i-th pixel predicted as background from the lower branch of MedSAM. i,2 =1-P i,1In equations (13) and (14), || represents the Euclidean distance;
[0090] 5.2 For Ub ROI Each pixel in the model has a feature value including the output m of the lower branch of the MedSAM model. 1 The standardized weighted eigenvalues M = {M i |i∈Ub ROI}, the predicted probability value P = {P i,j |i∈Ub ROI j = 1, 2, ..., K} and the uncertainty after fusion Get Ub ROI The corresponding characteristic matrix X:
[0091]
[0092] Among them, v μ It is the lower branch output m of MedSAM 1 The average value, v α It is m 1 The variance;
[0093] Step Six: Connect the work area USB ROI The adjacency matrix A and feature matrix X of the corresponding graph are input into a two-layer graph convolutional neural network (GCN) to obtain the output predicted probability Z. The parameters W = {W} of the GCN are updated by minimizing the joint optimization loss function L. (0) W (1)}, until the maximum number of iterations is reached:
[0094]
[0095] L = L CE +L Dice
[0096]
[0097] Among them W (0) and W (1) This is the weight parameter matrix of the first and second layers of the GCN. Sigmoid and ReLU are both activation functions. That is, adding an identity matrix I to the adjacency matrix A. N L CE and L Dice These are cross-entropy loss and Dice loss, respectively. Let y be the label corresponding to pixel i. i =1, where N is the predicted probability of the working area Ub. ROI The total number of pixels with labels, F = {F i |i∈Ub ROI} is the segmentation result after the output predicted probability Z of GCN is transformed into binary value through the sign function sgn (threshold is 0.5), F i ∈{0,1}, where 1 represents the foreground and 0 represents the background;
[0098] Step 7: Connect the work area USB ROI The adjacency matrix A and feature matrix X of the corresponding graph are input into a trained two-layer graph convolutional neural network (GCN) to obtain a binarized segmentation result F' = {F i '|i∈Ub ROI}, in the work area Ub ROI The original labeled pixel categories are preserved, while the categories of unlabeled pixels are replaced with the GCN segmentation results, finally yielding the final segmentation result F. * ={F i * |i∈Ub ROI}:
[0099]
[0100] This embodiment also provides a medical image segmentation post-processing system based on bi-branch evidence fusion and graph convolutional neural networks, including a preprocessing module, a bi-branch output module, an evidence fusion module, an ROI labeling module, a graph construction module, a GCN training module, and a GCN segmentation module. The above modules correspond to steps one through seven of the method of the present invention, respectively.
[0101] This embodiment utilizes a pre-trained basic segmentation model, MedSAM, combined with evidence theory and subjective logic. The output of MedSAM is represented as the degree of trust in different categories of the segmented target and an overall uncertainty. The segmentation results of MedSAM are post-processed through a graph convolutional neural network with uncertainty perception through dual-branch evidence fusion. This solves the segmentation error caused by MedSAM's overconfidence in medical image segmentation and improves the segmentation accuracy of high uncertainty regions in the MedSAM segmentation results.
[0102] As described above, the specific implementation steps of this invention make the invention clearer. Any modifications and alterations made to this invention within the spirit and scope of the claims fall within the protection scope of this invention.
Claims
1. A post-processing method for medical image segmentation based on bi-branch evidence fusion and graph convolutional neural networks, characterized in that, The method includes the following steps: Step 1: Medical image preprocessing. First, process the intensity value v of each pixel i in the input medical image x. i Normalize: Where v min v is the minimum pixel intensity value. max The adjusted pixel intensity value v is the maximum pixel intensity value. i ∈[0,255], and then the size of x is adjusted to H×W×3 using bicubic interpolation, where H and W are the height and width of the processed medical image x, respectively; Step 2: Obtain the dual-branch output. Input the processed medical image x into the MedSAM encoder, and then obtain the lower branch output m through the MedSAM decoder mask decoder 1. 1 The lower branch output m is obtained through the decoder mask decoder 2. 2 Mask decoder 2 shares network parameters with mask decoder 1. Mask decoder 2 is obtained by dropout operation based on mask decoder 1. That is, the output after two transposed convolutions and upsampling in mask decoder 1 is spatially multiplied by the output of the fully connected MLP layer to obtain m. 1 The output of the mask decoder 2 after two transposed convolution upsampling operations is spatially multiplied by the output of the last MLP layer after dropout to obtain m. 2 ; Step 3: Fuse the two branches of evidence to obtain evidence from the upper and lower branches. Calculate the trust value of the upper and lower branches. and uncertainty Where K is the total number of categories of the segmented target, n=1 represents the upper branch, n=2 represents the lower branch, and the uncertainty of the upper and lower branches is fused based on evidence theory to obtain the fused uncertainty. Step 4: Mark the pixels according to their level of uncertainty after fusion, and determine the working area Ub. ROI And Ub ROI Pixel labels with low internal uncertainty; Step 5: Using the work area Ub ROI Each pixel in the graph is treated as a node, and a weighted graph with semi-labeled nodes is constructed by connecting these nodes to obtain Ub. ROI The connection matrix A and characteristic matrix X of the corresponding graph; Step Six: Connect the work area USB ROI The adjacency matrix A and feature matrix X of the corresponding graph are input into a two-layer graph convolutional neural network (GCN) to obtain the output predicted probability Z. The parameters W = {W} of the GCN are updated by minimizing the joint optimization loss function L. (0) W (1) }, until the maximum number of iterations is reached: L=L CE +L Dice Among them W (0) and W (1) This is the weight parameter matrix of the first and second layers of the GCN. Sigmoid and ReLU are both activation functions. That is, adding an identity matrix I to the adjacency matrix A. N L CE and L Dice These are cross-entropy loss and Dice loss, respectively. Let y be the label corresponding to pixel i. i =1, where N is the predicted probability of the working area Ub. ROI The total number of pixels with labels, F = {F i |i∈Ub ROI } is the segmentation result after the output predicted probability Z of GCN is transformed into binary value through the sign function sgn (threshold is 0.5), F i ∈{0,1}, where 1 represents the foreground and 0 represents the background; Step 7: Connect the work area USB ROI The adjacency matrix A and feature matrix X of the corresponding graph are input into a trained two-layer graph convolutional neural network (GCN) to obtain a binarized segmentation result F' = {F i '|i∈Ub ROI }, in the work area Ub ROI The original labeled pixel categories are preserved, while the categories of unlabeled pixels are replaced with the GCN segmentation results, finally yielding the final segmentation result F. * ={F i * |i∈Ub ROI }:
2. The medical image segmentation post-processing method based on bi-branch evidence fusion and graph convolutional neural networks as described in claim 1, characterized in that, The process of step three is as follows: 3.1 Convert the output m of the dual branches 1 and m 2 Two-branch evidence e is obtained using the Softplus activation function. n : e n =Softplus(m n ), n=1,2 (2) 3.2 Based on the two-branch evidence e n Taking a uniform distribution, the parameters of the bi-branch Dirichlet distribution are obtained. and Dirichlet distribution intensity in a n =e n +1, n=1.2 (3) 3.3 Calculate the classification confidence value of each pixel i in the dual-branch output. and uncertainty for: 3.4 The uncertainties of the upper and lower branches are fused to obtain the fused uncertainty.
3. The medical image segmentation post-processing method based on bi-branch evidence fusion and graph convolutional neural networks as described in claim 1 or 2, characterized in that, The process of step four is as follows: 4.1 Based on the uncertainty after fusion Uncertainty level labeling Ub = {Ub i |i=1,2,...,H×W}: Where τ is the uncertain threshold, if Ub i =1 indicates that the i-th pixel is a pixel with high uncertainty, while Ub i =0 indicates that the i-th pixel is a pixel with low uncertainty; 4.2 Determine the working area Ub ROI : Ub ROI =Ub∪Eb (8) P=Sigmoid(m 1 ) (9) Where, P = {P} i,j |i=1,2,...,H×W,j=1,2,...,K} is the lower branch output m of MedSAM. 1 After applying the Sigmoid activation function, the predicted probability value for pixel classification is obtained, Eb={Eb i |i=1,2,...,H×W} is the segmentation result of P after being transformed into a binary value by the sign function sgn, where Eb i ∈{0,1}, where 1 represents the foreground and 0 represents the background. ∪ indicates that the two are first added together, and then the sign function sgn (with a threshold of 1) is used to convert the result into a binary mask, where 0 represents the background with low uncertainty, 1 represents pixels with high uncertainty and the foreground with low uncertainty, and finally, the region of all 1s is used as Ub. ROI ; 4.3 Determine Ub based on the level of uncertainty. ROI Labels for pixels with low uncertainty: The "unlabeled" option indicates that there is no label.
4. The medical image segmentation post-processing method based on bi-branch evidence fusion and graph convolutional neural networks as described in claim 1 or 2, characterized in that the process of step five is as follows: 5.1 For any i-th pixel, create connections to its four nearest neighbors (up, down, left, and right). Additionally, in Ub... ROI Randomly select some nodes and connect them, using the pixel values v = {v} of the original medical image. i The predicted probability value P corresponding to the output of the MedSAM lower branch and |i=1,2,...,H×W} is {P} i,j We use the values of |i=1,2,...,H×W,j=1,2,...,K} to calculate the weights of the edges, and obtain Ub. ROI The connection matrix A of the corresponding graph is {w} i,j |i∈Ub ROI ,j∈Ub ROI }: weight2(x i ,x j )=||v i -v j || 2 (13) weight3(x i ,x j )=||x i -x j || 2 (14) in, x i ,x j They are Ub ROI The positions of the i-th and j-th pixels are given, λ is a balance factor, weight1(·) is determined by the diversity between the two pixels, and weight2(x) is determined by the diversity between the two pixels. i ,x j ) and weight3(x i ,x j ) are respectively x i ,x j The pixel weights and distance weights between two points, σ1 and σ2, are respectively weight2(x i ,x j ) and weight3(x i ,x j The variance of ) is given in equation (12). When the segmented target consists only of foreground and background, K = 2, where P i,1 Let P be the probability value of the i-th pixel predicted as foreground and the probability value of the i-th pixel predicted as background from the lower branch of MedSAM. i,2 =1-P i,1 In equations (13) and (14), || represents the Euclidean distance; 5.2 For Ub ROI Each pixel in the model has a feature value including the output m of the lower branch of the MedSAM model. 1 The standardized weighted eigenvalues M = {M i |i∈Ub ROI }, the predicted probability value P = {P i,j |i∈Ub ROI j = 1, 2, ..., K} and the uncertainty after fusion Get Ub ROI The corresponding characteristic matrix X: Among them, v μ It is the lower branch output m of MedSAM 1 The average value, v α It is m 1 The variance.
5. A system implementing the medical image segmentation post-processing method based on bi-branch evidence fusion and graph convolutional neural network as described in claim 1, characterized in that: The system includes an image preprocessing module, a dual-branch output module, an evidence fusion module, an ROI labeling module, a graph construction module, a GCN training module, and a GCN segmentation module. Each module sequentially implements the processing steps one through seven of the method.
Citation Information
Patent Citations
Graffiti supervision maxillary sinus segmentation method and device based on superpixel and double-branch graph convolution network
CN118096812A
End-to-end hyperspectral image classification method, system, device and terminal
CN118262244A