A semi-supervised medical image segmentation method and device based on anatomical connectivity
By introducing anatomical connectivity judgment and pixel-level comparison learning in the medical image segmentation model, the non-connection segmentation problem caused by the existing model ignoring the connectivity between pixels is solved, and a more accurate anatomical structure segmentation is achieved.
Patent Information
- Application Number
- CN202510180763.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The existing medical image segmentation model ignores the intrinsic connectivity and geometric properties between pixels, resulting in unreasonable non-connected areas in the segmentation results.
A semi-supervised medical image segmentation method based on anatomical connectivity is proposed. By introducing the concept of connectivity judgment, pseudo-tagging of the largest connected area is obtained, and pixel-level comparison learning is carried out in combination with labeled data to improve the model's ability to extract features.
This method strengthens the model's understanding of the target structure, improves the model's accuracy in maintaining the continuity of anatomical structure, and significantly reduces the non-connected areas in the segmentation results.
Smart Images

Figure CN119649037B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image segmentation, and particularly to a semi-supervised medical image segmentation method and device based on anatomical connectivity. Background Art
[0002] In precise medical image analysis, the tolerance for minute geometric deviations is extremely low because these seemingly insignificant errors can trigger global topological changes, thereby inducing functional judgment errors in subsequent clinical decisions.
[0003] With the rapid development of deep neural networks, deep learning-based segmentation techniques have been able to capture information about a group of pixels from each local receptive field of an image through an encoder-decoder architecture and learn from it. Existing semi-supervised medical image segmentation algorithms ignore the connectivity of organs in anatomical structures and frame the task as a purely pixel-level classification problem. These models often have topological errors in predicting organs such as the heart, such as outlier predictions, disconnected regions, and non-simply connected predictions, resulting in discontinuous segmentation results. For example, when using an image segmentation model to segment a cardiac MRI dataset, the data distribution is complex and variable, and due to the similarity of gray levels and textures, it is easy to misidentify redundant regions as target regions, as Figure 2 shown, in the cardiac segmentation result, there are unreasonable fractures in the segmentation result, and non-connected regions should not appear in cardiac tissue.
[0004] This point-to-point segmentation method ignores the inherent connectivity and geometric properties between pixels, which is particularly obvious in the scenario of medical data with limited sample size and uneven quality. Low-quality medical images are often accompanied by high levels of noise and artifacts, which reduce spatial consistency and may cause topological problems. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the defect that existing medical image segmentation models ignore the inherent connectivity and geometric properties between pixels, resulting in unreasonable non-connected regions in the segmentation result. A semi-supervised medical image segmentation method and device based on anatomical connectivity are proposed, introducing the concept of connectivity judgment, strengthening the model's understanding of the target structure, and also improving the accuracy of the model in maintaining the continuity of anatomical structures.
[0006] The present invention adopts the following technical solutions:
[0007] A semi-supervised medical image segmentation method based on anatomical connectivity, comprising:
[0008] Segmenting the input labeled data and unlabeled data using a medical image segmentation model to obtain a first feature map and a second feature map;
[0009] Obtain the pseudo-labels of the largest connected regions of the first feature map and the second feature map through discrimination, and obtain the deterministic pseudo-labels according to the pseudo-labels of the largest connected regions and the uncertainty map;
[0010] Combine the deterministic pseudo-labels and the labeled data as the supervision signal to guide the segmentation model to perform pixel-level contrast learning to improve the feature extraction ability.
[0011] The medical image segmentation model includes a first segmentation unit and a second segmentation unit that use the U-Net network as the backbone network; the first segmentation unit segments the input labeled data and unlabeled data to obtain the first feature map. The first segmentation unit includes an encoder and a first decoder. The first decoder uses a transposed convolutional layer and uses a skip connection with one of the encoders to transfer information between different layers; the second segmentation unit segments the input labeled data and unlabeled data to obtain the second feature map. The second segmentation unit includes an encoder and a second decoder. The second decoder uses a linear interpolation layer and uses a skip connection with the encoder to transfer information between different layers.
[0012] For the labeled data, use the CE function and the Dice function to jointly calculate the supervision loss of the labeled data, as follows:
[0013] L a = CE(P a , y) + Dice(P a , y);
[0014] L b = CE(P b , y) + Dice(P b , y);
[0015] L sup = L a + L b ;
[0016] Among them, P a represents the first feature map, P b represents the second feature map, y represents the label of the labeled data, L sup represents the supervision loss of the labeled data, L a represents the supervision loss of the first feature map, L b represents the supervision loss of the second feature map, Dice(·) represents the Dice loss calculation function, and CE(·) represents the cross-entropy loss calculation function.
[0017] For the unlabeled data, by selecting the first feature map P a and the second feature map Pb The largest connected region of each category in [object] is used to further improve the spatial consistency of the prediction. The selection process of the connected region is as follows:
[0018]
[0019]
[0020] where C a,c and C b,c are the largest connected regions of category c in the first feature map P a and the second feature map P b ConnComp(Pa == c) represents returning an array that contains all the connected regions belonging to category c, and |i| represents the number of pixels in the connected region i; the argmax function represents selecting the connected region with the largest number of pixels;
[0021] Take the union of the largest connected regions C a,c and C b,c to obtain the pseudo-label C′ of the largest connected region, which is represented by the following formula:
[0022]
[0023] where N is the total number of categories, and ∪ represents the union of regions.
[0024] If there are conflicts among some pixels on the largest connected regions C a,c and C b,c , the conflicts are resolved by comparing the prediction probabilities at the corresponding positions: Let P a,c and P b,c represent the probabilities of the first feature map P a and the second feature map P b for category c respectively. Then the category label of each pixel can be determined in the following way:
[0025]
[0026] where (i, j) represents the position of the pixel, and C′(i, j) is the value of the pseudo-label C′ at the position (i, j); if the position (i, j) is in C a,c and P a,c (i, j) is greater than P b,c (i, j), then this position is labeled as category c; if the position (i, j) is in C b,c and P b,c (i, j) is greater than or equal to P a,cIf it is (i, j), it is also marked as category c; if the position (i, j) is not in the largest connected region of any category, it is marked as background or unlabeled in the pseudo-label C' and represented by 0.
[0027] The cross-entropy loss is used to measure the consistency loss L between the predicted output of the segmentation model and the pseudo-label, as follows: mcr , specifically as follows:
[0028]
[0029] where (i, j) represents the position of the pixel, c represents the category, and C' i,j,c is the value of the pseudo-label at position (i, j) and category c, and P a,i,j,c and P b,i,j,c are the probability values of the pixels in the first feature map and the second feature map at the same position and category, respectively.
[0030] Combining the deterministic pseudo-label and the labeled data as the supervision signal to guide the pixel-level contrast learning of the segmentation model specifically includes:
[0031] Design the positive sample z+ and negative sample n- in the contrast learning. For each category c, select the pixel index from the predicted output results of the medical image segmentation model with the same anchor point as the positive sample index, and filter out the pixels whose predicted results output by the medical image segmentation model are different from the current category c as the negative samples;
[0032] Set a memory buffer M with a fixed size to store samples from different categories, use the randomly selected anchor points from M to calculate the contrast loss, and then update the memory buffer;
[0033] The loss function in the contrast learning part uses infoNCE to calculate the pixel-level contrast loss between the retained feature tensors and the categories:
[0034]
[0035] where L pcl represents the pixel-level contrast loss, V n represents the nth anchor vector, and represent the nth positive sample vector and the mth negative sample vector respectively, K represents the number of negative samples, N' represents the number of anchor vectors, and τ is the temperature parameter.
[0036] The total loss function Loss of the medical image segmentation model is as follows:
[0037] Loss = L sup + λ 1 L fc + λ2 L mcr + λ 3 L pcl ;
[0038] Among them, L sup is the supervision loss, L pcl is the pixel-level contrast learning loss, L mcr is the maximum connected region consistency loss based on, L fc is the encoder-decoder consistency loss, and λ 1 , λ 2 and λ 3 are preset coefficients.
[0039] A semi-supervised medical image segmentation device based on anatomical connectivity, comprising
[0040] a segmentation module that uses a medical image segmentation model to segment the input labeled data and unlabeled data to obtain a first feature map and a second feature map;
[0041] a deterministic pseudo-label acquisition module that obtains pseudo-labels of the maximum connected regions of the first feature map and the second feature map through discrimination, and obtains deterministic pseudo-labels according to the pseudo-labels of the maximum connected regions and the uncertainty map;
[0042] a pixel-level contrast learning module that combines the deterministic pseudo-labels and the labeled data as a supervision signal to guide the segmentation model to perform pixel-level contrast learning to improve the feature extraction ability.
[0043] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects:
[0044] The present invention proposes a connectivity pseudo-label processing method for the common pixel connectivity problem in multi-class heart segmentation tasks. For unlabeled data, after integrating connectivity judgment in the prediction results of the model and then selecting the deterministic pseudo-labels with higher confidence and combining them with the labeled data as a supervision signal to guide the segmentation model to perform pixel-level contrast learning, it not only strengthens the model's understanding of the target structure, but also improves the accuracy of the model algorithm in maintaining the continuity of the anatomical structure.
[0045] The segmentation model of the present invention uses a first segmentation unit and a second segmentation unit composed of an encoder-decoder. Based on the feature consistency loss between the encoder and the decoder, it strengthens the transmission of feature information and improves the generalization ability of the model. Combined with the consistency module based on the maximum connected region. By discriminating to obtain the maximum connected region of the prediction results of the two decoders, further screening reliable and logical pseudo-labels for semi-supervised learning, obtaining a more accurate shape mapping from the unlabeled data, and avoiding discontinuous errors in the prediction results.
[0046] In the contrast learning part of the invention, the labeled data and high-quality pseudo-labels generated based on the largest connected component are combined as the supervision signals. This strategy not only optimizes the accuracy of the model for classifying each pixel, but also enhances the semantic relationship between pixels, promoting the in-depth learning of the model for the features of different regions. And a contrast learning loss function that combines labeled data and pseudo-labels is defined, aiming to maximize the amount of information learned by the model from complex image data. This loss function takes into account the direct supervision information from the labeled data and the unlabeled data guided by the pseudo-labels, thus comprehensively improving the training efficiency and generalization ability of the model. Brief Description of the Drawings
[0047] Figure 1 It is the model framework diagram of the present invention;
[0048] Figure 2 It is the visualization diagram of the unconnected region situation;
[0049] Figures 3(a) and 3(b) are respectively the comparison diagrams of different method metrics on the 5% and 10% labeled ACDC datasets;
[0050] Figure 4 It is the result comparison diagram under different labeled data;
[0051] Figure 5 It is the comparison diagram of the qualitative results of the model on the ACDC dataset.
[0052] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments. Detailed Embodiments
[0053] The present invention will be further described below through specific embodiments.
[0054] See Figure 1 , a semi-supervised medical image segmentation method based on anatomical connectivity, which mainly includes three parts, feature consistency constraint (FC) between the encoder and the decoder, pseudo-label generation strategy MCR based on the largest connected component, and pixel-level contrastive learning (PCL) based on pseudo-labels. The specific description is as follows:
[0055] S1 Use a medical image segmentation model to segment the input labeled data and unlabeled data to obtain a first feature map and a second feature map.
[0056] The medical image segmentation model includes a first segmentation unit and a second segmentation unit that use the U-Net network as the backbone network; the first segmentation unit segments the input labeled data and unlabeled data to obtain a first feature map. The first segmentation unit includes an encoder and a first decoder. The first decoder uses a transposed convolutional layer and uses a skip connection with one of the encoders to transfer information between different layers. The second segmentation unit segments the input labeled data and unlabeled data to obtain a second feature map. The second segmentation unit includes an encoder and a second decoder. The second decoder uses a linear interpolation layer and uses a skip connection with the encoder to transfer information between different layers.
[0057] The encoders of the first segmentation unit and the second segmentation unit of the present invention can be the same, while the decoders are different. The encoder is responsible for capturing image features and gradually reducing the resolution. The encoder contains four downsampling stages, each stage consisting of two convolutional layers and a pooling layer. After each convolutional layer, batch normalization and activation function operations are performed to enhance the stability of the network. When the input image passes through each stage, the size of the output feature map is reduced to half of the original, and the channel dimension is doubled. The decoder is responsible for gradually restoring the resolution and generating a prediction result of the same size as the input image.
[0058] The invention strengthens the transfer of feature information and improves the generalization ability of the model by designing a feature consistency loss between the encoder and the decoder. First, for the labeled data X1 and its label y, and the unlabeled data Xu, the architecture of the double decoder provides two different perspectives on the same data, obtaining the output predictions P a and P b .
[0059] Specifically, for the labeled data, the CE function and the Dice function are used to jointly calculate the supervised loss of the labeled data, as follows:
[0060] L a = CE(P a , y) + Dice(P a , y);
[0061] L b = CE(P b , y) + Dice(P b , y);
[0062] L sup = L a + L b ;
[0063] Among them, P a represents the first feature map, P b represents the second feature map, y represents the label of the labeled data, L sup represents the supervised loss of the labeled data, La The supervision loss representing the first feature map, L b The supervision loss representing the second feature map, Dice(·) represents the Dice loss calculation function, and CE(·) represents the cross-entropy loss calculation function.
[0064] The feature gap loss between the encoder and the decoder includes attention map generation and gap optimization, which are as follows:
[0065] Attention map generation: First, compress the feature map by applying global average pooling (GAP) to each layer of the encoder to obtain the corresponding attention map. For the feature map of any layer in the network The corresponding attention map can be generated through the following mapping function:
[0066]
[0067] where p = 2 to emphasize the sum of squares of elements within the channel, thus highlighting important features more prominently, C m ×H m ×W m represents the number of channels, height, and width of the feature map, H m ×W m represents the height and width of the attention map.
[0068] Gap optimization: After obtaining the attention maps corresponding to the encoder and the decoder, calculate the gap loss between the two attention maps:
[0069]
[0070] where, A d,i and A e,i respectively represent the i-th pixel of the m-th attention map in the decoder and the encoder, and λ(t) is a time-based Gaussian temperature function, defined as:
[0071]
[0072] t is the current training iteration number, and t max is the maximum iteration number. Due to the characteristics of the linear interpolation upsampling layer, the second decoder with a linear interpolation upsampling layer is used to calculate the feature gap loss. By this method, the model is encouraged to maintain feature consistency between the encoder and the decoder, thus promoting more accurate learning and prediction of anatomical structures.
[0073] S2 obtains the pseudo-labels of the largest connected regions of the first and second feature maps through discrimination, and obtains the deterministic pseudo-labels based on the pseudo-labels of the largest connected regions and the uncertainty map.
[0074] For unlabeled data, by selecting the first feature map P a and the second feature map P b to further improve the spatial consistency of the prediction by selecting the largest connected region for each category in them. The process of selecting the connected region is as follows:
[0075]
[0076]
[0077] where C a,c and C b,c are the largest connected regions of category c in the first feature map P a and the second feature map P b ; ConnComp(Pa == c) represents returning an array that contains all the connected regions belonging to category c, and |i| represents the number of pixels in connected region i; the argmax function represents selecting the connected region with the largest number of pixels;
[0078] Take the union of the largest connected regions C a,c and C b,c to obtain the pseudo-label C' for constructing the consistency of the largest connected regions of the two. If there are valid predictions at a certain position in both regions, select the category with a higher prediction probability as the label at that position, which is represented by the following formula:
[0079]
[0080] where N is the total number of categories, and ∪ represents the union of regions.
[0081] If there are conflicts for some pixels on the largest connected regions C a,c and C b,c (that is, these pixels are labeled as category c in both predictions), resolve the conflicts by comparing the prediction probabilities at the corresponding positions:
[0082] Let P a,c and P b,c represent the probabilities of the first feature map P a and the second feature map P b for category c respectively. Then the category label of each pixel can be determined as follows:
[0083]
[0084] where (i, j) represents the position of the pixel, and C'(i, j) is the value of the pseudo-label C' at position (i, j); if the position (i, j) is in C a,c and P a,c (i, j) is greater than P b,c(i, j), then this position is marked as category c; if the position (i, j) is in C b,c and P b,c (i, j) is greater than or equal to P a,c(i,j) , then it is also marked as category c; if the position (i, j) is not in the largest connected region of any category, then it is marked as background or unlabeled in the pseudo-label C' and represented by 0.
[0085] The pseudo-label C' retains the probability information by selecting the category with a higher probability among the two prediction results at each pixel position, and maintains the largest connectivity of the category through the union operation.
[0086] In this step, the uncertainty map is used to optimize and screen the pseudo-labels of the obtained largest connected region to obtain deterministic pseudo-labels, which can improve the segmentation accuracy of the model to a certain extent and reduce the situation of mis-segmentation.
[0087] The present invention uses cross-entropy loss to measure the consistency loss L between the prediction output of the segmentation model and the pseudo-labels mcr , specifically as follows:
[0088]
[0089] where (i, j) represents the position of the pixel, c represents the category, and C' i,j,c is the value of the pseudo-label at the position (i, j) and category c, and P a,i,j,c and P b,i,j,c are the probability values of the pixels in the same position and category in the first feature map and the second feature map respectively. The overall consistency loss is calculated by summing over all positions, categories, and predictions, and then semi-supervised learning is carried out using the unlabeled data.
[0090] S3 combines the deterministic pseudo-labels and the labeled data as the supervision signal to guide the segmentation model to perform pixel-level contrast learning to improve the feature extraction ability, and realizes medical image segmentation based on the optimized medical image segmentation model. The labeled data in this step can be different from the labeled data in step S1 or the same data.
[0091] The loss function in the contrast learning part of the present invention uses infoNCE to calculate the pixel-level contrast loss between the retained feature tensors and the categories, aiming to maximize the amount of information learned by the model from complex image data. This loss function takes into account the direct supervision information from the labeled data and the unlabeled data guided by the pseudo-labels, thus comprehensively improving the training efficiency and generalization ability of the model. First, the similarity between the anchor sample and other samples is calculated, and then the loss is calculated according to the temperature parameter. The loss function is expressed in the following form:
[0092]
[0093] Among them, L pcl represents the pixel-level contrast loss, V n represents the nth anchor vector, and represent the nth positive sample vector and the mth negative sample vector respectively, K represents the number of negative samples, N' represents the number of anchor vectors, and τ is the temperature parameter.
[0094] The total loss function Loss of the medical image segmentation model of the present invention is as follows:
[0095] Loss = L sup + λ 1 L fc + λ 2 L mcr + λ 3 L pcl ;
[0096] Among them, L sup is the supervision loss, L pcl is the pixel-level contrast learning loss, L mcr is the maximum connected region consistency loss based on, L fc is the encoder-decoder consistency loss, λ 1 , λ 2 and λ 3 are preset coefficients, and can be set to change according to the Gaussian temperature function with the number of iterations, and their values increase from 0 to 0.1.
[0097] A semi-supervised medical image segmentation device based on anatomical connectivity, which uses the above-mentioned semi-supervised medical image segmentation method based on anatomical connectivity to implement the segmentation of medical images, including:
[0098] A segmentation module that uses a medical image segmentation model to segment the input labeled data and unlabeled data to obtain a first feature map and a second feature map; this segmentation module is used to execute step S1 of the above method.
[0099] A deterministic pseudo-label acquisition module that obtains the pseudo-labels of the maximum connected regions of the first feature map and the second feature map through discrimination, and obtains the deterministic pseudo-labels according to the pseudo-labels of the maximum connected regions and the uncertainty map. This deterministic pseudo-label acquisition module is used to execute step S2 of the above method.
[0100] A pixel-level contrast learning module that combines the deterministic pseudo-labels and the labeled data as a supervision signal to guide the segmentation model to perform pixel-level contrast learning to improve the feature extraction ability. This pixel-level contrast learning module is used to execute step S3 of the above method.
[0101] To verify the effectiveness of the method of the present invention, the following comparative experiments are carried out for analysis: the performance of the proposed method is evaluated on the ACDC dataset. The human body images are mainly included, covering two types of segmentation tasks, 2D and 3D, such as: the automated cardiac diagnosis challenge dataset, the left atrial magnetic resonance image segmentation dataset of atrial fibrillation patients, etc.
[0102] The experiments selected advanced methods such as UAMT of the uncertainty-guided segmentation model, DTC proposed for multi-task consistency in medical image segmentation, and MC-Net+ for mutual consistency learning of cyclic pseudo-labels for comparative analysis, and used U-Net as the backbone network. In addition, the metrics of U-Net under 5% and 10% labeled data are reported as the reference performance baseline for comparison results.
[0103] To ensure the fairness of the experiment, the comparison methods used in this experiment are all publicly available official codes and hyperparameters. And the results of a single fully supervised U-Net network are selected as the benchmark for comparison. The present invention has made further improvements in the multi-class cardiac segmentation task. Table 1 shows that the comparative experiments carried out on the ACDC dataset clearly demonstrate the influence of different algorithms and label data ratios on the cardiac segmentation task.
[0104] Table 1
[0105]
[0106] The experimental results show that as the proportion of labeled data increases, the performance of all models has been significantly improved. When comparing different semi-supervised learning strategies, the method of the present invention has far exceeded the existing methods in terms of the improvement of the four evaluation metrics of Dice, Jaccard, 95HD, and ASD. Especially under the low label data ratio (5%), the method has achieved the best performance in all evaluation metrics. In the case of 5% label data, the Dice coefficient reached 83.39%. Compared with the baseline model U-Net, the four metrics of the method of the present invention have increased by 37.06%, 35.45%, 18.83, and 8.54; far exceeding 52.38% of MT and the performance of other algorithms. This result highlights the superior performance of the method of the present invention in dealing with the situation of scarce labels.
[0107] In addition, as the labeled data increases, the performance difference between algorithms gradually decreases. Under the condition of 10% label data, the Dice coefficient of the method of the present invention reached 89.68%. Compared with the baseline model U-Net, the four metrics of the method of the present invention have increased by 10.31%, 14.96%, 9.45, and 2.56. After adding connectivity discrimination, the method of the present invention obtained the best segmentation result. This proves the effectiveness of the method in comprehensively improving the segmentation accuracy and reducing the prediction error.
[0108] Figures 3(a) and 3(b) more intuitively show the result changes of all methods with different numbers of labels and the gaps between each model and the fully supervised U-Net model. In addition, Figure 4 it shows the changes of the method of the present invention with the growth of labeled data. It can be observed that not only good results are achieved in the case of extremely few labeled data, but when the amount of labeled data is increased to 20%, the gaps in the Dice and Jc metrics with the fully supervised model are reduced to 1.57% and 2.14%.
[0109] Figure 5 It shows the visual comparison results of 2D slices of the ACDC dataset under 5% and 10% labeled data. The model of the present invention has fewer redundant regions, effectively reduces the prediction of non-connected regions, strengthens the feature map transmission between the encoder and decoder, and significantly improves the segmentation accuracy. Especially in the case of extremely few labeled data, the model of the present invention can accurately identify and correctly label the target region, while the comparison models often produce incorrect segmentations in these challenging regions. The results under 10% labeled data, in which the present model shows its advantage in preserving the integrity of the cardiac ventricular septum structure. Compared with other methods, it shows a significant improvement in segmentation coherence. In addition, by comparing the performance of the models under low-proportion labeled data, it can be seen that the present model has a stronger adaptability to unlabeled data. Even with only 5% labeled data, the present model can still maintain the coherence and consistency of the segmentation results, thus highlighting its robustness and practicality.
[0110] The present invention should conduct ablation experiment analysis, and conduct ablation experiments on the ACDC dataset with 5% and 10% labeled data to test the effectiveness of each module of the model of the present invention. The experiments are divided into two groups, using 5% and 10% labeled samples respectively, and the results are shown in Table 2. For the case of using 3 labeled samples, the improvement in overall performance is more significant. Introducing the model L with the encoding-decoding consistency constraint module FC After that, compared with the baseline model trained only with labeled data, the Dice coefficient and Jaccard index are increased by 5.19% and 4.84% respectively, while 95HD and ASD are reduced by 2.14 and 0.48 respectively.
[0111] Further adding the pseudo-label generation strategy L based on the largest connected region MCR After that, compared with the baseline model trained only with labeled data, the Dice coefficient and Jaccard index are increased by 12.41% and 15.97% again, and 95HD and ASD are reduced by 4.20 and 1.49 respectively. It can be seen the importance of connectivity discrimination for cardiac segmentation. When combining all modules of the present invention (including pixel-level contrast learning L based on pseudo-labels PCL ), compared with only using LSeg For the baseline model, the Dice coefficient increased by 22.35%, the Jaccard index increased by 25.29%, and 95HD and ASD decreased by 6.58 and 3.00 respectively, achieving a significant improvement.
[0112] For the case of using 7 label samples, the improvement trend of the model is similar to that of using 3 samples. After adding L FC the Dice coefficient increased by 2.11%, while after adding L MCR it only increased by 0.30%. However, after combining all modules, compared with the baseline model, the Dice coefficient increased by 3.48%, the Jaccard index increased by 5.38%, and 95HD and ASD decreased by 9.45 and 2.67 respectively, showing a more comprehensive performance improvement. These results clearly show that it can be seen that the addition of different modules has a significant positive impact on the model performance, especially in maintaining the anatomical structure continuity and reducing prediction errors, demonstrating the complementarity and effectiveness of each module.
[0113] Table 2
[0114]
[0115] Parameter analysis was carried out for the maximum connectivity region pseudo-label strategy of the present invention, and only the contrast learning module was retained. Table 3 shows the comparison results of different contrast learning methods. By comparing the performance of three different contrast learning methods on the ACDC dataset, the following results can be observed: The pseudo-label method provides a baseline performance, showing the effect of training solely with pseudo-labels without further optimization, with a Dice coefficient of 87.48%, a Jaccard coefficient of 78.40%, a 95HD of 10.50, and an ASD of 2.84.
[0116] Table 3
[0117]
[0118] When combined with the use of uncertainty maps for screening reliable pixels, it shows an improvement compared to the baseline pseudo-labeling method. The Dice coefficient increases to 87.99%, the Jaccard coefficient is 79.28%, 95HD decreases to 7.58, and ASD decreases to 1.91. This indicates that by leveraging uncertainty maps to optimize pseudo-labels, the segmentation accuracy of the model can be improved to a certain extent and the situation of mis-segmentation can be reduced. Moreover, the uncertainty map of the present invention combined with connected component discrimination for deterministic pseudo-labels performs optimally, with the Dice coefficient further increased to 88.73%, the Jaccard coefficient being 80.38%, 95HD significantly reduced to 4.85, and ASD also reduced to the lowest 1.43. This result shows that by adding connected component discrimination in each category, the reliability of pseudo-labels is improved, the prediction of non-anatomical structures is reduced, and the model's understanding of anatomical structures is enhanced. It effectively reduces the noise and errors of pseudo-labels during the training process, thereby improving the performance of the model in a semi-supervised learning environment.
[0119] In summary, the present invention proposes a connectivity pseudo-label processing method for the common pixel connectivity problem in multi-class cardiac segmentation tasks. Based on the anatomical connectivity common sense, this method optimizes the generation process of pseudo-labels, thus ensuring the anatomical consistency of the segmentation results. In addition, the consistency of the feature space is combined with the connectivity of the prediction results. This method not only ensures high segmentation accuracy but also significantly improves the accuracy of the model for multi-class cardiac segmentation. The experimental results show that the semi-supervised learning strategy combining anatomical connectivity with deep learning technology can not only improve the accuracy of specific organ segmentation but also greatly enhance the overall performance and reliability of medical image analysis tasks by ensuring the continuity and consistency of predictions.
[0120] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0121] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0122] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure.
[0123] The above are only the specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection of the present invention.
Claims
1. A semi-supervised medical image segmentation method based on anatomical connectivity, characterized in that: include: Using a medical image segmentation model to segment the input labeled data and unlabeled data to obtain a first feature map and a second feature map; Obtaining a pseudo label of a maximum connected area of the first feature map and the second feature map by discrimination, and obtaining a deterministic pseudo label according to the pseudo label of the maximum connected area and an uncertainty map; Combining the deterministic pseudo-labels with the labeled data as supervisory signals to guide the medical image segmentation model to perform pixel-level contrast learning to improve the feature extraction capability; For the unlabeled data, by selecting the first feature map P a and the second feature map P b The maximum connected area of each category in the prediction is further improved to improve the spatial consistency of the prediction. The selection process of the connected area is as follows: Among them C a,c and C b,c is the first feature map P a and the second feature map P b The maximum connected area of category c in the graph, ConnComp(Pa==c) means returning an array containing the first feature graph P a All connected regions of category c in , |i| represents the number of pixels in connected region i; the argmax function represents selecting the connected region with the maximum number of pixels; The maximum connected area C of the two a,c and C b,c Take the union and get the pseudo label C' of the largest connected area, which is expressed by the following formula: Where N is the total number of categories, ∪ represents the union of regions; If some pixels are in the largest connected area C a,c and C b,c If there is a conflict, the conflict is resolved by comparing the predicted probabilities at the corresponding positions: Let P a,c and P b,c Respectively represent the first feature map P a And the second feature map P b The probability of being in category c, the category label of each pixel is determined as follows: Where (i, j) represents the position of the pixel, C′(i, j) is the value of the pseudo label C′ at position (i, j); if position (i, j) is in C a,c Medium and P a,c (i, j) is greater than P b,c (i, j), then the position is marked as category c; if the position (i, j) is in C b,c Medium and P b,c (i, j) is greater than or equal to P a,c (i, j), it is also marked as category c; if the position (i, j) is not in the maximum connected area of any category, it is marked as background or unlabeled in the pseudo-label C′ and represented by 0.
2. The semi-supervised medical image segmentation method based on anatomical connectivity as claimed in claim 1, characterized in that: The medical image segmentation model includes a first segmentation unit and a second segmentation unit using a U-Net network as a backbone network; the first segmentation unit segments the input labeled data and unlabeled data to obtain the first feature map, the first segmentation unit includes an encoder and a first decoder, the first decoder uses a transposed convolutional layer and uses a jump connection with one of the encoders to transfer information between different layers; The second segmentation unit segments the input labeled data and the unlabeled data to obtain the second feature map. The second segmentation unit includes an encoder and a second decoder. The second decoder uses a linear interpolation layer and uses a jump connection with the encoder to transfer information between different layers.
3. The semi-supervised medical image segmentation method based on anatomical connectivity as claimed in claim 1, characterized in that: For the labeled data, the CE function and the Dice function are used to jointly calculate the supervision loss of the labeled data, as follows: L a =CE(P a ,y)+Dice(P a ,y); L b =CE(P b ,y)+Dice(P b ,y); L sup =L a +L b ; Among them, P a represents the first feature map, P b represents the second feature map, y represents the label of the labeled data, L sup represents the supervision loss of label data, L a represents the supervision loss of the first feature map, L b represents the supervised loss of the second feature map, Dice(·) represents the Dice loss calculation function, and CE(·) represents the cross entropy loss calculation function.
4. The semi-supervised medical image segmentation method based on anatomical connectivity as claimed in claim 1, characterized in that: The cross entropy loss is used to measure the consistency loss L between the predicted output of the segmentation model and the pseudo label. mcr , as follows: Among them, (i, j) represents the position of the pixel, c represents the category, and C′ i,j,c is the value of the pseudo label at position (i, j) and category c, P a,i,j,c and P b,i,j,c are the probability values of pixels in the same position and category in the first feature map and the second feature map, respectively.
5. The semi-supervised medical image segmentation method based on anatomical connectivity as claimed in claim 1, characterized in that: The method of combining the deterministic pseudo-labels and the labeled data as a supervisory signal to guide the segmentation model to perform pixel-level contrast learning specifically includes: Design positive samples z+ and negative samples n in contrastive learning - , for each category c, select pixel indices from the prediction output results of the medical image segmentation model with the same anchor point as positive sample indices, and filter out pixels whose prediction results output by the medical image segmentation model are different from the current category c as negative samples; Set a fixed-size memory buffer M to store samples from different categories, use anchor points randomly selected from M to calculate the contrast loss, and then update the memory buffer; The loss function of the contrastive learning part uses infoNCE to calculate the pixel-level contrastive loss of the retained feature tensor and category: Among them, L pcl represents the pixel-level contrast loss, v n represents the nth anchor point vector, and They represent the nth positive sample vector and the mth negative sample vector respectively, K represents the number of negative samples, N' represents the number of anchor point vectors, and τ is the temperature parameter.
6. The semi-supervised medical image segmentation method based on anatomical connectivity as claimed in claim 1, characterized in that: The total loss function Loss of the medical image segmentation model is as follows: Loss=L sup +λ1L fc +λ2L mcr +λ3L pcl ; Among them, L sup To monitor the loss, L pcl is the pixel-level contrastive learning loss, L mcr is based on the maximum connected region consistency loss, L fc is the encoder-decoder consistency loss, and λ1, λ2, and λ3 are preset coefficients.
7. A semi-supervised medical image segmentation device based on anatomical connectivity, characterized in that: The semi-supervised medical image segmentation method based on anatomical connectivity according to any one of claims 1 to 6 comprises: A segmentation module, which uses a medical image segmentation model to segment the input labeled data and unlabeled data to obtain a first feature map and a second feature map; A deterministic pseudo label acquisition module, which obtains the pseudo label of the maximum connected area of the first feature map and the second feature map by discrimination, and obtains the deterministic pseudo label according to the pseudo label of the maximum connected area and the uncertainty map; The pixel-level contrastive learning module combines deterministic pseudo-labels and labeled data as supervisory signals to guide the segmentation model to perform pixel-level contrastive learning to improve the feature extraction capability.
Citation Information
Patent Citations
Medical image segmentation method based on weak supervision
CN117830332A
Semi-supervised medical image segmentation method based on mutual correction and pixel-level contrast learning
CN118587438A