Semi-supervised three-dimensional medical image segmentation method, system and device based on dual correction mutual learning, and storage medium

By adopting the dual correction mutual learning method in semi-supervised medical image segmentation, the problems of model cognitive bias and low pseudo-label quality are solved, and more efficient three-dimensional medical image segmentation performance is achieved.

CN120182599AActive Publication Date: 2025-06-20BEIHUA UNIV

Patent Information

Application Number
CN202510258701.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing semi-supervised medical image segmentation methods have problems such as model cognitive bias and low pseudo-label quality, resulting in limited segmentation performance.

Method used

Using a dual correction mutual learning method, the prediction differences between two networks of different structures are corrected and the generated preliminary prediction pseudo-mask is corrected to obtain reliable pseudo-labels to achieve mutual learning and parameter updates between networks.

Benefits of technology

It effectively reduces cognitive bias during model training, improves the quality of pseudo-labels and the performance of semi-supervised three-dimensional medical image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182599A_ABST
    Figure CN120182599A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised three-dimensional medical image segmentation method, system and equipment based on dual correction mutual learning, and a storage medium, belongs to the technical field of medical image processing, and solves the problems that an existing method has model cognition deviation and cannot obtain reliable pseudo labels. Training two networks with different structures by the three-dimensional medical image with the label, calculating the supervision loss of the three-dimensional medical image with the label, and correcting the prediction difference between the two networks with different structures to obtain two pre-training networks with different structures; a label-free three-dimensional medical image generates a preliminary prediction pseudo mask corresponding to the label-free three-dimensional medical image through two pre-training networks with different structures, and the preliminary prediction pseudo mask is corrected to obtain a reliable pseudo label corresponding to the preliminary prediction pseudo mask; obtaining two final networks with different structures based on the cross false supervision loss of the label-free three-dimensional medical image; and inputting a to-be-processed three-dimensional medical image into the final two networks with different structures to realize a three-dimensional medical image segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a semi-supervised three-dimensional medical image segmentation method, system, device, and storage medium based on dual-correction mutual learning. Background Art

[0002] The segmentation of anatomical structures and lesions in medical images plays an important role in clinical disease diagnosis, radiotherapy planning, surgical guidance, and prognosis assessment. In recent years, supervised learning techniques based on deep convolutional neural networks have achieved remarkable results in the field of medical image segmentation. However, the full-supervised learning method requires a large number of pixel-level labeled images to train the network model, which has certain limitations for the clinical three-dimensional medical image segmentation task with a small amount of pixel-level annotation data. Semi-supervised learning can use a small amount of labeled data and a large amount of unlabeled data to train the network model, effectively alleviating the dependence of the model on pixel-level labeled data, and better meeting the needs of clinical application scenarios in three-dimensional medical image segmentation.

[0003] In semi-supervised medical image segmentation, the mainstream methods are: pseudo-label-based methods and consistency regularization methods. The idea of pseudo-labels is to pre-train a model using labeled data, and then use this model to generate pseudo-labels for unlabeled data, and re-train the model with the unlabeled data. The idea of consistency regularization is to encourage the network to generate consistent predictions for an image under different perspectives or perturbations. In addition, in order to improve the segmentation performance of the model, some methods combine the two methods by introducing a cross-pseudo-supervision strategy between networks. There are also some methods, such as the Chinese patent CN118710912A, which discloses "a semi-supervised image segmentation method", and adopts a mutual learning strategy in the segmentation process, enabling the teacher model to learn knowledge from the labeled data like the student model, thus alleviating the problem of error accumulation in the teacher model.

[0004] Although these methods have achieved certain results, they still face the same challenges. On the one hand, the segmentation performance of the model depends on the quality of the pseudo-labels. Therefore, how to improve the quality of the pseudo-labels and obtain reliable pseudo-labels is a key problem to be solved. On the other hand, due to the cognitive bias in network feature learning, this bias will continuously intensify and be difficult to correct during the model training process. Therefore, how to use the differential information between network predictions to guide the network to self-correct is another key problem to be solved. Summary of the Invention

[0005] The present invention solves the problems that the existing methods have model cognitive bias and cannot obtain reliable pseudo-labels.

[0006] The semi-supervised three-dimensional medical image segmentation method based on dual-correction mutual learning described in the present invention is specifically as follows:

[0007] Preprocess the three-dimensional medical image, where the three-dimensional medical image includes a labeled three-dimensional medical image and an unlabeled three-dimensional medical image;

[0008] Use the labeled three-dimensional medical image to train two networks with different structures respectively. While calculating the supervised loss of the labeled three-dimensional medical image, correct the prediction differences between the two networks with different structures to obtain two pre-trained networks with different structures respectively;

[0009] The unlabeled three-dimensional medical image passes through two pre-trained networks with different structures respectively to generate corresponding preliminary predicted pseudo-masks. The preliminary predicted pseudo-masks generated by the two pre-trained networks with different structures are both corrected to obtain corresponding reliable pseudo-labels;

[0010] Based on the cross pseudo-supervised loss of the unlabeled three-dimensional medical image, realize the mutual learning between the two pre-trained networks with different structures, update the parameters of the two pre-trained networks with different structures, and obtain two final networks with different structures respectively;

[0011] The three-dimensional medical image to be processed is input into the two final networks with different structures respectively to implement the three-dimensional medical image segmentation task.

[0012] Furthermore, in an embodiment of the present invention, the two networks with different structures are respectively an uncertainty-aware three-dimensional segmentation network and a structure-enhanced three-dimensional segmentation network;

[0013] The decoder end of the uncertainty-aware three-dimensional segmentation network adopts an uncertainty-aware module. The uncertainty-aware module is specifically:

[0014]

[0015] where, u i is the overall uncertainty of sample i, is the trust quality of class k in sample i;

[0016] The encoder end of the structure-enhanced three-dimensional segmentation network adopts a structure-enhanced module. The structure-enhanced module is specifically:

[0017]

[0018] where, f ψ 、f x and are 1×1×1 convolutional layers respectively, x i and are the feature maps of the i-th layer encoder and decoder respectively, σ is the sigmod activation function, The i-th layer encoder feature map after enhancing the boundary contour information of the target region, and α is the feature map containing structural enhancement information.

[0019] Further, in an embodiment of the present invention, the correction of the prediction differences between two networks with different structures is specifically as follows:

[0020] Perform an exclusive OR operation on the binary mask maps respectively predicted and output by two networks with different structures to obtain the inconsistent partial mask between the predictions of the two networks with different structures;

[0021] Based on the inconsistent partial mask between the predictions of the two networks with different structures, calculate the prediction difference regions of the two networks with different structures;

[0022] Calculate the correction loss between the prediction difference regions of the two networks with different structures and their corresponding true mask regions to guide the two networks with different structures to correct wrong predictions.

[0023] Further, in an embodiment of the present invention, the obtaining of the inconsistent partial mask between the predictions of the two networks with different structures is specifically as follows:

[0024]

[0025] Wherein, and are respectively the softmax outputs of the labeled three-dimensional medical image passing through two networks with different structures, is the exclusive OR operation, BINA is the binarization operation, and M diff is the inconsistent partial mask after the labeled three-dimensional medical image is predicted by two networks with different structures;

[0026] The calculation of the prediction difference regions of the two networks with different structures is specifically as follows:

[0027]

[0028] Wherein, is the prediction difference region of the two networks with different structures, includes and CLIP is the operation of obtaining the prediction difference region corresponding to the inconsistent partial mask region;

[0029] The calculation of the correction loss between the prediction difference regions of the two networks with different structures and their corresponding true mask regions is specifically as follows:

[0030]

[0031] Wherein, For predicting the true mask image corresponding to the difference region, MSE is the mean squared error loss function, is the calibration loss between the predicted difference regions of two networks with different structures and their corresponding true mask regions.

[0032] Furthermore, in an embodiment of the present invention, the preliminary predicted pseudo-masks generated by two pre-trained networks with different structures are both calibrated to obtain their corresponding reliable pseudo-labels, specifically:

[0033] The two pre-trained networks with different structures are a pre-trained uncertainty-aware 3D segmentation network and a pre-trained structure-enhanced 3D segmentation network respectively;

[0034] The preliminary predicted pseudo-mask generated by the pre-trained uncertainty-aware 3D segmentation network is calibrated through uncertainty awareness to obtain its corresponding reliable pseudo-label;

[0035] The preliminary predicted pseudo-mask generated by the pre-trained structure-enhanced 3D segmentation network is calibrated through dynamically generating a high-confidence mask to obtain its corresponding reliable pseudo-label.

[0036] Furthermore, in an embodiment of the present invention, the uncertainty awareness calibration is specifically:

[0037]

[0038] Wherein, is the predicted probability map of the unlabeled 3D medical image is the binary map obtained through the argmax function, P A is the reliable pseudo-label generated after uncertainty awareness calibration, is the uncertainty map, is the indicator function, T = 0.2 is the threshold;

[0039] The dynamically generating high-confidence mask calibration is specifically:

[0040]

[0041] Wherein, C A and C B are the confidence score maps predicted by the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network respectively, ∈ is a very small constant, and are both the probabilities that a pixel point belongs to the k-th class, is the high-confidence mask map, is the predicted probability map of the unlabeled 3D medical image The binary image obtained through the argmax function, P B is the reliable pseudo-label generated after correction by dynamically generating a high-confidence mask.

[0042] Furthermore, in an embodiment of the present invention, the supervision loss is specifically:

[0043]

[0044] where is the supervision loss, is the supervision loss of the pre-trained uncertainty-aware 3D segmentation network, is the supervision loss of the pre-trained structure-enhanced 3D segmentation network, is the correction loss between the prediction difference region of two networks with different structures and its corresponding true mask region, g is the operation for calculating uncertainty, is the preliminary prediction probability map obtained by the uncertainty-aware 3D segmentation network for the labeled 3D medical image, is the preliminary prediction probability map obtained by the structure-enhanced 3D segmentation network for the labeled 3D medical image, and Y is the label map corresponding to the labeled 3D medical image;

[0045] The cross pseudo-supervision loss is specifically:

[0046]

[0047] where is the cross pseudo-supervision loss, P A is the reliable pseudo-label generated after correction by uncertainty awareness, P B is the reliable pseudo-label generated after correction by dynamically generating a high-confidence mask, is the cross-entropy loss function.

[0048] The semi-supervised 3D medical image segmentation system based on dual-correction mutual learning of the present invention is specifically:

[0049] Preprocess the 3D medical image, where the 3D medical image includes labeled 3D medical images and unlabeled 3D medical images;

[0050] Use the labeled 3D medical image to train two networks with different structures respectively. While calculating the supervision loss of the labeled 3D medical image, correct the prediction difference between the two networks with different structures to obtain two pre-trained networks with different structures respectively;

[0051] The unlabeled 3D medical images are respectively passed through two pre-trained networks with different structures to generate corresponding preliminary predicted pseudo-masks, and the preliminary predicted pseudo-masks generated by the two pre-trained networks with different structures are both corrected to obtain corresponding reliable pseudo-labels;

[0052] Based on the cross pseudo-supervision loss of the unlabeled 3D medical images, mutual learning between two pre-trained networks with different structures is realized, the parameters of the two pre-trained networks with different structures are updated, and finally two networks with different structures are obtained respectively;

[0053] The 3D medical images to be processed are respectively input into the finally obtained two networks with different structures to implement the 3D medical image segmentation task.

[0054] An electronic device according to the present invention includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0055] The memory is used to store a computer program;

[0056] The processor is used to implement the method steps described in the above-mentioned method 1 when executing the program stored on the memory.

[0057] A computer-readable storage medium according to the present invention stores a computer program therein, and when the computer program is executed by a processor, it implements the method steps described in any of the above-mentioned methods.

[0058] The present invention solves the problems that the existing methods have model cognitive biases and cannot obtain reliable pseudo-labels. The specific beneficial effects include:

[0059] 1. For the semi-supervised 3D medical image segmentation method based on dual correction and mutual learning of the present invention, the segmentation performance of existing models depends on the quality of pseudo-labels. Therefore, it is necessary to improve the quality of pseudo-labels. To solve the above technical problems, the present invention corrects the generated preliminary predicted pseudo-masks to obtain reliable pseudo-labels of unlabeled 3D medical images passing through two networks with different structures;

[0060] 2. For the semi-supervised 3D medical image segmentation method based on dual correction and mutual learning of the present invention, there are cognitive biases in the existing network feature learning, resulting in the continuous aggravation and difficulty in correcting this bias during the model training process. Therefore, how to guide the network to self-correct is another key problem to be solved. To solve the above technical problems, the present invention extracts partial information of the prediction differences of labeled images passing through two networks with different structures by using a difference correction module, and guides the two networks to self-correct based on the differences generated by this module, reducing the cognitive bias during the model training process.

[0061] 3. The semi-supervised 3D medical image segmentation method based on dual-correction mutual learning according to the present invention improves the performance of semi-supervised 3D medical image segmentation through dual-correction and mutual learning between sub-networks using the supervised loss of labeled images and the cross pseudo-supervised loss of unlabeled images. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the following description of embodiments in conjunction with the accompanying drawings, where:

[0063] Figure 1 is a schematic flowchart of the semi-supervised 3D medical image segmentation method based on dual-correction mutual learning described in Embodiment 1;

[0064] Figure 2 is a schematic diagram of the network structure described in Embodiment 2;

[0065] Figure 3 is a comparison schematic diagram with other methods after training the model with 20% labels on the LA dataset described in Embodiment 5. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The following will clearly and completely describe various embodiments of the present invention in conjunction with the accompanying drawings. The embodiments described by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0067] Embodiment 1. The semi-supervised 3D medical image segmentation method based on dual-correction mutual learning described in this embodiment is specifically as follows:

[0068] Preprocess the 3D medical images, where the 3D medical images include labeled 3D medical images and unlabeled 3D medical images;

[0069] Use the labeled 3D medical images to train two different-structured networks respectively. While calculating the supervised loss of the labeled 3D medical images, correct the prediction differences between the two different-structured networks to obtain two different-structured pre-trained networks respectively;

[0070] The unlabeled 3D medical images respectively pass through two different-structured pre-trained networks to generate corresponding preliminary prediction pseudo-masks. The corresponding preliminary prediction pseudo-masks generated by the two different-structured pre-trained networks are both corrected to obtain corresponding reliable pseudo-labels;

[0071] Based on the cross pseudo-supervised loss of the unlabeled 3D medical images, realize the mutual learning between the two different-structured pre-trained networks, update the parameters of the two different-structured pre-trained networks, and obtain two different-structured networks respectively finally;

[0072] The three-dimensional medical images to be processed are respectively input into two networks with different structures at the end to implement the three-dimensional medical image segmentation task.

[0073] In the existing technology, on the one hand, the segmentation performance of the model depends on the quality of the pseudo-labels. Therefore, how to obtain reliable pseudo-labels is the key problem to be solved. On the other hand, due to the existence of cognitive bias in network feature learning, this bias will continue to intensify and be difficult to correct during the model training process. Therefore, how to guide the network to self-correct is another key problem to be solved.

[0074] To solve the above technical problems, as Figure 1 shown, this embodiment proposes a semi-supervised three-dimensional medical image segmentation method based on dual-correction mutual learning, specifically:

[0075] Load the three-dimensional medical images and preprocess the three-dimensional medical images, specifically: normalize the three-dimensional medical image volume to zero mean and unit variance, and crop the three-dimensional medical image volume to expand the edge according to the target region position so that its size is 112×112×80;

[0076] The three-dimensional medical images described include labeled three-dimensional medical images and unlabeled three-dimensional medical images;

[0077] Use the labeled three-dimensional medical images to train two networks with different structures respectively. While calculating the supervised loss of the labeled three-dimensional medical images, correct the prediction differences between the two networks with different structures to obtain two pre-trained networks with different structures respectively;

[0078] The unlabeled three-dimensional medical images respectively pass through two pre-trained networks with different structures to generate corresponding preliminary predicted pseudo-masks. The preliminary predicted pseudo-masks generated by the two pre-trained networks with different structures are both corrected to obtain corresponding reliable pseudo-labels;

[0079] Based on the cross pseudo-supervised loss of the unlabeled three-dimensional medical images, realize the mutual learning between the two pre-trained networks with different structures, update the parameters of the two pre-trained networks with different structures, and obtain two networks with different structures at the end respectively;

[0080] Send the three-dimensional medical images into the two networks with different structures at the end that have been trained. By automatically extracting the features in the images, finally obtain the segmentation result map of the three-dimensional medical images.

[0081] In this embodiment, by using labeled 3D medical images to correct the prediction differences of two networks with different structures, the self-correction of the network is realized to reduce the cognitive bias problem of the model. And by correcting the initially generated prediction pseudo-masks, reliable pseudo-labels of unlabeled 3D medical images through two networks with different structures are obtained, realizing reliable mutual learning between two networks with different structures to improve the performance of semi-supervised 3D medical image segmentation.

[0082] Embodiment 2: This embodiment further limits the semi-supervised 3D medical image segmentation method based on dual-correction mutual learning described in Embodiment 1. The two networks with different structures are respectively an uncertainty-aware 3D segmentation network and a structure-enhanced 3D segmentation network;

[0083] At the decoder end of the uncertainty-aware 3D segmentation network, an uncertainty-aware module is adopted. The uncertainty-aware module is specifically:

[0084]

[0085] where u i is the overall uncertainty of sample i, is the trust quality of class k in sample i;

[0086] At the encoder end of the structure-enhanced 3D segmentation network, a structure-enhanced module is adopted. The structure-enhanced module is specifically:

[0087]

[0088] where f ψ 、f x and are respectively 1×1×1 convolutional layers, x i and are respectively the feature maps of the i-th layer encoder and decoder, σ is the sigmod activation function, is the i-th layer encoder feature map after enhancing the boundary contour information of the target region, and α is the feature map containing structure-enhanced information.

[0089] In this embodiment, as Figure 2 shown, the two networks with different structures are respectively an uncertainty-aware 3D segmentation network 3D-ResVnet and a structure-enhanced 3D segmentation network Vnet;

[0090] The uncertainty-aware 3D segmentation network is specifically:

[0091] At the decoder end of the uncertainty-aware 3D segmentation network, an uncertainty-aware module is adopted to simultaneously output the predicted segmentation probability map of the network and the corresponding uncertainty map;

[0092] The described uncertainty perception module is specifically as follows:

[0093] For an image segmentation task with K classes (including background), it is described using trust quality and overall uncertainty, specifically as follows:

[0094]

[0095] Among them, u i is the overall uncertainty of sample i (u i ≥0), is the trust quality of class k in sample i

[0096] The described u i and are specifically as follows:

[0097]

[0098] Among them, is the evidence vector e i of the generated class k, S i is the total evidence of the K-class predictions of sample i;

[0099] For evidence, which refers to the degree of support for classifying a certain pixel in the image into a certain class, the calculated total evidence is inversely proportional to the overall uncertainty. If there is no evidence, that is then the trust quality of each class prediction is 0, and the overall uncertainty is 1. On the contrary, if the total evidence is large, then the overall uncertainty u i will be small, and the prediction of the network has a higher confidence;

[0100] The described evidence vector e i is specifically as follows:

[0101]

[0102] Among them, is the preliminary prediction result obtained by passing sample i through the uncertainty perception three-dimensional segmentation network. τ is a scaling parameter, and 0 < τ < 1. Tanh is the hyperbolic tangent function, and the value range is specified as [-1, 1];

[0103] For the trust quality it corresponds to the parameter of the Dirichlet distribution function respectively, specifically as follows:

[0104]

[0105] Among them,

[0106] The described structure - enhanced three - dimensional segmentation network is specifically as follows:

[0107] In the encoder of the structure - enhanced three - dimensional segmentation network, a structure - enhanced module is adopted to enhance the structural information in the feature map and highlight the boundary contour of the target area in the feature map. Specifically:

[0108]

[0109] Among them, f ψ 、f x and are respectively 1×1×1 convolutional layers, x i and are respectively the feature maps of the i - th layer of the encoder and decoder, σ is the sigmod activation function, is the feature map of the i - th layer of the encoder after enhancing the boundary contour information of the target area, and α is the feature map containing structure - enhanced information.

[0110] Embodiment 3: This embodiment further limits the semi - supervised three - dimensional medical image segmentation method based on dual - correction mutual learning described in Embodiment 1. The correction of the prediction differences between two networks with different structures is specifically as follows:

[0111] Perform an exclusive - OR operation on the binary mask maps respectively predicted and output by two networks with different structures to obtain the inconsistent part mask between the predictions of the two networks with different structures;

[0112] Based on the inconsistent part mask between the predictions of the two networks with different structures, calculate the prediction difference region of the two networks with different structures;

[0113] Calculate the correction loss between the prediction difference region of the two networks with different structures and its corresponding true mask region to guide the two networks with different structures to correct wrong predictions.

[0114] In this embodiment, the obtaining of the inconsistent part mask between the predictions of the two networks with different structures is specifically as follows:

[0115]

[0116] Among them, and are respectively the softmax outputs of the labeled three - dimensional medical image passing through two networks with different structures, is the exclusive - OR operation, BINA is the binarization operation, and M diff is the inconsistent part mask after the labeled three - dimensional medical image is predicted by two networks with different structures;

[0117] Calculating the predicted difference regions of two networks with different structures specifically includes:

[0118]

[0119] Among them, is the predicted difference region of two networks with different structures, including and CLIP is an operation to obtain the predicted difference region corresponding to the inconsistent partial mask region;

[0120] Calculating the calibration loss between the predicted difference region of two networks with different structures and its corresponding true mask region specifically includes:

[0121]

[0122] Among them, is the true mask map corresponding to the predicted difference region, and MSE is the mean square error loss function, is the calibration loss between the predicted difference region of two networks with different structures and its corresponding true mask region.

[0123] In this embodiment, calibrating the predicted difference between two networks with different structures specifically includes:

[0124] Performing an exclusive OR operation on the binarized mask maps output by the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network to obtain the inconsistent partial mask M diff , specifically including:

[0125]

[0126] Among them, and are respectively the softmax outputs of the labeled 3D medical image passing through the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network, is the exclusive OR operation, BINA is the binarization operation, and M diff is the inconsistent partial mask after the labeled 3D medical image is predicted by the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network;

[0127] Calculating the predicted difference region according to the inconsistent partial mask predicted by the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network specifically includes:

[0128]

[0129] Among them, For the prediction difference region between the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network, including and CLIP is an operation to obtain the prediction difference region corresponding to the inconsistent partial mask region;

[0130] Correct the prediction difference between the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network, and calculate the correction loss between the prediction difference region and its corresponding ground truth mask region Guide the network to correct potential wrong predictions, specifically:

[0131]

[0132] wherein, is the ground truth mask map corresponding to the prediction difference region, and MSE is the mean square error loss function.

[0133] Embodiment 4. This embodiment further limits the semi-supervised 3D medical image segmentation method based on dual-correction mutual learning described in Embodiment 1. The preliminary prediction pseudo-masks generated by two pre-trained networks with different structures are both corrected to obtain their corresponding reliable pseudo-labels, specifically:

[0134] The two pre-trained networks with different structures are the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network respectively;

[0135] The preliminary prediction pseudo-mask generated by the pre-trained uncertainty-aware 3D segmentation network is corrected by uncertainty awareness to obtain its corresponding reliable pseudo-label;

[0136] The preliminary prediction pseudo-mask generated by the pre-trained structure-enhanced 3D segmentation network is corrected by dynamically generating a high-confidence mask to obtain its corresponding reliable pseudo-label.

[0137] In this embodiment, the uncertainty awareness correction is specifically:

[0138]

[0139] wherein, is the prediction probability map of the unlabeled 3D medical image The binary map obtained by passing through the argmax function, P A is the reliable pseudo-label generated after uncertainty awareness correction, is the uncertainty map, is the indicator function, and T = 0.2 is the threshold;

[0140] The so-called dynamic generation of high-confidence mask correction is specifically as follows:

[0141]

[0142] Among them, C A and C B are the confidence score maps predicted by the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network respectively. ∈ is a very small constant. and are both the probabilities that a pixel belongs to the k-th class. is the high-confidence mask map. is the predicted probability map of the unlabeled 3D medical image The binary map obtained through the argmax function, P B is the reliable pseudo-label generated after dynamic generation of high-confidence mask correction.

[0143] In this embodiment, after the uncertainty-aware 3D segmentation network and the structure-enhanced 3D segmentation network are trained with labeled 3D medical images, the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network are generated respectively;

[0144] The pre-trained uncertainty-aware 3D segmentation network generates its corresponding preliminary predicted pseudo-mask, which is corrected through uncertainty awareness to obtain its corresponding reliable pseudo-label;

[0145] The uncertainty map generated through the pre-trained uncertainty-aware module filters out the high-uncertainty part to improve the quality of the pseudo-labels generated by the network. Specifically:

[0146]

[0147] Among them, is the predicted probability map of the unlabeled 3D medical image The binary map obtained through the argmax function, P A is the reliable pseudo-label generated after uncertainty awareness correction. is the uncertainty map. is the indicator function, and T = 0.2 is the threshold;

[0148] The pre-trained structure-enhanced 3D segmentation network generates its corresponding preliminary predicted pseudo-mask, which is corrected through dynamic generation of high-confidence mask to obtain its corresponding reliable pseudo-label. Specifically:

[0149] Calculate the confidence of the predicted output of the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network respectively, and dynamically generate a high-confidence mask map of the pre-trained structure-enhanced 3D segmentation network by comparing the confidence scores between the two. Specifically:

[0150]

[0151]

[0152] Among them, C A and C B are the confidence score maps predicted by the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network respectively. ∈ is a very small constant. and are both the probabilities that the pixel belongs to the k-th class. is the high-confidence mask map. is the predicted probability map of the unlabeled 3D medical image. The binary map obtained through the argmax function, P B is the reliable pseudo-label generated after being corrected by the dynamically generated high-confidence mask.

[0153] Therefore, for the unlabeled images in this embodiment, two preliminary predicted pseudo-masks are generated by two pre-trained networks with different structures, and are corrected by uncertainty-aware correction and dynamically generated high-confidence mask correction respectively to obtain reliable pseudo-labels.

[0154] Embodiment 5: This embodiment further limits the semi-supervised 3D medical image segmentation method based on dual-correction mutual learning described in Embodiment 1. The supervised loss is specifically:

[0155]

[0156] Among them, is the supervised loss. is the supervised loss of the pre-trained uncertainty-aware 3D segmentation network. is the supervised loss of the pre-trained structure-enhanced 3D segmentation network. is the correction loss between the predicted difference region of the two networks with different structures and its corresponding true mask region. g is the operation for calculating uncertainty. is the preliminary predicted probability map of the labeled 3D medical image obtained by the uncertainty-aware 3D segmentation network. is the preliminary predicted probability map of the labeled 3D medical image obtained by the structure-enhanced 3D segmentation network. Y is the label map corresponding to the labeled 3D medical image.

[0157] The described cross pseudo-supervision loss is specifically as follows:

[0158]

[0159] Among them, is the cross pseudo-supervision loss, and P A is the reliable pseudo-label generated after uncertainty-aware calibration, and P B is the reliable pseudo-label generated after dynamic generation of a high-confidence mask calibration, is the cross-entropy loss function.

[0160] In this embodiment, the total loss of the calculation model includes two parts, the supervision loss and the cross pseudo-supervision loss (consistency loss), specifically as follows:

[0161]

[0162] Among them, λ is the parameter for balancing the supervision loss and the consistency loss The weight is gradually adjusted using the Gaussian ramp function , λ max = 0.1, t is the current training epoch number, and t max is the total number of training epochs;

[0163] The described supervision loss is specifically as follows:

[0164]

[0165] Among them, is the supervision loss of the pre-trained uncertainty-aware 3D segmentation network, is the supervision loss of the pre-trained structure-enhanced 3D segmentation network, is the calibration loss of the prediction difference region between two networks with different structures and its corresponding true mask region, g is the operation for calculating uncertainty, is the preliminary prediction probability map obtained by the uncertainty-aware 3D segmentation network for the labeled 3D medical image, is the preliminary prediction probability map obtained by the structure-enhanced 3D segmentation network for the labeled 3D medical image, and Y is the label map corresponding to the labeled 3D medical image;

[0166]

[0167] Among them, is the Dice loss function, is the evidence extraction loss;

[0168] The loss is calculated based on cross - entropy according to the Dirichlet distribution function The loss, and encourages the network to generate evidence for positive samples of different categories by minimizing the loss, while punishing the evidence of negative samples through the Kullback–Leibler (KL) divergence loss. Specifically:

[0169]

[0170] where \(B(\alpha i )\) is the k - dimensional polynomial beta function with parameter \(\alpha i \), \(S k \) is the k - dimensional simplex, \(\psi(\cdot)\) is the digamma function, \(\Gamma(\cdot)\) is the gamma function, \(\beta\) is used to balance the two losses, \(t\) is the current iteration index, \(t max \) is the maximum number of iterations, \(D(p i |1)\) is a Dirichlet distribution with uniform parameters;

[0171] The trust quality loss is specifically:

[0172]

[0173] where, \(\) is the simplex transformed from the belief quality \(b i \) through softmax;

[0174] The cross - pseudo - supervision loss described above is specifically:

[0175]

[0176] where, \(\) is the cross - pseudo - supervision loss, \(P A \) is the reliable pseudo - label generated after uncertainty - aware correction, \(P B \) is the reliable pseudo - label generated after dynamic generation of high - confidence mask correction, \(\) is the cross - entropy loss.

[0177] Therefore, this embodiment realizes the mutual learning between two networks by calculating the supervised loss of labeled three - dimensional medical images and the cross - pseudo - supervision loss of unlabeled three - dimensional medical images.

[0178] To better illustrate the semi - supervised three - dimensional medical image segmentation method based on dual - correction mutual learning described in any one of Embodiments 1 - 5, it is described in detail through the following examples:

[0179] To verify the effectiveness and superiority of the method of this embodiment, additional methods such as MT, UA-MT, SASSNet, DTC, MCNet, and MCF were selected for parallel comparison. The methods were quantitatively compared under 20% labeled images of the LA left atrial segmentation dataset, and the segmentation evaluation metrics of Dice, Jaccard, 95HD, and ASD were calculated for each method. The segmentation results are as Figure 3 shown. The comparison results are shown in Table 1. It can be seen from the charts that in the method of this embodiment, both the Dice and Jaccard metrics are higher than those of the comparable methods in the same category. It can be seen that the prediction result of the segmentation method in this embodiment is closest to the true result. The 95HD and ASD metrics of the method of this embodiment are lower than those of the comparable methods in the same category. It can be seen that the distance between the prediction result and the true result of the segmentation method in this embodiment is the smallest and the similarity is the highest.

[0180] Table 1

[0181]

[0182] In summary, this embodiment designs a semi-supervised 3D medical image segmentation scheme with dual-correction mutual learning, that is, using network prediction difference correction and pseudo-label correction, effectively alleviating the problems of model cognitive bias and low-quality pseudo-label filtering in existing semi-supervised medical image segmentation. Through the advantages of dual correction, reliable pseudo-masks are generated to guide network mutual learning, thereby improving the performance of semi-supervised 3D medical image segmentation.

[0183] Embodiment 6. The semi-supervised 3D medical image segmentation system based on dual-correction mutual learning described in this embodiment is specifically:

[0184] Preprocess the 3D medical images, where the 3D medical images include labeled 3D medical images and unlabeled 3D medical images;

[0185] Use the labeled 3D medical images to train two networks with different structures respectively. While calculating the supervised loss of the labeled 3D medical images, correct the prediction differences between the two networks with different structures to obtain two pre-trained networks with different structures respectively;

[0186] The unlabeled 3D medical images respectively generate corresponding preliminary prediction pseudo-masks through two pre-trained networks with different structures. The preliminary prediction pseudo-masks generated by the two pre-trained networks with different structures are both corrected to obtain corresponding reliable pseudo-labels;

[0187] Based on the cross pseudo-supervised loss of the unlabeled 3D medical images, realize the mutual learning between the two pre-trained networks with different structures, update the parameters of the two pre-trained networks with different structures, and obtain the final two networks with different structures respectively;

[0188] The three-dimensional medical images to be processed are respectively input into two different-structured networks at the end to implement the three-dimensional medical image segmentation task.

[0189] Embodiment 7. An electronic device according to this embodiment includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0190] The memory is used to store computer programs;

[0191] When the processor is used to execute the program stored on the memory, it implements the method steps described in any one of Embodiments 1-5.

[0192] Embodiment 8. A computer-readable storage medium according to this embodiment stores a computer program in the computer-readable storage medium. When the computer program is executed by a processor, it implements the method steps described in any one of Embodiments 1-5.

[0193] The semi-supervised three-dimensional medical image segmentation method, system, device, and storage medium based on dual-correction mutual learning proposed by the present invention are introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A semi-supervised 3D medical image segmentation method based on dual-correction mutual learning, characterized in that: Specifically: Preprocessing the three-dimensional medical image, wherein the three-dimensional medical image includes a labeled three-dimensional medical image and an unlabeled three-dimensional medical image; Two networks with different structures are trained respectively using labeled three-dimensional medical images, and while calculating the supervision loss of the labeled three-dimensional medical images, the prediction difference between the two networks with different structures is corrected to obtain two pre-trained networks with different structures; Unlabeled 3D medical images are respectively passed through two pre-trained networks with different structures to generate corresponding preliminary predicted pseudo masks. The preliminary predicted pseudo masks generated by the two pre-trained networks with different structures are corrected to obtain corresponding reliable pseudo labels. Based on the cross pseudo-supervision loss of unlabeled 3D medical images, mutual learning between two pre-trained networks with different structures is achieved, the parameters of the two pre-trained networks with different structures are updated, and the final two networks with different structures are obtained respectively; The three-dimensional medical images to be processed are respectively input into the final two networks with different structures to realize the three-dimensional medical image segmentation task.

2. The semi-supervised three-dimensional medical image segmentation method based on dual correction mutual learning according to claim 1, characterized in that: The two networks with different structures are respectively an uncertainty-aware three-dimensional segmentation network and a structure-enhanced three-dimensional segmentation network; The decoder end of the uncertainty-aware three-dimensional segmentation network adopts an uncertainty-aware module, and the uncertainty-aware module is specifically: Among them, u i is the overall uncertainty of sample i, is the trust quality of class k in sample i; The encoder end of the structure-enhanced three-dimensional segmentation network adopts a structure-enhanced module, and the structure-enhanced module is specifically: Among them, f ψ 、f x and They are 1×1×1 convolutional layers, x i and are the feature maps of the i-th layer encoder and decoder respectively, σ is the sigmoid activation function, is the feature map of the i-th layer encoder after enhancing the boundary contour information of the target area, and α is the feature map containing structural enhancement information.

3. The semi-supervised 3D medical image segmentation method based on dual correction mutual learning according to claim 1, characterized in that: The correction of the prediction difference between two networks with different structures is specifically: The binary mask images predicted and output by two networks with different structures are XORed to obtain the inconsistent partial masks predicted by the two networks with different structures. Based on the inconsistent partial masks predicted between the two networks with different structures, the predicted difference regions of the two networks with different structures are calculated; The correction loss between the predicted difference area of ​​two networks with different structures and their corresponding true mask area is calculated to guide the two networks with different structures to correct the wrong predictions.

4. The semi-supervised three-dimensional medical image segmentation method based on dual correction mutual learning according to claim 3, characterized in that: The method of obtaining the inconsistent partial masks predicted by two networks with different structures is specifically as follows: in, and They are the softmax outputs of labeled 3D medical images after two networks with different structures. is an XOR operation, BINA is a binary operation, M diff Mask the inconsistent parts of labeled 3D medical images after being predicted by two networks with different structures; The predicted difference area of ​​the two networks with different structures is calculated as follows: in, is the predicted difference region of two networks with different structures, include and CLIP is an operation to obtain the predicted difference region corresponding to the inconsistent partial mask region; The calculation of the correction loss between the predicted difference area of ​​the two networks with different structures and the corresponding true mask area is specifically: in, To predict the true mask map corresponding to the difference area, MSE is the mean square error loss function. It is the correction loss between the predicted difference area of ​​two networks with different structures and their corresponding true mask area.

5. The semi-supervised three-dimensional medical image segmentation method based on dual correction mutual learning according to claim 1, characterized in that: The two pre-trained networks with different structures generate the corresponding preliminary predicted pseudo masks and correct them to obtain the corresponding reliable pseudo labels, specifically: The two pre-trained networks with different structures are respectively a pre-trained uncertainty-aware three-dimensional segmentation network and a pre-trained structure-enhanced three-dimensional segmentation network; The pre-trained uncertainty-aware 3D segmentation network generates a preliminary predicted pseudo-mask corresponding to it, and then undergoes uncertainty-aware correction to obtain a reliable pseudo-label corresponding to it. The pre-trained structure-enhanced 3D segmentation network generates a preliminary predicted pseudo-mask corresponding to it, which is then corrected by dynamically generating a high-confidence mask to obtain a reliable pseudo-label corresponding to it.

6. The semi-supervised three-dimensional medical image segmentation method based on dual correction mutual learning according to claim 5, characterized in that: The uncertainty perception correction is specifically as follows: in, is the predicted probability map of unlabeled 3D medical images The binary image obtained by the argmax function, P A is a reliable pseudo-label generated after uncertainty-aware correction, A is the uncertainty diagram, is the indicator function, T = 0.2 is the threshold; The dynamic generation of high confidence mask correction is specifically as follows: Among them, C A and C B are the confidence score maps predicted by the pre-trained uncertainty-aware 3D segmentation network and the pre-trained structure-enhanced 3D segmentation network, ∈ is a very small constant, and are the probabilities that the pixel belongs to the kth class, is a high confidence mask map, is the predicted probability map of unlabeled 3D medical images The binary image obtained by the argmax function, P B Reliable pseudo labels generated after dynamic generation of high-confidence mask correction.

7. The semi-supervised three-dimensional medical image segmentation method based on dual correction mutual learning according to claim 1, characterized in that: The supervision loss is specifically: in, To monitor losses, The supervision loss for pre-training uncertainty-aware 3D segmentation networks, The supervision loss of the 3D segmentation network enhanced by the pre-trained structure, is the correction loss between the predicted difference area of ​​two networks with different structures and their corresponding true mask area, g is the computational uncertainty operation, It is the preliminary prediction probability map obtained by the uncertainty-aware 3D segmentation network for labeled 3D medical images. is the preliminary predicted probability map obtained by the structure-enhanced three-dimensional segmentation network of the labeled three-dimensional medical image, and Y is the label map corresponding to the labeled three-dimensional medical image; The cross pseudo-supervision loss is specifically: in, is the cross pseudo-supervision loss, P A is a reliable pseudo-label generated after uncertainty-aware correction, P B is a reliable pseudo-label generated after dynamic high-confidence mask correction. is the cross entropy loss function.

8. A semi-supervised 3D medical image segmentation system based on dual-correction mutual learning, characterized in that: Specifically: Preprocessing the three-dimensional medical image, wherein the three-dimensional medical image includes a labeled three-dimensional medical image and an unlabeled three-dimensional medical image; Two networks with different structures are trained respectively using labeled three-dimensional medical images, and while calculating the supervision loss of the labeled three-dimensional medical images, the prediction difference between the two networks with different structures is corrected to obtain two pre-trained networks with different structures; Unlabeled 3D medical images are respectively passed through two pre-trained networks with different structures to generate corresponding preliminary predicted pseudo masks. The preliminary predicted pseudo masks generated by the two pre-trained networks with different structures are corrected to obtain corresponding reliable pseudo labels. Based on the cross pseudo-supervision loss of unlabeled 3D medical images, mutual learning between two pre-trained networks with different structures is achieved, the parameters of the two pre-trained networks with different structures are updated, and the final two networks with different structures are obtained respectively; The three-dimensional medical images to be processed are respectively input into the final two networks with different structures to realize the three-dimensional medical image segmentation task.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Semi-supervised image segmentation method

    CN118710912A

  • Semi-supervised medical image segmentation method based on mutual correction and pixel-level contrast learning

    CN118587438A

  • Semi-supervised medical image segmentation method based on single-cycle regularization

    CN119478419A

  • Object region segmentation device and object region segmentation method thereof

    US20240338934A1

  • Mutual learning-based semi-supervised medical image segmentation method and system

    WO2023116635A1

Cited By

  • Semi-supervised medical image segmentation method based on asymmetric deformation guide mutual learning

    CN122023791A