Medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-tag
Through the method of progressive regional exchange and multi-dimensional collaborative pseudo-label, mixed images and labels with different semantic complexities are generated, which solves the problems of data distribution mismatch and pseudo-label quality in semi-supervised medical image segmentation, and achieves high-precision medical image segmentation.
Patent Information
- Application Number
- CN202510415853.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing semi-supervised medical image segmentation methods face the problems of data distribution mismatch, unstable pseudo-label quality and difficulty in evaluating samples, resulting in low segmentation accuracy of the model in medical image segmentation.
Using a method based on progressive regional exchange and multi-dimensional collaborative pseudo-label, a hybrid image and label with different semantic complexity is generated through two-way copy and paste operations, a multi-level difficulty structure and multi-dimensional supervision are constructed, and a tagged and unlabeled data are fused to improve the model's adaptability to different anatomical features.
With only 10% of the annotated data, the accuracy of dice indexes of up to 90.60% is achieved, which significantly improves the accuracy and robustness of medical image segmentation, and solves the problems of data distribution gap and insufficient pseudo-label quality.
Smart Images

Figure CN120259667A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of medical artificial intelligence and computer vision, and particularly relates to an optimization method for medical image segmentation based on progressive region exchange and multi - dimensional collaborative pseudo - labels. Background Art
[0002] Medical image segmentation is one of the key technologies in computer - aided diagnosis systems, and is widely used in clinical tasks such as tumor detection, organ structure analysis, and lesion area localization. By segmenting medical images such as Magnetic Resonance Imaging (MRI), doctors can more intuitively observe the lesion area, improve the early diagnosis rate of diseases, and assist in formulating precise treatment plans.
[0003] The research on medical image segmentation is gradually shifting from traditional fully - supervised methods to semi - supervised methods. Fully - supervised methods, such as U - Net[1], V - Net[2], UNet++[3], 3D U - Net[4], and Attention U - Net[5], etc., although performing excellently in tumor detection and organ segmentation, highly rely on high - quality data labeled by professional physicians, resulting in high clinical application costs. Therefore, semi - supervised learning methods, by integrating a small amount of labeled data and a large amount of unlabeled data, have become the current research hotspot.
[0004] Mainstream semi - supervised methods focus on solving the problem of data utilization efficiency: 1) Consistency regularization methods (MixMatch[6], FixMatch[7]) improve robustness by forcing the model to have consistent predictions under different data augmentations; 2) Data augmentation strategies (Copy - Paste series[8,9]) alleviate the distribution difference between labeled and unlabeled data through local image recombination; 3) Multi - task collaboration methods (Dual - task Consistency
[10] ) utilize cross - task constraints to enhance feature expression ability. However, these methods still face key challenges: the quality of pseudo - labels is significantly affected by the distribution difference between labeled / unlabeled data, and existing augmentation strategies mostly adopt a single difficulty mode and fail to simulate the human progressive learning process.
[0005] To effectively address this issue, some studies have proposed the "curriculum learning" strategy
[11] , which aims to enhance the model's learning and generalization abilities by starting with simple samples for training and gradually introducing more complex ones, mimicking the human learning process from easy to difficult. The core idea of the curriculum learning strategy is to design the training order based on the difficulty of the samples. Ionescu et al. proposed that the difficulty of an image can be defined by evaluating the human time required for a visual search task
[12] . According to the difficulty level of the samples, the training set can be dynamically adjusted, enabling the model to gradually master tasks from simple to complex. This step-by-step learning approach can effectively avoid the training difficulties caused by directly facing complex samples in the initial stage of training and can also improve the model's recognition ability for samples with higher difficulty to a certain extent. However, in the field of semi-supervised learning of medical images, it is clearly infeasible to evaluate the difficulty of all data using this method of human time.
[0006] In summary, although existing semi-supervised learning methods have made some progress in improving segmentation accuracy and utilizing unlabeled data, they still face problems such as data distribution mismatch, unstable pseudo-label quality, and difficulty in evaluating sample difficulty, namely, they mainly face three major challenges:
[0007] 1) There is an empirical distribution gap between labeled data and unlabeled data
[0008] Generally speaking, in semi-supervised medical image segmentation, the labeled data and unlabeled data come from the same distribution. However, in real-world scenarios, it is difficult to estimate the exact distribution from the labeled data because their quantity is small. Therefore, there is always an empirical distribution mismatch between a large amount of unlabeled data and a very small amount of labeled data.
[0009] 2) The dimension of data augmentation difficulty is single
[0010] Existing methods adopt a mixed strategy with a fixed difficulty, resulting in the model being exposed to samples of similar difficulty for a long time, with a single type of augmented data, which hinders the progressive learning of semantic information.
[0011] 3) The quality of traditional pseudo-label supervision is insufficient
[0012] Direct combined pseudo-labels are vulnerable to noise interference and lack a multi-dimensional supervision mechanism.
[0013]
References
[0014] [1]O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th
[0015] [2]International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pp. 234–241, Springer, 2015. F. Milletari, N. Navab, and S.-A. Ahmadi, “V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation,” in 2016 Fourth International Conference on 3D Vision (3DV), pp. 565–571, Ieee, 2016.
[0016] [3]Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, “UNet++: Redesigning Skip Connections to Exploit Multiscale Features in Image Segmentation,” IEEE Transactions on Medical Imaging, vol. 39, no. 6, pp. 1856–1867, 2019.
[0017] [4] A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger, “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2016: 19th International Conference, Athens, Greece, October 17 - 21, 2016, Proceedings, Part II 19, pp. 424–432, Springer, 2016.
[0018] [5]O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, et al., “Attention U-Net: Learning Where to Look for the Pancreas,” arXiv preprint arXiv:1804.03999, 2018.
[0019] [6]D. Berthelot, N. Carlini, I. Goodfellow, N. Papernot, A. Oliver, and C. A. Raffel, “MixMatch: A Holistic Approach to Semi-Supervised Learning,” Advances in Neural Information Processing Systems, vol. 32, 2019.
[0020] [7]K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. A. Raffel, E. D. Cubuk, A. Kurakin, and C.-L. Li, “FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence,” Advances in Neural Information Processing Systems, vol. 33, pp. 596–608, 2020.
[0021] [8]G. Ghiasi, Y. Cui, A. Srinivas, R. Qian, T.-Y. Lin, E. D. Cubuk, Q. V. Le, and B. Zoph, “Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 2918–2928, 2021.
[0022] [9]Y. Bai, D. Chen, Q. Li, W. Shen, and Y. Wang, “Bidirectional Copy-Paste for Semi-Supervised Medical Image Segmentation,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 11514–11524, 2023.
[0023]
[10] X. Luo, J. Chen, T. Song, and G. Wang, “Semi-supervised Medical Image Segmentation through Dual-task Consistency,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 8801–8809, 2021.
[0024]
[11] Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum Learning,” in Proceedings of the 26th Annual International Conference on Machine Learning, pp. 41–48, 2009.
[0025]
[12] R. T. Ionescu, B. Alexe, M. Leordeanu, M. Popescu, D. P. Papadopoulos, and V. Ferrari, “How Hard Can It Be?Estimating the Difficulty of Visual Search in an Image,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2157–2166, 2016. Summary of the Invention
[0026] The object of the present invention is to provide an optimized method for medical image segmentation based on progressive region exchange and multi-dimensional collaborative pseudo-labels. This method generates mixed images and labels with different semantic complexities through bidirectional copy and paste operations combined with size conversion. This method can not only promote the fusion of labeled and unlabeled data domains, but also improve the adaptability of the model to different anatomical features by constructing a multi-level difficulty structure and multi-dimensional supervision.
[0027] To achieve the above object, the technical solution of the present invention is: an optimized method for medical image segmentation based on progressive region exchange and multi-dimensional collaborative pseudo-labels, which generates mixed images and labels with different semantic complexities through bidirectional copy and paste operations combined with size conversion, and can not only promote the fusion of labeled data domain and unlabeled data domain, but also construct a multi-level difficulty structure and multi-dimensional supervision.
[0028] Furthermore, the method includes the following steps:
[0029] Step S1, construct a progressive composite mask;
[0030] Step S2, perform bidirectional region exchange and fusion on labeled data and unlabeled data using the mask constructed in Step S1;
[0031] Step S3, perform multi-dimensional collaborative supervision on the model using direct supervision labels and sharpened soft pseudo-labels.
[0032] Furthermore, step S1 is specifically implemented as follows:
[0033] S11. Generate an all-zero initial mask M R ∈{0} W×H×L , where W, H, and L respectively represent the width, height, and depth of the input image;
[0034] S12. Randomly select a region with a size satisfying β1W × β1H × β1L and fill it with a mask of value 1 where β1 represents the parameter of the first layer;
[0035] S13. Randomly select a region with a size satisfying β2(n)W × β2(n)H × β2(n)L in the mask M1 and cover it with a 0 mask ; where the parameter of the second layer β2(n) = β1 × λ × n, and λ represents the progressive parameter λ = 1 / z, z ∈ {z ∈ N + |z ≥ 2}, N + is the set of positive integers, n is an integer and n ∈ [0, z);
[0036] S14. Obtain the composite mask
[0037]
[0038] where ⊙ represents pixel-wise multiplication, represents pixel-wise addition.
[0039] Furthermore, in step S2, a teacher-student network framework is adopted. The teacher network is denoted as The student network is denoted as where, respectively represent the p-th and q-th unlabeled data. When the labeled data X l is a foreground image and the unlabeled data X u is a background image, the mixed image is X for . When the labeled data X l is a background image and the unlabeled data X u is a foreground image, the mixed image is X bac . The parameters Θ t of the teacher network and the parameters Θ s of the student network will be dynamically updated during the entire training process.
[0040] Furthermore, the parameters Θ t of the teacher network are updated by performing exponential moving average EMA on the parameters Θ s of the student network.
[0041] Furthermore, step S2 is specifically implemented as follows:
[0042] S21. For the teacher network
[0043] The formula for generating the training data used is as follows:
[0044]
[0045] where and respectively represent the i-th and j-th labeled data in the labeled data set, and i≠j, N is the total number of labeled data in the labeled data set ; is the operation of setting the 0 region of the mask to 1 and the 1 region to 0; the enhanced data input of the teacher network is one-way region exchange and consists only of labeled data;
[0046] S22. For the student network
[0047] The training enhanced data used is as follows:
[0048]
[0049] where and respectively represent the p-th and q-th unlabeled data in the unlabeled data set, and p≠q, M is the total number of labeled data in the unlabeled data set, and the student network is generated by mixing labeled data and unlabeled data through two-way region exchange. ;
[0050] Furthermore, step S3 includes:
[0051] S31. Multi-dimensional collaborative supervision
[0052] S311. Direct supervision label
[0053] The supervision signal in the training process of the teacher model consists entirely of real labels:
[0054]
[0055] The used is the same mask used when generating the corresponding group of data, represents the corresponding label;
[0056] is to generate a corresponding supervision signal for the student network, and the unlabeled data will enter the teacher network to generate corresponding pseudo-labels:
[0057]
[0058] and respectively represent and the results after the teacher network prediction, combined with the true labels through the same 0-1 mask used when generating the corresponding group of data combined into the direct supervision label Q for (n) and Q bac (n):
[0059]
[0060] S312, sharpened soft pseudo-labels
[0061] For the prediction results P for (n) and P bac (n) of the student network, use a temperature-controlled sharpening function to convert the prediction results into soft pseudo-labels, and the function equation is as follows:
[0062]
[0063] where T is a hyperparameter used to control the sharpening temperature, X represents the input 3D medical image, is expressed as the prediction result after the image X passes through the student network, is the result after the student network prediction result is calculated by the temperature-controlled sharpening function, and the prediction result P for (n) and P bac (n) are solved for the corresponding and Both of them will be used as soft pseudo-label supervision signals to jointly supervise the student network with the direct signal.
[0064] Furthermore, step S3 also includes:
[0065] S32, loss function calculation
[0066] S321, direct supervision label loss function
[0067] Adopt the linear combination of Dice loss and Cross-entropy loss Its expression is as follows:
[0068]
[0069] The loss calculation formula for the foreground and background combined image is as follows:
[0070]
[0071] Where α is a hyperparameter used to adjust the calculation ratio of the loss between the true label and the pseudo-label;
[0072] S322. Sharpened soft pseudo-label loss function
[0073] For the sharpened soft pseudo-label, the following function is used to calculate the loss of the prediction result
[0074]
[0075] Where N0 represents the total number of pixels in the image, is the linear sum of the foreground and background losses:
[0076]
[0077] S323. Total loss function expression
[0078] Total loss The specific expression is as follows:
[0079]
[0080] Where is the weight adjustment factor between the direct simple label and the sharpened soft pseudo-label.
[0081] The present invention also provides a medical image segmentation optimization system based on progressive region exchange and multi-dimensional collaborative pseudo-labels, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps described above can be implemented.
[0082] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method steps described above can be implemented.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] (1) Construct multi-difficulty-level training data: Generate mixed images with different semantic complexities through multi-level data augmentation, construct a learning scenario from easy to difficult, enhance the learning ability and robustness of the model, and solve the problem of single difficulty in data augmentation;
[0085] (2) Further reduce the data distribution gap: Well integrate the labeled and unlabeled data through the progressive region exchange method, and alleviate the problem of data distribution mismatch in semi-supervised learning;
[0086] (3) High segmentation accuracy: In the case of only 10% labeled data in the ACDC dataset, the dice index accuracy is as high as 90.60%; in the case of 10% labeled data in the LA dataset, the dice index accuracy is as high as 90.53%. The test results show that this method is significantly superior to existing mainstream models in terms of segmentation accuracy;
[0087] (4) Modules are flexible and plug-and-play: The proposed method is a highly flexible and easy-to-plug-and-play solution that can be seamlessly integrated into existing models to effectively improve model performance.
[0088] The present invention is mainly used in the field of 2D and 3D medical image segmentation to help doctors analyze medical images more accurately. For example, in the segmentation of cardiac MRI images, it can more accurately locate and analyze the cardiac structure. Brief Description of the Drawings
[0089] Figure 1 is the image segmentation process in the embodiment of the present invention;
[0090] Figure 2 is the schematic diagram of generating masks under different n values;
[0091] Figure 3 is the way of two-way regional exchange and fusion between labeled data and unlabeled data;
[0092] Figure 4 is the experimental result of MRI cardiac image segmentation in the embodiment of the present invention; among them, (a) is the original image, (b) is the true label of the segmentation target, and (c) is the result image segmented by the present invention. Detailed Embodiment
[0093] The technical solution of the present invention will be specifically described below with reference to the drawings.
[0094] The present invention provides an optimization method for medical image segmentation based on progressive regional exchange and multi-dimensional collaborative pseudo-labels, which is used to optimize the training and supervision of the model. This method generates mixed images and labels with different semantic complexities through two-way copy and paste operations combined with size conversion. This method can not only promote the fusion of labeled and unlabeled data domains, but also improve the model's adaptability to different anatomical features by constructing a multi-level difficulty structure and multi-dimensional supervision.
[0095] The following is the specific implementation process of the present invention.
[0096] 1. Symbol Explanation:
[0097] Define the 3D medical image as where W, H, and L respectively represent the width, height, and depth of the input image. The goal of model training is to generate pixel-level prediction samples where K is the number of categories.
[0098] Training set contains N labeled data and M unlabeled data, where is denoted as the labeled dataset, denotes the unlabeled dataset. The labeled data is far less than the unlabeled data (N << M). represents the i-th labeled data, represents its corresponding label. Similarly, represents the unlabeled data.
[0099] For the mixed image, it is defined that when the labeled data X l is the foreground image and the unlabeled data X u is the background image, the mixed image is X for ; conversely, when the labeled data X l is the background image and the unlabeled data X u is the foreground image, the mixed image is X bac .
[0100] 2. Overall framework
[0101] As Figure 1 shown, the method of the present invention mainly consists of a progressive region exchange module and a sharpening operation in the multi-dimensional collaborative supervision module. The present invention adopts a teacher-student network framework as the basis of our method, where the teacher network is denoted as and the student network is denoted as The parameters Θ t of the teacher network and the parameters Θ s of the student network will be dynamically updated during the entire training process. Specifically, the parameters Θ t of the teacher network are updated by the exponential moving average (EMA) of the parameters Θ s of the student network to ensure the stability and consistency of the learning process.
[0102] Since the student network needs to rely on pseudo-labels for supervision, when formally training the student network, first, the teacher network is pre-trained using the labeled data. Secondly, in the self-training stage, the labeled data and the unlabeled data are processed together through a multi-level step-by-step exchange module to generate combined data X for , X bac , and then input into the student network. The teacher network generates corresponding pseudo-labels for the unlabeled data, which are processed together with the true labels through the progressive region exchange to generate combined labels. These combined labels and the prediction sharpening results of the model together constitute the multi-dimensional collaborative supervision signal.
[0103] The method of the present invention specifically includes the following steps and features:
[0104] Step 1) Construct a progressive composite mask;
[0105] Step 2) Use the mask to perform two-way regional exchange and fusion on labeled and unlabeled data;
[0106] Step 3) Adopt direct and sharpened labels to perform multi-dimensional collaborative supervision on the model.
[0107] In a preferred embodiment, in Step 1), the method for constructing the progressive composite mask is as follows:
[0108] 1.1) Generate an initial mask M of all zeros R ∈{0} W×H×L
[0109] 1.2) Randomly select a region with a size satisfying β1W×β1H×β1L and fill it with a mask with a value of 1
[0110] 1.3) Randomly select a region in the mask M1 with a size satisfying β2(n)W×β2(n)H×β2(n)L and cover it with a 0 mask (where the second-level parameter β2(n) = β1×λ×n, λ = 1 / z, z∈{z∈N + |z≥2}, n is an integer and n∈[0,z))
[0111] 1.4) Finally, obtain the composite mask:
[0112] The overall process can be expressed by the following formula:
[0113]
[0114] where ⊙ represents pixel-level multiplication, represents pixel-level addition. Figure 2 shows the schematic diagram of the mask styles for different n values, where the white area represents the mask filled with a value of 0, and the shaded part is the mask filled with a value of 1
[0115] In a preferred embodiment, in Step 2), the method for performing two-way regional exchange and fusion on labeled and unlabeled data using the mask is as follows:
[0116] 2.1) For the teacher network
[0117] The formula for generating the training data used is as follows:
[0118]
[0119] where i and j respectively represent two different labeled input images, and i ≠ j, n is a fixed integer, To set the 0 regions of the mask to 1 and the 1 regions to 0. The enhanced data input of the teacher network is one-way region exchange and consists only of labeled data.
[0120] 2.2) For the student network
[0121] The training enhanced data used is as follows:
[0122]
[0123] Where i ≠ j, p and q represent two different unlabeled input images respectively, and p ≠ q. The student network is generated by mixing labeled and unlabeled data through two-way region exchange, where n is an integer and n ∈ [0, z).
[0124] Figure 3 is the way of two-way region exchange and fusion of labeled and unlabeled data.
[0125] In a preferred embodiment, in step 3), the multi-dimensional collaborative supervision signal and the loss calculation method are as follows:
[0126] 3.1) The multi-dimensional collaborative supervision signal consists of two parts: direct supervision labels and sharpened soft pseudo-labels
[0127] 3.1.1) Direct labels
[0128] The supervision signal during the training process of the teacher model is completely composed of real labels:
[0129]
[0130] The same mask used when generating this set of data.
[0131] To generate corresponding supervision signals for the student network, unlabeled data will enter the teacher network to generate corresponding pseudo-labels:
[0132]
[0133] And respectively represent And the results after prediction by the teacher network. Combining the real label through the same 0-1 mask used when generating this set of data to form the direct label Q for(n) and Q bac (n)
[0134]
[0135]
[0136] 3.1.2) Sharpen the soft pseudo - labels
[0137] For the prediction result P of the student network for (n) and P bac (n), use a temperature - controlled sharpening function to convert the prediction result into soft pseudo - labels. The function equation is as follows:
[0138]
[0139] where T is a hyperparameter used to control the sharpening temperature. The prediction result P of the student network for (n) and P bac (n) are solved through the formula to obtain the corresponding and Both of them will be used as soft pseudo - label supervision signals to jointly supervise the student network with the direct signal.
[0140] 3.2) Loss function calculation
[0141] When training the teacher network, use the direct label loss. When training the student network, use the direct label and the sharpened soft pseudo - labels for multi - dimensional collaborative supervision.
[0142] 3.2.1) Direct label loss function
[0143] The present invention adopts a linear combination of Dice loss and Cross - entropy loss Its expression is as follows:
[0144]
[0145] The loss calculation formula for the foreground and background combined image is as follows:
[0146]
[0147] where α is a hyperparameter used to adjust the calculation ratio of the loss between the true label and the pseudo - label.
[0148] 3.2.2) Sharpened soft pseudo - label loss function
[0149] For the soft pseudo - labels, use the following function to calculate the loss of the prediction result
[0150]
[0151] Where N0 represents the total number of pixels in the image. is the linear sum of the foreground and background losses:
[0152]
[0153] 3.2.3) Total loss function expression
[0154] Total loss The specific expression is as follows:
[0155]
[0156] Where is the weight adjustment factor between the direct label and the soft pseudo-label to better balance the learning direction of the model.
[0157] 3. Examples
[0158] In this example, pixel-level segmentation of the left ventricle, right ventricle, and myocardium will be performed on the diastolic and systolic frames in the cine-MRI (cardiac magnetic resonance imaging) of the ACDC (Automatic Cardiac Diagnosis Challenge) in the MICCAI 2017 Challenge. The specific steps are as follows:
[0159] 1) Construct a progressive composite mask:
[0160] 1.1) Generate an initial mask M of all zeros R ∈{0} W×H×L
[0161] 1.2) Randomly select a region with a size that meets β1W×β1H×β1L and fill it with a mask of value 1
[0162] 1.3) Randomly select a region in the mask M1 with a size that meets β2(n)W×β2(n)H×β2(n)L and cover it with a 0 mask (where the parameter β2(n) of the second layer = β1×λ×n, λ = 1 / z, z∈{z∈N + |z≥2}, n is an integer and n∈[0,z))
[0163] 1.4) Finally, obtain the composite mask:
[0164] The overall process can be expressed by the following formula:
[0165]
[0166] Where ⊙ represents pixel-level multiplication, Represents pixel - level addition. In this example, W = H = 256 and β1 = 2 / 3.
[0167] 2) Use a mask to perform two - way regional exchange and fusion on labeled and unlabeled data:
[0168] For the teacher network, the training data generation formula is as follows:
[0169]
[0170] where i and j respectively represent two different labeled input images, and i ≠ j, n is a fixed integer, is the operation of setting the 0 region of the mask to 1 and the 1 region to 0. The enhanced data input of the teacher network is one - way regional exchange and consists only of labeled data. In this example, z = 3, n = 2
[0171] For the student network, the training data generation formula is as follows:
[0172]
[0173] where i ≠ j, p and q respectively represent two different unlabeled input images, and p ≠ q. The student network is generated by mixing labeled and unlabeled data through two - way regional exchange, where n is an integer and n ∈ [0, z).
[0174] In this example, z = 3.
[0175] 3) Adopt direct and sharpened labels to perform multi - dimensional collaborative supervision on the model:
[0176] 3.1) Supervision signal
[0177] 3.1.1) Direct label
[0178] The supervision signal in the training process of the teacher model consists entirely of real labels:
[0179]
[0180] The is the same mask used when generating this set of data.
[0181] To generate the corresponding supervision signal for the student network, unlabeled data will enter the teacher network to generate corresponding pseudo - labels:
[0182]
[0183] and respectively represent and the results after being predicted by the teacher network. Combining the true labels with the same 0-1 mask used when generating this set of data to combine into the direct label Q for (n) and Q bac (n)
[0184]
[0185] 3.1.2) Sharpened soft pseudo-labels
[0186] For the prediction results P for (n) and P bac (n) of the student network, use the temperature-controlled sharpening function to convert the prediction results into soft pseudo-labels, and the function equation is as follows:
[0187]
[0188] where T is a hyperparameter used to control the sharpening temperature. The prediction results P for (n) and P bac (n) of the student network are solved through the formula to obtain the corresponding and Both of them will be used as soft pseudo-label supervision signals to jointly supervise the student network with the direct signal. In this example, T = 0.1.
[0189] 3.2) Loss calculation
[0190] When training the teacher network, use the direct label loss. When training the student network, use the direct label and the sharpened soft pseudo-labels for multi-dimensional collaborative supervision.
[0191] 3.2.1) Direct label loss function
[0192] The present invention adopts a linear combination of Dice loss and Cross-entropy loss:
[0193]
[0194] The loss calculation formula for the foreground and background combined image is as follows:
[0195]
[0196] where α is a hyperparameter used to adjust the calculation ratio of the loss between the true label and the pseudo-label. In this example, α = 0.5.
[0197] 3.2.2) Sharpened soft pseudo-label loss function
[0198] For soft pseudo-labels, the following function is used to calculate the loss of the prediction results
[0199]
[0200] where N0 represents the total number of pixels in the image. is the linear sum of the foreground and background losses:
[0201]
[0202] 3.2.3) Expression of the total loss function
[0203] Total loss The specific expression is as follows:
[0204]
[0205] where is the weight adjustment factor between the direct label and the soft pseudo-label to better balance the learning direction of the model. In this example
[0206] Figure 1 is the image segmentation process in the embodiment. Figure 4 is the comparison chart of the segmentation results of the embodiment. From Figure 4 it can be seen that the present invention can accurately complete the segmentation of the target image.
[0207] The present invention also provides a medical image segmentation optimization system based on progressive region exchange and multi-dimensional collaborative pseudo-labels, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the method steps described in any of the above can be implemented.
[0208] The present invention also provides a computer-readable storage medium, on which computer program instructions executable by the processor are stored. When the processor runs the computer program instructions, the method steps described in any of the above can be implemented.
[0209] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention and whose functional effects do not exceed the scope of the technical solutions of the present invention belong to the protection scope of the present invention.
Claims
1. An optimization method for medical image segmentation based on progressive region exchange and multi-dimensional collaborative pseudo-labels, characterized in that, By means of bidirectional copy and paste operations combined with size conversion, hybrid images and labels with different semantic complexities are generated, which can not only promote the fusion of the labeled data domain and the unlabeled data domain, but also construct a multi-level difficulty structure and multi-dimensional supervision.
2. The medical image segmentation optimization method based on progressive regional exchange and multi-dimensional collaborative pseudo-labels according to claim 1, wherein The method includes the following steps: Step S1, construct a progressive composite mask; Step S2, perform bidirectional regional exchange and fusion on the labeled data and the unlabeled data using the mask constructed in Step S1; Step S3, perform multi-dimensional collaborative supervision on the model using direct supervision labels and sharpened soft pseudo-labels.
3. The medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-labels according to claim 2, wherein, The specific implementation of Step S1 is as follows: S11. Generate an all-zero initial mask M R ∈ {0} W×H×L , where W, H, and L respectively represent the width, height, and depth of the input image; S12. Randomly select a region with dimensions β1W × β1H × β1L and fill it with a mask with a value of 1 where β1 represents a parameter of level one; S13. Randomly select a region in the mask M1 with a size satisfying β2(n)W × β2(n)H × β2(n)L and use a 0 mask for covering; where the level-two parameter β2(n) = β1 × λ × n, and λ is the asymptotic parameter λ = 1 / z, z ∈ {z ∈ N | z ≥ 2}, N + is the set of positive integers, n is an integer and n ∈ [0, z); + S14. Obtain a composite mask Among them, ⊙ represents pixel-level multiplication, represents pixel-level addition.
4. The medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-labels according to claim 3, characterized in that In step S2, a teacher-student network framework is adopted. The teacher network is denoted as and the student network is denoted as where represent the p-th and q-th unlabeled data respectively. When the labeled data X l is a foreground image and the unlabeled data X u is a background image, the mixed image is X for . When the labeled data X l is a background image and the unlabeled data X u is a foreground image, the mixed image is X bac . The parameters Θ t of the teacher network and the parameters Θ s of the student network will be dynamically updated during the entire training process.
5. The medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-labels according to claim 4, characterized in that, The parameters Θ of the teacher network t By performing an exponential moving average EMA update on the parameters Θ of the student network s for updating.
6. The medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-labels according to claim 4, characterized in that The specific implementation of Step S2 is as follows: S21. For the teacher network The formula for generating the training data used is as follows: Among them, and respectively represent the i-th and j-th labeled data in the labeled data set, where i ≠ j, and N is the total number of labeled data in the labeled data set ; is the operation of setting the 0 region of the mask to 1 and the 1 region to 0; the enhanced data input of the teacher network is one-way region exchange and consists only of labeled data; S22. For the student network The training enhanced data used is as follows: Among them, and respectively represent the p-th and q-th unlabeled data in the unlabeled dataset , where p≠q, and M is the total number of labeled data in the unlabeled dataset . The student network is generated after two-way regional exchange and mixing of labeled data and unlabeled data.
7. The medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-labels according to claim 6, wherein, Step S3 includes: S31. Multi-dimensional collaborative supervision S311. Direct supervision labels The supervision signal during the training process of the teacher model consists entirely of real labels: The one used is the same mask used when generating the corresponding group of data, which represents the corresponding label; To generate corresponding supervision signals for the student network, unlabeled data will enter the teacher network to generate corresponding pseudo-labels: and respectively represent and the results after prediction by the teacher network, combined with the true labels through the same 0-1 mask used when generating the corresponding group of data combined into the direct supervision label Q for (n) and Q bac (n): S312. Sharpened soft pseudo-labels For the prediction result P of the student network for (n) and P bac (n), use a temperature-controlled sharpening function to convert the prediction result into soft pseudo-labels. The function equation is as follows: where T is a hyperparameter used to control the sharpening temperature, and X represents the input 3D medical image. is the prediction result after the image X passes through the student network. is the result after the prediction result of the student network is calculated by the temperature-controlled sharpening function. The prediction result P for (n) and P bac (n) are solved for the corresponding and Both of them will be used as soft pseudo-label supervision signals to jointly supervise the student network with the direct signal.
8. The medical image segmentation optimization method based on progressive region exchange and multi-dimensional collaborative pseudo-labels according to claim 7, wherein Step S3 also includes: S32. Loss function calculation S321. Direct supervision label loss function Using Dice loss And a linear combination of Cross-entropy loss Its expression is as follows: The loss calculation formula for the foreground and background combined images is as follows: where α is a hyperparameter used to adjust the calculation ratio of the loss between the real label and the pseudo-label; S322. Sharpened soft pseudo-label loss function For the sharpened soft pseudo-label, the following function is used to calculate the loss of the prediction result where N0 represents the total number of pixels in the image, is the linear sum of the foreground and background losses: S323. Total loss function expression Total loss The specific expression is as follows: Among them is the weight adjustment factor for the direct simple label and the sharpened soft pseudo-label.
9. An optimized system for medical image segmentation based on progressive region exchange and multi-dimensional collaborative pseudo-labels, characterized in that, It includes a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps described in any one of claims 1-8 can be implemented.
10. A computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method steps described in any one of claims 1-8 can be implemented.