Semi-supervised medical image segmentation method and system
By distinguishing between learning reliable and unreliable pseudo labels in semi-supervised medical image segmentation and leveraging the encoder to learn decoder perspective differences, we address the problem that existing methods fail to fully utilize pseudo label information and achieve higher segmentation accuracy.
Patent Information
- Application Number
- CN202510739621.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Existing semi-supervised medical image segmentation methods fail to fully exploit the full potential of pseudo-labels, especially ignoring the additional information in low-quality pseudo-labels, resulting in insufficient segmentation accuracy.
By adopting the pseudo-label differentiation and full utilization (PLD-FU) method, reliable and unreliable pseudo-labels are distinguished and learned, and the difference information between the decoder perspectives is learned by the encoder to enhance the feature extraction capability.
The information of all pseudo labels is effectively utilized, which improves the accuracy of medical image segmentation and the training effect of the model, especially significantly improving the segmentation performance by utilizing low-quality pseudo labels.
Smart Images

Figure CN120672768A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical image segmentation, and in particular to a semi-supervised medical image segmentation method and system. Background Art
[0002] In recent years, semi-supervised medical image segmentation has become a research hotspot, significantly reducing its reliance on densely labeled data. The combination of consistency regularization and pseudo-labeling techniques has achieved remarkable results. However, traditional methods often focus solely on high-confidence pseudo-labels, ignoring the potential additional information contained in the pseudo-labels and thus failing to fully utilize their potential. Therefore, overcoming these technical issues and drawbacks has become a key challenge. Summary of the Invention
[0003] In order to overcome the above problems existing in the prior art, this application provides a semi-supervised medical image segmentation method and system, which adopts the following technical solutions:
[0004] In a first aspect, the present application provides a semi-supervised medical image segmentation method, comprising:
[0005] Obtain unlabeled medical images and labeled medical images; and train a semi-supervised medical image segmentation model based on the unlabeled medical images and the labeled medical images.
[0006] The medical image to be segmented is used as input to the trained semi-supervised medical image segmentation model to obtain a segmentation result of the medical image to be segmented.
[0007] Furthermore, a semi-supervised medical image segmentation model is trained based on unlabeled medical images and labeled medical images, including:
[0008] The encoder extracts features of the unlabeled medical image to obtain features of the unlabeled medical image.
[0009] Calculate the prediction confidence values of the two decoders for the feature maps, and obtain reliable pseudo-label masks and unreliable pseudo-label masks based on the prediction confidence values.
[0010] Based on the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder, the cross entropy loss of reliable labels and unreliable labels is obtained respectively.
[0011] Based on the reliable pseudo-label mask, the unreliable pseudo-label mask and the corresponding cross-entropy loss, the unsupervised loss is obtained. The supervised loss is obtained based on Diceloss.
[0012] The predicted labels of the labeled medical images are obtained in the last convolutional layer of the decoder, and the weights of the supervised loss are updated by comparing the predicted labels with the true labels.
[0013] The overall loss function of the semi-supervised medical image segmentation model is obtained by weighted summation to complete the training of the semi-supervised medical image segmentation model.
[0014] Furthermore, based on the encoder, the features of the unlabeled medical images are extracted to obtain the features of the unlabeled medical images, including: the encoder is composed of a Mamba module and a CNN in parallel, the Mamba module receives the feature map output by the CNN layer as input, and the feature map enhanced by the Mamba module is further passed to other CNN layers to extract the features of the unlabeled medical images.
[0015] Furthermore, the decoder-based acquisition of a reliable pseudo-label mask and an unreliable pseudo-label mask for a feature map includes: using two decoders to respectively acquire a probability value that the feature map is a certain category, acquiring prediction confidence values of the feature map by two encoders based on the probability values, comparing the prediction confidence values based on a comparison network, and respectively acquiring a reliable pseudo-label mask and an unreliable pseudo-label mask.
[0016] Furthermore, the cross entropy loss of obtaining a reliable label based on the features of the penultimate layer of the decoder and a prototype of a certain category of the corresponding decoder includes:
[0017] The features of the penultimate layer of the decoder and the cosine similarity of the prototype of a certain category of the corresponding decoder are obtained, and the features of the penultimate layer of the decoder and the cross entropy loss of the pixel level of the prototype of a certain category of the corresponding decoder are obtained based on the preset formula.
[0018] Furthermore, the cross entropy loss of unreliable labels is obtained based on the features of the penultimate layer of the decoder and a prototype of a certain category of the corresponding decoder, including:
[0019] Obtain the cosine similarity of the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder, obtain the cross-entropy loss of the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder based on a preset formula, and evaluate the difference information between the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder based on the cross-entropy loss.
[0020] Furthermore, the difference information is encoded into a feature vector and fused with the input data of the encoder to obtain a fused feature, which is used as the final input of the encoder.
[0021] In a second aspect, the present application also provides a semi-supervised medical image segmentation system, comprising:
[0022] The medical image segmentation model training module is used to obtain unlabeled medical images and labeled medical images and train a semi-supervised medical image segmentation model based on the unlabeled medical images and labeled medical images.
[0023] The medical image segmentation module to be segmented is used to take the medical image to be segmented as the input of the trained semi-supervised medical image segmentation model to obtain the segmentation result of the medical image to be segmented.
[0024] In a third aspect, the present application provides an electronic device, comprising:
[0025] One or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to perform the method as described in the first aspect.
[0026] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first aspect.
[0027] In a fifth aspect, the present application provides a computer program, which, when executed by a computer, is used to execute the method described in the first aspect.
[0028] In one possible design, the program in the fifth aspect may be stored in whole or in part on a storage medium packaged with the processor, or may be stored in whole or in part on a memory not packaged with the processor.
[0029] This application has the following beneficial effects:
[0030] 1. The present application obtains unlabeled medical images and labeled medical images; trains a semi-supervised medical image segmentation model based on the unlabeled medical images and labeled medical images; uses the medical image to be segmented as the input of the trained semi-supervised medical image segmentation model to obtain the segmentation result of the medical image to be segmented. The present application distinguishes pseudo-labels through the comparison network in the semi-supervised medical image segmentation model and uses a diversity strategy to learn pseudo-labels of different degrees; secondly, the encoder is added to learn the difference information of the same unlabeled data revealed by different decoder perspectives; finally, the different rules adopted by the classification layer in the decoder when learning pseudo-labels are used to fully utilize the various information of the pseudo-labels. The present application solves the problem of not being able to fully utilize unlabeled data and low quality of pseudo-labels in semi-supervised medical image segmentation by increasing the utilization of important information in low-quality pseudo-labels, thereby improving the segmentation accuracy of medical images.
[0031] 2. This application uses a mutual comparison network to distinguish pseudo labels and ensures that all pixels participate in the learning process. Not only pixels with high confidence, but even pixels with low confidence can be effectively learned. This strategy used in this application not only helps to improve the quality of pseudo labels, but also improves the model's utilization of pseudo labels.
[0032] 3. This application uses the encoder to supplement the decoder differences for learning and optimization. In this way, this application can more effectively train the encoder to extract more representative and discriminative features, thereby improving the effect and performance of semi-supervised learning.
[0033] 4. This application introduces an encoder composed of Mamba and CNN in parallel, and this structure performs well in semi-supervised learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is an exemplary system architecture diagram to which the embodiments of the present application can be applied;
[0035] Figure 2 This is an overall framework diagram of the semi-supervised medical image segmentation method according to an embodiment of the present application;
[0036] Figure 3 This is a flow chart of the semi-supervised medical image segmentation method according to an embodiment of the present application;
[0037] Figure 4 This is a flowchart of the training of a semi-supervised medical image segmentation model according to an embodiment of the present application;
[0038] Figure 5 Comparison visualization of the embodiment of the present application with the SOTA method on the left atrial dataset;
[0039] Figure 6 Comparison visualization of the embodiments of this application with the SOTA method on the Pancreas-NIH dataset;
[0040] Figure 7 This is a visualization of the segmentation effect of each component of the LA dataset under 10% annotation in the embodiment of this application;
[0041] Figure 8 This is a visualization of the segmentation effect of each component of the Pancreas-NIH dataset under 10% annotation in the embodiment of this application;
[0042] Figure 9 This is a system flow chart of an embodiment of the present application;
[0043] Figure 10 It is a schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0045] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0046] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0047] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0048] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0049] Terminal devices 101, 102, and 103 can be various electronic devices with display screens and support web browsing, including but not limited to smartphones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4) players, laptop computers, desktop computers, etc.
[0050] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .
[0051] It should be noted that the semi-supervised medical image segmentation method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the semi-supervised medical image segmentation system is generally set in the server / terminal device.
[0052] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0053] Continue to refer Figure 2 and Figure 3 , Figure 2 This is the overall framework diagram of the embodiment of this application. Figure 3 A flow chart of a semi-supervised medical image segmentation method of the present application is shown in FIG. , and the method comprises the following steps:
[0054] Step 201: Obtain an unlabeled medical image and a labeled medical image; and train a semi-supervised medical image segmentation model based on the unlabeled medical image and the labeled medical image.
[0055] In one possible implementation, a semi-supervised medical image segmentation model is trained based on unlabeled medical images and labeled medical images. Figure 4 , specifically including:
[0056] Step 41 : extracting features of the unlabeled medical image based on the encoder to obtain features of the unlabeled medical image.
[0057] In a possible implementation, the encoder extracts features of the unlabeled medical image to obtain features of the unlabeled medical image, such as Figure 2As shown in Figure 3, the encoder consists of a Mamba module and a CNN in parallel. The Mamba module receives the feature map output by the CNN layer as input. The feature map enhanced by the Mamba module is further passed to other CNN layers to extract features of unlabeled medical images. It should be noted that in this embodiment of the application, the Mamba module can enhance the CNN's ability to handle long-term dependencies.
[0058] Step 42: Calculate the prediction confidence values of the two decoders for the feature map, and obtain a reliable pseudo-label mask and an unreliable pseudo-label mask based on the prediction confidence values.
[0059] In a possible implementation, the two decoders include a decoder Da and a decoder Db, wherein the decoder Da uses a trilinear interpolation method to upsample the feature map, and the decoder Db uses a nearest neighbor interpolation method to upsample the feature map.
[0060] In one possible implementation, the present application calculates the prediction confidence values of the two decoders for the feature maps, and obtains a reliable pseudo-label mask and an unreliable pseudo-label mask based on the prediction confidence values, including:
[0061] First, the prediction confidence values of the two decoders for the unlabeled image are calculated. Then, by comparing these confidence values, high-confidence pseudo-labels and low-confidence pseudo-labels are screened out. High-confidence pseudo-labels generally indicate that the model is more confident in its prediction of the pixel and may be closer to the true label. Therefore, these pseudo-labels are considered more reliable. Conversely, low-confidence pseudo-labels are considered less reliable. Specifically, the process of calculating the confidence value can be expressed as: in are the prediction confidence values of decoder Da and decoder Db for unlabeled medical images, respectively. are respectively the probability values of decoder Da and decoder Db predicting that the class is C. ε is a small positive number close to zero but not zero, and in the embodiment of the present application, the value is 1e-6.
[0062] In one possible implementation, the decoder-based acquisition of a reliable pseudo-label mask and an unreliable pseudo-label mask for a feature map includes: using two decoders to respectively acquire a probability value that the feature map is a certain category, acquiring prediction confidence values of the feature map by two encoders based on the probability values, comparing the prediction confidence values based on a comparison network, and respectively acquiring a reliable pseudo-label mask and an unreliable pseudo-label mask.
[0063] Based on the comparison network, the predicted confidence values are compared to obtain reliable pseudo label masks and unreliable pseudo label masks, which can be expressed as:
[0064]
[0065] in, are the masks of reliable pseudo labels for decoder Da and decoder Db, respectively, are the masks of unreliable pseudo labels for decoder Da and decoder Db, respectively.
[0066] In step 43 , based on the features of the penultimate layer of the decoder and a prototype of a certain category of the corresponding decoder, the cross entropy losses of the reliable labels and the unreliable labels are obtained respectively.
[0067] In one possible implementation, obtaining a cross-entropy loss for a reliable label based on the features of the penultimate layer of the decoder and a corresponding prototype of a certain category of the decoder includes: obtaining the cosine similarity of the features of the penultimate layer of the decoder and a corresponding prototype of a certain category of the decoder, and obtaining a pixel-level cross-entropy loss for the features of the penultimate layer of the decoder and a corresponding prototype of a certain category of the decoder based on a preset formula. The purpose of obtaining the cross-entropy loss for the reliability label in this application is to optimize the prediction accuracy of the model at the pixel level.
[0068] It should be noted that this application proposes a mutual learning method that can fully utilize all pixels, which is achieved by distinguishing between reliable and unreliable pseudo-labels. Specifically, for reliable pseudo-labels, the similarity of reliable pseudo-labels is first calculated, where the cosine similarity of the features of the penultimate layer of the decoder and the corresponding prototype of a certain category of the decoder is obtained, including: in is the feature from the penultimate layer of decoder Da, is the prototype of the c-th category from the decoder Da, is the cosine similarity between the two.
[0069] It should be noted that the features of the penultimate layer of the decoder and the cross entropy loss of the prototype pixel level of a certain category of the decoder are obtained based on the preset formula, including: in It is calculated and The pixel-level cross entropy loss is , C is the total number of categories, to optimize the prediction accuracy of the model at the pixel level. The reliability of the generated decoder Db is similar, specifically: in It is calculated and The pixel-level cross entropy loss is
[0070] In one possible embodiment, the cross entropy loss of unreliable labels based on the features of the penultimate layer of the decoder and a prototype of a corresponding decoder category is obtained, including: obtaining the cosine similarity of the features of the penultimate layer of the decoder and a prototype of a corresponding decoder category, obtaining the cross entropy loss of the features of the penultimate layer of the decoder and a prototype of a corresponding decoder category based on a preset formula, and evaluating the difference information between the features of the penultimate layer of the decoder and a prototype of a corresponding decoder category based on the cross entropy loss. The greater the difference, the lower the reliability of the pseudo-label. Specifically, it can be expressed as: in It is calculated and The pixel-level cross entropy loss is used to evaluate the difference between the two. The unreliability of the generated decoder Db is similar, specifically: in It is calculated and The pixel-level cross entropy loss is
[0071] In one possible implementation, the difference information is encoded as a feature vector and fused with the encoder input data to obtain a fused feature, which is then used as the final input to the encoder. By encoding this difference information as a feature vector, the feature maps of each layer are enhanced with the corresponding difference information, thereby improving their representation capabilities.
[0072] This application uses decoder difference optimization coding to integrate the difference information obtained in the decoder into the encoder training process, encode the difference information into a feature vector, and add the feature map of each layer of the encoder to the difference information of the corresponding layer to enhance the expression of the feature map for further learning of the encoder. This difference information is repeatedly used in the loop of each training sample. This process can be expressed as: in Represent the feature maps of each layer of decoder Da and decoder Db respectively, is the input feature map of each layer of the encoder, and β is a scaling factor used to adjust the scale of the difference. idx It is the final input that combines the features of the original data and the difference information. idx As an input encoder, it can more effectively capture and utilize subtle changes in the data during the learning process, improving the ability to recognize complex patterns and structures.
[0073] In step 44, an unsupervised loss is obtained based on the reliable pseudo-label mask, the unreliable pseudo-label mask, and the corresponding cross-entropy loss; and a supervised loss is obtained based on Diceloss.
[0074] In one possible implementation, based on the reliable pseudo-label mask, the unreliable pseudo-label mask, and the corresponding cross-entropy loss, the unsupervised loss is obtained. That is, the process of the decoder Db learning the reliable pseudo-label can be expressed as: in is the index of the most probable category in each sample, and α is a Gaussian warm-up function that varies over time. The weighted cross entropy loss function is calculated between the predictions generated by the decoder Db and the reliable pseudo-labels generated by the decoder Da. The process of learning reliable pseudo-labels by the decoder Da is similar, specifically: in A weighted cross-entropy loss function is calculated between the predictions generated by decoder Da and the reliable pseudo-labels generated by decoder Db.
[0075] The process of learning unreliable pseudo labels by the decoder Db can be expressed as:
[0076] in A weighted cross-entropy loss function is calculated between the predictions generated by the decoder Db and the unreliable pseudo labels generated by the decoder Da. It can effectively guide the decoder Db to learn, thereby maintaining high performance in the face of data noise and label inconsistency. It enables the decoder Db to extract useful information from unreliable pseudo-labels while reducing over-reliance on reliable labels. The process of decoder Da learning unreliable pseudo-labels is similar, specifically: By effectively handling unreliable pseudo labels, the stability and prediction accuracy of the model can be improved. A weighted cross-entropy loss function is calculated between the predictions generated by decoder Da and the unreliable pseudo labels generated by decoder Db.
[0077] In one possible implementation, the supervised loss is obtained based on Diceloss and can be expressed as: Where P is the predicted score, T is the corresponding target label, and θ is a smoothing term.
[0078] In step 45, different layers of the decoder follow different gradient update rules.
[0079] It should be noted that a new gradient update rule is introduced for different layers of the decoder. During the training process, the optimization of the last layer only relies on the supervised loss, that is, the weights are updated by comparing the predicted labels with the true labels, thereby improving the segmentation accuracy. Figure 2 As shown in (b), the last convolutional layer of the decoder focuses on fine-tuning the model’s output through labeled data to ensure the accuracy of prediction and the segmentation performance of the model.
[0080] Step 46 , obtaining the overall loss function of the semi-supervised medical image segmentation model through weighted summation, thereby completing the training of the semi-supervised medical image segmentation model.
[0081] It should be noted that the overall loss function of the semi-supervised medical image segmentation model can be expressed as: L total =γL sup +L unsup , where L sup is the supervision loss, L unsup is the unsupervised loss, and γ is the weight hyperparameter of the supervised loss, which is used to control the proportion of supervised loss in the loss function.
[0082] For unlabeled medical images, this application emphasizes the importance of full-pixel pseudo-labels in model optimization. The unsupervised loss can be defined as: in is the consistency loss calculated under reliable pseudo-label filtering, is the consistency loss calculated under unreliable pseudo-label filtering.
[0083] It should be noted that this application solves the problem of insufficient utilization of unlabeled data and low quality of pseudo labels in semi-supervised medical image segmentation by adopting Pseudo label differentiation and full utilization (PLD-FU), that is, adopting the method of pseudo label differentiation and full utilization, and increasing the use of important information in low-quality pseudo labels. The overall architecture of this application is as follows Figure 2 As shown in (a), this application mainly consists of two parts. “PLD” can effectively distinguish reliable pseudo labels from unreliable pseudo labels, and “FU” can adopt different learning strategies to make full use of all pseudo labels. The combination of the two can improve the segmentation accuracy and convergence speed of the model.
[0084] Step 202 : Using the medical image to be segmented as input to the trained semi-supervised medical image segmentation model to obtain a segmentation result of the medical image to be segmented.
[0085] The experimental part of this application is as follows:
[0086] 1. Dataset
[0087] Left Atrium (LA) Dataset
[49] : The left atrium segmentation dataset contains 100 3D enhanced MRI scan images. Each sample represents the detailed image data of the left atrium of a patient at a certain moment.
[0088] Pancreas-NIH
[50] : This dataset contains 82 contrast-enhanced 3D CT scans, where each slice of the pancreas is manually annotated. The slices have a fixed resolution of 512×512 and thicknesses ranging from 1.5 mm to 2.5 mm.
[0089] Brats-2019
[51] : This dataset includes preoperative MRI images of 335 patients with glioma, covering samples from multiple medical institutions. In this dataset, 259 patients were diagnosed with high-grade glioma (HGG) and 76 patients were diagnosed with low-grade glioma (LGG). Each patient's MRI images contain four different imaging modes: T1, T1Gd, T2, and T2-FLAIR.
[0090] 2. Comparison with the most advanced methods
[0091] Table 1 Comparison with SOTA methods on LA dataset
[0092]
[0093] LADataset: This application evaluates the performance of this application's method and various competitors on the LA dataset. The methods involved in the comparison include UA-MT[7], SASSNet
[38] , DTC
[37] , MC-Ne
[38] , URPC
[33] , SS-Net
[35] , BCP
[36] , and MIA
[16] . The experiment examines the performance of these methods by using different proportions of labeled data. As shown in Table 1, this application's method achieves the best performance on all four evaluation indicators when trained with 5% and 10% labeled images. Figure 5 As shown in Figure 1, this application compares the qualitative results obtained by different methods on the LA dataset. These results are obtained by training with only 10% of the labeled data. It can be clearly seen from the figure that the method of this application is closer to the true label in the segmentation results.
[0094] Table 2 Comparison with SOTA methods on Pancreas-NIH dataset
[0095]
[0096] Pancreas-NIH: To verify the versatility of our model on other datasets, we conducted comparative experiments on the Pancreas-NIH dataset. Our model was compared with the following methods: UA-MT[7], SASSNet
[38] , DTC
[37] , MC-Net
[34] , URPC
[33] , SS-Net
[35] , Co-BioNet
[53] , BCP
[36] , and MIA
[16] . As shown in Table 2, our model outperformed other methods in all four evaluation metrics. It is worth noting that in the pancreas segmentation task, the performance of the model declined compared to that of the left atrium segmentation. This is mainly due to the more complex shape and texture of the pancreas. However, compared with other competing models, our model has a relatively small performance degradation. This shows that our model still performs well in dealing with complex tasks and can effectively cope with more challenging segmentation problems.
[0097] like Figure 6 As shown, this application obtains visualization results on the Pancreas-NIH dataset, which has only 10% labeled data.
[0098] Table 3 Comparison with SOTA methods on the Brats-2019 dataset
[0099]
[0100] Brats-2019: The whole brain tumor segmentation task is more challenging than organ segmentation. In order to evaluate the performance of our model in brain tumor segmentation, we conducted further comparative experiments on the Brats-2019 dataset. We used DAN
[54] , CPS
[11] , EM
[55] , ICT
[56] , DTC
[37] , CCT
[12] , URPC
[33] , AC-MT
[57] , and MIA
[16] for comparison.
[0101] 3. Ablation Experiment
[0102] Table 4 Ablation study of each component, where "MB" represents the addition of Mamaba-block in the encoder part. "PF" represents the "PLD-FU" method. "EI" represents Encoders learn and optimize through decoderdifferences. "DG" represents Decoder gradient update rules.
[0103]
[0104] To evaluate the impact of each component of our approach, we conducted ablation experiments on the LA and Pancreas-NIH datasets. We examined the effects of the Mamba module, the PLD-FU strategy, the encoder learning strategy, and the decoder gradient update rule on the segmentation task. The experimental results show that all four components significantly improve the model's segmentation performance.
[0105] The method of this application performs well in fully utilizing pseudo-labels and handling the additional potential of pseudo-labels. Figure 7 The significant contribution of each component of the proposed method to the segmentation task is intuitively demonstrated on the LA dataset with 10% labeled data. Figure 8 The visualization of the segmentation effect on the Pancreas-NIH dataset with 10% labeled data is shown.
[0106] 4. Conclusion
[0107] In this study, the present application proposes that full and accurate utilization of pseudo-label information is the key to improving semi-supervised image segmentation. Inadequate utilization of pseudo-labels will hinder the ability to achieve optimal performance. To solve this problem, the present application analyzes how to distinguish and learn pseudo-labels of different degrees and accurately utilize the potential information of pseudo-labels. The present application proposes a new semi-supervised learning framework. Specifically, in order to learn all pseudo-labels, the present application proposes to classify reliable and unreliable pseudo-labels and learn these pseudo-labels separately. Then, in order to accurately utilize the potential information of pseudo-labels, the present application enables the encoder to learn the differences between the decoder perspectives. At the same time, the present application optimizes the gradient update rules of the Mamba encoder and decoder, thereby improving the training process of the model. Experiments on the LA, Pancreas-NIH and Brats-2019 datasets show that the method of the present application effectively solves these problems and exhibits optimal performance.
[0108] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0109] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0110] Continue to refer Figure 9 , as a response to the above Figure 2 The present application provides an embodiment of a semi-supervised medical image segmentation system. Figure 2 Corresponding to the method embodiment shown, the system can be specifically applied to various electronic devices, including: a medical image segmentation model training module 901, and a medical image segmentation module 902.
[0111] The medical image segmentation model training module 901 is used to obtain unlabeled medical images and labeled medical images; and train a semi-supervised medical image segmentation model based on the unlabeled medical images and labeled medical images;
[0112] The medical image segmentation module 902 is configured to use the medical image to be segmented as input to the trained semi-supervised medical image segmentation model to obtain a segmentation result of the medical image to be segmented.
[0113] To solve the above technical problems, the present application also provides a computer device. Figure 10 , Figure 10 This is a basic structural block diagram of the computer device in this embodiment.
[0114] The computer device 10 includes a memory 10a, a processor 10b, and a network interface 10c that are interconnected and communicate with each other through a system bus. It should be noted that the figure only shows a computer device 10 having components 10a-10c, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0115] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0116] The memory 10a includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10a may be an internal storage unit of the computer device 10, such as a hard disk or memory of the computer device 10. In other embodiments, the memory 10a may also be an external storage device of the computer device 10, such as a plug-in hard disk equipped on the computer device 10, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Of course, the memory 10a may also include both the internal storage unit of the computer device 10 and its external storage device. In this embodiment, the memory 10a is generally used to store an operating system and various application software installed on the computer device 10, such as program code of a semi-supervised medical image segmentation method. In addition, the memory 10a can also be used to temporarily store various data that has been output or is about to be output.
[0117] In some embodiments, the processor 10b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 10b is generally used to control the overall operation of the computer device 10. In this embodiment, the processor 10b is used to execute program code stored in the memory 10a or process data, such as executing the program code of the semi-supervised medical image segmentation method.
[0118] The network interface 10c may include a wireless network interface or a wired network interface. The network interface 10c is generally used to establish a communication connection between the computer device 10 and other electronic devices.
[0119] The present application also provides another embodiment, namely, providing a non-volatile computer-readable storage medium, which stores a program of a semi-supervised medical image segmentation method, and the semi-supervised medical image segmentation can be executed by at least one processor so that the at least one processor performs the steps of the semi-supervised medical image segmentation method as described above.
[0120] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0121] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A semi-supervised medical image segmentation method, characterized in that: include: Obtaining unlabeled medical images and labeled medical images; Training a semi-supervised medical image segmentation model based on unlabeled medical images and labeled medical images; The medical image to be segmented is used as input to the trained semi-supervised medical image segmentation model to obtain a segmentation result of the medical image to be segmented.
2. The semi-supervised medical image segmentation method according to claim 1, characterized in that: Training a semi-supervised medical image segmentation model based on unlabeled and labeled medical images, including: Extracting features of unlabeled medical images based on the encoder to obtain features of the unlabeled medical images; Calculate the prediction confidence values of the two decoders for the feature map, and obtain reliable pseudo-label masks and unreliable pseudo-label masks based on the prediction confidence values; Based on the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder, the cross entropy loss of reliable labels and unreliable labels is obtained respectively; Obtain unsupervised loss based on reliable pseudo-label mask, unreliable pseudo-label mask and corresponding cross-entropy loss; obtain supervised loss based on Diceloss; The predicted labels of the labeled medical images are obtained in the last convolutional layer of the decoder, and the weight of the supervised loss is updated by comparing the predicted labels with the true labels; The overall loss function of the semi-supervised medical image segmentation model is obtained by weighted summation to complete the training of the semi-supervised medical image segmentation model.
3. The semi-supervised medical image segmentation method according to claim 2, characterized in that: The encoder extracts features of unlabeled medical images to obtain features of the unlabeled medical images, including: the encoder is composed of a Mamba module and a CNN in parallel, the Mamba module receives the feature map output by the CNN layer as input, and the feature map enhanced by the Mamba module is further passed to other CNN layers to extract features of the unlabeled medical images.
4. The semi-supervised medical image segmentation method according to claim 2, characterized in that: The decoder-based method for obtaining a reliable pseudo-label mask and an unreliable pseudo-label mask for a feature map includes: using two decoders to respectively obtain a probability value that the feature map is a certain category, obtaining prediction confidence values of the feature map by two encoders based on the probability values, comparing the prediction confidence values based on a comparison network, and respectively obtaining a reliable pseudo-label mask and an unreliable pseudo-label mask.
5. The semi-supervised medical image segmentation method according to claim 2, characterized in that: The cross entropy loss of obtaining reliable labels based on the features of the penultimate layer of the decoder and a prototype of a certain category of the corresponding decoder includes: The features of the penultimate layer of the decoder and the cosine similarity of the prototype of a certain category of the corresponding decoder are obtained, and the features of the penultimate layer of the decoder and the cross entropy loss of the pixel level of the prototype of a certain category of the corresponding decoder are obtained based on the preset formula.
6. The semi-supervised medical image segmentation method according to claim 2, characterized in that: The cross entropy loss of unreliable labels is obtained based on the features of the penultimate layer of the decoder and a prototype of a certain category of the corresponding decoder, including: Obtain the cosine similarity of the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder, obtain the cross-entropy loss of the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder based on a preset formula, and evaluate the difference information between the features of the penultimate layer of the decoder and the prototype of a certain category of the corresponding decoder based on the cross-entropy loss.
7. The semi-supervised medical image segmentation method according to claim 6, characterized in that: The difference information is encoded into a feature vector and fused with the input data of the encoder to obtain the fused feature, which is used as the final input of the encoder.
8. A semi-supervised medical image segmentation system for implementing the semi-supervised medical image segmentation method of claims 1-7, characterized in that: include: Medical image segmentation model training module, used to obtain unlabeled medical images and labeled medical images; Training a semi-supervised medical image segmentation model based on unlabeled medical images and labeled medical images; The medical image segmentation module to be segmented is used to take the medical image to be segmented as the input of the trained semi-supervised medical image segmentation model to obtain the segmentation result of the medical image to be segmented.
9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Training method of medical image diagnosis model, and medical image diagnosis method and system
CN121147213A
Training methods for medical image diagnostic models, medical image diagnostic methods and systems
CN121147213B
Methods of training a recognition model to recognize medical site image integrity and related products
CN122492705A