Semi-supervised medical image segmentation method based on double-disturbance consistency learning

By introducing confidence-guided adaptive perturbation and mixed feature perturbation modules in semi-supervised medical image segmentation, the problem of existing methods ignoring sample complexity and model learning progress when adding perturbation is solved, achieving more efficient label-free data utilization and performance improvement.

CN120107581APending Publication Date: 2025-06-06MINJIANG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510158936.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing semi-supervised medical image segmentation method based on consistency regularization ignores sample complexity and model learning progress when adding perturbations, resulting in limited performance improvement.

Method used

A method based on dual perturbation consistency learning is proposed, including confidence-guided adaptive perturbation (CAP) module and hybrid feature perturbation (MFP) module. The CAP module adaptively applies image-level perturbation based on the difficulty of the input image and the learning progress of the model; the MFP module applies feature-level perturbation at multiple levels of the feature pyramid, and promotes the simplicity and efficiency of the learning process through the mixing of features of weak and strong views.

Benefits of technology

Through adaptive image-level and feature-level perturbations, effectively utilize label-free data, improve the performance and stability of the model, reduce the impact of noise, and avoid overfitting and feature distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107581A_ABST
    Figure CN120107581A_ABST
Patent Text Reader

Abstract

The invention relates to a semi-supervised medical image segmentation method based on double-disturbance consistency learning, and belongs to the field of image segmentation. The method is based on weak-to-strong consistency learning, and comprises two modules: a confidence coefficient guided adaptive disturbance (CAP for short) module and a mixed feature disturbance (MFP for short) module. Specifically, the CAP module adaptively applies image-level perturbations according to the difficulty of inputting an image and the learning progress of a model. The MFP module applies feature-level disturbance on multiple levels of a feature pyramid, and promotes the learning process to become simpler, more convenient and more efficient through feature mixing of amblyograms and strong views.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image segmentation, and in particular relates to a semi-supervised medical image segmentation method based on double perturbation consistency learning. Background Art

[0002] Medical image segmentation is a key component of computer-aided diagnosis, which significantly improves the accuracy of medical diagnosis and treatment. It can help clinicians quickly and accurately identify the type and severity of diseases, and facilitate the formulation of optimal treatment plans by locating abnormal areas in medical images. In recent years, with the development of deep learning, convolutional neural networks (CNNs) have received increasing attention in medical image segmentation. CNNs have driven the success of segmentation algorithms with their powerful feature extraction and learning capabilities, especially in terms of accuracy and efficiency. Despite the remarkable results, training CNNs requires a large amount of densely labeled image data, which is expensive to obtain.

[0003] In order to reduce the dependence on labeled data, some semi-supervised medical image segmentation methods have emerged recently. The mainstream semi-supervised medical image segmentation methods are mainly divided into two directions: pseudo-labeling and consistency regularization. Specifically, pseudo-labeling methods usually use pre-trained models to generate pseudo-labels for unlabeled samples, and then combine these pseudo-labels with labeled samples for model training. Although this method is simple and effective, directly using the pseudo-labels predicted by the model may introduce noise, thereby reducing the performance of the model [1]. In order to reduce the impact of noise, some studies [2,3] use threshold screening methods to filter noise. Other studies [4-6] design special modules to optimize pseudo-labels to improve model performance. On the other hand, consistency regularization methods [7,8] impose perturbations on unlabeled samples to force the model to output consistent outputs under different perturbations. Recent studies [8,9] have shown that powerful data augmentation perturbation techniques such as CutMix

[10] can significantly improve the performance of semi-supervised learning methods.

[0004] At present, the ways of adding perturbations to semi-supervised medical image segmentation methods based on consistency regularization are divided into image perturbations and feature perturbations. However, both ways of adding perturbations have their limitations: (1) At the image perturbation level, strong enhancement is applied to all unlabeled samples indiscriminately, ignoring the sample complexity, the difficulty of training unlabeled samples, and the learning progress of the model at different stages. For example, in the early stages of learning, the model may find it difficult to effectively learn from strongly enhanced unlabeled samples, thereby hindering performance improvement. On the contrary, in the later stages of learning, repeatedly training unlabeled samples that are easy to learn may lead to overfitting; (2) At the feature perturbation level, feature perturbations are applied to weakly enhanced images. Due to randomness, the changes after feature perturbation may not be significant. On the contrary, only perturbing features on strongly enhanced images may distort the feature distribution, resulting in incorrect decision boundaries for the model. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a semi-supervised medical image segmentation method based on dual perturbation consistency learning, which is based on weak to strong consistency learning and includes two modules: a confidence-guided adaptive perturbation (CAP) module and a mixed feature perturbation (MFP) module. Specifically, the CAP module adaptively applies image-level perturbations according to the difficulty of the input image and the learning progress of the model. The MFP module applies feature-level perturbations at multiple levels of the feature pyramid and promotes the learning process to become simpler and more efficient through the mixing of features of weak views and strong views.

[0006] To achieve the above object, the technical solution of the present invention is: a semi-supervised medical image segmentation method based on double perturbation consistency learning, comprising:

[0007] A confidence-guided adaptive perturbation CAP module is proposed to adaptively apply image-level perturbations according to the difficulty of the input image and the current learning progress of the model.

[0008] A mixed feature perturbation (MFP) module is proposed to apply feature-level perturbations on multiple levels of the feature pyramid and promote the learning process to become simpler and more efficient by mixing features from weak and strong views.

[0009] In one embodiment of the present invention, the CAP module adaptively applies image-level perturbations according to the difficulty of the input image and the current learning progress of the model, that is, according to the current learning progress of the model, the confidence map of the unlabeled image is used to infer its difficulty and the confidence of the correct segmentation, so as to adaptively apply strong perturbations to each unlabeled image.

[0010] In one embodiment of the present invention, strong perturbations are adaptively applied to each unlabeled image. Specifically, the current learning progress of the model is tracked by an exponential moving average of the image difficulty, and the relationship between the current learning progress of the model and the predetermined image difficulty is used to determine whether strong perturbations need to be applied to the corresponding unlabeled image.

[0011] In one embodiment of the present invention, the method is implemented as follows:

[0012] Given a small amount of labeled data and a large unlabeled dataset in are the i-th and j-th images in the labeled and unlabeled datasets, respectively, and S is the set of image locations, i.e. | S | =D×W×H, where D, W, and H are the three spatial dimensions of the input image, respectively. i ∈C |S| is the image x i , where C is the category space of the label, M and N are the sizes of the labeled dataset and the unlabeled dataset, respectively. A semi-supervised algorithm FixMatch that integrates consistency regularization and pseudo-labels is used to constrain the prediction of the strongly enhanced unlabeled images using the prediction of the weakly enhanced unlabeled images as pseudo-labels. The overall loss function is the supervised loss function and the unsupervised loss function A combination of Defined as:

[0013]

[0014] in, and i denote the prediction and label of the model respectively. In this work, They are cross entropy loss and Dice loss respectively. For unlabeled images, the prediction of the model is first obtained by weak enhancement and strong enhancement input. The formula is as follows:

[0015]

[0016] Among them, H w (·),H s (·) indicates weak enhancement and strong enhancement, respectively. represents the corresponding model output probability, ★={w,s}, and in addition, the predicted label under weak enhancement is used As a pseudo-label to supervise the prediction of the strongly enhanced version of the unlabeled image, from FixMatch, we know that A binary mask is required To retain the pixel positions where the model has high confidence in the prediction; for The element at image position s∈S, To predict the model at image position s under weak enhancement, the following threshold rule is used to obtain the mask at image position s:

[0017]

[0018] in, is the indicator function, the mask It is 1 only at the image positions of the most likely class whose predicted probability exceeds the threshold τ. Therefore, the unsupervised loss function Written as:

[0019]

[0020] in, is the mask Dice loss, that is, only Apply Dice loss to image locations where the value is 1

[0021] Finally, the overall loss function is defined as:

[0022]

[0023] Among them, λ u It is a hyperparameter used to control the ratio of supervised loss and unsupervised loss in the overall loss.

[0024] In one embodiment of the present invention, each image and its label is a two-dimensional slice or a three-dimensional volume.

[0025] In one embodiment of the present invention, the confidence-guided adaptive perturbation CAP module proposes an adaptive strategy to decide whether to apply copy-paste enhancement (abbreviated as CutMix), which is as follows:

[0026] Two vectors are introduced in the training process. In each given mini-batch (a mini-batch refers to a small number of samples used in each training iteration for training. Specifically, a mini-batch refers to a small batch of data processed each time during the training process, where a "mini-batch" refers to a group of data with a small number of samples.), they are r = [ r 1 ,r 1 ,…,r K] and c = [ c 1 ,c 1 ,…,c K], where K is the number of unlabeled images in the mini-batch, and r records the proportion of correctly predicted pixels in the weakly enhanced unlabeled images, written as:

[0027]

[0028] Among them, p k w ,s represents the model prediction at position s of the kth weakly enhanced unlabeled image in the mini-batch, τ is a manually set threshold used to filter out pixels with incorrect predictions; r K The higher the value, the lower the training difficulty; use l t to denote the proportion of correctly predicted pixels of all weakly enhanced unlabeled images in a mini-batch, where the subscript t corresponds to the time step in training, and l t Written as:

[0029]

[0030] Because l t It may fluctuate between different batches, and is used to update a more stable estimate of the overall average difficulty of unlabeled images R t ; Assuming that each mini-batch is a time step t, as the training progresses, at the beginning of the training, the initial value R 0 Set to 0, then R t Updated by Exponentially Weighted Average EMA:

[0031] R t =αR t-1 +(1-α)l t (8)

[0032] Among them, α is a hyperparameter used to control the update speed. By comparing r k and R t , quantifies the training difficulty of a given unlabeled image and compares it to the overall average difficulty of all unlabeled images, R t can be regarded as the current learning progress of the model; in particular, r k and R t The relative relationship between them forms the basis for the decision on whether to apply CutMix enhancement to the unlabeled image.

[0033] In one embodiment of the present invention, r k and R t The relative relationship between them forms whether CutMix enhancement is applied to the unlabeled image, which is specifically implemented as follows:

[0034] Using r k and R tThe relative relationship between them, calculate c k , it will adaptively divide the images in the mini-batch into two parts. Specifically, c k Indicates whether CutMix enhancement should be applied to each image in the current mini-batch. Each element of the vector c has only two values ​​0 and 1. Here, 0 means that the conventional strong perturbation H should be applied to the corresponding image. s (·), and 1 means that the conventional strong perturbation H should be applied to the corresponding images simultaneously s (·) and CutMix enhancement, this image adaptive perturbation is represented as H cap (·), c k The update is written as:

[0035]

[0036] From formula (9), we can see that only when the difficulty of the image r k Higher than or equal to the current average difficulty R t When the corresponding image is enhanced, CutMix is ​​applied. k The update is performed in each mini-batch. In other words, whether to apply CutMix enhancement to the unlabeled image depends on the learning progress of the model and the training difficulty of the unlabeled image. In particular, for the same unlabeled image, c k The value of may change during the training process. For images with low training difficulty and that the model can handle, it is necessary to apply CutMix enhancement to make the learning task more challenging, thereby driving learning further to handle more difficult cases. On the other hand, for images that already have high training difficulty, CutMix enhancement may destroy the data distribution of the image, thereby negatively affecting model learning, so CutMix enhancement is not applied. In summary, the unsupervised loss of the proposed confidence-guided adaptive perturbation is expressed as:

[0037]

[0038] Among them, H cap (·)and They represent the model output after adaptive perturbation and perturbation respectively.

[0039] In one embodiment of the present invention, the mixed feature perturbation MFP module is specifically implemented as follows:

[0040] Segmentation Model Decomposed into encoder ε and decoder Let l∈{l 1 ,l 2 ,…,l L} is the L level of feature mixing in the encoder ε; therefore, after passing through the encoder, the intermediate features obtained from the weakly enhanced and strongly enhanced images are expressed as and As shown below:

[0041]

[0042] The subscript l is used to represent the intermediate features of the lth layer in the encoder ε, and image mixing enhancement (abbreviated as MixUp) is used to mix the features of weakly enhanced and strongly enhanced unlabeled images to enhance the richness of the features; for a pair of features, the feature mixing coefficient λ' is calculated as follows:

[0043] λ∽Beta(β,β) (13)

[0044] λ′=max(λ,1-λ) (14)

[0045] Among them, λ is obtained by sampling using Beta distribution (Beta(β,β) for short), β is the shape parameter of Beta distribution, which determines the shape and concentration of distribution, and λ' is the larger of λ and 1-λ, which ensures that λ' is always greater than 0.5, thus avoiding excessive bias towards one sample during the mixing process and maintaining the diversity of mixed features; use λ' to generate a feature mixing matrix Among them C l ×W l ×H l Represents intermediate features and Dimensions; then, using M l To mix the features from strongly enhanced and weakly enhanced unlabeled images, it is expressed as:

[0046]

[0047] Among them, ⊙ is the element-by-element multiplication operation; next, the output after the decoder is expressed as:

[0048]

[0049] in, is the last layer feature output by the encoder, l L is the last layer in the encoder ε; this feature is then fed into the decoder is processed; finally, the unsupervised loss of the mixed feature perturbation is defined as:

[0050]

[0051] In one embodiment of the present invention, in the method, the unlabeled image has three forward flows:

[0052] (i) Image-level weak perturbation flow:

[0053] (ii) Image-level CAP flow:

[0054] (iii) Feature-level MFP flow:

[0055] The final loss function Calculated as:

[0056]

[0057] Among them, μ cap , μ mfp are the corresponding loss weights respectively.

[0058] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.

[0059] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention proposes a framework for semi-supervised medical image segmentation. The framework aims to add reasonable perturbations at the image level and feature level to better utilize unlabeled data. To this end, the present invention proposes two novel perturbation modules: confidence-guided adaptive perturbation (CAP) and mixed feature perturbation (MFP). The CAP module performs adaptive image-level perturbations on unlabeled samples based on the confidence map obtained by network prediction according to the learning progress of the model and the training difficulty of the image. In addition, the MFP module mixes the features of strongly enhanced and weakly enhanced versions of unlabeled samples in multiple network layers to increase the diversity of feature representations and improve the gradient flow of cross-view consistency learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 The conventional perturbation method (a) and the perturbation method (b) of the present invention are shown.

[0061] Figure 2 This is the network model architecture of the method of the present invention.

[0062] Figure 3 To visualize the comparison results of different methods on the ACDC dataset, 10% labeled data is used. DETAILED DESCRIPTION

[0063] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0064] The present invention proposes a semi-supervised medical image segmentation method based on double perturbation consistency learning, comprising:

[0065] A confidence-guided adaptive perturbation CAP module is proposed to adaptively apply image-level perturbations according to the difficulty of the input image and the current learning progress of the model.

[0066] A mixed feature perturbation (MFP) module is proposed to apply feature-level perturbations on multiple levels of the feature pyramid and promote the learning process to become simpler and more efficient by mixing features from weak and strong views.

[0067] The following is the specific implementation process of the present invention.

[0068] The present invention first proposes a confidence-guided adaptive perturbation module (CAP for short). The goal of the CAP module is to adaptively apply strong perturbations to each unlabeled image according to the current learning progress of the model, that is, to avoid learning from samples that are too difficult or too simple. To achieve this goal, the present invention uses the confidence map of the unlabeled image to infer its difficulty and the confidence of the correct segmentation, and this inference is based on the state of the model currently being learned. Specifically, this is achieved by estimating the proportion of correctly predicted pixels based on a high confidence threshold. Then, the difficulty of a single image needs to be compared with a reference value, which represents the overall learning progress of the model. In order not to introduce additional parameters, the present invention can use the average difficulty of all images as a reference. Therefore, the present invention tracks the learning progress of the model through the exponential moving average of the image difficulty, and uses the relationship between the learning progress and the difficulty of a specific image to decide whether it is necessary to apply strong perturbations to the image. In addition, the present invention also proposes a mixed feature perturbation (MFP for short) module, which uses features from two views for perturbation at the same time, and improves model performance by exploring richer feature interactions. Different from previous feature perturbation methods [11,12], the MFP module considers the encoded features of both weakly enhanced and strongly enhanced unlabeled samples. It should be noted that the features of these two views are complementary in nature. For example, when the feature perturbation in the weakly enhanced view is not strong enough, the perturbation in the strongly enhanced view can serve as a supplement. Therefore, mixing these features has great potential to enhance feature representation and ultimately improve model performance. More importantly, the present invention proposes to mix these features at multiple levels of the feature pyramid, thereby providing more paths for the cross-view gradient flow in the optimization. Figure 1The perturbation method proposed in the present invention is different from other perturbation methods. The left side of Figure (a) shows the feature perturbation (WFP) of the weakly enhanced image and the indifferent random perturbation (RP) of the image. The right side of Figure (a) shows the feature perturbation (SFP) of the strongly enhanced image. (b) is the CAP and MFP modules proposed in the present invention. CAP replaces RP by applying image adaptive input perturbation. MFP replaces WFP and SFP by performing multi-level feature mixing on weakly enhanced and strongly enhanced inputs. Here, x w and x s Represent weakly enhanced and strongly enhanced images respectively. As pseudo labels to supervise the outputs of different perturbations, they are and

[0069] 1. Brief Introduction

[0070] In semi-supervised medical image segmentation, the present invention gives a small amount of labeled data set and a large unlabeled dataset in are the i-th and j-th images in the labeled and unlabeled datasets, respectively. Here, S is the set of image locations, i.e., |S| = D×W×H, where D, W, and H are the three spatial dimensions of the input image, respectively. i ∈C |S| is the image x i where C is the class space of labels. M and N are the sizes of labeled and unlabeled datasets, respectively. Each image and its label can be a two-dimensional slice or a three-dimensional volume. The present invention chooses to use the latter to improve the versatility of representation. More specifically, FixMatch

[13] It is a semi-supervised algorithm that integrates consistency regularization and pseudo-labeling. Originally designed for semi-supervised classification, FixMatch uses the predictions of weakly enhanced unlabeled images as pseudo-labels to constrain the predictions of strongly enhanced unlabeled images. The overall loss function is the supervised loss function and the unsupervised loss function A combination of Defined as:

[0071]

[0072] in, and i denote the model’s prediction and label respectively. In this work, in addition to the standard cross entropy loss The present invention also uses Dice loss To minimize the difference between the prediction and the label. For unlabeled images, the present invention first obtains the prediction of the model through weak enhancement and strong enhancement input, and the formula is as follows:

[0073]

[0074] Among them, H ★ (·), ★={w,s} represents weak enhancement and strong enhancement respectively. Represents the corresponding model output probability. In addition, the present invention uses the predicted label under weak enhancement As pseudo-labels to supervise the prediction of strongly enhanced versions of unlabeled images.

[13] You can know, A binary mask is required To retain pixel locations where the model has high confidence in the prediction. for The element at image position s∈S, For model prediction of image position s under weak enhancement, the present invention uses the following threshold rule to obtain the mask at image position s:

[0075]

[0076] in, is the indicator function, the mask It is 1 only at the image positions of the most likely class whose predicted probability exceeds the threshold τ. Therefore, the unsupervised loss function It can be written as:

[0077]

[0078] in, is the mask Dice loss, that is, only Apply Dice loss to image locations where the value is 1 Finally, the overall loss function can be defined as:

[0079]

[0080] Among them, u It is a hyperparameter used to control the ratio of supervised loss and unsupervised loss in the overall loss.

[0081] 2 Confidence-Guided Adaptive Perturbation

[0082] In semi-supervised image segmentation, CutMix

[13] It is widely used as a strong input augmentation method. Specifically, CutMix generates new training data by randomly copying a region of an unlabeled image and pasting it to the corresponding position of another unlabeled image. [8,9,13] We show that CutMix is ​​able to improve the performance of semi-supervised algorithms, but it still has some major limitations. In particular, the decision to apply CutMix augmentation to an unlabeled image is usually made in a random manner. However, this random augmentation does not take into account the training difficulty of a particular image and the current learning progress of the model. Furthermore, applying CutMix to images that are already very challenging for the model, or at an early stage in the training process, may have a negative impact on model performance.

[0083] In order to solve the above problems, the present invention proposes an adaptive strategy to decide whether to apply CutMix enhancement. To this end, the present invention introduces two vectors in the training process, which are r = [r 1 ,r 1 ,…,r K ] and c=[c 1 ,c 1 ,…,c K ], where K is the number of unlabeled images in the mini-batch. Here, r records the proportion of correctly predicted pixels in the weakly enhanced unlabeled images. It can be written as:

[0084]

[0085] in, represents the model prediction at position s of the kth weakly enhanced unlabeled image in the mini-batch, and τ is a manually set threshold used to filter out pixels that are mispredicted. K The higher the value, the lower the training difficulty. t to denote the proportion of correctly predicted pixels of all weakly enhanced unlabeled images in a mini-batch. The subscript t here corresponds to the time step in training. More specifically, l t It can be written as:

[0086]

[0087] Because l t It may fluctuate between different batches, and the present invention uses it to update a more stable estimate R of the overall average difficulty of unlabeled images t The present invention assumes that each mini-batch is a time step t. As the training progresses, at the beginning of the training, the initial value R 0 Set to 0, then R tThrough Exponential Moving Average (EMA)

[14] renew:

[0088] R t =αR t-1 +(1-α)l t (8)

[0089] Among them, α is a hyperparameter used to control the update speed, and the present invention sets it to α = 0.999. By comparing r k and R t , the present invention can quantify the training difficulty of a given unlabeled image and compare it with the overall average difficulty of all unlabeled images. In this context, R t It can also be seen as the current learning progress of the model. In particular, r k and R t The relative relationship between them forms the decision basis of the present invention on whether to apply CutMix enhancement to unlabeled images. Specifically, using r k and R t The relative relationship between the two, the present invention calculates c k , it will adaptively divide the images in the mini-batch into two parts. Specifically, c k Indicates whether CutMix enhancement should be applied to each image in the current mini-batch. Each element of the vector c has only two values ​​0 and 1. Here, 0 means that the conventional strong perturbation H should be applied to the corresponding image. s (·), and 1 means that the conventional strong perturbation H should be applied to the image at the same time s (·) and CutMix enhancement. Figure 2 This image adaptive perturbation is represented as H cap (·). c k The update can be written as:

[0090]

[0091] From formula (9), we can see that only when the difficulty of the image r k Higher than or equal to the current average difficulty R t CutMix enhancement is applied to the image only when k The update of is performed in each mini-batch. In other words, whether to apply CutMix enhancement to the unlabeled image depends on the learning progress of the model and the training difficulty of the unlabeled image. In particular, for the same unlabeled image, c kThe value of may change during the training process. For images with low training difficulty and that the model is able to handle, it is necessary to apply CutMix enhancement to make the learning task more challenging, thereby pushing the learning further to handle more difficult cases. On the other hand, for images that already have high training difficulty, CutMix enhancement may destroy the data distribution of the image, thereby negatively affecting model learning, so CutMix enhancement is not applied. In summary, the unsupervised loss of the proposed confidence-guided adaptive perturbation can be expressed as:

[0092]

[0093] Among them, H cap (·)and denote the model output after adaptive perturbation and perturbation, respectively. The main difference between Eq. (11) and Eq. (4), i.e., the traditional unsupervised loss in FixMatch, lies in the image motion used.

[0094] 3 Mixed feature perturbations

[0095] This section discusses feature-level perturbations and introduces the mixed feature perturbations proposed in this paper. Existing feature perturbations are usually performed only on weakly enhanced or strongly enhanced unlabeled images. Unlike these feature perturbations, the present invention allows the interaction between features in strongly and weakly enhanced unlabeled images to achieve communication between encoding features from different perspectives. Specifically, the present invention performs multi-level feature mixing on multiple different network layers, each with different spatial resolution and receptive field. Subsequently, the deeply intertwined features from weak and strong views are input into a common decoder to obtain model predictions. Mixing features at multiple levels can provide more cross-view gradient flow channels, making it easier to optimize the consistency learning objective. Formally, the segmentation model F can be decomposed into an encoder ε and a decoder Let l∈{l 1 ,l 2 ,…,l L} is the L levels of feature mixing in the encoder ε. Therefore, after passing through the encoder, the intermediate features obtained from the weakly enhanced and strongly enhanced images of the present invention are respectively expressed as and As shown below:

[0096]

[0097] The present invention uses the subscript l to represent the intermediate features of the lth layer in the encoder ε.

[15] Inspired by this, the present invention uses MixUp

[16] To mix the features of weakly enhanced and strongly enhanced unlabeled images, thereby enhancing the richness of the features. For a pair of features, the feature mixing coefficient λ' can be calculated as follows:

[0098] λ∽Beta(β,β) (13)

[0099] λ′=max(λ,1-λ) (14)

[0100] Among them, λ is obtained by sampling using Beta distribution (Beta(β,β) for short), β is the shape parameter of Beta distribution, which determines the shape and concentration of distribution, and λ' is the larger of λ and 1-λ, which ensures that λ' is always greater than 0.5, thus avoiding excessive bias towards one sample during the mixing process and maintaining the diversity of mixed features; The present invention uses λ' to generate a feature mixing matrix Among them C l ×W l ×H l Represents intermediate features and Then, the present invention uses M l To mix the features from strongly enhanced and weakly enhanced unlabeled images, it can be expressed as:

[0101]

[0102] Among them, ⊙ is the element-by-element multiplication operation. Next, the output after the decoder is expressed as:

[0103]

[0104] in, is the last layer feature output by the encoder, l L is the last layer in the encoder ε. This feature is then fed into the decoder Finally, the unsupervised loss of mixed feature perturbation can be defined as:

[0105]

[0106] 4 Overall framework

[0107] So far, the present invention has introduced two important components of the method: confidence-guided adaptive perturbation (CAP) and mixed feature perturbation (MFP). In the overall framework, there are three forward flows for unlabeled images: (i) image-level weak perturbation flow: (ii) Image-level CAP flow: (iii) Feature-level MFP flow: The final loss function Calculated as:

[0108]

[0109] During training, image-level and feature-level perturbations each have their own advantages and effects. Therefore, their loss weights μ cap and μ mfp are all set to 0.5. and In , the confidence threshold τ is set to 0.95.

[0110] 5 Experiments

[0111] In order to evaluate the effectiveness of the algorithm of the present invention on medical image data, the publicly available multi-category cardiac segmentation (ACDC) dataset has been used to extensively evaluate the method of the present invention and compared with several recently published algorithms. In order to quantitatively evaluate the segmentation performance of the algorithm on these three datasets, the present invention selects the following four evaluation indicators: Dice score, Jaccard score, 95% Hausdorff distance (95HD) and average surface distance (ASD). Among them, the Dice score and Jaccard score are used to measure the similarity between the network prediction and the true label. Higher Dice and Jaccard scores indicate a higher similarity between the predicted value and the true value. 95HD and ASD represent the distance between the network prediction and the boundary of the true label. Lower 95HD and ASD represent that the predicted value overlaps more with the boundary of the true value.

[0112] 5.1 Quantitative comparison

[0113] On the ACDC dataset, in order to verify the effectiveness of the proposed method, the present invention compares it with the most advanced semi-supervised methods, including URPC

[10] UAMT

[17] 、CLD [5] 、CPS

[18] and SCP-Net [7] Among them, URPC

[10] UAMT

[17] 、CLD [5]] and SCP-Net [7] It is a method specially designed for semi-supervised medical image segmentation, while CPS

[18] It is a general semantic segmentation algorithm. In order to ensure fair use of labeled data and unlabeled data for all methods, the present invention conducts experiments at labeled data ratios of 10% and 5%, respectively. Table 1 shows the experimental results of the method of the present invention and other comparative methods on the ACDC dataset. It can be seen from the results in Table 1 that the method of the present invention is significantly better than other methods under different experimental settings. This result shows that the method of the present invention is more effective than other methods in utilizing unlabeled data. It is worth noting that as the proportion of labeled data decreases, the performance improvement brought by the method of the present invention is more significant, which shows that it has good applicability in the low labeled data scenario of the ACDC dataset.

[0114] Table 1 Performance on the ACDC dataset

[0115]

[0116]

[0117] 5.2 Qualitative comparison

[0118] In order to compare the state-of-the-art semi-supervised algorithms with the algorithm of the present invention, the present invention outputs the predicted segmentation images of different methods in the test set of the ACDC dataset, such as Figure 3 As shown in the figure, UAMT

[17] 、CLD [5] and CPS

[18] The model predictions on the two images of different styles have obvious semantic errors.

[10] and SCP-Net [7] Although the models of these two methods predict fewer semantic errors, the boundary prediction of the predicted image is inaccurate compared with the true value. The image predicted by the method of the present invention is closest to the true value, and can well predict the category in the uncertain boundary area and the area where semantic errors are easily confused, showing the best performance.

[0119] 5.3 Ablation Experiment

[0120] The present invention conducted an ablation experiment on the ACDC dataset, using 10% labeled data to evaluate the effectiveness of the proposed module. Specifically, the present invention compared the following models: (1) the original FixMatch, (2) CAP module + FixMatch, (3) MFP module + FixMatch, and (4) CAP module + MFP module + FixMatch (the complete method of the present invention). The results are shown in Table 2. The CAP and MFP modules brought 1.12% and 1.73% improvements, and 1.27% and 1.91% improvements in the Dice and Jaccard indicators, respectively. This shows that each module can bring performance improvements alone, and when the two modules work together, the performance improvement is more significant.

[0121] Table 2 Performance on the ACDC dataset

[0122]

[0123] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.

[0124] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0125] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0126] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0128] References:

[0129] [1]Qin,Xuebin,et al."U2-Net:Going deeper with nested U-structure forsalient object detection."Pattern recognition 106(2020)

[0130] [2]Wang,X,et al.:Ssa-net:Spatial self-attention network for covid-19pneumonia infection segmentation with semi-supervised few-shotlearning. Medical Image Analysis 79,102459(2022)

[0131] [3]Yao, H., Hu, X., Li,

[0132] [4]Wang,T.,Huang,Z.,Wu,J.,Cai,Y.,Li,Z.:Semi-supervised medical imagesegmentation with co-distribution alignment.Bioengineering 10(7),869(2023)

[0133] [5]Lin,Y.,Yao,H.,Li,Z.,Zheng,G.,Li,X.:Calibrating label distributionfor classimbalanced barely-supervised knee segmentation.In:InternationalConference on Medical Image Computing and Computer-Assisted Intervention,pp.109–118(2022).

[0134] [6]Shi,Y.,Zhang,J.,Ling,T.,Lu,J.,Zheng,Y.,Yu,Q.,Qi,L.,Gao,Y.:Inconsistency-aware uncertainty estimation for semi-supervised medical imagesegmentation.IEEE Transactions on Medical Imaging 41(3),608–620(2021)

[0135] [7]Zhang,Z.,Ran,R.,Tian,C.,Zhou,H.,Li,X.,Yang,F.,Jiao,Z.:Self-awareand cross-sample prototypical learning for semi-supervised medical imagesegmentation.In:International Conference on Medical Image Computing andComputer-Assisted Intervention,pp.192–201(2023)

[0136] [8]Yang,L.,Qi,L.,Feng,L.,Zhang,W.,Shi,Y.:Revisiting weak-to-strongconsistency in semi-supervised semantic segmentation.In:Proceedings of theIEEE / CVF Conference on Computer Vision and Pattern Recognition,pp.7236–7246(2023)

[0137] [9]Bai,Y.,Chen,D.,Li,Q.,Shen,W.,Wang,Y.:Bidirectional copy-paste forsemi-supervised medical image segmentation.In:Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition,pp.11514–11524(2023)

[0138]

[10] Luo,X.,Wang,G.,Liao,W.,Chen,J.,Song,T.,Chen,Y.,Zhang,S.,Metaxas,D.N.,Zhang,S.:Semi-supervised medical image segmentation via uncertaintyrectified pyramid consistency.Medical Image Analysis 80,102517(2022)

[0139]

[11] Liu,Y.,Tian,Y.,Chen,Y.,Liu,F.,Belagiannis,V.,Carneiro,G.:Perturbed and strict mean teachers for semi-supervised semanticsegmentation.In:Proceedings ofthe IEEE / CVF Conference on Computer Vision andPattern Recognition,pp.4258–4267(2022)

[0140]

[12] Ouali,Y.,Hudelot,C.,Tami,M.:Semi-supervised semantic segmentationwith cross-consistency training.In:Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition,pp.12674–12684(2020)

[0141]

[13] Yun,S.,Han,D.,Oh,S.J.,Chun,S.,Choe,J.,Yoo,Y.:Cutmix:Regularization strategy to train strong classifiers with localizablefeatures.In:Proceedings of the IEEE / CVF International Conference on ComputerVision,pp.6023–6032(2019)

[0142]

[14] A.Tarvainen,H.Valpola,Mean teachers are better role models:Weight-averaged consistency targets improve semi-supervised deep learningresults,Advances in neural information processing systems 30(2017)

[0143]

[15] D.Berthelot,N.Carlini,I.Goodfellow,N.Papernot,A.Oliver,C.A.Raffel,Mixmatch:A holistic approach to semi-supervised learning,Advancesin Neural Information Processing Systems 32(2019)

[0144]

[16] H.Zhang,M.Cisse,YNDauphin,D.Lopez-Paz,mixup:Beyond empiricalrisk minimization,arXiv preprint arXiv:1710.09412(2017)

[0145]

[17] L.Yu, S.Wang,

[0146]

[18]

[0147] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A semi-supervised medical image segmentation method based on double perturbation consistency learning, characterized in that: include: A confidence-guided adaptive perturbation CAP module is proposed to adaptively apply image-level perturbations according to the difficulty of the input image and the current learning progress of the model. A mixed feature perturbation (MFP) module is proposed to apply feature-level perturbations on multiple levels of the feature pyramid and promote the learning process to become simpler and more efficient by mixing features from weak and strong views.

2. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 1, characterized in that: The CAP module adaptively applies image-level perturbations according to the difficulty of the input image and the current learning progress of the model, that is, according to the current learning progress of the model, the confidence map of the unlabeled image is used to infer its difficulty and the confidence of the correct segmentation to adaptively apply strong perturbations to each unlabeled image.

3. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 2, characterized in that: Adaptively apply strong perturbations to each unlabeled image. Specifically, the current learning progress of the model is tracked through the exponential moving average of the image difficulty, and the relationship between the current learning progress of the model and the predetermined image difficulty is used to decide whether strong perturbations need to be applied to the corresponding unlabeled image.

4. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 1, characterized in that: The method is implemented as follows: Given a small amount of labeled data and a large unlabeled dataset in are the i-th and j-th images in the labeled and unlabeled datasets, respectively. S is the set of image locations, i.e., |S| = D × W × H, where D, W, and H are the three spatial dimensions of the input image, respectively. i ∈C |S| is the image x i , where C is the category space of the label, M and N are the sizes of the labeled dataset and the unlabeled dataset, respectively. A semi-supervised algorithm FixMatch that integrates consistency regularization and pseudo-labels is used to constrain the prediction of the strongly enhanced unlabeled images by using the prediction of the weakly enhanced unlabeled images as pseudo-labels. The overall loss function is the supervised loss function and the unsupervised loss function A combination of Defined as: in, and i denote the prediction and label of the model respectively. In this work, They are cross entropy loss and Dice loss respectively. For unlabeled images, the prediction of the model is first obtained by weak enhancement and strong enhancement input. The formula is as follows: Among them, H w (·),H s (·) represent weak enhancement and strong enhancement, respectively. represents the corresponding model output probability, * = {w, s}, and in addition, the predicted label under weak enhancement is used As a pseudo-label to supervise the prediction of the strongly enhanced version of the unlabeled image, from FixMatch, we know that A binary mask is required To retain the pixel positions where the model has high confidence in the prediction; for The element at image position s∈S, To predict the model at image position s under weak enhancement, the following threshold rule is used to obtain the mask at image position s: in, is the indicator function, the mask It is 1 only at the image positions of the most likely class whose predicted probability exceeds the threshold τ. Therefore, the unsupervised loss function Written as: in, is the mask Dice loss, that is, only Apply Dice loss to image locations where the value is 1 Finally, the overall loss function is defined as: Among them, λ u It is a hyperparameter used to control the ratio of supervised loss and unsupervised loss in the overall loss.

5. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 4, characterized in that: Each image and its label is either a 2D slice or a 3D volume.

6. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 4, characterized in that: The confidence-guided adaptive perturbation CAP module proposes an adaptive strategy to decide whether to apply copy-paste enhancement CutMix, as follows: Two vectors are introduced in the training process. In each given mini-batch, they are r = [r1, r1, …, r K ] and c=[c1,c1,…,c K ], where K is the number of unlabeled images in the mini-batch, and r records the proportion of correctly predicted pixels in the weakly enhanced unlabeled image, written as: in, represents the model prediction at position s of the kth weakly enhanced unlabeled image in the mini-batch, τ is a manually set threshold used to filter out pixels with incorrect predictions; r K The higher the value, the lower the training difficulty; use l t to denote the proportion of correctly predicted pixels of all weakly enhanced unlabeled images in a mini-batch, where the subscript t corresponds to the time step in training, and l t Written as: Because l t It may fluctuate between different batches, and is used to update a more stable estimate of the overall average difficulty of unlabeled images R t ; Assuming that each mini-batch is a time step t, as the training progresses, at the beginning of the training, the initial value R0 is set to 0, and then R t Updated by Exponentially Weighted Average EMA: R t =αR t-1 +(1-a)l t (8) Among them, α is a hyperparameter used to control the update speed. By comparing r k and R t , quantifies the training difficulty of a given unlabeled image and compares it to the overall average difficulty of all unlabeled images, R t can be regarded as the current learning progress of the model; in particular, r k and R t The relative relationship between them forms the basis for the decision on whether to apply CutMix enhancement to the unlabeled image.

7. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 6, characterized in that: r k and R t The relative relationship between them forms whether CutMix enhancement is applied to the unlabeled image, which is specifically implemented as follows: Using r k and R t The relative relationship between them, calculate c k , it will adaptively divide the images in the mini-batch into two parts. Specifically, c k Indicates whether CutMix enhancement should be applied to each image in the current mini-batch. Each element of the vector c has only two values ​​0 and 1. Here, 0 means that the conventional strong perturbation H should be applied to the corresponding image. s (·), and 1 means that the conventional strong perturbation H should be applied to the corresponding image simultaneously s (·) and CutMix enhancement, this image adaptive perturbation is represented as H cap (·), c k The update is written as: From formula (9), we can see that only when the difficulty of the image r k Higher than or equal to the current average difficulty R t CutMix enhancement will be applied to the corresponding image only when k The update is performed in each mini-batch. In other words, whether to apply CutMix enhancement to the unlabeled image depends on the learning progress of the model and the training difficulty of the unlabeled image. In particular, for the same unlabeled image, c k The value of may change during the training process. For images with low training difficulty and that the model can handle, it is necessary to apply CutMix enhancement to make the learning task more challenging, thereby driving learning further to handle more difficult cases. On the other hand, for images that already have high training difficulty, CutMix enhancement may destroy the data distribution of the image, thereby negatively affecting model learning, so CutMix enhancement is not applied. In summary, the unsupervised loss of the proposed confidence-guided adaptive perturbation is expressed as: Among them, H cap (·)and They represent the model output after adaptive perturbation and perturbation respectively.

8. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 7, characterized in that: The mixed feature perturbation MFP module is implemented as follows: Segmentation Model Decomposed into encoder ε and decoder Let l∈{l1,l2,…,l L } is the L level of feature mixing in the encoder ε; therefore, after passing through the encoder, the intermediate features obtained from the weakly enhanced and strongly enhanced images are expressed as and As shown below: Among them, the subscript l is used to represent the intermediate features of the lth layer in the encoder ε, and the image mixing enhancement MixUp is used to mix the features of weakly enhanced and strongly enhanced unlabeled images to enhance the richness of the features; for a pair of features, the feature mixing coefficient λ' is calculated as follows: λ∽Beta(β,β) (13) λ=max(λ,1-λ) (14) Among them, λ is obtained by sampling using the Beta distribution Beta(β,β), β is the shape parameter of the Beta distribution, which determines the shape and concentration of the distribution, and λ' is the larger of λ and 1-λ. Ensure that λ' is always greater than 0.5 to avoid being too biased towards one sample during the mixing process and maintain the diversity of the mixed features; use λ' to generate a feature mixing matrix Among them C l ×W l ×H l Represents intermediate features and Dimensions; then, using M l To mix the features from strongly enhanced and weakly enhanced unlabeled images, it is expressed as: Among them, ⊙ is the element-by-element multiplication operation; next, the output after the decoder is expressed as: in, is the last layer feature output by the encoder, l L is the last layer in the encoder ε; this feature is then fed into the decoder is processed; finally, the unsupervised loss of the mixed feature perturbation is defined as:

9. The semi-supervised medical image segmentation method based on double perturbation consistency learning according to claim 8, characterized in that: In the method described, there are 3 forward flows for unlabeled images: (i) Image-level weak perturbation flow: (ii) Image-level CAP flow: (iii) Feature-level MFP flow: The final loss function Calculated as: Among them, μ cap , μ mfp are the corresponding loss weights respectively.

10. A computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, and when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.

Citation Information

Cited By

  • Federal learning-oriented anti-backdoor attack image segmentation model training method and device

    CN120563969A