Active learning method and system for medical image data annotation with combination of uncertainty and representativeness
Through an active learning method combining uncertainty and representativeness, using variational autoencoder and gradient entropy divergence evaluation, the problems of high annotation cost and improper sample selection in medical image segmentation are solved, and efficient sample screening and model performance improvement are achieved.
Patent Information
- Application Number
- CN202510326568.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
AI Technical Summary
Existing medical image segmentation methods rely on large-scale fine labeling data sets, which have high labeling costs and low efficiency. Existing active learning algorithms are prone to ignore decision boundary areas or select redundant samples when selecting samples, affecting model performance.
An active learning method combining uncertainty and representativeness is adopted to construct a latent spatial generative model through a variational autoencoder, combining gradient and model entropy divergence evaluation strategy, screen candidate samples and optimize model parameters to ensure the comprehensive coverage of data distribution and the refined segmentation of decision boundaries.
It significantly improves the comprehensiveness and pertinence of sample selection, optimizes segmentation performance and clinical applicability, reduces the impact of overconfidence in neural networks, and improves the accuracy and efficiency of medical image segmentation.
Smart Images

Figure CN120236134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image data, and particularly to an active learning method and system for medical image data annotation combining uncertainty and representativeness. Background Art
[0002] In recent years, deep learning segmentation networks (such as U-Net, UNet++, DeepLabV3+, etc.) have achieved remarkable success in medical image segmentation tasks, and these methods have been greatly improved in terms of accuracy and generality. However, the performance of these methods largely depends on large-scale finely annotated datasets. Although the scale of medical image datasets has been continuously expanding in recent years, it is still difficult to meet the urgent needs of the rapid development of multi-fields of intelligent medicine.
[0003] Constructing a large-scale medical image dataset requires high annotation costs, which are mainly reflected in two aspects: First, deep learning-based medical image segmentation networks require a large number of annotated samples to achieve good segmentation effects; second, the annotation accuracy requirements for medical images are relatively high, the annotation difficulty is large, and the annotation process usually takes a long time. Especially for some images with complex structures and difficult-to-identify lesions, it often requires multiple experts to discuss together to obtain a consistent annotation result. To reduce the data annotation cost, there are usually two ways: one is to optimize the learning algorithm to reduce the demand for the number of samples; the other is to develop intelligent annotation tools to improve the annotation efficiency.
[0004] Currently, transfer learning, semi-supervised learning, etc. are to reduce the dependence on data samples from the learning method, while active learning starts from the data samples, optimizes the sample selection strategy to reduce the demand for training samples, and improves the training efficiency of sample data.
[0005] The paper "One-shot active learning for image segmentation via contrastive learning and diversity-based sampling》 proposes a clustering-based representative active learning method, which mainly selects samples according to the representative index. Although this method can better cover the global distribution of medical images, it is easy to ignore the training needs of difficult samples in the decision boundary region. These regions usually correspond to the key edges of lesions or complex anatomical structures, and the lack of segmentation accuracy may significantly affect the clinical applicability of disease diagnosis and treatment.
[0006] Chinese Patent Document CN117173701B discloses an active learning method for semantic segmentation based on superpixel feature representation learning. This method mainly selects samples based on information entropy, but does not fully consider the representativeness of data distribution, resulting in the sampling result being biased towards complex or edge regions. In addition, during the clustering process, although the K-means algorithm is used to cluster the superpixel feature vectors, only the sample with the highest information entropy is selected within the cluster, and the global representativeness of the sample to the cluster is not comprehensively considered. If only relying on the uncertainty index to select samples, it may lead to the algorithm selecting overly redundant samples. The decision boundary regions (such as fuzzy regions and edge regions) in medical images usually have high uncertainty, but they often lack global representativeness. Selecting only these data as the training set will lead to model overfitting and ignore the diversity and distribution of the entire dataset.
[0007] The paper "D2ADA: Dynamic Density-aware Active Domain Adaptation for Semantic Segmentation" proposes a strategy that dynamically combines uncertainty and representativeness. However, this method does not consider the drawback that the uncertainty index is prone to selecting redundant samples. The better the effect of an uncertainty algorithm, the easier it is to select redundant samples. If this defect is not constrained, after combining with the representativeness index, it may even occur that the combined performance is lower than that of only using the representativeness index. In the later rounds of active learning of this algorithm, the weight of the uncertainty index becomes larger, and at this time, the algorithm will select more redundant samples, thus affecting the model performance.
[0008] The paper "Deep Batch Active Learning by Diverse, Uncertain Gradient Lower Bounds" proposes an active learning method BADGE based on gradient embedding, which directly calculates the gradient using pseudo-labels. However, the predictions of neural networks are often overconfident in some wrong categories, resulting in a small gradient magnitude of the pseudo-labels. In this case, gradient embedding will underestimate the uncertainty of these samples, and may thus miss the samples crucial for the decision boundary, affecting the effectiveness of sample selection. Summary of the Invention
[0009] The purpose of the present invention is to propose an active learning method and system for medical image data annotation that combines uncertainty and representativeness. Through the effective combination of these two strategies, the comprehensiveness and pertinence of sample selection are achieved, providing a solid foundation for improving the performance of active learning algorithms.
[0010] According to the first aspect of the embodiments of the present disclosure, an active learning method for medical image data annotation that combines uncertainty and representativeness is provided, including the following steps:
[0011] Train a variational autoencoder infoVAE on the image pool X pool to construct a latent space generation model;
[0012] Randomly select M samples from the image pool X pool and label them to construct an initial labeled image set X an , and train a segmentation model on the labeled image set X an ;
[0013] Perform T rounds of active learning loops. In each round t of the active learning loop, the following steps are specifically executed:
[0014] Screen candidate samples based on a representative method;
[0015] Screen final samples based on an uncertainty method;
[0016] Update the labeled and unlabeled data sets;
[0017] Retrain the segmentation model on the updated labeled image set X an to optimize the model parameters;
[0018] Obtain the final model parameters.
[0019] According to the second aspect of the embodiments of the present disclosure, an active learning system for medical image data annotation combining uncertainty and representativeness is provided, including:
[0020] A latent space generation model, which trains a variational autoencoder infoVAE on the image pool X pool ;
[0021] An initial model training module, which randomly selects M samples from the image pool X pool and labels them to construct an initial labeled image set X an , and trains a segmentation model on the labeled image set X an ;
[0022] An active learning module, which performs T rounds of active learning loops. In each round T of the active learning loop, it includes:
[0023] Screen candidate samples based on a representative method;
[0024] Screen final samples based on an uncertainty method;
[0025] Update the labeled and unlabeled data sets;
[0026] Retrain the segmentation model on the updated labeled image set X an to optimize the model parameters;
[0027] A parameter acquisition module that obtains the final model parameters.
[0028] According to the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program running on the memory. When the processor executes the program, the active learning method for medical image data annotation combining uncertainty and representativeness is implemented.
[0029] According to the fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the active learning method for medical image data annotation combining uncertainty and representativeness is implemented.
[0030] The above technical solutions adopted by the present invention, compared with the prior art, have the following advantages:
[0031] (1) The comprehensiveness and pertinence of sample selection are improved, and the segmentation performance and clinical applicability are significantly optimized. The present invention adopts a dual strategy combining comprehensive uncertainty and representativeness in sample selection. Among them, the representativeness method based on the latent space preferentially selects samples that can best cover the global data distribution, ensuring a comprehensive coverage of the data distribution; the uncertainty method focuses on difficult samples near the decision boundary, strengthening the model's segmentation ability for fuzzy boundaries and complex anatomical structures in medical images. Through the combination of this dual strategy, the performance limitations brought by single-dimensional consideration are effectively solved.
[0032] (2) The influence of the overconfidence of the neural network on sample selection is significantly reduced. The present invention designs an uncertainty evaluation strategy based on gradient and model entropy divergence, avoiding the defects of over-reliance on pseudo-labels and being affected by overconfidence.
[0033] (3) The effect of combining uncertainty and representativeness is optimized. The present invention adopts a progressive selection strategy. During the sample selection process, a batch of candidate samples that cover the global data distribution are first selected from the unlabeled image pool by the representativeness method, and then the final training samples are screened from the candidate samples based on the uncertainty method. Through this combination method, the defect that the uncertainty index in the existing active learning algorithm is prone to select redundant samples is effectively overcome, and the respective advantages of the uncertainty and representativeness methods are fully exerted. Description of the Drawings
[0034] The specification drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application.
[0035] Figure 1Flowchart of an active learning method for medical image data annotation combining uncertainty and representativeness;
[0036] Figure 2 Flowchart for screening candidate samples by a representativeness-based method;
[0037] Figure 3 Flowchart for screening final samples by an uncertainty-based method. Specific implementation manners
[0038] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0039] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs.
[0040] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0041] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Similarly, it should be noted that each block in the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0042] According to the task requirements of medical image segmentation, the main purpose of medical image segmentation is still to assign labels to the special semantic information in the image (such as tumors, organs, blood vessels, etc.). However, the number of categories in medical image segmentation is generally not as large as that in natural image semantic segmentation, and often there are more binary classification problems. Therefore, representativeness and uncertainty are used as the principles for sample selection. Among them, the representative active learning method is committed to selecting samples that can represent the distribution of the entire image pool to ensure the model's learning ability for the global data distribution; the uncertainty active learning method is committed to selecting samples on the decision boundary, which can effectively find difficult examples that the model has not mastered. Therefore, after combining the two selection strategies, the utilization efficiency of data samples is improved and the model learning is balanced.
[0043] Embodiment 1:
[0044] This embodiment provides an active learning method for medical image data annotation that combines uncertainty and representativeness, including the following steps:
[0045] S1. Train a variational autoencoder infoVAE on the image pool X pool to construct a latent space generation model;
[0046] Specifically, the variational autoencoder infoVAE includes an encoder and a decoder. The goal of the encoder is to model the latent structure of the input medical image data and normalize it into a prior distribution (such as a Gaussian distribution) for optimization and sampling; the goal of the decoder is to use the latent variable z to reconstruct samples as close as possible to the input medical image data. Through the loss function L infoVAE , infoVAE is trained on the complete image pool X pool . The loss function L infoVAE is shown as follows:
[0047] L infoVAE = L AE + L MMD
[0048] where L AE is the reconstruction error, ensuring that the samples reconstructed through the latent variable z can be as close as possible to the input data. The loss function L AE is shown as follows:
[0049]
[0050] where q φ (z|x) represents the approximate posterior distribution learned by the encoder, mapping the input data x to the distribution of the latent variable. p θ (x|z) represents the conditional generation distribution learned by the decoder, indicating the probability distribution of generating x given the latent variable z.
[0051] L MMD is the maximum mean discrepancy (MMD) loss, which is used to measure the distance between the latent variable distribution q φ (z) and the prior distribution p(z), ensuring that the distribution in the latent space can match the prior distribution p(z). The loss function L MMD is shown as follows:
[0052]
[0053] where q φ (z) is the latent variable distribution obtained by the encoder from the input data x. p(z) is the set prior distribution, which is a standard multivariate normal distribution. k(z, z') is the Gaussian kernel, defined as k(z, z') = exp(-||z - z'|| / 2τ 2 ), where σ = 1.
[0054] S2. Randomly select M samples from the image pool X pool and label them to construct the initial labeled image set X an , and train the segmentation model on the labeled image set X an ;
[0055] Specifically, during the training process, the cross-entropy loss function is used to minimize the error between the prediction of the model f(x; θ) and the true label y.
[0056] It should be noted that the present invention does not select the state-of-the-art segmentation model, but adopts the classic U-Net as the segmentation network. The reason for such a design is to select the optimal samples through the active learning algorithm, thereby improving the model performance. By using a non-state-of-the-art segmentation model, the impact of different active learning algorithms on the improvement of the model performance can be more intuitively compared, so as to verify the superiority of the algorithm of the present invention in sample selection.
[0057] S3. Conduct T rounds of active learning loops. In each round t of the active learning loop, the following steps are specifically executed:
[0058] S31. Screen candidate samples based on a representative method;
[0059] The purpose of the current step is to select R non candidate samples that can best represent the complete image pool X t from the unlabeled data set X pool through the variational autoencoder infoVAE trained in S1, while ensuring that these candidate samples are different from the labeled data set X anThe sample distribution in it is made as non-redundant as possible. To achieve the above goal, first, use the trained variational autoencoder infoVAE in S1 to encode the samples of X pool and X an respectively, map them to the latent space and normalize them to a multivariate Gaussian distribution.
[0060] For each sample x i ∈X pool , the encoder of infoVAE outputs the distribution parameters of the sample in the latent space through the learned parameter φ, including the mean vector and the diagonal covariance matrix Specifically: where q φ (z|x i ) is the distribution of the sample x i in the latent space. Then, integrate the latent distributions of all samples in X pool to fit the global distribution of X pool : where μ pool and are the mean and variance of the samples in Z pool respectively, describing the global distribution characteristics of X pool .
[0061] Using a similar method, project the samples in the labeled dataset X an to the latent space Z an , and obtain its global distribution: where μ an and are the mean and variance of the samples in Z an respectively.
[0062] Through the above steps, the sample distributions of X pool and X an have been normalized to Gaussian distributions in the latent space. Next, select R non samples from X pool = X an \X t to maximize representativeness and avoid redundancy. Specifically, for each unlabeled sample x i ∈X non , obtain its representativeness score through the following formula:
[0063]
[0064] where, q φ (z|x i ,X pool) represents the sample x i For the complete dataset X pool representativeness; q φ (z|x i ,X an ) measures the redundancy of the sample x i with respect to the labeled dataset X an .
[0065] Using the previously fitted Gaussian distribution and to obtain the distribution value q i for each sample x φ (z|x i ,X). Specifically, it is estimated through the cumulative distribution function (CDF) of the Gaussian distribution, as shown in the following formula:
[0066]
[0067] where erf(x) is the error function; μ X and σ X are the mean and standard deviation of the Gaussian distribution of the dataset X, respectively.
[0068] Through this calculation method, R non candidate samples that best represent the complete image pool X t are successfully selected from X pool , providing high-quality preliminary screening results for subsequent uncertainty screening.
[0069] S32. Screening the final samples based on the uncertainty method;
[0070] The purpose of the current step is to further select S t samples for annotation from the R t candidate samples screened in S3.1. For this purpose, the present invention designs an uncertainty active learning method that combines gradient and model entropy divergence, aiming to more accurately quantify the uncertainty of samples and mitigate the overconfidence problem of neural networks. The specific steps are as follows.
[0071] First, input a sample x i to be screened into the current segmentation network f(x; θ), and obtain the output result O = f(x i ; θ). Subsequently, for the sample x i , calculate the entropy loss of each pixel of it, and take the average of the entropy losses of all pixels as the entropy loss of the entire image x i . The formula for calculating the entropy loss of each pixel in the sample x i is shown in the following formula:
[0072]
[0073] where p j represents the j-th pixel in the sample x i , and P(y c |p j ) represents the predicted probability that the pixel p j belongs to the class y c . C is the total number of classes. Then, calculate the gradient of the entropy loss of the entire image (i.e., the average of the entropy losses of all pixels) with respect to the parameters of the last layer of the segmentation network, and calculate its L2 norm, denoted as the uncertainty metric 1. This metric reflects the importance of the sample for updating the model parameters through the magnitude of the gradient, avoiding the error accumulation directly relying on the pseudo-labels.
[0074] To further evaluate the prediction uncertainty of the model for the sample, record the set L of pixel positions with relatively low current predicted entropy values. Low-entropy pixels usually represent that the model is more confident in predicting these pixels. However, neural networks often have the problem of "overconfidence", that is, showing high confidence in wrong predictions. Therefore, it is necessary to perform secondary uncertainty verification on these low-entropy pixels.
[0075] Based on the current model parameters θ, add Gaussian noise K times to generate perturbed model parameters θ k = θ + ∈ k , where σ is the hyperparameter of the noise intensity. For each perturbed model f(x; θ k ), perform forward propagation on the candidate sample x i to obtain the output O k = f(x i ; θ k ). Subsequently, obtain the difference in the output entropy values of the pixels in the low-entropy pixel set L before and after perturbation, and obtain the entropy difference set {Δ1, Δ2, …, Δ K}. Finally, by calculating the L2 norm of the entropy difference set, obtain the uncertainty metric 2, as shown in the following formula:
[0076] U(x i ) = ||{Δ1, Δ2, …, Δ K}||2
[0077] This metric quantifies the instability of low-entropy pixels under perturbation, reflecting whether there is an overconfidence problem with these pixels.
[0078] Finally, fuse the uncertainty metric 1 and the uncertainty metric 2 as the comprehensive uncertainty metric of the sample x i . After calculating the uncertainty scores of all candidate samples, sort the R t candidate samples in descending order of the uncertainty scores, and select the top S tThese samples are used as the final labeled sample set.
[0079] Compared with the current state-of-the-art active learning algorithms (such as BADGE), this method has been improved in two aspects. The first aspect is to reduce the dependence of the uncertainty active learning algorithm on pseudo-labels. The BADGE method calculates the cross-entropy loss and Dice loss through pseudo-labels, and uses the loss value to calculate the gradient of the parameters of the last layer of the model. This method is too dependent on the accuracy of pseudo-labels, which may lead to error accumulation. To reduce the influence of pseudo-labels, this method uses entropy loss as an alternative metric, only using the sample prediction distribution itself, avoiding the interference of pseudo-labels. The second aspect is to alleviate the overconfidence problem of neural networks. Neural networks have the problem of overconfidence during prediction, especially for incorrect predictions in complex tasks. This method further validates the low-entropy pixel positions, uses the entropy difference of the perturbed model output to quantify the impact of overconfidence on uncertainty, and further improves the reliability of sample selection.
[0080] S33. Update the labeled and unlabeled data sets;
[0081] Specifically, label the finally selected S t samples to obtain their true labels y. Remove these newly labeled samples from the unlabeled data set X non and add them to the labeled image set X an . The updated data set X an contains all the labeled samples for subsequent model training, while the unlabeled data set X non continues to retain the remaining unlabeled samples, providing a candidate set for the next round of active learning.
[0082] S34. Retrain the segmentation model on the updated labeled image set X an and optimize the model parameters;
[0083] Specifically, by minimizing the cross-entropy loss function, optimize the model parameters θ t to improve the fitting ability of the model to the labeled data and provide a more reliable prediction basis for the next round of active learning.
[0084] S5. Obtain the final model parameters.
[0085] After completing T rounds of active learning cycles, obtain the finally trained model parameters. Apply this model to an independent test set to evaluate its performance on the key performance indicators of Dice coefficient and ASD coefficient, thereby verifying the improvement effect of the active learning method on the segmentation performance.
[0086] This embodiment uses the ACDC dataset (Bernard et al., 2018), which contains short-axis cardiac MR images from 100 patients. Only the end-diastolic frames of each patient are used for evaluation, resulting in a total of 100 scans. Each scan corresponds to a manual segmentation mask of the left ventricle (LV), myocardium (MYO), and right ventricle (RV). The dataset is divided into a training set (70 scans, 656 slices), a validation set (10 scans), and a test set (20 scans) according to the division method of Luo et al. (2022). Considering the large spacing in the z-axis direction, a 2D segmentation model is used for training in the experiment, and the evaluation is performed through a 3D volume according to the method of Bai et al. (2017).
[0087] The experiment runs on a workstation equipped with an NVIDIA Quadro RTX 8000 graphics card (48GB video memory), with the operating system being Ubuntu 20.04.6 LTS (Focal Fossa) and the CUDA version being 12.0. The model is implemented based on the PyTorch framework, relying on the large video memory and high-performance computing capabilities of the graphics card, effectively supporting the training of the medical image segmentation model and the iterative execution of the progressive hybrid active learning algorithm.
[0088] The present invention is verified on the ACDC left ventricle medical dataset. Samples selected by the active learning method proposed by the present invention are used to train the medical image segmentation model. The obtained model is significantly higher than other active learning algorithms in terms of the key evaluation indicators Dice coefficient and average surface distance (ASD). The uncertainty evaluation strategy focuses on optimizing the ASD index, which measures the refinement degree of the segmentation edge. The uncertainty method designed by the present invention optimizes the ASD index by an average of 18.47% compared to the BADGE method, significantly improving the model's ability to handle complex boundaries in the medical image segmentation task.
[0089] Compared with using only the uncertainty method, in 5 rounds of active learning, the average improvement of the Dice coefficient is 10.15 percentage points (from 79.15% using only the uncertainty method to 89.30% using the progressive combination strategy). At the same time, the average optimization amplitude of the ASD coefficient is 72.8% (from 5.88 using only the uncertainty method to 1.60 using the progressive combination strategy). This result fully demonstrates the effectiveness of the combination strategy of the present invention in suppressing the defects of the uncertainty method. Compared with the representative method using only the latent space, in 5 rounds of active learning, the average improvement of the Dice coefficient is 1.23 percentage points (from 88.07% using only the representative method to 89.30% using the progressive combination strategy). At the same time, the average optimization amplitude of the ASD coefficient is 31.03% (from 2.32 using only the representative method to 1.60 using the progressive combination strategy). This result fully demonstrates that the combination strategy of the present invention enables both uncertainty and representative AL to fully exert their own advantages.
[0090] Example Two:
[0091] This example provides an active learning system for medical image data annotation that combines uncertainty and representativeness, including:
[0092] A latent space generation model that trains a variational autoencoder infoVAE on the image pool X pool ;
[0093] An initial model training module that randomly selects M samples from the image pool X pool for annotation and constructs an initial labeled image set X an , and trains a segmentation model on the labeled image set X an ;
[0094] An active learning module that performs T rounds of active learning loops. In each round t of the active learning loop, it includes:
[0095] Screening candidate samples based on the representative method;
[0096] Screening final samples based on the uncertainty method;
[0097] Updating the labeled and unlabeled data sets;
[0098] Retraining the segmentation model on the updated labeled image set X an to optimize the model parameters;
[0099] A parameter acquisition module that obtains the final model parameters.
[0100] In a representative aspect of the present invention, based on the global data distribution, through the estimation of the Gaussian distribution in the latent space, samples with global representativeness are preferentially selected, thus ensuring comprehensive coverage of the data distribution. In terms of uncertainty, the present invention focuses on difficult samples near the decision boundary, strengthening the refined segmentation ability of the model when dealing with complex boundaries and fuzzy anatomical structures in medical images. Through the effective combination of these two strategies, the present invention achieves comprehensiveness and pertinence in sample selection, providing a solid foundation for improving the performance of active learning methods.
[0101] Existing active learning methods usually rely on prediction probabilities, such as prediction entropy or pseudo-label loss, in uncertainty assessment. However, the phenomenon of "overconfidence" is prevalent in the neural network training process, that is, high confidence is given to mispredicted samples, resulting in sample selection errors. This problem is particularly significant in medical images with complex anatomical structures and fuzzy boundaries. In response, the present invention proposes an uncertainty active learning method that combines gradient and model entropy divergence to mitigate the impact of overconfidence in neural networks. First, calculate the entropy loss of each image and its gradient with respect to the model parameters, and quantify the importance of samples for updating the model parameters through the norm of the gradient. By using the entropy loss of the image, the dependence on pseudo-labels is reduced. Second, for pixels with high current prediction confidence (low entropy), perturb the model parameters by adding Gaussian noise multiple times to verify the stability of their prediction results. Quantify the potential instability by calculating the entropy difference of low-entropy pixels before and after perturbation. Finally, fuse the gradient information and entropy difference into a comprehensive uncertainty index to accurately screen out the difficult samples that are most valuable for model optimization. Through the above technical means, the present invention effectively alleviates the problem of sample selection errors caused by overconfidence in medical image segmentation tasks, significantly improving the sample selection quality and segmentation performance of active learning.
[0102] Some existing active learning methods attempt to combine uncertainty and representativeness metrics to select samples, but lack effective constraints on the disadvantage that redundant samples are easily selected for uncertainty metrics, resulting in limited sample selection effects. In response, the present invention introduces an effective representative active learning method. By using the latent space to map all samples into a Gaussian distribution, samples that can best represent the global data distribution are preferentially selected, while avoiding redundancy with the selected samples. In terms of the combination method, the present invention adopts a progressive selection strategy: first, select a batch of candidate samples using the representativeness method to ensure coverage of the global data distribution; then, based on the uncertainty method, screen out the final training samples from the candidate samples, focusing on the difficult samples at the decision boundary. Through this sample selection framework of the progressive hybrid strategy, the present invention effectively suppresses the defect of the uncertainty method in selecting redundant samples, fully leveraging the advantages of uncertainty and representativeness methods, and improving the comprehensive performance of active learning in medical image segmentation.
[0103] Example 3:
[0104] An electronic device includes a memory, a processor, and a computer program running on the memory. When the processor executes the program, it implements the active learning method for medical image data annotation combining uncertainty and representativeness, including:
[0105] Train a variational autoencoder infoVAE on the image pool X pool to construct a latent space generation model;
[0106] Randomly select M samples from the image pool X pool and label them to construct an initial labeled image set X an , and train a segmentation model on the labeled image set X an ;
[0107] Perform T rounds of active learning loops. In each round t of the active learning loop, specifically execute the following steps:
[0108] Screen candidate samples based on a representativeness-based method;
[0109] Screen final samples based on an uncertainty-based method;
[0110] Update the labeled and unlabeled data sets;
[0111] Retrain the segmentation model on the updated labeled image set X an to optimize the model parameters;
[0112] Obtain the final model parameters.
[0113] Example 4:
[0114] A computer-readable storage medium stores a computer program. When the program is executed by a processor, it implements the active learning method for medical image data annotation combining uncertainty and representativeness, including:
[0115] Train a variational autoencoder infoVAE on the image pool X pool to construct a latent space generation model;
[0116] Randomly select M samples from the image pool X pool and label them to construct an initial labeled image set X an , and train a segmentation model on the labeled image set X an ;
[0117] Perform T rounds of active learning loops. In each round t of the active learning loop, specifically execute the following steps:
[0118] Screen candidate samples based on representative methods;
[0119] Screen final samples based on uncertainty methods;
[0120] Update the labeled and unlabeled datasets;
[0121] On the updated labeled image set X an Retrain the segmentation model and optimize the model parameters;
[0122] Obtain the final model parameters.
[0123] Those skilled in the art should understand that the above-mentioned modules or steps of the present disclosure can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present disclosure is not limited to any specific combination of hardware and software.
[0124] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0125] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present disclosure.
Claims
1. An active learning method for labeling medical image data combining uncertainty and representativeness, characterized in that: The following steps are involved: In Image Pool X pool Train a variational autoencoder infoVAE to build a latent space generation model; From Image Pool X pool Randomly select M samples from and annotate them to construct the initial annotated image set X an , in the labeled image set X an Train the segmentation model on Perform T rounds of active learning cycles. In each round of active learning cycle, perform the following steps: Screen candidate samples based on representative methods; The final sample was selected based on uncertainty-based methods; Update labeled and unlabeled datasets; In the updated labeled image set X an Retrain the segmentation model and optimize the model parameters; Get the final model parameters.
2. The active learning method for medical image data annotation combining uncertainty and representativeness according to claim 1, characterized in that: The variational autoencoder infoVAE includes an encoder and a decoder, wherein the encoder aims to model the latent structure of the input medical imaging data and normalize it into a prior distribution; The decoder goal is to use the latent variables z to reconstruct samples that are as close as possible to the input medical image data.
3. The active learning method for medical image data annotation combining uncertainty and representativeness according to claim 1, characterized in that: The specific implementation method of screening candidate samples based on representativeness is as follows: Using the variational autoencoder infoVAE, the image pool X pool and the labeled image set X an Encode the samples and map them to the latent space and normalized to a multivariate Gaussian distribution; For each sample x i ∈X pool , the infoVAE encoder outputs the distribution parameters of the sample in the latent space through the learned parameters φ, including the mean vector and the diagonal covariance matrix Specifically: where q φ (z|x i ) is the sample x i distribution in the latent space; then, X pool Integrate the potential distribution of all samples in and fit X pool Global distribution of: where μ pool and Z pool The mean and variance of the samples in X describe pool The global distribution characteristics of The labeled dataset X an The samples in are projected into the latent space Z an , and obtain its global distribution: where μ an and Z an The mean and variance of the samples in ; From X non =X pool \X an Select R t samples.
4. The active learning method for labeling medical image data combining uncertainty and representativeness according to claim 3, characterized in that: For each unlabeled sample x i ∈X non , and its representative score is obtained by the following formula: Among them, q φ (z|x i ,X pool ) represents the sample x i For the complete dataset X pool The representativeness of q φ (z|x i ,X an ) measure sample x i For the labeled dataset X an Redundancy; Using Gaussian distribution and Get each sample x i The distribution value q φ (z|x i ,X): Where erf(x) is the error function; μ X and σ X are the mean and standard deviation of the Gaussian distribution of data set X, respectively.
5. The active learning method for medical image data annotation combining uncertainty and representativeness according to claim 1, characterized in that: The specific implementation method of screening the final sample based on the uncertainty method is: First, a sample x to be screened i Input the current segmentation network f(x;θ) and get the output result O=f(x i ;θ); Then, for the sample x i Get the entropy loss of each pixel and take the average of all pixel entropy losses as the entire image x i The entropy loss of ; the entropy loss is obtained as follows: Among them, p j Represents sample x i The jth pixel in P(y c |p j ) represents pixel p j Belongs to category y c The predicted probability of , C is the total number of categories; Next, obtain the gradient of the entropy loss of the entire image relative to the parameters of the last layer of the segmentation network, and calculate its L2 norm, which is recorded as the uncertainty index 1.
6. The active learning method for medical image data annotation combining uncertainty and representativeness according to claim 5, characterized in that: Based on the current model parameter θ, add K times Gaussian noise Generate perturbed model parameters θ k =θ+∈ k , where σ is the noise intensity hyperparameter; for each perturbed model f(x; θ k ), for candidate sample x i Perform forward propagation and get the output O k =f(x i θ k ); Then, the difference in the output entropy value of the pixels in the low entropy pixel set L before and after the disturbance is obtained to obtain the entropy difference set {Δ1, Δ2, …, Δ K }; Finally, by calculating the L2 norm of the entropy difference set, the uncertainty index 2 is obtained, as shown in the following formula U(x i )=||{Δ1,Δ2,…,Δ K }||2 This indicator quantifies the instability of low entropy pixels under perturbations, reflecting whether these pixels have overconfidence problems. The uncertainty index 1 and uncertainty index 2 are combined as sample x i After all the candidate samples have their uncertainty scores calculated, R t The candidate samples are sorted from high to low according to the uncertainty scores, and the top S are selected. t samples as the final labeled sample set.
7. The active learning method for medical image data annotation combining uncertainty and representativeness according to claim 1, characterized in that: The method for updating labeled and unlabeled datasets is: For the selected S t samples and obtain their true labels y; these newly labeled samples are extracted from the unlabeled dataset X non Removed from the image set X and added to the labeled image set X an ; Updated labeled dataset X an Includes all labeled samples for subsequent model training, while the unlabeled dataset X non The remaining unlabeled samples are retained to provide candidate sets for the next round of active learning.
8. An active learning system for labeling medical image data combining uncertainty and representativeness, characterized in that: include: Latent space generation model, in image pool X pool Train a variational autoencoder infoVAE on ; Initial model training module, from image pool X pool Randomly select M samples from and annotate them to construct the initial annotated image set X an , in the labeled image set X an Train the segmentation model on The active learning module performs T rounds of active learning cycles. In each round of t active learning cycles, it includes: Screen candidate samples based on representative methods; The final sample was selected based on uncertainty-based methods; Update labeled and unlabeled datasets; In the updated labeled image set X an Retrain the segmentation model and optimize the model parameters; Parameter acquisition module, obtains the final model parameters.
9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the active learning method for medical image data annotation combining uncertainty and representativeness is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the active learning method for labeling medical imaging data combining uncertainty and representativeness is implemented.
Citation Information
Patent Citations
An active learning method for semantic segmentation based on superpixel feature representation learning
CN117173701B
Cited By
Graph-guided data annotation treatment method for curved surface shell defect detection
CN120726425A
Multi-dimensional anti-fraud and risk control auditing method and system for large transaction
CN120996940A
Crack detection model training method and device based on dynamic multi-strategy active learning
CN121121343A
Layered simulation geological modeling method under constraint of anisotropic tensor and uncertainty
CN121479912A
Layered simulation geologic modeling method under anisotropic tensor and uncertainty constraint
CN121479912B