Medical intelligent labeling method based on adaptive density clustering and active learning
By using medical comparison encoder and adaptive weight evaluation methods in high-dimensional medical data, the problems of density estimation failure and category imbalance are solved, efficient and robust medical active learning is achieved, and sample recall and model performance are improved.
Patent Information
- Application Number
- CN202510578848.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art density estimation failure, dynamic data distribution and clustering drift, density deviation under category imbalance and information-diversity trade-off dilemma in high-dimensional data, resulting in a degradation in the performance of medical active learning models in high-dimensional data and category imbalance scenarios.
The pre-trained medical contrast encoder converts unlabeled samples into embedded vectors, and obtains high-density core sets and boundary candidate sets through local density calibration. The samples are extracted from them using an adaptive weight evaluation method for annotation. The weight parameters are dynamically adjusted to balance the amount of information and diversity, and iterative training is optimized.
It significantly improves the performance of medical active learning, especially in pulmonary nodule detection and pulmonary hypertension classification, reducing the annotation volume and improving the recall of a few categories of samples, breaking through the performance bottleneck of traditional methods under high-dimensional space and category imbalance.
Smart Images

Figure CN120451715A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence active learning, and specifically relates to a medical intelligent labeling method based on adaptive density clustering and active learning. Background Art
[0002] With the in-depth application of artificial intelligence technology in fields such as medical image analysis, disease risk prediction, and personalized treatment, the scarcity of high-quality labeled data and the inherent imbalance of medical data have become core bottlenecks restricting the clinical practical application of models. In medical scenarios, data labeling is not only costly (for example, it takes radiologists several hours to label a single 3D tumor segmentation image), but there is also significant class imbalance (for example, the proportion of positive samples for rare diseases is less than 0.1%), making it difficult for traditional supervised learning to build reliable diagnostic models. The current mainstream strategy of active learning in medicine has a dual optimization orientation. The first is the uncertainty-driven strategy, which uses entropy or confidence to filter samples with ambiguous model decisions (such as the selection of borderline lesion areas in breast cancer subtype classification), which can effectively capture diagnostic boundary cases; the second is the diversity-driven strategy, which selects representative slices covering anatomical structural variability through density clustering (such as sampling of ventricular morphological differences in cardiac MRI), ensuring that the model adapts to the characteristic distribution of different patient groups.
[0003] In recent years, active learning methods that combine clustering and diversity-driven strategies in general fields have gradually attracted attention. Density clustering can more effectively identify boundary points and representative samples in the data by measuring the density of point distribution in the area where the samples are located. In particular, methods based on density peak clustering can accurately locate the decision boundary of the data, providing key support for subsequent sample selection. In addition, the diversity-driven strategy helps to more comprehensively characterize the distribution characteristics of the data by optimizing the distribution coverage of the samples, and thus improve the representativeness of active learning sampling after combining with density clustering. On this basis, entropy, as an important indicator for measuring the amount of information, can quantify the uncertainty of the sample and provide a unified evaluation standard for sample selection. However, the current density clustering active learning has the following difficulties:
[0004] (1) Density estimation failure in high-dimensional data: Density clustering (such as DPC) relies on distance metrics between samples (such as Euclidean distance), but in high-dimensional space, all samples tend to be equidistantly distributed ("curse of dimensionality"), resulting in distortion of density calculation.
[0005] (2) Dynamic data distribution and cluster drift: During the active learning iteration process, model updates will change the decision boundary, causing the initial clustering results to not match the current data distribution, resulting in distribution drift.
[0006] (3) Density bias under class imbalance: The density of the majority class region is significantly higher than that of the minority class, which causes the samples selected by DPC to be biased towards the majority class, exacerbating the underfitting of the model for the minority class.
[0007] (4) The information-diversity trade-off dilemma: Density clustering guarantees diversity but ignores uncertainty, while entropy maximization may select outliers. Simply weighting the two will lead to Pareto frontier conflicts.
[0008] These challenges indicate that active learning based on density clustering needs to improve the uncertainty and representativeness of sampling during the iteration process to improve the accuracy of the model. This is a complex problem that requires overcoming many technical difficulties. Summary of the Invention
[0009] To solve the above problems, the present invention provides a medical intelligent labeling method based on adaptive density clustering and active learning, comprising the following steps:
[0010] S1. Obtain an unlabeled sample set, where each unlabeled sample includes unlabeled image samples and text data of the same patient;
[0011] S2. Use a pre-trained medical contrast encoder to convert all unlabeled samples into embedding vectors;
[0012] S3. Randomly select some unlabeled samples from the unlabeled dataset and label them to form a labeled dataset to train the classification model and update the unlabeled sample set;
[0013] S4. Perform local density calibration based on the embedding vector to obtain a high-density core set and boundary candidate set;
[0014] S5. Extract unlabeled samples from the high-density core set and the boundary candidate set according to the adaptive weight evaluation method, label them, and add them to the labeled dataset to update the unlabeled sample set;
[0015] S6. Use the labeled dataset to train the classification model. If the classification model converges, the training ends. Otherwise, return to step S4 for the next iteration.
[0016] S7. Use the trained classification model to classify medical images.
[0017] Beneficial effects of the present invention:
[0018] This invention significantly improves the performance of medical active learning through boundary-sensitive screening of density clustering and a dynamic balance mechanism of adaptive weights:
[0019] 1) Density clustering enhances sample representativeness: Based on density peaks, it accurately locates anatomical boundary regions (such as tumor infiltration margins). This reduces the amount of annotation required for lung nodule detection by 80% while increasing the recall rate of malignant samples by 58%, overcoming the boundary under-detection problem caused by the failure of high-dimensional distances in traditional methods.
[0020] 2) Adaptive weight optimization iterative strategy: Dynamically reconciles data coverage and model uncertainty through stage-aware weight parameters. In the classification of pulmonary arterial hypertension (positive proportion 0.08%), only 200 labeled samples can achieve a minority class recall rate of 78% (compared to 32% of traditional entropy sampling), breaking through the selection bias bottleneck of static strategies in data imbalance scenarios, and providing a new paradigm of efficient and robust active learning for medical AI. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] In view of the technical difficulties and challenges in the existing technology, the present invention intends to solve them through density clustering and adaptive weighting, wherein:
[0024] In terms of density clustering, the present invention preprocesses the data in each round of iteration based on density peak clustering to extract boundary point samples, which usually have higher information entropy.
[0025] In terms of adaptive weights, the present invention designs an adaptive weight formula method for the unlabeled sample pool after each round of iteration to dynamically select samples for iterative training when the data is constantly changing.
[0026] Specifically, the present invention provides a medical intelligent labeling method based on adaptive density clustering and active learning, such as Figure 1 As shown, the following steps are included:
[0027] S1. Obtain an unlabeled sample set, where each unlabeled sample includes unlabeled image samples and text data of the same patient.
[0028] Specifically, an unlabeled sample may be a CT image of the same patient and its corresponding pathology report.
[0029] S2. Use a pre-trained medical contrast encoder to convert all unlabeled samples into embedding vectors.
[0030] Specifically, to address the high-dimensional heterogeneity of medical data, the present invention first uses a medical contrast encoder to map the medical data into a 128-dimensional unified embedding space. The medical contrast encoder is referenced in "Medclip: Contrastive learning from unpaired medical images and text." The pre-training process for the medical contrast encoder in step S2 includes:
[0031] Using a medical contrast encoder, each unlabeled sample is mapped into a 128-dimensional embedding space, generating corresponding image and text embedding vectors. A cross-modal contrast loss function is then used to calculate the loss based on these embedding vectors, forcing matched pairs of unlabeled images and text data to be close together in the embedding space and mismatched pairs to be far apart. Compared to traditional PCA, mapping unlabeled samples into the embedding space for distance calculation reduces the distance variance between similar anatomical structures by 73%, addressing the failure of traditional Euclidean distance in high-dimensional spaces.
[0032] Among them, the cross-modal contrast loss function L cross Expressed as:
[0033]
[0034] Among them, N represents the number of unlabeled samples, represents the image embedding vector of the unlabeled image in the i-th unlabeled sample, Represents the text embedding vector of the text data in the i-th unlabeled sample, sim() represents the similarity calculation, and τ is a hyperparameter used to adjust the similarity scaling ratio in the cross-modal contrast loss function.
[0035] S3. Randomly select some unlabeled samples from the unlabeled dataset and label them to form a labeled dataset to train the classification model and update the unlabeled sample set.
[0036] S4. Perform local density calibration based on the embedding vector to obtain high-density core sets and boundary candidate sets.
[0037] Specifically, the present invention performs density peak clustering based on embedded vectors, improves traditional density peak clustering, further integrates density index and distance index to divide samples, and finally obtains two sets: a high-density core set and a boundary candidate set.
[0038] In this paper, we define a high-density core set to represent the dense core region of the data distribution, covering the main body of the data distribution. The samples it contains have high density (large ρ) and are close to the center of higher density (small δ). The boundary candidate set is defined to represent the high uncertainty region, corresponding to the classification boundary or rare patterns, and the samples it contains may be located in the density transition region (medium ρ) or far away from the center of higher density (large δ).
[0039] Based on the above, in the embedding space mapped by the medical contrast encoder, local density calibration is performed according to the embedding vector, including:
[0040] S41. For each unlabeled sample, calculate its distance to each labeled sample in the labeled dataset based on the embedding vector. Arrange all distances in ascending order and take the median of the first K distances as the bandwidth density of the unlabeled sample. Dynamic adjustment of the bandwidth density can significantly improve the sensitivity of the boundary area.
[0041] S42. According to the bandwidth density, calculate the adaptive bandwidth kernel density of each unlabeled sample, expressed as
[0042]
[0043] σ i =median(||v i -v kNN ||)
[0044] Among them, ρ(x i ) represents the i-th unlabeled sample x i Adaptive bandwidth kernel density, n represents the number of unlabeled samples, v i Represents unlabeled sample x i The embedding vector, σ i Represents the i-th unlabeled sample x i The bandwidth density of the unlabeled sample x i The median of the distances to the K nearest neighbor samples; median() is the median function; ||·|| represents the Euclidean distance (L2 norm), which is used to measure the geometric distance between two embedded vectors in the embedding space; v kNN Indicates that in the embedding space, the unlabeled sample x i The embedding vector of the kth nearest neighbor sample;
[0045] S43. For each unlabeled sample x i , find the unlabeled samples with a closer distance among the unlabeled samples with higher density than its adaptive bandwidth kernel, and calculate the distance δ between the two i ; For the unlabeled sample x with the largest adaptive bandwidth kernel density j , define δ j is an unlabeled sample xj The maximum distance to all other unlabeled samples;
[0046] S44. For each unlabeled sample x i , calculate ρ(x i ) and δ i All unlabeled samples are sorted in ascending order according to the product value, the first 30% of the unlabeled samples in the sequence form a high-density core set, and the last 15% of the unlabeled samples in the sequence form a boundary candidate set.
[0047] S5. Extract unlabeled samples from the high-density core set and boundary candidate set according to the adaptive weight evaluation method, label them, and then add them to the labeled dataset to update the unlabeled sample set.
[0048] Specifically, step S5 includes:
[0049] S51. Give density gain to the minority class unlabeled samples in the high-density core set and the boundary candidate set, and update the adaptive bandwidth kernel density of the minority class unlabeled samples in the high-density core set and the boundary candidate set.
[0050] Specifically, the public statement for updating the density gain of unlabeled samples in the minority class is:
[0051]
[0052] Among them, ρ minority (x) represents the adaptive bandwidth kernel density of the minority class unlabeled sample x after being given density gain, ρ(x) represents the adaptive bandwidth kernel density of the minority class unlabeled sample x before being given density gain, N major Represents the number of majority class samples, N minor Represents the number of minority class samples.
[0053] S52. Calculate the value score of each unlabeled sample in the high-density core set and the boundary candidate set based on the weight parameter of the current iteration.
[0054] Specifically, in the tth iteration, any unlabeled sample x in the high-density core set and the boundary candidate set a The value score S(x a ,t) is calculated as
[0055] S(x a ,t)=α(t)·Entropy(f θ (x a ))+(1-α(t))·ρ(x a )
[0056] Among them, α(t) represents the weight parameter of the tth round iteration; f θ (xa ) represents the unlabeled sample x based on the current model parameters θ a The predicted output is usually a probability distribution vector; Entropy(f θ (x a )) indicates that the θ (x a ) The information entropy of the probability distribution obtained is used to quantify the model's response to the sample x a The classification uncertainty of ρ(x a ) represents the unlabeled sample x a Adaptive bandwidth kernel density.
[0057] Specifically, the weight parameter α(t) is dynamically adjusted according to the number of iterations, and its calculation formula is:
[0058] α(t)=0.2+0.6·e -βt
[0059] Wherein, β is the calculation factor, β=0.3.
[0060] S53. Arrange all unlabeled samples in the high-density core set and the boundary candidate set in descending order of their value scores; calculate the number of samples X required for the current iteration according to a fixed ratio (i.e., multiply the fixed ratio by the current number of unlabeled samples to obtain the number of samples X required for the current iteration); extract the first b × X unlabeled samples from the high-density core set, and extract the middle (1-b) × X unlabeled samples from the boundary candidate set; label the extracted X unlabeled samples and add them to the labeled dataset.
[0061] Specifically, the sampling ratio b of unlabeled samples is different in different iteration periods. In the early iterative training, the unlabeled samples in the high-density core set are mainly obtained to quickly cover the distribution of the main data (such as the common morphology of lung nodules); a small number of unlabeled samples in the boundary candidate set are sampled to preliminarily explore potential boundary samples; in the later iterative training, the sampling ratio of unlabeled samples in the high-density core set is reduced, and the stability of the classification model for common cases is maintained by a small number of unlabeled samples in the high-density core set; the sampling ratio of unlabeled samples in the boundary candidate set is increased, and high entropy samples (such as atypical calcifications, multimodal contradictory cases) are prioritized to be labeled to optimize the sensitivity of the classification model to the boundary. In the embodiment of the present invention, it is set that when t≤3, b=0.7, when t>5, b=0.3, and when t=4 / 5, b=0.5.
[0062] This method uses value scoring as the core link between density peak clustering and active learning. Dynamic weight parameters balance information quantity and diversity, guiding the classification model to annotate the most effective samples at different learning stages. The sampling ratio is adjusted according to the iteration stage. Early iterations focus on density diversity, relying on a high-density core set for rapid modeling. Later iterations focus on uncertainty, optimizing boundaries through boundary candidate sets.
[0063] At the same time, a compensation mechanism is used to grant density gain to minority unlabeled samples. The amplified adaptive bandwidth kernel density value directly increases the value score of minority unlabeled samples, making them more likely to be selected. In breast cancer subtype classification, this compensation mechanism can increase the selection rate of LCIS (lobular carcinoma in situ) samples from 12% to 34%.
[0064] S6. Use the labeled dataset to train the classification model. If the classification model converges, the training ends. Otherwise, return to step S4 for the next iteration.
[0065] S7. Use the trained classification model to classify medical images.
[0066] In one embodiment, incremental training and validation are performed. In incremental training, the top 15% high-value unlabeled samples (about 100-200 cases) are selected from the unlabeled sample set for additional annotation in each iteration; training is stopped when the AUC improvement of the validation set is less than 0.5% for three consecutive rounds or the total number of annotations reaches a preset upper limit (e.g., 1000 cases).
[0067] Specifically, the embodiment of the present invention targets an extremely unbalanced breast cancer subtype classification scenario and is validated based on the TCGA-BRCA dataset, which contains 1,092 samples in total, of which LCIS accounts for 0.8%.
[0068] Implementation effect: The number of annotations was reduced from 800 in traditional AL to 220 (a decrease of 72.5%); the minority class recall rate LCIS was increased from 12% to 34%, and the F1-score was increased from 0.21 to 0.63; the computing efficiency was compressed from 58 minutes to 9 minutes per round of iteration.
[0069] In the present invention, unless otherwise clearly stipulated and limited, the terms "installation", "setting", "connection", "fixation", "rotation" and the like should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.
[0070] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A medical intelligent labeling method based on adaptive density clustering and active learning, characterized by: The following steps are involved: S1. Obtain an unlabeled sample set, where each unlabeled sample includes unlabeled image samples and text data of the same patient; S2. Use a pre-trained medical contrast encoder to convert all unlabeled samples into embedding vectors; S3. Randomly select some unlabeled samples from the unlabeled dataset and label them to form a labeled dataset to train the classification model and update the unlabeled sample set; S4. Perform local density calibration based on the embedding vector to obtain a high-density core set and boundary candidate set; S5. Extract unlabeled samples from the high-density core set and the boundary candidate set according to the adaptive weight evaluation method, label them, and add them to the labeled dataset to update the unlabeled sample set; S6. Use the labeled dataset to train the classification model. If the classification model converges, the training ends. Otherwise, return to step S4 for the next iteration. S7. Use the trained classification model to classify medical images.
2. A medical intelligent labeling method based on adaptive density clustering and active learning according to claim 1, characterized in that: The pre-training process of the medical contrast encoder in step S2 includes: Each unlabeled sample is mapped to a 128-dimensional embedding space through a medical contrast encoder to obtain the corresponding image embedding vector and text embedding vector. Based on the image embedding vector and text embedding vector, the loss is calculated through a cross-modal contrast loss function, forcing matched pairs of unlabeled images and text data to be close in the embedding space, and mismatched pairs of unlabeled images and text data to be far away in the embedding space.
3. The medical intelligent labeling method based on adaptive density clustering and active learning according to claim 1 is characterized in that: Perform local density calibration based on the embedding vector, including: S41. For each unlabeled sample, calculate its distance to each labeled sample in the labeled dataset based on the embedding vector, sort all distances in ascending order, and take the median of the first K distances as the bandwidth density of the unlabeled sample; S42. According to the bandwidth density, calculate the adaptive bandwidth kernel density of each unlabeled sample, expressed as σ i =median(||v i -v kNN ||) Among them, ρ(x i ) represents the i-th unlabeled sample x i Adaptive bandwidth kernel density, n represents the number of unlabeled samples, v i Represents unlabeled sample x i The embedding vector, σ i Represents the i-th unlabeled sample x i Bandwidth density, median() is the median function, ||·|| represents the Euclidean distance, v kNN Indicates that in the embedding space, the unlabeled sample x i The embedding vector of the kth nearest neighbor sample; S43. For each unlabeled sample x i , find the unlabeled samples with a closer distance among the unlabeled samples with higher density than its adaptive bandwidth kernel, and calculate the distance δ between the two i ; For the unlabeled sample x with the largest adaptive bandwidth kernel density j , define δ j is an unlabeled sample x j The maximum distance to all other unlabeled samples; S44. For each unlabeled sample x i , calculate ρ(x i ) and δ i All unlabeled samples are sorted in ascending order according to the product value, the first 30% of the unlabeled samples in the sequence form a high-density core set, and the last 15% of the unlabeled samples in the sequence form a boundary candidate set.
4. The medical intelligent labeling method based on adaptive density clustering and active learning according to claim 1 is characterized in that: Step S5 specifically includes: S51. Give density gain to the minority class unlabeled samples in the high-density core set and the boundary candidate set, and update the adaptive bandwidth kernel density of the minority class unlabeled samples in the high-density core set and the boundary candidate set; S52. Calculate the value score of each unlabeled sample in the high-density core set and the boundary candidate set based on the weight parameter of the current iteration; S53. Arrange all unlabeled samples in the high-density core set and the boundary candidate set in descending order of their value scores; calculate the number of samples X required for the current iteration according to a fixed ratio, extract the first b × X unlabeled samples from the high-density core set, and extract the middle (1-b) × X unlabeled samples from the boundary candidate set; label the extracted X unlabeled samples and add them to the labeled dataset.
5. The medical intelligent labeling method based on adaptive density clustering and active learning according to claim 4 is characterized in that: The public statement for updating the density gain of unlabeled samples in the minority class is Among them, ρ minority (x) represents the adaptive bandwidth kernel density of the minority class unlabeled sample x after being given density gain, ρ(x) represents the adaptive bandwidth kernel density of the minority class unlabeled sample x before being given density gain, N major Represents the number of majority class samples, N minor Indicates the number of minority class samples.
6. The medical intelligent labeling method based on adaptive density clustering and active learning according to claim 4 is characterized in that: In the tth iteration, any unlabeled sample x in the high-density core set and the boundary candidate set a The value score S(x a ,t) is calculated as S(x a ,t)=α(t)·Entropy(f θ (x a ))+(1-α(t))·ρ(x a ) Among them, α(t) represents the weight parameter of the tth round iteration; f θ (x a ) represents the unlabeled sample x based on the current model parameters θ a The predicted output of Entropy(f θ (x a )) indicates that the θ (x a ) The information entropy of the probability distribution obtained is used to quantify the model's response to the sample x a The classification uncertainty of ρ(x a ) represents the unlabeled sample x a Adaptive bandwidth kernel density.
7. The medical intelligent labeling method based on adaptive density clustering and active learning according to claim 6, characterized in that: The weight parameter α(t) is dynamically adjusted according to the number of iterations, and its calculation formula is: α(t)=0.2+0.6·e -βt Wherein, β is the calculation factor, β=0.
3.
8. The medical intelligent labeling method based on adaptive density clustering and active learning according to claim 4 is characterized in that: When t≤3, b=0.7, when t>5, b=0.3, when t=4 / 5, b=0.5; t is the iteration round.
Citation Information
Cited By
Target feature intelligent identification system and method based on multi-source data
CN121365269A
Interpretable active learning method cooperatively driven by cell image attributes
CN121837152A