An image annotation method, device, electronic device, and medium for medical image model training

The method optimizes medical image model training by unsupervised clustering and selective annotation of representative samples, addressing inefficiencies in traditional methods and reducing costs.

CN119964739BActive Publication Date: 2025-07-15SHANGHAI PANORAMIC MEDICAL IMAGING DIAGNOSIS CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510445296.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-15
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the training of existing medical imaging models, repeated labeling of similar cases in lung nodules samples leads to waste of resources and insufficient labeling of special cases, which increases the workload and cost of manual labeling.

Method used

Through unsupervised classification, image samples are divided into multiple feature categories, representative samples are selected for manual annotation, and multi-scale feature extraction and density adaptive hierarchical clustering algorithm are used to dynamically determine the number of clusters, and optimize the annotation strategy with self-supervised pre-training models.

Benefits of technology

The number of manual labeling samples is reduced, the labeling efficiency is improved, the model development cycle is shortened, the cost is reduced, and the generalization ability and diagnostic accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964739B_ABST
    Figure CN119964739B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical image recognition. There is provided an image annotation method, apparatus, electronic device, and medium for training a medical image model. The method includes: obtaining a set of medical image samples to be annotated; performing unsupervised classification on the image samples in the set of medical image samples and dividing them into multiple feature categories, where each feature category contains samples with similar image features; selecting at least one representative sample from each feature category according to a preset strategy; only manually annotating the representative samples; and training a medical image recognition model based on the annotated representative samples. By means of the methods of unsupervised classification and representative sample selection, the present invention effectively improves the annotation efficiency of medical image samples, optimizes the allocation of annotation resources, thereby accelerating the model development process and reducing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image recognition, and particularly relates to an image annotation method, device, electronic device, and medium for training a medical image model. Background Art

[0002] In the field of medical image processing, especially in the recognition and analysis of pulmonary nodules, model training is a key link to improve the diagnostic accuracy. Traditional model training methods rely on a large amount of accurately annotated image data. Taking pulmonary nodule detection as an example, the current practice is usually to comprehensively annotate all provided image samples, and then use these annotated data to train the recognition model. However, this method has encountered significant challenges in actual operation: on the one hand, there are generally a large number of similar cases in pulmonary nodule samples, which are highly consistent in morphological features and have a high repeatability in contributing to the model training; on the other hand, the annotation of a few special or complex cases is crucial for improving the generalization ability of the model, but under the comprehensive annotation strategy, these special cases may be submerged in a large number of similar samples, resulting in inefficient use of annotation resources.

[0003] In addition, the practice of comprehensive annotation also greatly increases the workload of manual annotation, not only prolonging the model development cycle but also increasing the cost. Therefore, how to effectively reduce unnecessary repeated annotation while ensuring the model training effect has become an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides an image annotation method, device, electronic device, and medium for training a medical image model, aiming to achieve a more efficient and targeted annotation strategy through the classification of image samples in the preprocessing stage to optimize the model training process.

[0005] The present invention provides an image annotation method for training a medical image model, including:

[0006] S1. Obtain a set of medical image samples to be annotated;

[0007] S2. Perform unsupervised classification on the image samples in the set of medical image samples, and divide them into multiple feature categories, where each feature category contains samples with similar image features;

[0008] S3. Select at least one representative sample from each feature category according to a preset strategy;

[0009] S4. Only perform manual annotation on the representative samples;

[0010] S5. Train a medical image recognition model based on the annotated representative samples.

[0011] An image annotation method for medical image model training provided by the present invention. In step S2, the unsupervised classification includes:

[0012] S21. Use a multi-scale feature extraction model to extract the global feature vector and local feature vector of the image sample; the multi-scale feature extraction model is obtained through self-supervised pre-training; the global feature is used to capture the macroscopic features of the image sample, and the local feature is used to capture the detailed features of the image sample;

[0013] S22. Perform weighted fusion on the global feature vector and local feature vector of each image sample to generate a fused feature vector;

[0014] S23. Group the fused feature vectors using a density-adaptive hierarchical clustering algorithm to obtain multiple feature categories.

[0015] According to an image annotation method for medical image model training provided by the present invention, the weighted fusion includes: assigning higher weights to the feature parts with sparsity higher than the first threshold in the global feature vector and the local feature vector.

[0016] According to an image annotation method for medical image model training provided by the present invention, the grouping of the fused feature vectors using a density-adaptive hierarchical clustering algorithm includes:

[0017] The number of categories of the clustering algorithm is dynamically determined according to the complexity of the sample feature distribution; the complexity is obtained by the ratio of the coverage radius R of the samples in the feature space to the within-class distance D;

[0018] Among them, the calculation formula of the coverage radius R is as follows:

[0019]

[0020] Among them, C represents the sample set; is the sample and the distance in the feature space;

[0021] The calculation formula of the within-class distance D is as follows:

[0022] .

[0023] According to an image annotation method for medical image model training provided by the present invention, in step S3, the preset strategy includes:

[0024] For the feature categories with the number of samples greater than the second threshold, representative samples are proportionally extracted; for the feature categories with the number of samples less than or equal to the second threshold, all samples are extracted.

[0025] An image annotation method for medical image model training provided by the present invention, the preset strategy includes:

[0026] a) Determine the local density of samples within each of the feature categories;

[0027] b) Use a pre-trained initial model to determine the uncertainty of unlabeled samples within each of the feature categories;

[0028] c) Determine the comprehensive weight of a sample according to the local density and uncertainty of the sample;

[0029] d) For a category with the number of samples greater than the second threshold, extract a preset proportion of samples from high to low according to the comprehensive weight; for a category with the number of samples less than or equal to the second threshold, only extract samples with uncertainty higher than the second threshold; label all the extracted samples;

[0030] e) Update the initial model based on the labeled samples, and iteratively execute steps a)-e) until convergence.

[0031] An image annotation method for medical image model training provided by the present invention, after the step S5, further includes:

[0032] Use the trained medical image recognition model to predict the remaining unlabeled samples to determine the confidence of each sample, and add the samples with confidence lower than the third threshold to the manual annotation queue.

[0033] The present invention also provides an image annotation device for medical image model training, including:

[0034] An acquisition module, configured to acquire a set of medical image samples to be annotated;

[0035] A division module, configured to perform unsupervised classification on the image samples in the set of medical image samples, and divide them into multiple feature categories, where each feature category contains samples with similar image features;

[0036] A selection module, configured to select at least one representative sample from each feature category according to a preset strategy;

[0037] A manual annotation module, configured to perform manual annotation only on the representative samples;

[0038] A training module, configured to train a medical image recognition model based on the labeled representative samples.

[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image annotation method for medical image model training as described in any one of the above.

[0040] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image annotation method for medical image model training as described in any one of the above.

[0041] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the image annotation method for medical image model training as described in any one of the above.

[0042] The image annotation method, device, electronic device, and medium for medical image model training provided by the present invention have the following advantages:

[0043] 1. By unsupervised classification, the image samples are divided into multiple feature categories, and representative samples are selected from each category for annotation, greatly reducing the number of samples that need to be manually annotated, thus significantly improving the annotation efficiency.

[0044] 2. It can ensure that at least one sample in each feature category is annotated, while avoiding repeated annotation of a large number of similar samples. This enables the annotation resources to be more concentrated on annotating a few special or complex cases, which helps to improve the generalization ability of the model.

[0045] 3. Due to the reduction of the annotation workload, the annotation data required for model training can be prepared more quickly, thus shortening the model development cycle.

[0046] 4. The improvement of the annotation efficiency and the optimized allocation of resources mean that a large amount of manual annotation costs can be saved, which is of great significance for the practical application in the field of medical image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic flowchart of the image annotation method for medical image model training provided by the present invention;

[0048] Figure 2 It is a schematic structural diagram of the image annotation device for medical image model training provided by the present invention;

[0049] Figure 3 It is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Figure 1 The following is a schematic flowchart of an image annotation method for medical image model training provided by the present invention, including the following steps:

[0052] S1. Obtain a medical image sample set to be annotated.

[0053] Specifically, collect and organize a certain number of medical image samples, which will be used for subsequent unsupervised classification and model training.

[0054] S2. Perform unsupervised classification on the image samples in the medical image sample set, and divide them into multiple feature categories, where each feature category contains samples with similar image features.

[0055] Specifically, the present invention uses an unsupervised classification algorithm to group medical image samples, and the samples within each group have high similarity in image features. This can ensure that the representative samples selected from each category can better represent the features of the category.

[0056] S3. Select at least one representative sample from each feature category according to a preset strategy.

[0057] Specifically, through the preset strategy, ensure that at least one sample is selected from each feature category for annotation. The preset strategy can be formulated based on factors such as the number of samples and feature distribution to ensure that the selected representative samples can comprehensively reflect the features of each category.

[0058] S4. Manually annotate only the representative samples.

[0059] Specifically, the present invention manually annotates the selected representative samples, and the annotation content usually includes key information such as the location, size, and shape of the lesion. Since only the representative samples are annotated, the annotation workload can be greatly reduced.

[0060] S5. Train a medical image recognition model based on the annotated representative samples.

[0061] Specifically, use the annotated representative samples to train a medical image recognition model. Since these samples can better represent the feature distribution of the entire data set, the trained model usually has high accuracy and generalization ability.

[0062] The image annotation method for medical imaging model training provided by the present invention has the following advantages:

[0063] 1. By performing unsupervised classification to divide the imaging samples into multiple feature categories and selecting representative samples from each category for annotation, the number of samples that need to be manually annotated is greatly reduced, thus significantly improving the annotation efficiency.

[0064] 2. It can ensure that at least one sample in each feature category is annotated, while avoiding duplicate annotation of a large number of similar samples. This enables the annotation resources to be more concentrated on annotating a small number of special or complex cases, which helps to improve the generalization ability of the model.

[0065] 3. Due to the reduction of the annotation workload, the annotation data required for model training can be prepared more quickly, thereby shortening the model development cycle.

[0066] 4. The improvement of the annotation efficiency and the optimized allocation of resources mean that a large amount of manual annotation costs can be saved, which is of great significance for the practical application in the field of medical image processing.

[0067] In another embodiment of the present invention, the unsupervised classification includes:

[0068] S21. Using a multi-scale feature extraction model to extract the global feature vector and local feature vector of the imaging sample; this multi-scale feature extraction model is obtained through self-supervised pre-training; the global feature is used to capture the macroscopic features of the imaging sample, and the local feature is used to capture the detailed features of the imaging sample.

[0069] Specifically, in medical images, such as CT images of pulmonary nodules, they usually contain information at different levels. Global features may involve the position and overall shape of the nodules, while local features may include details such as edge texture and calcification points. For a CT image containing a pulmonary nodule, the multi-scale model captures the edge and texture information (local features) of the CT image through its shallow network, and captures the macroscopic features (global features) of the CT image through its deep network. Finally, the local features and global features are converted into corresponding feature vectors for fusion. For example, taking the classification of CT images of pulmonary nodules as an example, assume that the input CT image contains an isolated pulmonary nodule with a diameter of about 5 mm, whose edge is blurred and there are calcification points inside. In the embodiments of the present invention, the deep convolutional layer of the multi-scale feature extraction model (such as the last layer of ResNet-50) extracts the high-dimensional features of the entire CT image (i.e., global features) to capture the overall shape of the nodule (such as round, lobulated) and its position distribution in the lung lobe. The shallow convolutional layer of the model (such as the middle layer of ResNet-50) extracts the detailed features of the nodule edge region (i.e., local features), such as the texture of the calcification points and the degree of edge blurring. Finally, using the skip connection of the U-Net architecture, the global feature vector and the local feature vector are concatenated to achieve multi-scale fusion, and finally a fused feature vector is obtained.

[0070] In medical images, the global morphology of the lesion (such as nodule size, position) and local details (such as edge spiculation, internal density) are equally important for classification, but single-scale features are difficult to capture both at the same time, resulting in inaccurate classification. To solve this problem, this application distinguishes nodule and non-nodule regions through global features, identifies key details such as calcification and spiculation through local features, and finally through multi-scale fusion, enables the model to learn both macroscopic and microscopic features of the lesion, avoiding classification bias caused by a single scale.

[0071] In addition, since supervised learning relies on a large amount of labeled data, and the cost of medical image annotation is extremely high, and the samples of rare features (such as calcified nodules) are insufficient. Self-supervised pre-training is an unsupervised learning method, and its core idea is to automatically generate supervision signals (pseudo-labels) from the data itself by designing proxy tasks, enabling the model to learn useful feature representations for downstream tasks (such as classification, detection) without manual annotation of data. Therefore, the present invention adopts the method of self-supervised pre-training to learn general features using unlabeled data, such as making the model focus on the anatomical structure invariance (such as the shape of the lung lobe) and local perturbation sensitivity (such as changes in calcification points) of medical images through contrastive learning.

[0072] In an embodiment of the present invention, self-supervised pre-training can use a contrastive learning framework (such as SimCLR) to perform data enhancement (such as random cropping, rotation, and grayscale transformation) on unlabeled CT images to generate positive sample pairs (different enhanced versions of the same image) and negative sample pairs (enhanced versions of different images). The contrastive learning framework can shorten the feature distance of the positive sample pairs and push the distance of the negative sample pairs. Since medical image annotation requires the participation of professional doctors and is extremely costly, the present application only requires unlabeled data to learn common features through self-supervised pre-training, which greatly reduces the amount of annotation required for subsequent classification. For example: in the classification of lung nodules, traditional supervised learning requires the annotation of 10,000 images, while after self-supervised pre-training, only 1,000 images need to be labeled to achieve the same accuracy. Therefore, in the case of scarce medical image data annotation, diverse lesion morphology, and uneven data distribution, the use of self-supervised pre-training can make full use of a large amount of unlabeled data, alleviate the above problems, and thus reduce the cost of annotation.

[0073] The image annotation method for medical imaging model training provided by the present invention, on the one hand, reduces the dependence on annotation data through self-supervised pre-training; on the other hand, through multi-scale feature fusion, it can simultaneously capture the overall morphology and micro-texture differences of nodules, reducing the misclassification of similar samples; thus, through the combination of the two, the annotation cost is significantly reduced.

[0074] S22, performing weighted fusion on the global feature vector and the local feature vector of each image sample to generate a fused feature vector.

[0075] S23. Grouping the fused feature vectors using a density-adaptive hierarchical clustering algorithm to obtain multiple feature categories.

[0076] Specifically, the density-adaptive hierarchical clustering algorithm combines the advantages of density clustering and hierarchical clustering, can handle data distributions of different densities, and generate hierarchical clustering results. It does not require a preset number of clusters, but extracts the optimal clustering by analyzing the stability of the data. The core of the algorithm is: first define the core distance and mutual reach distance of the sample, then construct the minimum spanning tree of the mutual reach distance graph, and generate a hierarchical clustering structure through pruning. Finally, the optimal plane clustering is extracted based on stability analysis.

[0077] The image annotation method for medical image model training provided by the present invention obtains multiple feature categories by grouping fused feature vectors using a density-adaptive hierarchical clustering algorithm, thereby avoiding the limitation of the number of preset categories in traditional clustering and effectively improving the classification accuracy.

[0078] Furthermore, the global feature vector and the local feature vector of each image sample are weighted and fused to generate a fused feature vector, including: assigning higher weights to the feature parts with sparsity higher than the first threshold in both the global feature vector and the local feature vector.

[0079] Specifically, the so-called sparsity refers to that in the feature space, the number of non-zero elements of the feature vector is relatively small or the distribution of the eigenvalues is relatively sparse. Since the features with high sparsity often contain more unique or key information, which may be more important for distinguishing different lesions or conditions. Therefore, in the embodiments of the present invention, higher weights are assigned to the local regions with sparsity higher than the first threshold, so that the feature vectors of the local lesion regions with higher sparsity can be assigned greater weights. Suppose there is a lung CT image sample, which contains an obvious lung nodule lesion. First, the global feature map is extracted from the entire CT image, and this feature map may contain information such as the overall shape, size, position, and texture of the lungs; then, the local feature vector is extracted for the lung nodule lesion region. This feature vector may contain detailed information such as the size, shape, density, and edge features of the nodule. Suppose the sparsity of the global feature map is low (i.e., it contains more non-zero elements or the eigenvalue distribution is relatively dense), while the sparsity of the local lesion region feature vector is high (i.e., it contains fewer non-zero elements or the eigenvalue distribution is relatively sparse) and higher than the first threshold, then according to the principle of dynamically adjusting the weights based on sparsity in this application, a higher weight will be assigned to the local lesion region feature vector. Finally, the global feature map and the local lesion region feature vector are weighted and fused according to the adjusted weights to generate a fused feature vector. This fused feature vector contains both the global overall information and emphasizes the key features of the local lesion region.

[0080] The image annotation method for medical image model training provided by the present invention can more accurately capture and analyze the key information in medical images by performing feature fusion in the way of assigning higher weights to the feature parts with sparsity higher than the first threshold in both the global feature vector and the local feature vector, thereby improving the accuracy and reliability of diagnosis. At the same time, this method also has a certain degree of flexibility and generality and can adapt to different medical image tasks and data sets.

[0081] Furthermore, the following introduces how to group the fused feature vector using a density-adaptive hierarchical clustering algorithm, specifically including: the number of categories of the clustering algorithm is dynamically determined according to the complexity of the sample feature distribution; the complexity is obtained by the ratio of the coverage radius of the samples in the feature space to the intra-class distance.

[0082] Specifically, based on the fused feature vectors, the present invention uses a density - adaptive hierarchical clustering algorithm to group samples. The number of clusters in the clustering algorithm is not preset, but is dynamically determined according to the complexity of the sample feature distribution. This complexity is quantified by calculating the ratio of the covering radius R of the samples in the feature space to the within - class distance D. This method can automatically adapt to the feature distributions of different data sets, thus obtaining more reasonable and accurate clustering results.

[0083] Among them, the calculation formula for the covering radius R is as follows:

[0084]

[0085] Where C represents the sample set; is the sample and distance in the feature space.

[0086] The within - class distance D is the distance between samples within the same class. Its calculation formula is as follows:

[0087]

[0088] The complexity ratio C = R / D; the higher the ratio, the more complex the distribution of the samples in the feature space.

[0089] The present invention dynamically determines the number of clusters of the clustering algorithm according to the calculated complexity. The higher the complexity, the more clusters may be required to better capture the details and differences of the data. For example, suppose there is a set of fused feature vectors of medical image samples, and these samples contain different types of lesions (such as tumors, inflammations, etc.). First, global and local features are extracted from each sample, and fused feature vectors are generated. Then, the density - adaptive hierarchical clustering algorithm is applied to cluster the fused feature vectors. The algorithm automatically adjusts the clustering strategy according to the density distribution of the samples in the feature space. During the calculation process, the algorithm calculates the covering radius and within - class distance of each sample, and calculates their ratio to quantify the complexity. According to the calculated complexity, the algorithm automatically determines the optimal number of clusters. For example, if the distribution of the samples in the feature space is very complex (i.e., the ratio is high), the algorithm may choose a larger number of clusters to capture this complexity. Finally, the clustering results are analyzed to observe the features and differences between different classes. These classes may represent different types of lesions or conditions, thus providing useful information for subsequent medical analysis and diagnosis.

[0090] The image annotation method for medical image model training provided by the present invention can automatically adapt to the feature distributions of different data sets and obtain more reasonable and accurate clustering results by dynamically determining the number of clusters according to the complexity of the sample feature distribution. This is of great significance in fields such as medical image analysis, which can help doctors better understand and analyze the condition, and improve the accuracy and efficiency of diagnosis.

[0091] Further, the preset strategy provided by the embodiments of the present invention will be introduced in detail below, which specifically includes the following steps:

[0092] a) Determine the local density of samples within each feature category.

[0093] Specifically, for the samples within each feature category, calculate their local density in the feature space. The local density can reflect the degree of aggregation of samples in the feature space, and samples with high density may be more representative. In the embodiments of the present invention, the K-nearest neighbor algorithm can be used to count the number of neighbors within the radius around each sample to determine the local density.

[0094] b) Use the pre-trained initial model to determine the uncertainty of unlabeled samples within each feature category.

[0095] Specifically, uncertainty is an indicator to measure the prediction confidence of the model. Samples with high uncertainty may contain more new information and are more valuable for model training. In the embodiments of the present invention, the pre-trained initial model (such as a self-supervised pre-trained model) is used to predict the unlabeled samples, and calculate their prediction confidence (such as Softmax entropy or Monte Carlo Dropout variance). Among them, the lower the confidence, the higher the uncertainty.

[0096] c) Determine the comprehensive weight of the samples according to the local density and uncertainty of the samples.

[0097] Specifically, combine the local density and uncertainty to assign a comprehensive weight to each sample, which helps to comprehensively consider multiple factors when extracting samples. Among them, the comprehensive weight The calculation formula is as follows:

[0098]

[0099] Among them, and are adjustable hyperparameters used to balance the influence of density and uncertainty.

[0100] d) For the categories with the number of samples greater than the second threshold, extract a preset proportion of samples from high to low according to the comprehensive weight; for the categories with the number of samples less than or equal to the second threshold, only extract the samples with uncertainty higher than the second threshold.

[0101] Specifically, according to the comprehensive weight sort in descending order, and extract the top high-weight samples; if the uncertainty of the sample is higher than the second threshold, all are labeled; otherwise, only the samples with high uncertainty are labeled. In the embodiment of the present invention, for the category with the number of samples greater than the second threshold, a preset proportion of samples are extracted from high to low according to the comprehensive weight; for the category with the number of samples less than or equal to the second threshold, only the samples with uncertainty higher than the second threshold are extracted: this rule takes into account both the number of samples and the quality and representativeness of the samples.

[0102] e) Update the initial model based on the labeled samples, and iteratively execute steps a)-e) until convergence.

[0103] Specifically, add the labeled samples to the training set to update the initial model, repeat steps a)-e) for multiple rounds of iteration, gradually optimize the sample selection, and finally select the required representative samples.

[0104] In each round of iteration, in the embodiment of the present invention, a queue of samples to be labeled is extracted from each feature category according to the dynamic weight (local density + uncertainty). For example: sort by weight and extract the top k% samples (such as the top 20% with the highest weight) as high-density categories; and extract the samples with uncertainty higher than the second threshold (such as the top 30% of entropy values) as low-density categories. Then, the samples to be labeled are manually labeled (such as marking the benign and malignant of pulmonary nodules) to form labeled samples. Finally, these labeled samples are used to update the parameters of the model, and steps a)-e) are re-executed based on the updated model to generate a new sampling strategy. In this way, the labeled samples in each round are the most informative samples dynamically selected based on the current model state and sample distribution.

[0105] Since the initial model may have low prediction confidence (high uncertainty) for some samples (calcified nodule samples), and the labeled model learns to recognize calcification features, thereby reducing the uncertainty of similar nodules in subsequent rounds. At the same time, through iterative feedback, the model gradually focuses on difficult samples (high uncertainty) and key samples (low-density regions), avoiding redundant labeling of redundant samples.

[0106] The image annotation method for medical image model training provided by the present invention determines the local density and uncertainty of samples within each feature category, and extracts samples based on the comprehensive weight corresponding to the local density and uncertainty of the samples, so that the present invention can reduce the annotation of samples with high density and low uncertainty (such as a large number of similar nodules); give priority to the annotation of samples with low density and high uncertainty (such as rare calcified nodules), and through model iterative update and weight adjustment, form a "labeling-training-reselection" closed loop, adaptively optimize the annotation set, and achieve the goal of only annotating the samples most valuable for model evolution, thereby reducing the manual annotation cost.

[0107] Further, after training the medical image recognition model based on the labeled representative samples, the embodiments of the present invention further include: using the trained medical image recognition model to predict the remaining unlabeled samples to determine the confidence of each sample, and adding the samples with confidence lower than the third threshold to the manual annotation queue.

[0108] Specifically, the present invention uses the trained model to predict unlabeled medical image samples to generate pseudo-labels (i.e., model prediction results). For example: the model predicts an unlabeled CT image as "malignant lung nodule" with a confidence of 0.85. Then, according to the third threshold (such as 0.9), the pseudo-labeling results are filtered, that is: for samples with confidence ≥ 0.9, directly adopt the pseudo-labels and add them to the training set; for samples with confidence < 0.9, add them to the manual annotation queue and wait for manual review and annotation. After the samples in the manual annotation queue are annotated, they are added to the training set together with the high-confidence pseudo-labeled samples; the model is retrained using the updated training set, thus forming a closed loop of "training - pseudo-labeling - manual verification".

[0109] The image annotation method for medical image model training provided by the present invention forms a closed loop of "training - pseudo-labeling - manual verification" by retraining the model using the updated training set, enabling the manual annotation resources to be concentrated on key samples and avoiding waste on simple samples, thereby further reducing the manual annotation cost.

[0110] Next, the image annotation device for medical image model training provided by the present invention will be described. The image annotation device for medical image model training described below can be mutually corresponding and referred to with the image annotation method for medical image model training described above.

[0111] Figure 2 is a schematic structural diagram of the image annotation device for medical image model training provided by the present invention, as Figure 2 shown, the device includes:

[0112] An acquisition module 201, configured to acquire a set of medical image samples to be annotated;

[0113] A division module 202, configured to perform unsupervised classification on the image samples in the set of medical image samples and divide them into multiple feature categories, where each feature category contains samples with similar image features;

[0114] A selection module 203, configured to select at least one representative sample from each feature category according to a preset strategy;

[0115] A manual annotation module 204, configured to perform manual annotation only on the representative samples;

[0116] A training module 205 for training a medical image recognition model based on the labeled representative samples.

[0117] Figure 3 An exemplary physical structure diagram of an electronic device is shown as Figure 3 shown. The electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 310 may call the logical instructions in the memory 330 to execute an image annotation method for medical image model training. The method includes:

[0118] S1. Obtain a set of medical image samples to be annotated;

[0119] S2. Perform unsupervised classification on the image samples in the set of medical image samples, and divide them into multiple feature categories, where each feature category contains samples with similar image features;

[0120] S3. Select at least one representative sample from each feature category according to a preset strategy;

[0121] S4. Manually annotate only the representative samples;

[0122] S5. Train a medical image recognition model based on the labeled representative samples.

[0123] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disc that can store program codes.

[0124] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image annotation method for medical image model training provided by the above-mentioned various methods. The method includes:

[0125] S1. Obtain a set of medical image samples to be labeled;

[0126] S2. Perform unsupervised classification on the image samples in the set of medical image samples, and divide them into multiple feature categories, where each feature category contains samples with similar image features;

[0127] S3. Select at least one representative sample from each feature category according to a preset strategy;

[0128] S4. Only perform manual annotation on the representative samples;

[0129] S5. Train a medical image recognition model based on the labeled representative samples.

[0130] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the image annotation method for medical image model training provided by the above methods. The method includes:

[0131] S1. Obtain a set of medical image samples to be labeled;

[0132] S2. Perform unsupervised classification on the image samples in the set of medical image samples, and divide them into multiple feature categories, where each feature category contains samples with similar image features;

[0133] S3. Select at least one representative sample from each feature category according to a preset strategy;

[0134] S4. Only perform manual annotation on the representative samples;

[0135] S5. Train a medical image recognition model based on the labeled representative samples.

[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image annotation method for medical image model training, characterized in that, Including: S1. Obtain a medical image sample set to be labeled; S2. Perform unsupervised classification on the image samples in the medical image sample set, and divide them into multiple feature categories, where each feature category contains samples with similar image features; S3. Select at least one representative sample from each feature category according to a preset strategy; S4. Manually label only the representative samples; S5. Train a medical image recognition model based on the labeled representative samples; In step S2, the unsupervised classification includes: S21. Use a multi-scale feature extraction model to extract the global feature vector and local feature vector of the image sample; the multi-scale feature extraction model is obtained through self-supervised pre-training; the global feature is used to capture the macroscopic features of the image sample, and the local feature is used to capture the detailed features of the image sample; S22. Perform weighted fusion on the global feature vector and local feature vector of each image sample to generate a fused feature vector; S23. Group the fused feature vectors using a density-adaptive hierarchical clustering algorithm to obtain multiple feature categories; The grouping of the fused feature vectors using a density-adaptive hierarchical clustering algorithm includes: The number of categories of the clustering algorithm is dynamically determined according to the complexity of the sample feature distribution; the complexity is obtained by the ratio of the coverage radius R of the samples in the feature space to the within-class distance D; Among them, the calculation formula of the coverage radius R is as follows: ; where C represents the sample set; is the sample and distance in the feature space; The calculation formula of the within-class distance D is as follows: 。 2. The image annotation method for medical image model training according to claim 1, wherein, The weighted fusion includes: assigning higher weights to the feature parts with sparsity higher than the first threshold in the global feature vector and the local feature vector.

3. The image annotation method for medical image model training according to claim 1, wherein The preset strategy includes: a) Determine the local density of the samples in each feature category; b) Use a pre-trained initial model to determine the uncertainty of the unlabeled samples in each feature category; c) Determine the comprehensive weight of the samples according to the local density and uncertainty of the samples; d) For the categories with the number of samples greater than the second threshold, extract a preset proportion of samples from high to low according to the comprehensive weight; for the categories with the number of samples less than or equal to the second threshold, only extract the samples with uncertainty higher than the second threshold; label all the extracted samples; e) Update the initial model based on the labeled samples, and iteratively execute steps a)-e) until convergence.

4. The image annotation method for medical image model training according to claim 1, wherein, After step S5, it further includes: Use the trained medical image recognition model to predict the remaining unlabeled samples to determine the confidence of each sample, and add the samples with confidence lower than the third threshold to the manual labeling queue.

5. An apparatus for performing the image annotation method for medical image model training according to claim 1, characterized in that, Including: An acquisition module, used to obtain a medical image sample set to be labeled; A division module, used to perform unsupervised classification on the image samples in the medical image sample set, and divide them into multiple feature categories, where each feature category contains samples with similar image features; A selection module, used to select at least one representative sample from each feature category according to a preset strategy; A manual labeling module, used to manually label only the representative samples; A training module, used to train a medical image recognition model based on the labeled representative samples.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the image annotation method for medical image model training according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image annotation method for medical image model training according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Efficient medical image marking and learning system

    CN113314205A

  • Unsupervised industrial data classification method

    CN118154985A