A data labeling method and device and a disease classification model training method

By merging relevant classification labels and reducing label disagreements using entropy values ​​and greedy algorithms, the problem of label inconsistency caused by subjectivity in disease classification is solved, and the accuracy and consistency of model training is improved.

CN114140653BActive Publication Date: 2025-08-26BEIJING AIRDOC TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210004573.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-05
Publication Date
2025-08-26
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

In disease classification, due to inconsistent classification labels due to the subjectivity of different judges, it is difficult for the prior art to obtain universal label information, which affects the model training effect.

Method used

By annotating the sample data set, merging relevant classification labels, using entropy values ​​and greedy algorithms to reduce label disparity, using greedy algorithms to merge label pairs to obtain objective labels, and training convolutional neural network models.

Benefits of technology

Preprocessing of data with certain subjectivity is realized, universal labels are obtained, and the accuracy and consistency of disease classification models are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140653B_ABST
    Figure CN114140653B_ABST
Patent Text Reader

Abstract

The present invention provides a method for data labeling of a sample data set, comprising the following steps: S1, obtaining a sample data set, wherein each sample in the sample data set includes one or more classification labels respectively labeled by multiple annotators; S2, merging the label types of samples containing multiple classification labels to merge related classification label pairs and using one label in the label pair as the merged label; wherein the related classification label pair refers to a pair of different labels annotated to the same sample by different annotators; S3, re-labeling the samples in the sample data set based on the merged classification label. Compared with the existing technology, the method of the present invention can realize preprocessing of data with a certain degree of subjectivity to objectify the subjective evaluation using other relevant indicators to obtain universal labels to achieve data labeling, and then train the relevant classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, specifically, to the field of supervised machine learning in the field of artificial intelligence, and more specifically, to a data labeling method and model for multi-classification problems of data with some subjectivity, a disease classification model training method based on fundus images, and a disease classification method based on fundus images. Background Art

[0002] Supervised machine learning in the field of artificial intelligence involves training a machine using pre-existing labeled training samples. The model's output is then compared with the labeled information of the training samples, allowing the model to be trained and corrected during iterations using the existing information. Typical supervised learning problems can be categorized into regression, classification, and labeling, depending on the characteristics of the input and output.

[0003] Classification problems involve prediction problems where the output variable can take on a finite number of discrete values. Supervised learning involves learning a classification decision function from data, called a classifier, and then predicting the output for new inputs. This process is called classification. Multi-classification problems, on the other hand, involve multiple prediction categories, and are typically learned using a splitting strategy.

[0004] In machine learning, the categories set for general classification problems are relatively universal, and there are fewer disagreements. For simple classification problems, especially for common object classification, the category conclusions drawn from object classification pictures based on human cognition are roughly the same. Therefore, without refining the problem, this generalized solution is more acceptable to the public. However, for problems such as the classification of diseases and sub-classification within the classification of diseases, these are relatively detailed multi-classification problems, and different judges may have different classifications for the same object. For example, when doctors diagnose a disease based on the same specimen, different doctors may give different diagnostic results, and the diagnostic results given may be inconsistent. Taking disease diagnosis based on fundus images as an example, Figure 1 As shown, Figure 1 (a) is a "leopard-spot change," which means that clearly defined choroidal vessels can be observed around the fovea and the posterior pole vascular arch, indicating pathological myopia. Figure 1(b) is "reduced arterial elasticity," which can generally be identified by observing the reflection of the arteries, changes in diameter, and the pressure marks at the intersection of arteries and veins. These two types of diseases are somewhat related and have overlapping characteristics, making it possible for both conditions to coexist in the same patient. In comparison, these diseases are not very serious, so when diagnosing them, doctors may overlook other symptoms that may also occur, resulting in disagreements between different doctors. This shows that for classification problems with a certain degree of subjectivity or evaluation, different people may obtain different results on the same sample, resulting in disagreements. This makes it impossible for supervised learning classification problems to obtain universal label information under established or objective evaluations, making model training difficult. Therefore, how to preprocess these subjective labels is a prerequisite for solving classification problems and training universal classification models. Summary of the Invention

[0005] Therefore, the purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide a data labeling method and device as well as a disease classification model training method.

[0006] According to a first aspect of the present invention, a method for data labeling of a sample data set is provided, comprising the steps of: S1, obtaining a sample data set, wherein each sample in the sample data set includes one or more classification labels respectively labeled by multiple annotators; S2, merging the label types of samples containing multiple classification labels to merge associated classification label pairs and using one label in the label pair as the merged label; wherein the associated classification label pair refers to a paired combination of different labels annotated on the same sample by different annotators; S3, re-labeling the samples in the sample data set based on the merged classification labels.

[0007] In some embodiments of the present invention, the degree of disagreement in the annotations of a sample data set by different annotators is measured by the number of decreases in a preset target value reduction; wherein, one of the disagreement rate, zero-deleted entropy value, disagreement entropy value, total entropy value, and weight corresponding to each class label in the sample data set is used as the target value reduction. The disagreement rate corresponding to each class label is the proportion of samples with inconsistent annotations by all annotators for that class label in the total samples; the zero-deleted entropy value corresponding to each class label is the average entropy value of samples with consistent annotations by all annotators for that class label; the disagreement entropy value corresponding to each class label is the average entropy value of samples with inconsistent annotations by all annotators for that class label; and the total entropy value corresponding to each class label is the average entropy value of all samples for that class label. The degree of disagreement in the annotations of a sample data set by different annotators is measured by the number of decreases in a preset target value reduction; wherein, the average or weighted average of the disagreement rate, zero-deleted entropy, disagreement entropy, or total entropy values ​​corresponding to all class labels in the sample data set is set as the target value reduction. Preferably, the target reduction is set in the following manner: calculate the disagreement rate or zero-free entropy value or disagreement entropy value or total entropy value corresponding to each type of label, wherein the disagreement rate corresponding to each type of label is the proportion of samples with inconsistent annotations by all annotators for that type of label in the total samples; the zero-free entropy value corresponding to each type of label is the average entropy value of samples with consistent annotations by all annotators for that type of label; the disagreement entropy value corresponding to each type of label is the average entropy value of samples with inconsistent annotations by all annotators for that type of label; the total entropy value corresponding to each type of label is the average entropy value of all samples for that type of label; based on the calculated disagreement rate or zero-free entropy value or disagreement entropy value or total entropy value corresponding to each type of label, calculate the average or weighted average of the disagreement rate or zero-free entropy value or disagreement entropy value or total entropy value corresponding to all category labels in the sample data set.

[0008] The divergence rate or the zeroed entropy value or the divergence entropy value or the average value of the total entropy value is calculated as follows: Among them, W H is the average of the divergence rate or zero entropy or divergence entropy or total entropy corresponding to all category labels in the sample data set, H is the divergence rate or zero entropy or divergence entropy or total entropy corresponding to each category label, and N is the number of label categories; the weighted average of the divergence rate or zero entropy or divergence entropy or total entropy is calculated as follows: Among them, Q H It is the weighted average of the divergence rate or zero entropy or divergence entropy or total entropy corresponding to all category labels in the sample data set, and P is the sample frequency corresponding to each category label.

[0009] Preferably, in step S2, a greedy algorithm is used to iteratively merge the label types of samples containing multiple classification labels multiple times until the target value reduction is less than or equal to 0. Each sample in the sample dataset contains multiple related label pairs, and each time a merge is performed, the label pair that, after the merge, reduces the divergence of annotations of the sample dataset by different annotators the most is merged. Preferably, after merging the related label pairs, the label with the highest target value reduction is used as the merged label.

[0010] In some embodiments of the present invention, the entropy value corresponding to each sample of each class label is calculated as follows:

[0011] S=-plogp-(1-p)log(1-p)

[0012] Among them, S represents the entropy value of the current sample for the current class label, and p is the proportion of all annotators who have annotated the current sample with the current class label.

[0013] In some embodiments of the present invention, the tag pairs to be merged are selected according to a preset frequency threshold, wherein the preset frequency threshold is set to the tag pairs whose occurrences in the sample data set are in the top 50%.

[0014] According to a second aspect of the present invention, there is provided a data labeling device for implementing the method as described in the first aspect of the present invention, the device comprising: a data acquisition module for acquiring a sample data set to be processed, wherein each sample in the sample data set contains one or more classification labels respectively labeled by multiple annotators; a data processing module for counting the types of classification labels corresponding to all samples in the sample data set, merging the label types of samples containing multiple classification labels to merge related classification label pairs and using one label in the label pair as the merged label; and a data labeling module for re-labeling the samples in the sample data set based on the classification labels merged by the data processing module.

[0015] According to a third aspect of the present invention, a method for training a disease classification model based on fundus images is provided, the method comprising: F1, obtaining a fundus image dataset, each fundus image containing one or more disease classification labels annotated by multiple annotators; F2, using the fundus image dataset as a sample dataset, and annotating all fundus images in the fundus image dataset using the method described in the first aspect of the present invention to obtain a fundus image training dataset; F3, training a convolutional neural network using the fundus image training dataset in step F2 until convergence.

[0016] According to a fourth aspect of the present invention, a disease classification method based on fundus images is provided, the method comprising: P1, obtaining a fundus image to be processed; P2, performing disease classification on the image to be processed using a disease classification model trained using the method according to the third aspect of the present invention.

[0017] Compared with the existing technology, the method of the present invention can preprocess data with a certain degree of subjectivity to objectify the subjective evaluation using other relevant indicators to obtain universal labels to achieve data annotation, and then train relevant classification models. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:

[0019] Figure 1 is a schematic diagram of a fundus image according to an example of the present invention;

[0020] Figure 2 Schematic diagram of a data annotation method according to an embodiment of the present invention;

[0021] Figure 3 Schematic diagram of comparison results of average entropy value and average decrease of weight value thereof in experimental verification according to an embodiment of the present invention;

[0022] Figure 4 Schematic diagram of comparison results of zero entropy value and average decrease of weighted value verified by experiments according to an embodiment of the present invention;

[0023] Figure 5 Schematic diagram of comparison results of confirmation rate and average increase of weights thereof in experimental verification according to an embodiment of the present invention;

[0024] Figure 6 Schematic diagram of comparison results of divergence rate and average decrease of weights verified by experiments according to an embodiment of the present invention;

[0025] Figure 7 A schematic diagram of entropy comparison for verifying the effect of the present invention on different data sets according to an embodiment of the present invention;

[0026] Figure 8 Schematic diagram of the entropy weighted average comparison results of verifying the present invention and existing technologies on the first-fd16 dataset according to an embodiment of the present invention;

[0027] Figure 9 Schematic diagram of the annotation effect on the epiretinal membrane fundus image according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] As described in the background technology, when facing classification problems in which the identification content has a certain degree of subjectivity or evaluation, the existing technology cannot obtain universal labels for this type of classification problem because the same input (processing object) can have multiple labels in the judgment of a judge, and there are differences between different judges. Therefore, the present invention proposes a method for preprocessing data with a certain degree of subjectivity to objectify the subjective evaluation using other relevant indicators to obtain universal labels to achieve data labeling, and then train relevant classification models.

[0030] The technical idea of ​​the present invention can be summarized as follows:

[0031] First, the sample dataset is decomposed into label information to obtain a binary diagnostic result table for a calibrated label in the entire dataset. In this table, if a judge in the sample calibrated the label, its corresponding value is 1, otherwise it is 0. The total number of times the label appears in the entire table is counted as the sample frequency of the label.

[0032] Then, divide the statistical table data into three parts: all 1 (m1), containing both 0 and 1 (m2) and all 0 (m3), which are called all-one interval, uncertain interval and all-zero interval respectively, and calculate the proportion of each part.

[0033] Secondly, by calculating the entropy value, the entropy is used for comparison to achieve the purpose of quantification. Among them, the entropy value of the judgment result of each row in the statistical table is evaluated, and the average entropy value is calculated in the m1∪m2 interval, the m2 interval and the overall interval, and the zero entropy value, the uncertain interval entropy value e2 and the total entropy value e are obtained. They can reflect the disagreement of the judge in each input to a certain extent. If the disagreement is higher, the entropy value is larger, otherwise it is smaller. Based on the consideration of the entropy value, the present invention takes the above part of the indicators as the standard for judging the degree of disagreement, and then uses the mean and weighted average of these comparison standards as the target reduction, and merges different labels. First, the related labels are extracted, and then one of the comparison standards is selected as the target reduction standard, and the average value or weighted average of the standard is selected as the reduction indicator.

[0034] Finally, based on the greedy algorithm, the post-merger reduction of all associated label pairs is calculated. The label pair with the largest reduction is then selected as the label pair for this merge. After the merge, the combined reduction of the merged row and column is recalculated, and the merging process is repeated until the target reduction of the entire merged row and column is no greater than 0, that is, the difference is no less than 0. At this point, the greedy algorithm calculation is completed, and the merged objective label is obtained to annotate the input data.

[0035] On top of this, there are two mechanisms that can further adjust the merge: 1. Before using the greedy algorithm to select the merge pair with the largest reduction, the current sample frequency of the two merged parties is removed from the calculated target entropy value, thereby reducing the impact of sample frequency on the merged item; 2. Use multiple comparison criteria for merging, and then merge the merge pairs with the highest number of effective times, thereby reducing the impact of low-frequency samples on the merged item.

[0036] The method of the present invention is applicable to various types of data, including but not limited to image data, voice data, etc.

[0037] According to one embodiment of the present invention, Figure 2 As shown, the method for labeling sample data of the present invention includes steps S1, S2, and S3. Each step is described in detail below:

[0038] In step S1, a sample data set is obtained, wherein each sample in the sample data set includes classification labels annotated by multiple annotators, and each sample includes one or more classification labels annotated by different annotators;

[0039] In the embodiment of the present invention, fundus images are used as exemplary data, and disease diagnosis is performed through fundus images to explain the present invention in detail.

[0040] In fundus images, Did (Disease ID) is used to represent the labels of observable diseases or conditions in different fundus images, and it is used to mark fundus images to achieve a diagnostic record for a fundus image. Since the diagnostic results given by different doctors are not consistent, the diagnostic results of different doctors may differ, and it is impossible to obtain a unified label. To solve this problem, the present invention proposes an entropy-based fundus image labeling method. By merging different categories, the differences caused by labeling different classification labels are minimized. At the same time, the characteristics of the categories can still be retained after the merger, making the fundus image labeling method somewhat interpretable.

[0041] In order to better understand the present invention, the present invention takes the fundus image information shown in Table 1 as an example to explain the fundus image annotation method in detail.

[0042] Table 1

[0043]

[0044]

[0045] Among them, img_id represents the fundus image identifier, img_1, img_2, img_3, img_4, and img_5 represent 5 different fundus images respectively, A, B, C, D, and E represent different annotators (i.e., doctors), and the numbers in the table represent the annotation labels of fundus images by different annotators. Each number corresponds to a Did. Since an image may correspond to multiple annotation results, there can be multiple Dids in each square, separated by semicolons, as shown in Table 1. Table 1 contains three labels, namely Did = 1, and / or 2, and / or 3.

[0046] Preferably, firstly, Table 1 is processed, the labels are split, and the binary diagnosis result table for each label is counted. For the data in Table 1, the binary diagnosis result tables for Did=1, 2, and 3 are counted as shown in Table 2, Table 3, and Table 4.

[0047] Table 2

[0048] Did=1 A B C D E img_1 1 1 1 1 1 img_2 1 1 1 1 1 img_3 0 0 0 0 1 img_4 1 0 1 1 1 img_5 0 0 0 0 0

[0049] Table 3

[0050] Did=2 A B C D E img_1 1 0 0 1 0 img_2 0 0 0 0 0 img_3 0 0 1 0 0 img_4 0 0 1 0 0 img_5 1 1 1 1 1

[0051] Table 4

[0052] Did=3 A B C D E img_1 0 0 0 0 0 img_2 0 0 0 0 0 img_3 1 1 1 1 1 img_4 0 1 0 0 0 img_5 0 1 1 1 0

[0053] Count the total number of times different labels appear in the entire table, that is, count the number of 1s in the binary diagnosis result table of different labels, and use this as the sample frequency of the label. As shown in Tables 2, 3, and 4, the sample frequencies of Did = 1, 2, and 3 are 15, 9, and 9, respectively.

[0054] Then, based on the binary diagnosis result table, the data in each table is divided into three parts: fundus images with all 1s (m1), fundus images containing both 0s and 1s (m2), and images with all 0s (m3). These are called the all-one interval, the uncertain interval, and the all-zero interval, respectively. Taking the label Did = 1 as an example, the data in Table 2 after data division is shown in Table 5:

[0055] Table 5

[0056]

[0057] Calculate the proportion of each part of the image to the total image, marked as r1, r2 and r3, which represent the confirmed rate, uncertain rate and asymptomatic rate respectively. Taking Table 5 as an example, part m1 has two images in total, so when Did = 1 Using the same method, we obtain r2 = 0.4 and r3 = 0.2. As the names suggest, r1 is the proportion of images where all doctors agree that the disease is confirmed (Did = 1), r2 is the proportion of images where doctors disagree and the overall results make it uncertain whether the image is a diagnosis of the disease, and r3 is the proportion of images where all doctors agree that the disease is not confirmed. Similarly, when Did = 2, r1 = 0.2, r2 = 0.6, and r3 = 0.2; when Did = 3, r1 = 0.2, r2 = 0.4, and r3 = 0.4.

[0058] The idea of ​​the present invention is to achieve the purpose of objectively quantifying subjective judgments by comparing entropy values. For categories with low universality, information entropy can be used to statistically analyze the results given by different people to measure the amount of information. The so-called information entropy refers to the measure of the amount of information given to a certain output message. The smaller the possibility of a message appearing, the more information the message carries. Therefore, for messages with a small probability of appearing, the amount of information is large and the information entropy is also large. The present invention uses information entropy to measure the differences caused by the labels made by different labelers on a certain input, that is, the uncertainty of their judgments, and thereby analyzes the effectiveness of the method of merging categories. For example, in the exemplary data, since different doctors have different diagnostic result labels, when they have different diagnostic result labels for the same fundus image, they have differences, which increases the entropy value. Preferably, the entropy value of each image in each label (disease) is calculated by formula (1):

[0059] S=-plogp-(1-p)log(1-p)

[0060] Where p is the corresponding judgment probability, which is equivalent to the proportion of doctors who gave a 1 (i.e., confirmed diagnosis) for the fundus image in the binary diagnosis result table. For example, for img_3 in Table 2 above, a total of 1 doctor (E) gave a diagnosis result of Did = 1, p = 1 / 5 = 0.2, so when Did = 1, the entropy value of the fundus image img_3 is:

[0061]

[0062] Preferably, the entropy value of each doctor's judgment result for each fundus image is calculated using a formula for calculating entropy. This comparison can infer the difference in the doctors' disagreement on the Did for that image. The larger the entropy value, the greater the doctors' disagreement on the Did for that image. Based on the obtained results, the average entropy value is calculated for each image within the interval m1∪m2, the interval m2, and the overall interval, respectively, to obtain the zero-free entropy value e1, the uncertainty interval entropy value e2, and the total entropy value e. These values ​​can, to some extent, reflect the doctors' disagreement on Did. The higher the doctors' disagreement on Did, the greater the entropy value, and vice versa. Still taking the data of Did=1 in Table 2 as an example, the entropy values ​​of img_1, img_2, img_3, img_4, and img_5 are: 0, 0, 0.50040, 0.50040, and 0 respectively; therefore, the average entropy value of the interval m1∪m2 is the average of the entropy values ​​of img_1, img_2, img_3, and img_4, that is, (0+0+0.50040+0.50040) / 4=0.25020, then the zero-deleted entropy value e1 when Did=1 is 0.25020; the average entropy value of the interval m2 is The average entropy value between the two is the average of the entropy values ​​of img_3 and img_4, that is, (0.50040+0.50040) / 2=0.50040, so the uncertainty interval entropy value e2 when Did=1 is 0.50040; the total entropy value of the entire interval is the average of the entropy values ​​of img_1, img_2, img_3, img_4, and img_5, that is, (0+0+0.50040+0.50040+0) / 5=0.20016, so the total entropy value e when Did=1 is 0.20016. These entropy values ​​can reflect the disagreement of doctors on Did=1. If the disagreement of doctors on Did=1 is higher, the entropy value is larger, and vice versa. For the overall label combination, the average of its values ​​can be calculated as a measure of the disagreement of this label combination. In addition, for each obtained standard, the sample frequency of the label itself can be introduced for consideration. The specific method is to introduce the concept of weighted average, that is, the standard value obtained for each Did is multiplied by its frequency ratio. Taking r1 as an example, its average is (0.4+0.2+0.2) / 3=0.26667, and its weighted average is (0.4*15+0.2*9+0.2*9) / (15+9+9)=0.29091, and the prefix w is added, denoted as wr1. Similarly, the corresponding entropy values ​​of Did=2 and Did=3 can be obtained, and the entropy values, confirmation rate and uncertainty rate shown in Table 6 are obtained:

[0063] Table 6

[0064]

[0065]

[0066] Preferably, the present invention uses the average value of the confirmation rate r1, the uncertainty rate r2, the zero-deleted entropy value e1, the uncertainty interval entropy value e2 and the total entropy value e or the average of their corresponding weights as the standard for judging the degree of disagreement. It can be seen that the higher the confirmation rate r1, the smaller the disagreement; except for the confirmation rate r1, the smaller the values ​​of the other four indicators, the smaller the disagreement.

[0067] In step S2, the label types of the sample containing multiple classification labels are merged multiple times iteratively to merge the related classification label pairs and use one label in the label pair as the merged label, wherein the related classification label pair refers to a pair combination of different labels annotated by different annotators on the same sample.

[0068] According to one embodiment of the present invention, the present invention adopts a greedy algorithm to merge label pairs. The greedy algorithm is also called a greedy algorithm. Each time, the best choice is made according to the current state, regardless of the consequences, and the sub-problems are simplified from top to bottom. The principle of the greedy algorithm is well known in the field, and the present invention will not elaborate on it. It will only be explained from the perspective of the problem setting of the present invention and the benefit evaluation of the greedy algorithm. In the present invention, four comparison criteria (divergence rate r2, zero entropy value e1, divergence entropy value e2 and total entropy value e) and weighted average are used as target reduction values, and then five comparison criteria (plus the diagnosis rate r1) are used to measure the benefit evaluation of each fundus image annotation. Here, the target reduction value is used as the evaluation criterion used in problem processing. Since the four criteria used can give higher benefits (divergence reduction) when they are reduced, the core of the problem of the present invention is to reduce the target reduction value from the merging processing operation to the minimum, and the sub-problem is to find the best strategy for the next merging.

[0069] Still taking the fundus image exemplary data as an example, before the fundus image annotation and merging are performed, each pair of related Dids will be extracted first. The main processing benchmark is that two Dids have appeared simultaneously in the diagnostic results of the same row (not necessarily in the same grid) (such as in Table 1, 1 and 2 appeared simultaneously in the img_1 row, 1, 2, and 3 appeared simultaneously in the img_3 row, 1, 2, and 3 appeared simultaneously in the img_4 row, and 2 and 3 appeared simultaneously in the img_4 row). Only these Did pairs are merged. This is because the entropy value after merging Dids that do not appear in the same row cannot be less than the original entropy value. Based on Table 1, the Did-related pairs that may be merged (that is, the related label pairs) are 1 and 2, 2 and 3, and 1 and 3. The present invention selects one of the four comparison criteria mentioned above (divergence rate r2, zero entropy value e1, divergence entropy value e2 and total entropy value e) as the target reduction standard. Then, based on the greedy algorithm, the reduction of the reduction index of these related Dids or Did classes after merging is first calculated, and then the Did pair with the largest reduction is selected as the Did pair for this merger. For example, assuming that the divergence rate r2 is the target reduction, the reduction of all r2 after merging 1 and 2, merging 2 and 3, and merging 1 and 3 is calculated respectively. If the reduction of all r2 after merging 2 and 3 is the largest, then 2 and 3 are merged. The merged Did pair will take the Did of one of them as the label class name, and continue to exist in the mergeable row in the form of Did. Preferably, the label class name selected as the merged label pair is the one with the largest target reduction. For example, if after merging 2 and 3, the target reduction of 3 is the largest, then 3 is used as the merged label class name, and the label data shown in Table 7 is obtained.

[0070] Table 7

[0071] img_id A B C D E img_1 1;3 1 1 1;3 1 img_2 1 1 1 1 1 img_3 3 3 3 3 1,3 img_4 1 3 1;3 1 1 img_5 3 3 3 3 3

[0072] After merging, recalculate the merged reduction of the merged rows and columns, and repeat the merging steps until the target reduction of the entire merged rows and columns is no more than 0, that is, the difference is no less than 0. At this point, the calculation of the greedy algorithm is completed. Assuming that 1 and 2 are merged to 1, and 1 and 3 are merged to 1, the label data shown in Table 8 can be obtained:

[0073] Table 8

[0074] img_id A B C D E img_1 1 1 1 1 1 img_2 1 1 1 1 1 img_3 1 1 1 1 1 img_4 1 1 1 1 1 img_5 1 1 1 1 1

[0075] The label combinations obtained above are all classified into one category due to the simplicity of the examples, but they successfully reduce the differences for all fundus images.

[0076] In step S3, the samples in the sample dataset are relabeled based on the merged classification labels.

[0077] Taking the data labels in Table 8 as an example, the fundus images are labeled, and the labels of img_1, img_2, img_3, and img_4 are Did=1, and the label of img_5 is Did=3. The model obtained by training the neural network with the data set composed of these fundus images can be used to label fundus images and classify diseases based on fundus images.

[0078] After the above calculations, we obtain a new category combination after merging. If two Dids are merged into the same category, these two Dids are defined as a merged pair within this category combination. Because different merging methods have different effects on different indicators, the inventors hope that the fundus image annotation method can achieve overall improved performance across all indicators. Therefore, we can observe the frequency of different merging pairs (label pairs) under various target reductions. The so-called frequency of different merging pairs under various target reductions refers to the total number of times the same label pair is selected for merging under different target reductions. For example, taking the data in Table 1 above as an example, when the target reduction is r2, 1 and 2 are selected for merging twice, when the target reduction is e1, 1 and 2 are selected for merging once, when the target reduction is e2, 1 and 2 are selected for merging once, and when the target reduction is e, 1 and 2 are selected for merging zero times. Therefore, the frequency of label pair 1 and 2 is 2+1+1+0=4 times. Similarly, the frequencies of other label pairs can be obtained, thereby analyzing their merging trends and providing some explanation for their correlation. Generally, half of the total number of category combinations (i.e., the number of times the greedy algorithm is used) can be used as the frequency threshold, or further analysis can be performed on the merging pairs that appear in the top 50% of the frequency in the dataset. In addition, we can also select the merged pairs with a large number of merges, set a certain frequency threshold, and merge the Did pairs that meet the conditions to achieve a trade-off between balancing feature preservation and reducing divergence.

[0079] In order to verify the effect of the present invention, the method of the present invention is used to process different medical fundus images, and the entropy values ​​after merging Did with different annotations are compared. The comparison results are shown in the figure. Figure 3-7 As shown in the figure. The data used are: initial fundus image annotations, fundus image annotations for seven groups of symptoms based on medical data (categorized by severity), fundus image Did combinations after similar symptoms, and fundus image annotations after the greedy algorithm using different standards as target reductions. Different target reductions are calculated using the Did combinations obtained using the greedy algorithm when using different standards as target reductions. The prefix w in the data in the figure represents the target reduction as the weighted average of the Did, i.e., the weighted average calculated using the frequency of 1 in Did as its weight. The horizontal line is a reference line to the original classification at this standard value, used to determine the increase or decrease in the target value after fundus image annotation.

[0080] in, Figure 3 The average entropy value obtained by using the greedy algorithm with different standards as the target reduction ( Figure 3 (a)) and its weighted average ( Figure 3 (b) is a schematic diagram of the decrease comparison result. It can be seen that the average entropy value of the fundus image marked by the method of the present invention is significantly decreased.

[0081] Figure 4 The zero entropy value obtained by performing a greedy algorithm with different standards as the target reduction ( Figure 4 (a)) and its weighted average ( Figure 4 (b) is a schematic diagram of the decrease comparison result. It can be seen that the zero-free entropy value of the fundus image marked by the method of the present invention is significantly decreased.

[0082] Figure 5 The diagnosis rate obtained by using the greedy algorithm with different standards as the target reduction ( Figure 5 (a)) and its weighted average ( Figure 5 (b) is a schematic diagram of the comparison results. It can be seen that the diagnosis rate of the fundus images marked by the method of the present invention is significantly improved.

[0083] Figure 6 The divergence rate obtained by using the greedy algorithm with different standards as the target reduction ( Figure 6 (a)) and its weighted average ( Figure 6 (b) is a schematic diagram of the comparison results of the decrease. It can be seen that the divergence rate of the fundus image marked by the method of the present invention is significantly reduced.

[0084] The present invention also verifies the entropy value of the method of the present invention on different data sets, and the effect is as follows: Figure 7 As shown, it can be seen that the entropy value is significantly reduced when the method of the present invention is used on different data sets.

[0085] Finally, the present invention also uses four methods: initial fundus image annotation, fundus image annotation of seven groups of diseases based on medical data, fundus image Di d combination annotation after similar disease processing, and fundus image annotation using the greedy algorithm method of the present invention. The four methods are verified on the first-fd16 dataset, and the results are as follows: Figure 8 As shown, it can be seen that the method of the present invention makes the average of each weight decrease significantly.

[0086] At the same time, in the above-mentioned Did merging process, the inventors also found a large number of Did pairs that have certain reasons to reduce differences through merging. For example, among all Did merged pairs, the two Did pairs that appear the most times in fundus images are [67,68] and [123,166]. Figure 9As shown, in the Did pair [67,68], 67 represents the epiretinal membrane (e.g. Figure 9 (a)), and 68 corresponds to the macular epiretinal membrane_treatment (as shown Figure 9 (b) Doctors assign a result between the two based on severity, not both simultaneously. 123,166 Dids were classified as mild arteriosclerosis and decreased arterial elasticity, respectively. The former is an extension of the latter, so doctors also don't assign both results simultaneously. Because there's a transition period between one Did and the next, it can be difficult for doctors to make an accurate and consistent diagnosis. Therefore, combining these Dids improves fundus image annotation performance.

[0087] It should be noted that although the above describes the various steps in a specific order, it does not mean that the steps must be performed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0088] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0089] Computer-readable storage media can be a tangible device that holds and stores the instructions used by an instruction execution device. Computer-readable storage media can, for example, include, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, a punch card or a raised structure in a groove on which instructions are stored, for example, and any suitable combination thereof.

[0090] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A training method for a disease classification model based on fundus images, characterized in that: The method comprises: F1. Obtain a fundus image dataset, where each fundus image contains one or more disease classification labels annotated by multiple annotators. F2. Label all fundus images in the fundus image dataset to obtain a fundus image training dataset using the following method: F21. Merge the label types in a fundus image dataset containing multiple classification labels to merge related classification label pairs and use one label in the label pair as the merged label; wherein the related classification label pair refers to a pair of different labels annotated by different annotators on the same fundus image; F22. Re-label the fundus images in the fundus image dataset based on the merged classification labels; F3. Use the fundus image training dataset in step F2 to train the convolutional neural network until convergence.

2. The method according to claim 1, characterized in that In step F21 , a greedy algorithm is used to perform multiple iterations to merge the label types of the fundus image containing multiple classification labels.

3. The method according to claim 2, characterized in that Each fundus image in the fundus image dataset contains a plurality of associated label pairs, and each time a label pair is merged, the label pair that reduces the most the divergence of the annotations of the fundus image dataset by different annotators after the merger is merged.

4. The method according to claim 3, characterized in that The divergence of annotations of fundus image datasets by different annotators is measured by the number of decreases in a preset target value; wherein the divergence rate, or zeroed entropy value, or divergence entropy value, or the average or weighted average of the total entropy values ​​corresponding to all category labels in the fundus image dataset is set as the target value.

5. The method according to claim 4, characterized in that Set the target drop value as follows: Calculate the disagreement rate or zero-deleted entropy value or disagreement entropy value or total entropy value corresponding to each type of label, wherein the disagreement rate corresponding to each type of label is the proportion of fundus images with inconsistent annotations by all annotators of this type of label in the fundus image dataset; the zero-deleted entropy value corresponding to each type of label is the average entropy value of fundus images with consistent annotations by all annotators of this type of label; the disagreement entropy value corresponding to each type of label is the average entropy value of fundus images with inconsistent annotations by all annotators of this type of label; the total entropy value corresponding to each type of label is the average entropy value of all fundus images with this type of label; Based on the calculated divergence rate or zero-free entropy value or divergence entropy value or total entropy value corresponding to each class label, the average or weighted average of the divergence rate or zero-free entropy value or divergence entropy value or total entropy value corresponding to all class labels in the fundus image dataset is calculated, where: The divergence rate or the zeroed entropy or the divergence entropy or the average value of the total entropy is calculated as follows: ,in, is the average of the divergence rate or zero-deleted entropy or divergence entropy or total entropy corresponding to all category labels in the fundus image dataset, H is the divergence rate or zero-deleted entropy or divergence entropy or total entropy corresponding to each category of labels, and N is the number of label categories; The weighted average of the divergence rate or zero entropy or divergence entropy or total entropy is calculated as follows: ,in, It is the weighted average of the divergence rate or zero-entropy value or divergence entropy value or total entropy value corresponding to all category labels in the fundus image dataset, and P is the fundus image frequency corresponding to each category label.

6. The method according to claim 5, characterized in that The entropy value corresponding to each fundus image of each class label is calculated as follows: in, Represents the entropy value of the current fundus image for the current class label, The proportion of all annotators who have annotated the current fundus image with the current class label.

7. The method according to claim 3, characterized in that After merging the associated label pairs, the label with the largest target reduction value in the merged label pair is used as the merged label.

8. The method according to claim 3, characterized in that In step F21 , a greedy algorithm is used to iteratively merge the label types of the fundus image containing multiple classification labels until the target reduction value is less than or equal to 0.

9. The method according to claim 8, characterized in that In step F21 , the label pairs to be merged are selected according to a preset frequency threshold, wherein the preset frequency threshold is set to the label pairs whose occurrence frequency in the fundus image dataset ranks in the top 50%.

10. A disease classification method based on fundus images, characterized in that: The method comprises: P1. Acquire the fundus image to be processed; P2. Use the disease classification model trained by the method described in any one of claims 1 to 9 to classify the image to be processed.

Citation Information

Patent Citations

  • Target detection method based on full-automatic learning

    CN111191732A