A crowdsourcing labeling ensemble method based on instance redistribution
By using instance redistribution technology in the crowdsourcing tag integration method, calculating multi-noise tag distribution vectors and instance redistribution vectors, and updating the instance weight and integration tags, the problem of high complexity of tag integration methods in the existing technology is solved, and an efficient and simple tag inference process is realized, and the marking quality is improved.
Patent Information
- Application Number
- CN202210419838.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-04-21
AI Technical Summary
Existing tag integration methods increase complexity while improving performance, making it difficult to find a simple and efficient way to infer the true label of an instance from a multi-noise tag set.
The crowdsourcing tag integration method based on instance redistribution is adopted, and the instance redistribution vector is calculated by calculating the multi-noise tag distribution vector and the instance redistribution vector, and the instance weight and integration tag are updated to ensure the simplicity and efficiency of the tag inference process.
The method of mining the true marks of instances from both the marking perspective and the feature perspective is realized, which improves the marking quality, avoids time overhead, and verifies the effectiveness of the method through experiments.
Smart Images

Figure BDA0003607106770000025 
Figure BDA0003607106770000026 
Figure BDA0003607106770000027
Abstract
Description
Technical Field
[0001] The invention relates to a crowdsourcing tag integration method based on instance reallocation, and belongs to the field of image tag estimation. Background Art
[0002] Due to the rapid development of artificial intelligence, the demand for large amounts of labeled data has also increased accordingly. The emergence of crowdsourcing platforms provides an effective and low-cost way for instances to collect multiple labels from crowdsourcing workers. Therefore, on crowdsourcing platforms, each instance can obtain a multi-noise label set. Label ensemble methods are used to infer the true label of an instance from its multi-noise label set. For example, emotion label estimation is performed on pilot facial photos, species label estimation is performed on leaf images, and so on.
[0003] At present, a large number of label ensemble methods have been proposed, which use various machine learning techniques to optimize the label inference process. Early label ensemble methods only infer the true label of the instance from the label level. They build a flexible probabilistic model to find the generative relationship between a given multi-noise label set and the unknown true label. These methods completely ignore the feature information of the instance. With the development of label ensemble research, some label ensemble methods consider adding the feature information of the instance to participate in the construction of the model during the inference process. However, while these methods are committed to improving the performance of label ensemble methods, they also increase the complexity of the methods. Therefore, finding a simple and efficient label ensemble method has become a challenge. Summary of the invention
[0004] In order to address the shortcomings of the prior art, the present invention provides a crowdsourcing labeling integration method based on instance redistribution, which makes full use of the effective information contained in the crowdsourcing dataset, explores the hidden true labels of the instances from the labeling perspective and the feature perspective respectively, and ensures that the labeling inference process is not only simple and easy to implement but also efficient, which not only improves the labeling quality of the labeling integration method, but also avoids the time overhead of the labeling inference process.
[0005] The technical solution adopted by the present invention to solve the technical problem is: a crowdsourcing tag integration method based on instance redistribution is provided, comprising the following steps:
[0006] S1. For a given crowdsourcing dataset, each element of which consists of an instance and a multi-noise tag set corresponding to the instance; the multi-noise tag distribution vector of the instance is calculated according to the multi-noise tag set of each instance, and its instance weight is further assigned and the integration mark is initialized, and the crowdsourcing dataset is divided into multiple subsets according to the initialized integration mark;
[0007] S2. For each subset, generate the centroid of the subset and calculate the attribute value of the centroid;
[0008] S3. For each instance in the crowdsourcing data set, the distance between the instance and each centroid is calculated respectively, and the similarity between the instance and each centroid is calculated respectively according to the distance value;
[0009] S4. For each instance in the crowdsourcing dataset, the instance reallocation vector of the instance is obtained according to the similarity between the instance and each centroid, the weight of each instance is updated using the instance reallocation vector, and the integrated label of each instance is updated using the updated weight.
[0010] Step S1 specifically includes the following process:
[0011] S1.1. Given a crowdsourcing dataset Where N is the number of instances contained in D, x i is the i-th instance, instance Contains M attributes, x i The mth attribute of is represented by a im ; L i is x i The multi-noise label set, It consists of the marks provided by J workers, and the jth mark l ij Provided by the jth worker, l ij From the class label set {0,c1,c2,…,c Q}, where 0 means that the jth worker has no mark x i ;
[0012] S1.2. Calculate x using the following formula i The multi-noise label distribution vector
[0013]
[0014] Among them, δ(l ij ,c q ) represents a binary function, which takes the value 1 if its two arguments are the same and 0 otherwise;
[0015] S1.3, use the following formula to assign x i Instance weight:
[0016]
[0017] S1.4. Initialize x using the following formula i Integration standard:
[0018]
[0019] S1.5. According to example x i After initialization, the integrated Divide the data set D into Q subsets {D1,D2,…,D Q}.
[0020] Step S2 specifically includes the following process:
[0021] S2.1. Generate your own centroid z q ;
[0022] S2.2. Calculate the centroid attribute value using the average of the weighted numerical feature values or the mode of the weighted nominal feature values:
[0023]
[0024] where z qm It is a subset D q The center of mass z q The attribute value of the mth attribute, N q Yes D q The number of instances, w i is x i The weight of mv is the vth attribute value of the mth attribute.
[0025] Step S3 specifically includes the following process:
[0026] S3.1. Calculate the instance x using the following formula i and the centroid z q Distance:
[0027]
[0028] where d m (x i ,z q ) represents x i With z q The distance on the mth attribute, d m (x i ,z q ) is calculated by the following formula:
[0029]
[0030] Among them, overlap(a im ,a qm ) represents x i With z q The calculation method for noun attributes is calculated using the following formula:
[0031]
[0032] rn_diff(a im ,a qm) represents x i With z q The calculation method for numerical attributes is calculated using the following formula:
[0033]
[0034] S3.2, calculate x using the following formula i With z q The similarity S iq :
[0035]
[0036] Among them, exp(·) is the exponential kernel function, and σ is the hyperparameter of the exponential kernel function.
[0037] Step S4 specifically includes the following process:
[0038] S4.1. Calculate x using the following formula i The instance reallocation vector
[0039]
[0040] S4.2. Update x using the following formula i Weight:
[0041]
[0042] S4.3. Update instance x using the following formula i Integration standard:
[0043]
[0044] That is, instance x i Inferred integration mark.
[0045] The beneficial effects of the present invention based on its technical solution are: a crowdsourcing labeling integration method based on instance reallocation provided by the present invention makes full use of the effective information contained in the crowdsourcing data set, mines the true label of the instance from both the labeling perspective and the feature perspective, and estimates the probability that the instance belongs to each category. From the labeling perspective, the multi-noise label distribution vector reveals the subjective cognitive tendency of crowdsourcing workers towards the instance. From the feature perspective, the instance reallocation vector reveals the objective connection between the instance features and the representative features of each category. The step of fusing the two vectors to update the integrated label of the instance not only takes into account the complementarity of the advantages and disadvantages between the multi-noise label distribution vector and the instance reallocation vector, but also maintains the computational complexity and simplicity of the model. More importantly, the effectiveness of the new method provided by the present invention is verified by experimental results. DETAILED DESCRIPTION
[0046] The present invention will be further described below in conjunction with the embodiments.
[0047] Example:
[0048] The present invention provides a crowdsourcing tag integration method based on instance redistribution, comprising the following steps:
[0049] S1. For a given crowdsourcing dataset, each element of which consists of an instance and a multi-noise tag set corresponding to the instance; the multi-noise tag distribution vector of the instance is calculated based on the multi-noise tag set of each instance, and its instance weight is further assigned and the integration mark is initialized, and the crowdsourcing dataset is divided into multiple subsets according to the initialized integration mark. The specific process includes the following:
[0050] S1.1. Given a crowdsourcing dataset Where N is the number of instances contained in D, x i is the i-th instance, instance Contains M attributes, x i The mth attribute of is represented by a im ; L i is x i The multi-noise label set, It consists of the marks provided by J workers, and the jth mark l ij Provided by the jth worker, l ij From the class label set {0,c1,c2,…,c Q}, where 0 means that the jth worker has no mark x i ;
[0051] S1.2. Calculate x using the following formula i The multi-noise label distribution vector
[0052]
[0053] Among them, δ(l ij ,c q ) represents a binary function, which takes the value 1 if its two arguments are the same and 0 otherwise;
[0054] S1.3, use the following formula to assign x i Instance weight:
[0055]
[0056] S1.4. Initialize x using the following formula i Integration standard:
[0057]
[0058] S1.5. According to example x i After initialization, the integrated Divide the data set D into Q subsets {D1,D2,…,D Q}.
[0059] S2. For each subset, generate the centroid of the subset and calculate the attribute value of the centroid. The specific process includes the following:
[0060] S2.1. Generate your own centroid z q ;
[0061] S2.2. Calculate the centroid attribute value using the average of the weighted numerical feature values or the mode of the weighted nominal feature values:
[0062]
[0063] where z qm It is a subset D q The center of mass z q The attribute value of the mth attribute, N q Yes D q The number of instances, w i is x i The weight of mv is the vth attribute value of the mth attribute.
[0064] S3. For each instance in the crowdsourcing data set, the distance between the instance and each centroid is calculated, and the similarity between the instance and each centroid is calculated based on the distance value. The specific process includes the following:
[0065] S3.1. Calculate the instance x using the following formula i and the centroid z q Distance:
[0066]
[0067] where d m (x i ,z q ) represents x i With z q The distance on the mth attribute, d m (x i ,z q ) is calculated by the following formula:
[0068]
[0069] Among them, overlap(a im ,a qm ) represents x i With zq The calculation method for noun attributes is calculated using the following formula:
[0070]
[0071] rn_diff(a im ,a qm ) represents x i With z q The calculation method for numerical attributes is calculated using the following formula:
[0072]
[0073] S3.2, calculate x using the following formula i With z q The similarity S iq :
[0074]
[0075] Among them, exp(·) is the exponential kernel function, and σ is the hyperparameter of the exponential kernel function.
[0076] S4. For each instance in the crowdsourcing dataset, the instance redistribution vector of the instance is obtained according to the similarity between the instance and each centroid, the weight of each instance is updated using the instance redistribution vector, and the integrated label of each instance is updated using the updated weight. Specifically, the following process is included:
[0077] S4.1. Calculate x using the following formula i The instance reallocation vector
[0078]
[0079] S4.2. Update x using the following formula i Weight:
[0080]
[0081] S4.3. Update instance x using the following formula i Integration standard:
[0082]
[0083] That is, instance x i Inferred integration mark.
[0084] Experimental comparison:
[0085] The present invention proposes a crowd-sourced labeling ensemble method based on instance redistribution, and the resulting model is called crowd-sourced labeling ensemble based on instance redistribution (abbreviated as IRLI). In the following experimental part, the crowd-sourced labeling ensemble based on instance redistribution (abbreviated as IRLI) proposed in the present invention is compared with majority voting (abbreviated as MV), ZenCrowd method (abbreviated as ZC), Karger Oh and Shah method (abbreviated as KOS) and iterative weighted majority voting (abbreviated as IWMV).
[0086] MV selects the most frequently appearing tag in the multi-noise tag set as the integrated tag of the instance.
[0087] ZC iteratively estimates the reliability of crowdworkers and the probability of instances belonging to each class by using the expectation maximization technique.
[0088] KOS completes labeling integration by taking into account the reliability of workers using the singular value decomposition technique of low-rank matrices and the belief propagation-like technique.
[0089] IWMV searches for reliable crowdworkers by iteratively optimizing the error rate bound to improve the performance of majority voting.
[0090] In order to verify the effectiveness of the crowdsourcing labeling ensemble method based on instance reallocation proposed in this invention, the classification performance of IRLI, MV, ZC, KOS and IWMV is compared experimentally.
[0091] The experiment reproduces the IWMV method and uses the existing MV, ZC and KOS methods implemented by the CEKA platform. The hyperparameter σ of IRLI is set to 1. These methods are tested on two widely used crowdsourced real datasets, which come from different fields and represent different data characteristics. Table 1 describes the main features of these two datasets in detail. The specific data can be downloaded from the CEKA platform website.
[0092]
[0093] Table 1 Datasets used in the experiments
[0094] Table 2 shows the labeling quality obtained by each method on each dataset. Each row is marked in bold to indicate the labeling ensemble method with the best labeling quality on the dataset corresponding to that row.
[0095] Dataset MV ZC KOS IWMV IRLI Income 72.17 72.17 60.67 71.67 74.00 Leaves 64.06 64.58 64.32 64.06 72.66
[0096] Table 2 Comparison of labeling quality of MV, ZC, KOS, IWMV and IRLI
[0097] From the experimental results in Table 2, we can summarize the following advantages of this method:
[0098] (1) On the real dataset Income, the labeling quality of the IRLI method proposed in this paper (74.00%) is significantly better than that of all the comparison objects, among which the labeling quality of MV is 71.76%, the labeling quality of ZC is 72.17%, the labeling quality of KOS is 60.67%, and the labeling quality of IWMV is 71.67%.
[0099] (2) On the real dataset Leaves, the labeling quality of the IRLI method proposed in this paper (72.66%) is significantly better than that of all the comparison objects, among which the labeling quality of MV is 64.06%, the labeling quality of ZC is 64.58%, the labeling quality of KOS is 64.32%, and the labeling quality of IWMV is 64.06%.
[0100] (3) The labeling quality of the IRLI method proposed in this paper is much higher than that of all the comparison objects, which proves that the labeling integration method using instance redistribution is very effective. The fusion of the estimated multi-noise label distribution vector and the strength redistribution vector effectively complements the advantages and disadvantages of the two.
[0101] In summary, the crowdsourcing labeling integration method based on instance reallocation proposed in the present invention is simple and easy to implement and has the best labeling quality among all the compared methods. It can be used in the field of image label estimation and play an advantageous role.
Claims
1. A crowdsourcing labeling ensemble method based on instance redistribution, characterized in that The following steps are involved: S1. For a given image, a crowdsourced dataset, where each element consists of an instance and a set of multiple noise labels corresponding to the instance; The multi-noise label distribution vector of each instance is calculated based on the multi-noise label set of the instance, and the instance weight is further assigned and the integration mark is initialized. The crowdsourcing dataset is divided into multiple subsets according to the initialized integration mark. The process includes: S1.
1. Given a crowdsourcing dataset Where N is the number of instances contained in D, x i is the i-th instance, instance Contains M attributes, x i The mth attribute of is represented by a im ; L i is x i The multi-noise label set, It consists of the marks provided by J workers, and the jth mark l ij Provided by the jth worker, l ij From the class label set {0,c1,c2,…,c Q }, where 0 means that the jth worker has no mark x i ; S1.
2. Calculate x using the following formula i The multi-noise label distribution vector Among them, δ(l ij ,c q ) represents a binary function, which takes the value 1 if its two arguments are the same and 0 otherwise; S1.3, use the following formula to assign x i Instance weight: S1.
4. Initialize x using the following formula i Integration standard: S1.
5. According to example x i After initialization, the integrated Divide the data set D into Q subsets {D1,D2,…,D Q }; S2. For each subset, generate the centroid of the subset and calculate the attribute value of the centroid; The process includes: S2.
1. Generate their respective centroids z q ; S2.
2. Calculate the centroid attribute value using the average of the weighted numerical feature values or the mode of the weighted nominal feature values: where z qm It is a subset D q The center of mass z q The attribute value of the mth attribute, N q Yes D q The number of instances, w i is x i The weight of mv is the vth attribute value of the mth attribute; S3. For each instance in the crowdsourcing data set, the distance between the instance and each centroid is calculated respectively, and the similarity between the instance and each centroid is calculated respectively according to the distance value; S4. For each instance in the crowdsourcing dataset, the instance redistribution vector of the instance is obtained according to the similarity between the instance and each centroid, the weight of each instance is updated using the instance redistribution vector, and the integrated label of each instance is updated using the updated weight.
2. The crowdsourcing labeling integration method based on instance redistribution according to claim 1, characterized in that: Step S3 specifically The process includes: S3.
1. Calculate the instance x using the following formula i and the centroid z q Distance: where d m (x i ,z q ) represents x i With z q The distance on the mth attribute, d m (x i ,z q ) is calculated by the following formula: Among them, overlap(a im ,a qm ) represents x i With z q The calculation method for noun attributes is calculated using the following formula: rn_diff(a im ,a qm ) represents x i With z q The calculation method for numerical attributes is calculated using the following formula: S3.2, calculate x using the following formula i With z q The similarity S iq : Among them, exp(·) is the exponential kernel function, and σ is the hyperparameter of the exponential kernel function.
3. The crowdsourcing tagging integration method based on instance redistribution according to claim 2 is characterized in that: Step S4 specifically includes the following process: S4.
1. Calculate x using the following formula i The instance reallocation vector S4.
2. Update x using the following formula i Weight: S4.
3. Update instance x using the following formula i Integration standard: That is, instance x i Inferred integration mark.
Citation Information
Patent Citations
Tag noise correction based crowd-sourced tagging data quality improvement method
CN105426826A
Crowdsourcing repeat label-based deep learning target detection method and system
CN110580499A