A crowdsourcing labeling noise filtering method based on weighted relative density

By applying a weighted relative density noise filtering method in crowdsourcing data, using the information in multi-noise marking, the problem of poor noise filtering effect in the prior art is solved, and a higher quality clean data set is achieved.

CN114781519BActive Publication Date: 2025-05-23HUBEI WEILAN GENERAL AVIATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210432916.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-05-23
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

When processing crowdsourcing data, existing noise filtering methods fail to effectively utilize the implicit information in multi-noise markers, resulting in poor filtering effect.

Method used

Using a method based on weighted relative density, the data quality of the clean set is improved by calculating the weighted relative density of each sample under each data subset, and the noise samples are identified and filtered.

Benefits of technology

This method not only utilizes the attribute feature information of the sample, but also uses the information in crowdsourcing tags to filter out clean samples from the original data more accurately, significantly improving the data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003611663950000012
    Figure BDA0003611663950000012
  • Figure BDA0003611663950000023
    Figure BDA0003611663950000023
  • Figure BDA0003611663950000032
    Figure BDA0003611663950000032
Patent Text Reader

Abstract

The present invention provides a crowdsourcing labeling noise filtering method based on weighted relative density, generates data subsets according to multiple noise labels of crowdsourcing data, then calculates the relative density of each sample for these data subsets, and finally filters the original data according to the relative density. The crowdsourcing labeling noise filtering method based on weighted relative density provided by the present invention not only utilizes the attribute feature information of the sample, but also utilizes the information in the crowdsourcing label, more accurately filters out clean samples in the original data, and can verify the effectiveness of the present invention through experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a crowdsourcing labeling noise filtering method based on weighted relative density, and belongs to the field of image and text labeling estimation. Background Art

[0002] Due to the rapid development of artificial intelligence, the demand for large amounts of labeled data has also increased accordingly. The emergence of crowdsourcing platforms provides an effective and low-cost way for instances to collect multiple labels from crowdsourcing workers. Therefore, on crowdsourcing platforms, each instance can obtain a multi-noise label set. Label ensemble methods are used to infer the true label of an instance from its multi-noise label set. For example, the category label estimation of leaf images, the emotion label estimation of pilot facial photos, the category label estimation of news, emails and other texts, and so on.

[0003] In machine learning, obtaining high-quality sample labels often requires high costs. Crowdsourcing systems provide an efficient and convenient platform to collect labels. On the crowdsourcing platform, workers choose crowdsourcing tasks that they are interested in and label the samples. The task publisher will also give appropriate compensation. Crowdsourcing labels have received widespread attention because they are cheap and easy to obtain. However, crowdsourcing workers are not experts in the field. During the labeling process, they may make incorrect labels due to lack of professionalism or inherent personal bias. Each sample often requires multiple labels, and the labeling quality of each sample is improved through a suitable label integration algorithm. Therefore, in the crowdsourcing dataset, each sample has a multi-noise label.

[0004] A crowdsourced dataset with N samples <x n ,L n >, which consists of two parts: N sample attributes and a multi-noise label set. n = {l n1 ,l n2 ,…,l nN}, M is the number of workers. ij represents the mark given by worker j to sample i, l ij ∈{c 1 ,c 2 ,…,c P}, where c i represents the i-th category, and there are a total of P categories in the dataset. If the worker does not label the sample, the value is -1. The goal of label integration is to aggregate the multiple noise label sets of each sample to obtain an integrated label The integrated dataset is

[0005] Although the use of label integration algorithms can improve the quality of labeling, because each sample has a different number of labels and different workers have different understandings of the same sample, label integration has limited effect on improving the quality of crowdsourcing data. Experiments show that no matter which label integration algorithm is used, the final data set will inevitably contain a certain amount of noise. Therefore, noise filtering of the integrated data set is also a very important process. The goal of noise filtering is to filter out samples that may be clean as much as possible and clean the integrated data set. Get two clean data sets D c and noise set D n Among them, the samples that are considered to be correctly labeled are divided into the clean set, and the remaining samples are divided into the noise set. The clean set is the high-quality dataset we expect.

[0006] At present, some scholars have proposed some noise filtering methods, such as ClassifierFiler (classifier filtering), CompleteRandomForestFilter (complete random forest filtering), etc. However, these methods are designed for traditional data sets. They only use the attribute feature information of samples for noise filtering, and do not use the implicit information in the multiple noise labels in the crowdsourcing data set. They cannot give full play to the advantages of crowdsourcing labels, and the filtering effect is not good. Summary of the invention

[0007] In order to address the shortcomings of the prior art, the present invention provides a crowdsourcing labeling noise filtering method based on weighted relative density, which utilizes the implicit information in the crowdsourcing labels to identify and filter noise, can greatly improve the data quality of the clean set, and can verify the superiority of the method through experiments.

[0008] The technical solution adopted by the present invention to solve the technical problem is: a crowdsourcing labeling noise filtering method based on weighted relative density is provided, comprising the following steps:

[0009] S1. For a crowdsourcing dataset after labeling and integration, each sample consists of a sample attribute, a multi-noise label set corresponding to the sample attribute, and an integrated label. The class probability distribution of each sample in all categories is obtained according to the multi-noise label set of each sample.

[0010] S2, dividing the crowdsourcing dataset into multiple data subsets according to the class probability distribution, each data subset corresponds to a category;

[0011] S3, calculating the weighted absolute density of each sample under each data subset;

[0012] S4, calculating the weighted relative density of each sample relative to each data subset according to the weighted absolute density, and obtaining a relative density vector;

[0013] S5. Use the relative density vector to perform noise filtering, add the samples identified as noise samples in the crowdsourcing dataset to the noise set, and add the remaining samples in the crowdsourcing dataset to the clean set to complete the filtering.

[0014] The crowdsourcing dataset in step S1 consists of N samples. Represented by, where the i-th sample is represented by the sample attribute x i , multi-noise label set L i and integration tags It consists of three parts: i = {l i1 ,l i2 ,…,l iM}, M is the number of workers, l ij represents the mark given by the jth worker to sample i, l ij ∈{0,c 1 ,c 2 ,…,c P}, where c i represents the i-th category. There are P categories in the dataset. 0 represents that the j-th worker has no information about the sample attribute x. i Marking;

[0015] The following formula is used to calculate the class probability distribution of each sample over all categories:

[0016]

[0017] Among them, P(x i ,c j ) represents the i-th sample attribute x in the crowdsourcing dataset i Belongs to the jth category c j The probability of M is the number of workers, l im Indicates that the mth worker gives the sample attribute x i The labeled category, δ(a,b) is a binary function that returns 1 when the values ​​of a and b are the same, otherwise it returns 0; then the class probability of each sample satisfies

[0018] Step S2 divides the crowdsourcing dataset D into multiple data subsets, which specifically includes the following process: for any sample attribute x of the i-th sample i and a given category c j , if p(x i ,c j )≠0, then add the sample attributes of the sample to the data subset D j and p(x i ,c j ) as the sample attribute x i In Dj After processing each sample, we get the set D of data subsets. sub ={D 1 ,D 2 ,…,D P}, where each data subset consists of sample attributes and corresponding weights, and the jth data subset D j The corresponding category is c j , that is, the same data subset D j The samples in have the same category c j .

[0019] Step S3 calculates the weighted absolute density of each sample under each data subset and specifically includes the following process:

[0020] S3.1. For each sample attribute in the crowdsourcing dataset D, its relationship with each data subset D is calculated by the following formula: p The weighted distance of each sample attribute in :

[0021]

[0022] Among them, wd(x i ,x j ,D p ) j Represents the attribute x of the i-th sample in the crowdsourcing dataset D i With the pth data subset D p The j-th sample attribute x in j The weighted distance, w j is the sample attribute x j In the pth data subset D p The weights in , are assigned class probabilities calculated in step S2;

[0023] S3.2 Calculate the weighted absolute density of each sample in each data subset according to the following formula:

[0024]

[0025] Among them, AD(x i ,D p ) represents the i-th sample x in the crowdsourcing dataset D i In the pth data subset D p The weighted absolute density under kNeighbor(x i ,D p ) represents the sample attribute x i In the dataset D p The k nearest neighbors in .

[0026] Step S4 specifically includes the following process:

[0027] S4.1. Calculate the weighted relative density of each sample relative to each data subset according to the following formula:

[0028]

[0029] Among them, RD(x i ,D j ) represents the i-th sample x in the crowdsourcing dataset D i In the jth data subset D j The weighted relative density under label(i) is the integrated label of the i-th sample in the crowdsourcing dataset D The categories of the set D in the data subset sub If the corresponding data subset in Then D label(i) =D j ;

[0030] S4.2. After substituting the weighted absolute density into the formula of step S4.1, the weighted relative density is simplified to:

[0031]

[0032] Among them, x p Belong to x i In the data subset D label(i) k nearest neighbors in x q For x i In the data subset D j The k nearest neighbors in D are the sum of the weighted distances of the k nearest neighbors of the i-th sample attribute in D on the data subset corresponding to its integrated label divided by the sum of the weighted distances of the k nearest neighbors on each data subset.

[0033] Step S5 specifically includes the following process:

[0034] S5.1. For each sample x i , and the weighted relative density under P data subsets constitutes the sample relative density vector V i ={rd 1 ,rd 2 ,…,rd P}, where rd j =RD(x i ,D j );

[0035] S5.2. For each sample x i , sort the relative density vectors from large to small, and get the category of the data subset corresponding to the rearranged relative density vector:

[0036]

[0037] S5.3, traverse all samples, for the i-th sample, if Make Then the sample x i is treated as noise and filtered into the noise set D n middle;

[0038] S5.4. The remaining samples and their integrated labels form a clean set D c , the filtering is completed.

[0039] The beneficial effect of the technical solution of the present invention is that the crowdsourcing labeling noise filtering method based on weighted relative density provided by the present invention generates data subsets according to the multiple noise labels of crowdsourcing data, then calculates the relative density of each sample for these data subsets, and finally filters the original data according to the relative density. Not only the attribute feature information of the sample is used, but also the information in the crowdsourcing label is used to more accurately filter out clean samples in the original data. DETAILED DESCRIPTION

[0040] The present invention will be further described below in conjunction with the embodiments.

[0041] Example:

[0042] The present invention provides a crowdsourcing labeling noise filtering method based on weighted relative density, comprising the following steps:

[0043] S1. For a labeled integrated crowdsourcing dataset, each sample consists of sample attributes, a multi-noise label set corresponding to the sample attribute, and an integrated label. The class probability distribution of each sample in all categories is obtained according to the multi-noise label set of each sample.

[0044] The crowdsourcing dataset consists of N samples. Represented by, where the i-th sample is represented by the sample attribute x i , multi-noise label set L i and integration tags It consists of three parts: i = {l i1 ,l i2 ,…,l iM}, M is the number of workers, l ij represents the mark given by the jth worker to sample i, l ij ∈{0,c 1 ,c 2 ,…,c P}, where c i represents the i-th category. There are P categories in the dataset. 0 represents that the j-th worker has no information about the sample attribute x. iMarking;

[0045] The following formula is used to calculate the class probability distribution of each sample over all categories:

[0046]

[0047] Among them, P(x i ,c j ) represents the i-th sample attribute x in the crowdsourcing dataset i Belongs to the jth category c j The probability of M is the number of workers, l im Indicates that the mth worker gives the sample attribute x i The labeled category, δ(a,b) is a binary function that returns 1 when the values ​​of a and b are the same, otherwise it returns 0; then the class probability of each sample satisfies

[0048] S2. Divide the crowdsourcing dataset into multiple data subsets according to the class probability distribution, each data subset corresponds to a category. The specific process includes the following: For any sample attribute x of the i-th sample i and a given category c j , if p(x i ,c j )≠0, then add the sample attributes of the sample to the data subset D j and p(x i ,c j ) as the sample attribute x i In D j After processing each sample, we get the set D of data subsets. sub ={D 1 ,D 2 ,…,D P}, where each data subset consists of sample attributes and corresponding weights, and the jth data subset D j The corresponding category is c j , that is, the same data subset D j The samples in have the same category c j .

[0049] S3. Calculate the weighted absolute density of each sample in each data subset. Specifically, it includes the following process:

[0050] S3.1. For each sample attribute in the crowdsourcing dataset D, its relationship with each data subset D is calculated by the following formula: p The weighted distance of each sample attribute in :

[0051]

[0052] Among them, wd(x i ,x j ,D p ) j Represents the attribute x of the i-th sample in the crowdsourcing dataset D i With the pth data subset D p The j-th sample attribute x in j The weighted distance, w j is the sample attribute x j In the pth data subset D p The weights in , are assigned class probabilities calculated in step S2;

[0053] S3.2 Calculate the weighted absolute density of each sample in each data subset according to the following formula:

[0054]

[0055] Among them, AD(x i ,D p ) represents the i-th sample x in the crowdsourcing dataset D i In the pth data subset D p The weighted absolute density under kNeighbor(x i ,D p ) represents the sample attribute x i In the dataset D p The k nearest neighbors in .

[0056] S4. Calculate the weighted relative density of each sample relative to each data subset based on the weighted absolute density to obtain a relative density vector. Specifically, the process includes the following:

[0057] S4.1. Calculate the weighted relative density of each sample relative to each data subset according to the following formula:

[0058]

[0059] Among them, RD(x i ,D j ) represents the i-th sample x in the crowdsourcing dataset D i In the jth data subset D j The weighted relative density under label(i) is the integrated label of the i-th sample in the crowdsourcing dataset D The categories of the set D in the data subset sub If the corresponding data subset in Then D label(i) =D j ;

[0060] S4.2. After substituting the weighted absolute density into the formula of step S4.1, the weighted relative density is simplified to:

[0061]

[0062] Among them, x p Belong to x i In the data subset D label(i) k nearest neighbors in x q For x i In the data subset D j The k nearest neighbors in D are the sum of the weighted distances of the k nearest neighbors of the i-th sample attribute in D on the data subset corresponding to its integrated label divided by the sum of the weighted distances of the k nearest neighbors on each data subset.

[0063] S5. Use the relative density vector to filter noise, add the samples in the crowdsourcing dataset that are identified as noise samples to the noise set, and add the remaining samples in the crowdsourcing dataset to the clean set to complete the filtering. The specific process includes the following:

[0064] S5.1. For each sample x i , and the weighted relative density under P data subsets constitutes the sample relative density vector V i ={rd 1 ,rd 2 ,…,rd P}, where rd j =RD(x i ,D j );

[0065] S5.2. For each sample x i , sort the relative density vectors from large to small, and get the category of the data subset corresponding to the rearranged relative density vector:

[0066]

[0067] S5.3, traverse all samples, for the i-th sample, if Make Then the sample x i is treated as noise and filtered into the noise set D n middle;

[0068] S5.4. The remaining samples and their integrated labels form a clean set D c , the filtering is completed.

[0069] Experimental comparison:

[0070] In the experimental part, a weighted relative density based crowdsourcing labeling noise filtering method (abbreviated as WRDNC) proposed in this invention is compared with the benchmark method after labeling integration. In addition, it is also compared with an existing simple and efficient noise filtering method, classifier filtering (abbreviated as CF).

[0071] First, we briefly introduce the objects with which this method is compared:

[0072] 1. Majority voting (MV for short) is the most commonly used tag integration method. The category with the most votes is selected as the integrated tag by majority voting. This method is the starting point of the design of this scheme. This experiment will use this method to verify the effectiveness of the method of the present invention.

[0073] 2. Classifier filtering (abbreviated as CF) is used to filter noise through cross-validation. First, the entire data set is divided into k parts, and one part of the data set is removed each time. The remaining k-1 parts are used as training sets to train the classifier. After obtaining k classifiers, these samples are classified in turn. If their prediction results are different, they are filtered out, and the remaining data is the clean set data.

[0074] This experiment selects two real datasets that are widely used in crowdsourcing experiments. They come from different fields and represent different features. Income is a census data that records the age, gender, race, and work of 600 people. It can be used to predict whether a person's income reaches $50,000 per year. Leaves is a multi-classification image dataset that distinguishes 6 different types of leaves. Table 1 describes the main features of these two datasets. The specific data can be downloaded from the CEKA platform.

[0075] Dataset Number of samples Number of attributes Number of categories Income 600 52 2 Leaves 384 64 6

[0076] Table 1 Datasets used in the experiments

[0077] Table 2 shows the results of various methods on the two datasets. The first column under each method represents the noise rate of the clean set, and the second column represents the size of the clean set. Since MV is the starting point of our work and no filtering is performed, the MV column only gives the noise rate after integration.

[0078]

[0079] Table 2 Comparison results between WRDNC and classification

[0080] From the experimental results we can see that:

[0081] (1) Compared with the labeled integrated datasets, the proposed method WRDNC can filter out all noise and select clean samples on both real datasets.

[0082] (2) The WRDNC method of the present invention is significantly better than CF in selecting clean samples. Although the number of clean samples screened is less than that of CF, it can greatly improve the accuracy of the clean set.

Claims

1. A crowdsourcing labeling noise filtering method based on weighted relative density, Features The following steps are involved: S1. For a crowdsourced dataset of text or images after labeling and integration, each sample consists of a sample attribute, a multi-noise label set corresponding to the sample attribute, and an integrated label. The class probability distribution of each sample in all categories is obtained according to the multi-noise label set of each sample. The crowdsourcing dataset consists of N samples. Represented by, where the i-th sample is represented by the sample attribute x i , multi-noise label set L i and integration tags It consists of three parts: i = {l i1 ,l i2 ,…,l iM }, M is the number of workers, l ij represents the mark given by the jth worker to sample i, l ij ∈{0,c 1 ,c 2 ,…,c P }, where c i represents the i-th category. There are P categories in the dataset. 0 represents that the j-th worker has no information about the sample attribute x. i Marking; The following formula is used to calculate the class probability distribution of each sample over all categories: Among them, P(x i ,c j ) represents the i-th sample attribute x in the crowdsourcing dataset i Belongs to the jth category c j The probability of M is the number of workers, l im Indicates that the mth worker gives the sample attribute x i The labeled category, δ(a,b) is a binary function that returns 1 when the values ​​of a and b are the same, otherwise it returns 0; then the class probability distribution of each sample satisfies S2. Divide the crowdsourcing dataset into multiple data subsets according to the class probability distribution, each data subset corresponds to a category; specifically, the following process is included: For any sample attribute x of the i-th sample i and a given category c j , if p(x i ,c j )≠0, then add the sample attributes of the sample to the data subset D j and p(x i ,c j ) as the sample attribute x i In D j After processing each sample, we get the set D of data subsets. sub ={D 1 ,D 2 ,…,D P }, where each data subset consists of sample attributes and corresponding weights, and the jth data subset D j The corresponding category is c j , that is, the same data subset D j The samples in have the same category c j ; S3. Calculate the weighted absolute density of each sample in each data subset; specifically, the following process is included: S3.

1. For each sample attribute in the crowdsourcing dataset D, its relationship with each data subset D is calculated by the following formula: p The weighted distance of each sample attribute in : Among them, wd(x i ,x j ,D p ) represents the attribute x of the ith sample in the crowdsourcing dataset D i With the pth data subset D p The j-th sample attribute x in j The weighted distance, w j is the sample attribute x j In the pth data subset D p The weights in are assigned by the class probability distribution calculated in step S2; S3.2 Calculate the weighted absolute density of each sample in each data subset according to the following formula: Among them, AD(x i ,D p ) represents the i-th sample x in the crowdsourcing dataset D i In the pth data subset D p The weighted absolute density under kNeighbor(x i ,D p ) represents the sample attribute x i In the dataset D p The k nearest neighbors in ; S4. Calculate the weighted relative density of each sample relative to each data subset according to the weighted absolute density to obtain a relative density vector; specifically, the process includes the following steps: S4.

1. Calculate the weighted relative density of each sample relative to each data subset according to the following formula: Among them, RD(x i ,D j ) represents the i-th sample x in the crowdsourcing dataset D i In the jth data subset D j The weighted relative density under label(i) is the integrated label of the i-th sample in the crowdsourcing dataset D The categories of the set D in the data subset sub If the corresponding data subset in Then D label(i) =D j ; S4.

2. After substituting the weighted absolute density into the formula of step S4.1, the weighted relative density is simplified to: Among them, x p Belong to x i In the data subset D label(i) k nearest neighbors in x q For x i In the data subset D j The k nearest neighbors in D, that is, the sum of the weighted distances of the k nearest neighbors of the i-th sample attribute in D on the data subset corresponding to its integrated label divided by the sum of the weighted distances of the k nearest neighbors on each data subset; S5. Use the relative density vector to perform noise filtering, add the samples identified as noise samples in the crowdsourcing dataset to the noise set, and add the remaining samples in the crowdsourcing dataset to the clean set to complete the filtering.

2. The crowdsourcing labeling noise filtering method based on weighted relative density according to claim 1, Features: Step S5 specifically The process includes: S5.

1. For each sample x i , and the weighted relative density under P data subsets constitutes the sample relative density vector V i ={rd 1 ,rd 2 ,…,rd P }, where rd j =RD(x i ,D j ); S5.

2. For each sample x i , sort the relative density vectors from large to small, and get the category of the data subset corresponding to the rearranged relative density vector: S5.3, traverse all samples, for the i-th sample, if Make Then the sample x i is treated as noise and filtered into the noise set D n middle; S5.

4. The remaining samples and their integrated labels form a clean set D c , the filtering is completed.

Citation Information

Patent Citations

  • Method and system for improving quality of a dataset

    CA3070925A1

  • Label noise detection method based on multi-granularity relative density

    CN111178387A