Label correction crowdsourcing result convergence method and system based on deep clustering

By adopting a deep cluster-based label correction method in the crowdsourcing system, using feature extraction and variational deep embedding model, combining Gaussian hybrid model and worker labeling information, the label noise problem in the crowdsourcing system is solved, the accuracy of task labels is improved, and a more reliable data foundation is provided for supervised learning.

CN120030368APending Publication Date: 2025-05-23ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510147868.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In crowdsourcing systems, due to uneven worker abilities and task complexity, the accuracy of task tags is difficult to guarantee. Especially in the case of sparse source tags, the label noise problem is significant, affecting the truth value inference.

Method used

The label correction method based on deep clustering is adopted, and the task features are extracted using the feature extraction network, combined with the variational depth embedding model and the Gaussian hybrid model, hidden variables are clustered, and the label inference process is optimized through the worker's labeling information to identify and correct tasks with low clustering accuracy.

Benefits of technology

Effectively eliminate low-quality annotations, improve the accuracy of task labels, and provide a more reliable data foundation for supervised learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030368A_ABST
    Figure CN120030368A_ABST
Patent Text Reader

Abstract

The invention discloses a label correction crowdsourcing result aggregation method and system based on deep clustering. Firstly, task features are extracted through a pre-training model to serve as input data of the model; secondly, introducing a worker loss item, defining a comprehensive loss function of the model in combination with reconstruction loss and KL divergence in the variational depth embedding model, and training the model through a gradient descent method; secondly, in the training process, clustering the hidden variables by using a Gaussian mixture model to obtain clustering clusters, and mapping a specific class label for each clustering cluster in combination with task labeling information of a worker; and finally, the contour coefficient of each task is calculated, the task with low clustering accuracy is identified, and the precision of the clustering result is further improved through readjustment and correction of worker labels. According to the method, tasks with similar features are clustered, and tasks with low clustering quality are corrected by using worker tags, so that the interference of tag noise on truth value inference is effectively reduced, and the performance of a result convergence model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of truth-value reasoning of crowdsourcing, and specifically relates to a method and system for aggregating label-corrected crowdsourcing results based on deep clustering. Background Art

[0002] With the widespread application of crowdsourcing in data annotation, task distribution and collaborative work, it has become an important means to obtain large-scale and diverse data. However, due to the uneven ability of workers, the complexity of tasks and the influence of subjective factors, the accuracy of task labels is difficult to guarantee. The problem of label noise is particularly significant, especially in the case of sparse source labels, that is, each object only obtains labels from very few sources, which makes it difficult to provide sufficient basis for the inference of true labels. Compared with sparse labeled data, object features contain more valuable information. By clustering objects with similar features, the interference of label noise on true value inference can be effectively reduced. Summary of the invention

[0003] In order to solve the problem that workers in a crowdsourcing system are subject to subjective and objective influences and the accuracy of task labels is difficult to guarantee, the present invention provides a method and system for aggregating crowdsourcing results with label correction based on deep clustering.

[0004] In a first aspect, the present invention provides a method for aggregating label correction crowdsourcing results based on deep clustering, the method comprising:

[0005] Use the feature extraction network to extract the features of the crowdsourcing task as the input data of the model;

[0006] Construct a worker loss term based on worker label consistency, combine the reconstruction loss and KL divergence in the variational deep embedding model, define a comprehensive loss function, and train the model using the gradient descent method;

[0007] During the training process, the Gaussian mixture model is used to cluster the latent variables and map specific class labels to each cluster.

[0008] The silhouette coefficient of each task is calculated, tasks with low clustering accuracy are identified, and these tasks are readjusted and corrected using worker annotation information.

[0009] In a second aspect, the present invention provides a label correction crowdsourcing result aggregation system based on deep clustering, comprising:

[0010] The feature extraction module is used to extract the features of the crowdsourcing task using the feature extraction network as the input data of the model;

[0011] The loss function construction module is used to construct a worker loss term based on the consistency of worker labels, and combine the reconstruction loss and KL divergence in the variational deep embedding model to define a comprehensive loss function;

[0012] The training module is used to train the model using the gradient descent method;

[0013] The clustering module is used to cluster latent variables using a Gaussian mixture model during training and map specific class labels to each cluster.

[0014] The clustering quality assessment module is used to calculate the silhouette coefficient of each task and identify tasks with low clustering accuracy;

[0015] The label correction module is used to readjust and correct tasks with low clustering accuracy through worker annotation information.

[0016] Beneficial effects of the present invention: The present invention combines the variational deep embedding model and the characteristics of crowdsourcing tasks, and proposes a label correction result aggregation model based on deep clustering. The present invention can not only use the high-dimensional features of the task for deep modeling, but also optimize the label inference process in combination with worker behavior data. By integrating task characteristics and worker information, this method can effectively eliminate low-quality annotations, improve the accuracy of task labels, and provide a more reliable data foundation for supervised learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A neural network model diagram provided for an embodiment of the present application. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] like Figure 1 As shown, the embodiment of the present application provides a method for aggregating label correction crowdsourcing results based on deep clustering, and the method comprises the following steps:

[0020] (1) Use feature extraction networks (e.g., ResNet, BERT, etc.) to extract crowdsourcing task features as the input data x of the model.

[0021] (2) Using the workers’ annotation information on tasks, we construct a penalty term based on the consistency of worker labels, referred to as the worker loss term. Combining the reconstruction loss and KL divergence in the original variational deep embedding (VaDE) model, we define the comprehensive loss function of this model, and train the model using the gradient descent method to optimize the distribution of the latent variable z and the clustering results.

[0022] (3) During the training process, the Gaussian mixture model (GMM) is used to cluster the latent variable z to obtain clusters. The worker's annotation information on the task is further combined to map a specific class label to each cluster.

[0023] (4) Calculate the silhouette coefficient of each task and filter out tasks with low clustering accuracy. For these tasks, readjust and correct them based on the workers’ annotation information to further improve the accuracy of the clustering structure.

[0024] In one embodiment, step (2) includes the following steps:

[0025] (2-1) Use the stacked autoencoder (SAE) to pre-train the encoder and decoder to obtain the initial latent variable z, and use GMM to fit the latent variable z to obtain the initial prior parameters of GMM Where π is the weight of a certain category distribution in GMM, μ c and are the corresponding means and variances;

[0026] (2-2) Input the input data x into the entire model, and after being processed by the encoder and decoder, the reconstructed data is obtained The model is trained by optimizing the loss function, which can be expressed as the following formula:

[0027] Loss_model=Loss_rescon+α*Loss_KL+β*Loss_crowd

[0028] Where α and β represent the weights of Loss_KL and Loss_crowd in the overall loss.

[0029] Furthermore, Loss_rescon is the reconstruction error term, calculated using the mean square error (MSE):

[0030]

[0031] where L is the number of Monte Carlo sampling, x and Represent the input data and reconstructed data respectively. The purpose of this loss is to ensure that the latent variable z can effectively capture the main information of the input data x.

[0032] Furthermore, Loss_KL is the KL divergence between the approximate posterior distribution and the GMM prior distribution, calculated by the following formula:

[0033] Loss_KL = D KL (q(z,c|x)||p(z,c))=T 1 +T 2

[0034] in

[0035]

[0036] γ c =q(c|x) Calculation method In step (3), C represents the number of mixed distributions and clusters, and N is the mean μ of the latent variable I and variance And the mean μ of a specific distribution corresponding to GMM c and variance σ c 2 Dimension.

[0037] D KL (q(z,c|x)||p(z,c)) is the KL divergence between the prior distribution p(z,c) and the posterior distribution q(z,c|x). This loss is used to distribute the latent variable z on the mean field of a mixed Gaussian, ensuring the smoothness and separability of the latent variable z.

[0038] Furthermore, Loss_crowd is the worker loss term, which optimizes the latent variable by constraining the worker's judgment information on the task category. The worker loss is calculated as follows:

[0039]

[0040] in:

[0041]

[0042] R j = {i|A(j,i) = 1, i∈Ι} represents the set of tasks marked by worker j, S j ={(u,v)|u,v∈R j ,u≠v} represents the set R j The possible task pair combinations in ju represents the label that worker j annotates for task u, l jv represents the label annotated by worker j for task v, N represents the dimension of the latent variable, and d(u,v) calculates the Euclidean distance between task u and task v.

[0043] The worker loss is calculated by the category judgment between any two tasks marked by each crowdsourcing worker. Specifically, if the two tasks have the same label, they are considered to belong to the same category, and their latent variables z should be similar; conversely, if the labels are different, their latent variables z should be significantly different. Introducing the judgment bias of crowdsourcing workers on task categories can optimize the distribution of latent variables z, so that it can achieve better clustering effects in the future.

[0044] In one embodiment, step (3) includes the following steps:

[0045] (3-1) In GMM, assuming that all samples are divided into C categories, and each input data x obeys a Gaussian distribution, then the learning process of GMM is to estimate the probability density and corresponding weights of C Gaussian distributions. Substitute each input data x into the obtained C Gaussian distributions to calculate the probability that it belongs to each cluster;

[0046]

[0047] Where C represents the number of sub-distributions, π c is the weight of the mixture distribution, and the sum of all weights Ν(z|μ c ,Σ c ) represents the cth Gaussian distribution, μ c and Σ c are the mean and variance of the corresponding distribution, respectively, and q(c|x) is the calculated probability that the input data x belongs to cluster c.

[0048] Finally, the input data x is assigned to the cluster with the highest probability.

[0049] c = argmax c q(c|x)

[0050] By calculating the posterior probability q(c|x), we can get the probability that the input data x belongs to each cluster, and then select the cluster label with the highest probability to assign it to the input data x, thereby realizing the clustering of the input data x.

[0051] (3-2) Combined with the workers’ task labeling information, count the distribution of all task labels in each cluster c, and calculate the specific class label for each cluster c;

[0052]

[0053] For each cluster c obtained by GMM, create a vector τ (c) , used to record the tasks contained in the cluster, M (c) represents the set of all tasks in cluster c, T i represents the set of all workers who have marked task i, l ij represents the label that worker j annotates for task i, represents the number of tasks of category k in cluster c, K (c) Represents the specific class label of the calculated cluster c.

[0054] Through this step, the distribution of each class label of all tasks in each cluster can be obtained, and the class label with the largest number is assigned to the cluster. In this way, the specific class label of the cluster can be obtained, and then the specific class label of each task can be obtained.

[0055] In one embodiment, step (4) includes the following steps:

[0056] (4-1) Calculate the silhouette coefficient of each task and identify tasks with low clustering accuracy;

[0057] The calculation formula of task silhouette coefficient is as follows:

[0058]

[0059] S(i) represents the silhouette coefficient value, which ranges from -1 to 1. a(i) represents the average distance between task i and other sample points in the same cluster, that is, the intra-cluster distance, and b(i) represents the average distance between task i and the nearest cluster in other clusters, that is, the inter-cluster distance.

[0060] The silhouette coefficient is used as an indicator to measure the quality of clustering results because it comprehensively considers the closeness between the sample and the samples in the same cluster and the separation between the sample and the samples in other clusters.

[0061] (4-2) A portion of the tasks with low silhouette coefficients are readjusted and corrected through worker labels.

[0062]

[0063] in represents the number of workers who label task i as category k, w i represents the set of workers who provide labels for task i, l ij K represents the label that worker j annotates for task i. (i) represents the label reassigned to task i.

[0064] The core idea of ​​this process is to identify tasks with large ambiguity in clustering results by quantifying the uncertainty of task clustering results. For these tasks, the labels provided by workers are reused to correct them, thereby further improving the accuracy of clustering results.

[0065] Based on the concept of the above method, the embodiment of the present application also provides a label correction crowdsourcing result aggregation system based on deep clustering, including:

[0066] The feature extraction module is used to extract the features of the crowdsourcing task using the feature extraction network as the input data of the model;

[0067] The loss function construction module is used to construct a worker loss term based on the consistency of worker labels, and combine the reconstruction loss and KL divergence in the variational deep embedding model to define a comprehensive loss function;

[0068] The training module is used to train the model using the gradient descent method;

[0069] The clustering module is used to cluster latent variables using a Gaussian mixture model during training and map specific class labels to each cluster.

[0070] The clustering quality assessment module is used to calculate the silhouette coefficient of each task and identify tasks with low clustering accuracy;

[0071] The label correction module is used to readjust and correct tasks with low clustering accuracy through worker annotation information.

[0072] Furthermore, to verify the effectiveness of the present invention, three real data sets were used for comparative experiments, and the accuracy was used as the evaluation index, as shown in the following table:

[0073] Table 1 Comparison of the accuracy of the data set in this application embodiment and other crowdsourcing result aggregation methods

[0074] Method Labelme Music Cifar-10N MV 0.7693±0.0021 0.7171±0.0035 0.9111±0.0018 DS 0.7833±0.0037 0.7342±0.0023 0.8938±0.0035 GLAD 0.7770±0.0025 0.7912±0.0051 0.9096±0.0032 GTIC 0.7703±0.0029 0.7257±0.0032 0.9145±0.0025 Agg_net 0.8421±0.0062 0.7685±0.0055 0.8554±0.0046 CoNAL 0.8168±0.0032 0.7693±0.0036 0.8356±0.0065 CCC 0.8479±0.0047 0.7629±0.0043 0.8625±0.0065 ours 0.8506±0.0046 0.8055±0.0015 0.9176±0.0023

[0075] Table 1 intuitively compares the average accuracy of the present invention with the other seven methods. From the last column of Table 1, it can be seen that the performance of the present invention on the three data sets is better than that of the other methods.

[0076] The above is only a preferred embodiment of the present invention. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Any technician familiar with the art can make many possible changes and modifications to the technical solution of the present invention by using the above disclosed methods and technical contents without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.

Claims

1. A label correction crowdsourcing result aggregation method based on deep clustering, characterized in that: The following steps are involved: Use the feature extraction network to extract the features of the crowdsourcing task as the input data of the model; Construct a worker loss term based on worker label consistency, combine the reconstruction loss and KL divergence in the variational deep embedding model, define a comprehensive loss function, and train the model using the gradient descent method; During the training process, the Gaussian mixture model is used to cluster the latent variables and map specific class labels to each cluster. The silhouette coefficient of each task is calculated, tasks with low clustering accuracy are identified, and these tasks are readjusted and corrected using worker annotation information.

2. The label correction crowdsourcing result aggregation method based on deep clustering according to claim 1 is characterized by: The feature extraction network is a deep learning network, including ResNet, BERT, and CNN, which is used to extract high-dimensional features of the task and use the extracted features as input data of the model.

3. The label correction crowdsourcing result aggregation method based on deep clustering according to claim 1 is characterized by: The comprehensive loss function is composed of reconstruction loss, KL divergence and worker loss terms, and includes the following parts: The reconstruction loss is used to measure the difference between the input data and the reconstructed data, which is calculated by the mean square error; KL divergence is used to measure the difference between the approximate posterior distribution and the Gaussian mixture model prior distribution; The worker loss term constrains the distribution of latent variables through the consistency of workers' task annotations, which is achieved by calculating the Euclidean distance between tasks.

4. The method for aggregating label correction crowdsourcing results based on deep clustering according to claim 3, characterized in that: The worker loss term constrains the distribution of latent variables by calculating the consistency of workers' task annotations. Specifically, for any two tasks, if the labels annotated by the workers are the same, the latent variables should be closer; if the annotated labels are different, the latent variables should be more dispersed, thereby optimizing the clustering effect.

5. The label correction crowdsourcing result aggregation method based on deep clustering according to claim 1 is characterized by: Before model training, the encoder and decoder are pre-trained layer by layer using a stacked autoencoder to obtain the initial latent variables and initial prior parameters of the Gaussian mixture model, including the following steps: Pre-train the encoder and decoder layer by layer to obtain the initial latent variable distribution; The Gaussian mixture model is used to fit the initial latent variables to obtain the initial prior parameters, including the weight, mean and variance of each Gaussian distribution.

6. The label correction crowdsourcing result aggregation method based on deep clustering according to claim 1 or 5, characterized in that: When using the Gaussian mixture model to cluster latent variables, the probability of each input data belonging to each cluster is calculated, and the data is assigned to the cluster with the highest probability, specifically: For each input data, calculate the probability that it belongs to each Gaussian distribution; Assign the input data to the cluster with the highest probability.

7. The method for aggregating label correction crowdsourcing results based on deep clustering according to claim 6, characterized in that: By counting the label distribution of tasks in each cluster, the class label with the largest number is assigned to the cluster, so as to determine the specific class label for each task, specifically: Count the task labels in each cluster and calculate the number of each label; The label with the largest number is taken as the final label of the cluster.

8. The method for aggregating label correction crowdsourcing results based on deep clustering according to claim 1, characterized in that: The silhouette coefficient is used as an indicator to measure the quality of clustering results. The silhouette coefficient of a task is determined by calculating the average distance from the task to other sample points in the same cluster and the average distance to the nearest cluster in other clusters.

9. The label correction crowdsourcing result aggregation method based on deep clustering according to claim 1 or 8, characterized in that: For tasks with low silhouette coefficients, adjustments and corrections are made by reusing the labels provided by workers, specifically: For tasks whose silhouette coefficient is lower than a preset threshold, recalculate their labels; Reassign task labels based on the category with the largest number of labels annotated by statistical workers.

10. A label correction crowdsourcing result aggregation system based on deep clustering, characterized in that: include: The feature extraction module is used to extract the features of the crowdsourcing task using the feature extraction network as the input data of the model; The loss function construction module is used to construct a worker loss term based on the consistency of worker labels, and combine the reconstruction loss and KL divergence in the variational deep embedding model to define a comprehensive loss function; The training module is used to train the model using the gradient descent method; The clustering module is used to cluster latent variables using a Gaussian mixture model during training and map specific class labels to each cluster. The clustering quality assessment module is used to calculate the silhouette coefficient of each task and identify tasks with low clustering accuracy; The label correction module is used to readjust and correct tasks with low clustering accuracy through worker annotation information.

Citation Information

Cited By

  • Graph-guided data labeling governance method for curved shell defect detection

    CN120726425B