A Crowdsourcing Annotation Data Integration Method Based on Task Difficulty and Annotator Ability
A crowdsourced labeling and data integration technology, applied in the field of crowdsourced labeling data integration based on task difficulty and the ability of labelers, can solve problems such as lack of task difficulty, accuracy deviation, and labeler evaluation deviation, and achieve convenient difficulty Accurate evaluation, labeling results, and ability to evaluate objectively and accurately
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Publication Date
- 2017-08-08
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention belongs to the technical field of data labeling, and in particular relates to a crowdsourcing labeling data integration method based on task difficulty and labeler's ability. Background technique
[0002] High-quality labeled datasets are very important resources in the field of computer research and applications. Algorithms in the fields of computer vision, artificial intelligence, and machine learning are mostly trained and optimized based on corresponding labeled data sets. Obtaining high-quality and large-scale labeled datasets quickly and efficiently has always been a concern of various researchers. The traditional way to obtain labeled datasets is to hire experts to manually label the datasets. The annotation data obtained in this way is of high quality, but the annotation takes a long time, and the financial cost of hiring experts is also very large.
[0003] In recent years, with the development of crowdsourcing technology, the...
Examples
Embodiment Construction
[0040] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0041] The flow process of the inventive method is as figure 1 As shown, it specifically includes the following steps:
[0042] Step (1): The assessment of task difficulty is from the collected labeled data set Find the difficulty set of all tasks {D i |i∈[1,a]}; where is the tagging result of the i-th task by the w-th tagger, D i Indicates the difficulty of the i-th task, a is the total number of tasks, and W is the total number of annotators. The method is described below taking the difficulty of the i-th task as an example, and the steps are as follows:
[0043] 1-1: Collect the collected annotation data Perform statistics to obtain the number K of the types of labeling results made by all labelers for the i-th task i , and the proportion set ...