Data mining and package data marking method and system

A crowdsourcing labeling and data mining technology, applied in the field of data labeling, can solve the problems of low efficiency, small amount of data processing, large investment, etc., and achieve the effects of improving labeling quality, facilitating review and modification, and reducing labeling costs

Inactive Publication Date: 2017-03-08
SHENZHEN GOWILD ROBOTICS CO LTD
View PDF8 Cites 34 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

Disadvantages are large investment, low efficiency, and small amount of data processing
Moreover, the annotators are all ordin

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Data mining and package data marking method and system
  • Data mining and package data marking method and system

Examples

Experimental program
Comparison scheme
Effect test

Embodiment 1

[0041] Such as figure 1 As shown, a data labeling method based on data mining and crowdsourcing is disclosed in this embodiment, including:

[0042] S101. Acquiring raw data to be labeled;

[0043] S102. Use an integrated algorithm to classify and crowdsource distribution of the original data;

[0044] S103. Obtain the crowdsourcing labeling results, use the integrated algorithm to automatically review the crowdsourcing labeling results, screen out the problem labeling results, and mark the problem labeling results;

[0045] S104. Output the crowdsourcing labeling results that have been automatically reviewed, and the crowdsourcing labeling results include question labeling results.

[0046] Wherein, the marked data scope includes but not limited to text, image, audio, statistical data and other data.

[0047]In the existing crowdsourcing technology, the annotators are ordinary users from the Internet, and the quality of their annotations is not guaranteed, and the annotati...

Embodiment 2

[0079] Such as figure 2 As shown, a data labeling system based on data mining and crowdsourcing is disclosed in this embodiment, including:

[0080] Grabbing module 201, for obtaining the original data to be marked;

[0081] A distribution module 202, configured to use an integrated algorithm to classify and distribute the raw data through crowdsourcing;

[0082] The processing module 203 is used to obtain the crowdsourcing labeling results, use the integrated algorithm to automatically review the crowdsourcing labeling results, screen out the problem labeling results, and mark the problem labeling results;

[0083] The output module 204 is configured to output crowdsourcing annotation results that have been automatically reviewed, and the crowdsourcing annotation results include question annotation results.

[0084] The data labeling system disclosed in this embodiment includes: a capture module 201, used to obtain the original data to be labeled; a distribution module 202...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to view more

PUM

No PUM Login to view more

Abstract

The present invention provides a data mining and package data marking method, comprising: obtaining the original data to be marked; using the integration algorithm to perform classification and package distribution of the original data; obtaining the package marking result, and using the integration algorithm to perform automation checking of the packaging mark result, screening the problem marking result, and marking the problem marking result; outputting the package marking result through the checking, wherein the result includes a problem marking result. The problem marking result can be found in all the package marking results, and are marked to facliate checking and modulation of the problem marking result so as to greatly facilitate find the problem marking result, improve the ouput result marking quality and effectively reduce the marking cost.

Description

technical field [0001] The invention relates to the technical field of data labeling, in particular to a data labeling method and system based on data mining and crowdsourcing. Background technique [0002] In recent years, with the development of crowdsourcing technology, the use of crowdsourcing technology for data annotation has attracted the attention of researchers. Crowdsourcing technology is a distributed problem-solving method. The technology harnesses the wisdom and power of crowds to solve tasks that are difficult for computers, especially tasks that are easy for humans but very difficult for computers, such as data labeling and object recognition. Many labeling tasks, such as text labeling, image classification, etc., can be published on the Internet through crowdsourcing platforms, and are marked by ordinary users from the Internet. Ordinary users complete data labeling tasks and receive economic rewards from publishers. [0003] The advantage of the crowdsour...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to view more

Application Information

Patent Timeline
no application Login to view more
IPC IPC(8): G06F17/30
CPCG06F16/2465G06F40/166
Inventor 杨新宇王昊奋邱楠
Owner SHENZHEN GOWILD ROBOTICS CO LTD
Who we serve
  • R&D Engineer
  • R&D Manager
  • IP Professional
Why Eureka
  • Industry Leading Data Capabilities
  • Powerful AI technology
  • Patent DNA Extraction
Social media
Try Eureka
PatSnap group products