Data mining and package data marking method and system
A crowdsourcing labeling and data mining technology, applied in the field of data labeling, can solve the problems of low efficiency, small amount of data processing, large investment, etc., and achieve the effects of improving labeling quality, facilitating review and modification, and reducing labeling costs
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Image
Examples
Embodiment 1
[0041] Such as figure 1 As shown, a data labeling method based on data mining and crowdsourcing is disclosed in this embodiment, including:
[0042] S101. Acquiring raw data to be labeled;
[0043] S102. Use an integrated algorithm to classify and crowdsource distribution of the original data;
[0044] S103. Obtain the crowdsourcing labeling results, use the integrated algorithm to automatically review the crowdsourcing labeling results, screen out the problem labeling results, and mark the problem labeling results;
[0045] S104. Output the crowdsourcing labeling results that have been automatically reviewed, and the crowdsourcing labeling results include question labeling results.
[0046] Wherein, the marked data scope includes but not limited to text, image, audio, statistical data and other data.
[0047]In the existing crowdsourcing technology, the annotators are ordinary users from the Internet, and the quality of their annotations is not guaranteed, and the annotati...
Embodiment 2
[0079] Such as figure 2 As shown, a data labeling system based on data mining and crowdsourcing is disclosed in this embodiment, including:
[0080] Grabbing module 201, for obtaining the original data to be marked;
[0081] A distribution module 202, configured to use an integrated algorithm to classify and distribute the raw data through crowdsourcing;
[0082] The processing module 203 is used to obtain the crowdsourcing labeling results, use the integrated algorithm to automatically review the crowdsourcing labeling results, screen out the problem labeling results, and mark the problem labeling results;
[0083] The output module 204 is configured to output crowdsourcing annotation results that have been automatically reviewed, and the crowdsourcing annotation results include question annotation results.
[0084] The data labeling system disclosed in this embodiment includes: a capture module 201, used to obtain the original data to be labeled; a distribution module 202...
PUM
Abstract
Description
Claims
Application Information
- R&D Engineer
- R&D Manager
- IP Professional
- Industry Leading Data Capabilities
- Powerful AI technology
- Patent DNA Extraction
Browse by: Latest US Patents, China's latest patents, Technical Efficacy Thesaurus, Application Domain, Technology Topic, Popular Technical Reports.
© 2024 PatSnap. All rights reserved.Legal|Privacy policy|Modern Slavery Act Transparency Statement|Sitemap|About US| Contact US: help@patsnap.com