Corpus labeling method and device, server and storage medium

A corpus labeling and corpus technology, applied in the field of information processing, can solve problems such as single corpus labeling results, cognitive level and operating habits affecting the quality of corpus labeling, and difficulty in judging the accuracy of labeling results, so as to ensure high quality and accuracy Effect

CN110069602AActive Publication Date: 2019-07-30CHINANETCENT TECH
13 Cites 4 Cited by

Patent Information

Authority / Receiving Office
CN · China
Current Assignee / Owner
Publication Date
2019-07-30

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of information processing, in particular to a corpus labeling method and device, a server and a storage medium. The corpus annotation method comprises the steps of acquiring an even number of manual annotation results of an initial corpus and a model annotation result of the initial corpus, wherein the model annotation result of the initial corpus is obtained according to a preset annotation model, and the preset annotation model is obtained by training a plurality of manually annotated initial corpora; and obtaining a unique labeling result meeting a preset condition from all labeling results including the manual labeling result and the model labeling result, and taking the labeling result as a final labeling result of the initial corpus. According to the embodiment of the invention, the high-quality corpus annotation result of the initial corpus can be obtained, and the influence of a single annotator on the corpus annotation quality is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The embodiments of the present invention relate to the technical field of information processing, and in particular to a corpus tagging method, device, server and storage medium. Background technique

[0002] Natural language processing refers to the computer receiving input in the form of natural language, and internally processes and calculates the natural language through user-defined algorithms to return the results expected by the user. It can usually be applied to fields such as text retrieval, machine translation, and information question answering. Users usually define the algorithm by establishing an algorithm model, and the established algorithm model needs to be trained through a large number of labeled original language materials; labeling the original language materials refers to processing the original corpus, and combining various representations of language features The additional codes are marked on the corresponding language component...

Examples

Embodiment Construction

[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention more clear, various implementation modes of the present invention will be described in detail below in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that, in each implementation manner of the present invention, many technical details are provided for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following implementation modes, the technical solution claimed in this application can also be realized. The division of the following embodiments is for the convenience of description, and should not constitute any limitation to the specific implementation of the present invention, and the various embodiments can be combined and referred to each other on the premise of no contradiction.

[0023] The first embodi...