An automatic labeling system and method for cultural resource entity recognition training data
Through the automatic labeling system, prefix collection construction, keyword matching, interval merging, tool calling and format conversion modules are used to solve the problems of low efficiency and low accuracy of traditional manual labeling, and efficient and accurate generation of training data sets is achieved.
Patent Information
- Application Number
- CN202111279572.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-10-29
AI Technical Summary
The traditional manual annotation method is inefficient, low accuracy and high cost when labeling large-scale data, especially when labeling text content.
It provides an automatic annotation system for cultural resource entity recognition training data, including prefix collection construction module, keyword matching module, interval merging module, tool calling module and format conversion module. Through the collaborative work of these modules, it automatically recognizes and labels domain keywords and general proper nouns in text.
It significantly improves the labeling efficiency, reduces the error rate and labeling cost, and can efficiently generate high-quality training data sets.
Smart Images

Figure CN113886516B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer and artificial intelligence technology, and particularly relates to an automatic labeling system and method for cultural resource entity recognition training data. Background Art
[0002] In recent years, artificial intelligence technology has developed rapidly, and all walks of life have begun to integrate artificial intelligence technology for innovation and development. The core content of artificial intelligence lies in its algorithm or model. A model often requires a large amount of data to train the model in order to improve its intelligence. Therefore, data formatting and annotation is the task of the first stage of artificial intelligence application. At present, the formatting and annotation methods are mainly manual, one is full manual annotation, and the other is manual annotation with the assistance of annotation tools. Regardless of the annotation method, when encountering large-scale data that needs to be labeled, the labor cost will be very high according to the previous annotation method, and the efficiency will be low, and the accuracy cannot be guaranteed. This is a major problem faced in annotation.
[0003] The purpose of artificial intelligence technology is to enable machines to have human cognitive abilities. Human cognitive abilities are acquired through continuous learning. Similarly, machine cognitive abilities also need to be acquired through continuous learning, and labeled data is the learning material for machines. For example, if we want a machine to recognize a picture of a "dog", we can directly give it a picture of a puppy, and it will not be able to tell that it is a puppy. When we give a large number of labeled "dog" pictures to the machine and let it recognize and learn, the machine extracts the features of a large number of pictures and corresponds to the label of "dog". At this time, we give it another picture that the machine has never seen, and the machine will recognize the corresponding label based on the features of the picture. The training set and test set, collectively referred to as data sets, are the basis for training machine learning models. The accuracy of machine learning models is based on the scale of data sets and the accuracy of annotations. In addition to high-quality models, a high-performance artificial intelligence application also requires high-quality data sets to train the model. The higher the quality of the data set, the higher the accuracy of the model and the higher the value of the artificial intelligence application. Therefore, efficiently producing high-quality training sets is the basis of artificial intelligence. The traditional manual annotation method is very inefficient in producing training sets, especially when annotating text content. Summary of the invention
[0004] In order to solve the above problems, the present disclosure provides an automatic annotation system for cultural resource entity recognition training data, which includes a prefix set construction module, a keyword matching module, an interval merging module, a tool calling module and a format conversion module, wherein:
[0005] The prefix set construction module converts the keyword dictionary into a prefix set by using the prefix set construction algorithm through the keyword dictionary read in;
[0006] The keyword matching module receives the prefix set and the original text, identifies the domain keywords in the original text through the keyword matching algorithm, and records the position interval of the keywords in the original text into the information set;
[0007] The interval merging module receives the information set, solves the problem of keyword interval inclusion and intersection through the interval merging algorithm, finally generates a new information set, and saves the elements in the new information set into the analysis text;
[0008] The tool calling module is used to identify common proper nouns in the cultural field, add them to the new information set, and save the elements in the new information set into the analysis text;
[0009] The format conversion module converts the analysis text and the original text into mature annotated text through a format conversion algorithm.
[0010] The present disclosure also provides a method for automatically labeling cultural resource entity recognition training data, which comprises the following steps:
[0011] S100: converting the keyword dictionary into a prefix set by using a prefix set construction algorithm through the read keyword dictionary;
[0012] S200: receiving the prefix set and the original text, identifying the domain keywords in the original text through a keyword matching algorithm and recording their position intervals in the original text into an information set;
[0013] S300: receiving the information set, solving the problem of keyword interval inclusion and intersection by using interval merging algorithm, finally generating a new information set, and saving the elements in the new information set into the analysis text;
[0014] S400: Identify common proper nouns in the cultural field, add them to the new information set, and save the elements in the new information set into the analysis text;
[0015] S500: Convert the analyzed text and the original text into mature annotated text through a format conversion algorithm.
[0016] Through the above technical solution, this method can significantly improve the labeling efficiency and greatly reduce the error rate and labeling cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of an automatic labeling method for cultural resource entity recognition training data provided in one embodiment of the present disclosure;
[0018] Figure 2It is a structural diagram of an automatic labeling method for cultural resource entity recognition training data provided in one embodiment of the present disclosure;
[0019] Figure 3 is a flowchart of a prefix set construction algorithm in one embodiment of the present disclosure;
[0020] Figure 4 is a flowchart of a keyword matching algorithm in one embodiment of the present disclosure;
[0021] Figure 5 is a flow chart of an interval merging algorithm in one embodiment of the present disclosure;
[0022] Figure 6 It is a flow chart of a format conversion algorithm in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] The training of entity recognition models in the cultural field requires a large amount of training data. The traditional method of producing training data is mainly based on manual annotation, which has the problems of low annotation efficiency, high annotation error rate and high annotation cost. In order to solve this problem, we proposed a BIOES automated annotation system and method based on keyword matching and existing tool calls for the annotation task of resource entity recognition datasets in the cultural field.
[0024] In one embodiment, Figure 1 As shown, the present disclosure provides a method for automatically labeling cultural resource entity recognition training data, which includes the following steps:
[0025] S100: converting the keyword dictionary into a prefix set by using a prefix set construction algorithm through the read keyword dictionary;
[0026] S200: receiving the prefix set and the original text, identifying the domain keywords in the original text through a keyword matching algorithm and recording their position intervals in the original text into an information set;
[0027] S300: receiving the information set, solving the problem of keyword interval inclusion and intersection by using interval merging algorithm, finally generating a new information set, and saving the elements in the new information set into the analysis text;
[0028] S400: Identify common proper nouns in the cultural field, add them to the new information set, and save the elements in the new information set into the analysis text;
[0029] S500: Convert the analyzed text and the original text into mature annotated text through a format conversion algorithm.
[0030] As far as this embodiment is concerned, the method is based on the prefix set construction algorithm, keyword matching algorithm, interval merging algorithm and format conversion algorithm as well as existing natural language processing tools, and uses the BIOFS labeling system as the entity annotation scheme to realize the automatic annotation of entity recognition training data. Construct a prefix set through a keyword dictionary, and execute a keyword matching algorithm on the prefix set to realize the recognition of domain-specific nouns in raw corpus; realize the recognition of common proper nouns in raw corpus through the use of natural language processing tools; solve the problem of intersection and inclusion of keyword intervals through the interval merging algorithm; and convert the original text into annotated text through the format conversion algorithm combined with the analysis file. The structural diagram of this method is shown in the figure below. Figure 2 shown.
[0031] The purpose of converting to a prefix set is to facilitate keyword matching and improve keyword matching efficiency.
[0032] Previous keyword matching algorithms were mainly based on keyword lists. By traversing the keyword list, each keyword in the list is searched in the full text to determine whether the keyword exists in the original text. This method is highly complex and inefficient. This keyword matching algorithm is proposed based on the prefix set. With the prefix set, the original text can be matched by traversing the original text character by character only once, which improves the matching efficiency.
[0033] The previous interval merging algorithm first needs to sort the intervals, which requires additional sorting time and increases time overhead. However, this interval merging algorithm does not need to sort the intervals. The input of this interval merging algorithm comes from the output of the keyword matching algorithm. In the keyword matching algorithm, the original text is traversed once in order, so the keyword interval information obtained is already sorted. Therefore, the sorting step is eliminated in this interval merging algorithm, thereby reducing time overhead and improving the efficiency of interval merging.
[0034] In another embodiment, the prefix set construction algorithm puts keywords with the same prefix into the same group.
[0035] For this embodiment, W = {w0, w1, ..., w n} represents a keyword set, where w i ={c i,0 , c i,1 , ..., c i,len-1} represents the i-th keyword, where c i,j Represents the characters that make up the keyword, where 0≤j≤len-1, and len represents the length of the i-th keyword.
[0036] There may be keywords with a common prefix in the set W. In order to improve the matching efficiency of keywords, keywords with the same prefix are placed in the same group. Define a prefix set P to group keywords with a common prefix together. Its recursive definition is as follows:
[0037] P={{c0,p 0,1 ,f},{c1,p 0,2 ,f},...,{c n , p 0,n ,f}}
[0038] p i,j ={{c0, p i+1,0 ,f},{c1,p i+1,1 ,f},...,{c m , p i+1,m ,f}}
[0039] P represents the prefix set corresponding to all keywords, c i represents the i-th prefix set p 0,i The common prefix of , where 0≤i≤n, f represents the flag bit, which is used to determine the current character c i Whether it constitutes a keyword alone. i,j represents the jth prefix set of the i-th layer, c0,...,c n Indicates that there are n common prefixes in the i-th layer, and j represents the flag. For example, there are three keywords w1 = {a, b, c, d}, w2 = {a, b, e, g, h} and w3 = {h, i, j}, where the common prefix of w1 and w2 is 'ab'; w3 and other keywords have no common prefix. So the prefix set of these three keywords is expressed in the form of a dictionary as follows:
[0040]
[0041]
[0042] Algorithm 1 is the prefix set construction algorithm.
[0043] Input: keyword set W
[0044] Output: prefix set P
[0045]
[0046]
[0047] The flowchart of Algorithm 1 is as follows Figure 3 As shown. The initial input of Algorithm 1 is the keyword set W, and the output is the prefix set P. First, the set P is initialized, and the outermost loop reads each keyword w from the keyword set.i , the inner loop traverses the current keyword w i Each character c j Initially, let P′=P, which means the current prefix set is P. Each subsequent assignment of P′ indicates that it points to a new prefix set. If there is a prefix with c in the current prefix set P′, j For the prefix set p of the local prefix, let P′=p, indicating that the next character c j+1 As a local prefix, it can only exist in the prefix set p, or be added to p as a new local prefix; otherwise, it is determined whether w is read. i The last character of j If it is the last character, set the flag f = 1, otherwise set f = 0, and set p = {} to indicate that the character c j is the prefix set of the local prefix, p is empty, and c j , p and f are added to P'. When the loop ends, the prefix set P corresponding to the keyword set W is completely constructed, and finally the prefix set P is returned.
[0048] This module is the core module of the whole system. Before starting this module, we have collected and sorted out the keywords of various categories in the cultural field and stored them in different files by category. Different category keywords correspond to different prefix sets. By executing the prefix set construction algorithm multiple times, we can build prefix sets of all categories. Then, for each prefix set, we execute the keyword matching algorithm on the original text to obtain the information set.
[0049] In another embodiment, the elements in the information set are represented in the form of a tuple (b, e), where b represents the start index of the domain keyword in the original text, and e represents the end index of the domain keyword in the original text.
[0050] For this embodiment, the text is segmented into individual characters using a set C = {c0, c1, ..., c m}, represented by P = {p0, p1, ..., p n} represents the set of layer 0 prefixes after the flag bit is incorporated. Assume Indicates a keyword. If c x ∈p z , then it means that keyword w can only appear in the prefix set p z In, c x+1 Must be in p z In the next level prefix set, all the way to c y Entered p z In the prefix set of the yxth layer, of course, the prefix set of this layer contains the character c y, flag f = 1 and prefix set Indicates that this keyword has been matched successfully. In the same way, match all keywords, using the tuple (b, e) to represent the start index b and end index e of keyword w in the text. Use the set R = {r0, r1, ..., r n} represents the position information corresponding to all keywords in the text, where r i =(b,e).
[0051] Algorithm 2 is a keyword matching algorithm.
[0052] Input: prefix set P, text character set C
[0053] Output: Keyword matching result R
[0054]
[0055]
[0056] The flowchart of Algorithm 2 is as follows Figure 4 As shown in the figure, the initial input of Algorithm 2 is the prefix set P and the original text character set C. After the keywords are identified, a keyword information tuple (b, e) is constructed, and the information set R consisting of all keyword tuples is returned.
[0057] First, the algorithm is an outermost loop, traversing each character c in the character set C i ; Initialize the current prefix set p x,y For p 0,0 , determine the current character c i Whether it belongs to the prefix set p x,y , if c i does not belong to the prefix set, indicating that p x,y There is no character c i The keyword that starts with c continues to read the next character; if c i Belongs to the prefix set p x,y , and the current flag f value is 1, indicating that the current character c i A single keyword is constructed, and the binary group (i, i) corresponding to the current keyword is added to the result set R; if c i Belongs to the prefix set p x,y , indicating that there may be a i The keyword consisting of multiple characters is headed by p x,y =p x+1z , where p x+1,z Indicates that the current character c i The prefix set of the local prefix continues through a loop to read the current character c in the set C iThe following characters are represented by c k Indicates, get all the characters c i The keyword binary group (i, k) consisting of multiple characters is put into the information set R. When the outermost loop ends, all the keyword binary group information will be stored in the information set R, and finally R is returned.
[0058] In another embodiment, the elements in the new information set are represented by a triple (b, e, k), where b represents the start index of the domain keyword in the original text, e represents the end index of the domain keyword in the original text, and k represents the category information of the entity.
[0059] For this embodiment, the information set R = {r0, r1, .. ., r n}, r i =(b, e) means that the starting position index of the i-th keyword in the text is b and the ending position index is e. The intervals composed of position indexes may contain and intersect. For example, there are two intervals (b1, e1) and (b2, e2). If e1>b2 and e1<e2, it means If b1≤b2 and e2<e1, or b1<b2 and e2≤e1, or b1<b2 and e2<e1, it means For the former, we need to take the union of the two For the latter, only the large interval (b1, e1) needs to be retained.
[0060] Algorithm 3 is the interval merging algorithm.
[0061] Input: Information set R
[0062] Output: Information set R′ after interval merging
[0063] Add (b, e)0 to the result set R'
[0064]
[0065]
[0066] The flowchart of Algorithm 3 is as follows Figure 5 As shown in Figure 3, the initial input of Algorithm 3 is the information set R obtained by Algorithm 2, and the output is the information set R′ after interval merging.
[0067] The information set R obtained by Algorithm 2 has all elements sorted in ascending order, that is, all elements are sorted in ascending order according to the value of b. If the values of b are the same, they are sorted in ascending order according to the value of e. Because of this natural property, no additional sorting is required during the interval merging process. The purpose of sorting is to reduce judgment branches. For example, there are two intervals (b1, e1) and (b2, e2). It is possible that b1<b2 and e2<e1, indicating It is also possible that b2<b1 and e1<e2, indicating that Reducing the possible cases by sorting is an optimization of interval merging. First, initialize the information set R' with the first element (b, e) 0 in the information set R, use a loop to traverse all elements in R, set a flag value f = 0, initialize the variable r' with the last element in the information set R' and remove the last element from the information set R'. Then perform intersection judgment. If r'(e)>r(b) and r'(e)<r(e), it means that the two intervals intersect. Let the variable At the same time, the flag value f is set to 1.
[0068] If the two intervals do not intersect, then determine whether the two intervals are in a containment relationship. If r(b)≥r′(b) and r′(e)≥r(e), it means Set the flag value f to 1. Regardless of whether the above two situations are true or not, r' needs to be added to the information set R'. Finally, determine whether the flag value is 0. If it is 0, it means that the two intervals neither intersect nor contain each other. At this time, add the new element r to the information set R'. At the end of the loop, a new information set R' will be obtained, which can be returned.
[0069] The core of this module is the interval merging algorithm. The elements in the information set obtained in Algorithm 2 are tuples (b, e). After interval merging, the problems of interval intersection and inclusion are solved. The purpose of this method is to obtain annotated text, that is, to classify and label the entities found in the original text. Therefore, it is not enough to have only the location information of the entity in the original text. Here, the variable k is introduced to represent the category information of the entity, and the tuple (b, e) is expanded to a triple (b, e, k). Finally, the triple information corresponding to all entities is written into the analysis text to prepare for the subsequent format conversion.
[0070] In another embodiment, a natural language processing tool is called to identify the common entities involved in the cultural field, and the start and end position information of these entities in the original text and the entity category information are recorded and represented by a triple (b, e, k), where b represents the index of the start position of the entity in the original text, and e represents the index of the end position of the entity in the original text. In the recognition process, the category information k of the domain keyword is obtained at the same time, indicating the category to which the entity belongs. Finally, the triple is written into the analysis text.
[0071] As far as this embodiment is concerned, there are many relatively mature natural language processing tools at present, which can be directly used to perform operations such as sentence segmentation, word segmentation, part-of-speech tagging and general domain named entity recognition on the original text. The general domain named entity categories that these natural language processing tools can recognize include names, place names, organization names, institution names, time, location, works, places, etc. For the general entities involved in the cultural field, existing tools can be used directly to perform entity recognition. This method recognizes the general proper nouns involved in the cultural field by calling existing natural language processing tools.
[0072] In order to facilitate the study of automatic annotation methods, this method determines seven entity categories in the field of Shaanxi Province's food culture, and automatically annotates the obtained Shaanxi Province food culture raw corpus according to these seven entity categories. Entity categories include: food name, food category, raw materials, related figures, origin, spread region, and origin dynasty. Food names, such as honey pumpkin, rose mirror cake, Hengshan mutton, etc.; food categories, such as special dishes, innovative dishes, antique dishes, etc.; raw materials, such as oil, salt, sauce, vinegar, sugar, etc.; related figures, such as Kou Zhun, Zhang Caifeng, Nai Jiangti, etc., who are related to the production and development of crystal cakes; origin places, such as Liquan, the birthplace of baked noodles, Binxian, the birthplace of imperial noodles, and Zhonglou, the birthplace of small milk cakes; spread regions, such as southern Shaanxi, Guanzhong, northern Shaanxi, etc.; origin dynasties, such as crystal cakes originated in the Song Dynasty, water basin mutton originated in the Shang and Zhou Dynasties, and Mashi originated in the Yuan Dynasty.
[0073] For example, in the cultural field, there are three special categories of general entities, namely, "character", "place of origin" and "dynasty of origin". "Character" can be converted to "name" category, "place of origin" can be converted to "place" category, and "dynasty of origin" can be converted to "time" category. These three categories belong to general entity categories. For the first two categories of entities, the tool is directly called to identify them. When calling the tool for testing, if the category of "dynasty of origin" is only identified as the "time" category, the identification range will be expanded. For example, words such as "today", "tomorrow", "2015", "this month" and "this quarter" will be identified as the "time" category, but they are not dynasties. Another problem found during the test is that when the "dynasty" keyword is identified, sometimes the irrelevant words before and after it are also identified as the "dynasty" keyword. Therefore, for dynasties, we still need to maintain a keyword dictionary and the prefix set corresponding to the dictionary. Only the "time" keywords that belong to the keyword dictionary or can be successfully matched in the prefix set can be marked as the "dynasty of origin" category.
[0074] In another embodiment, the format conversion algorithm uses the BIOES tag system.
[0075] As far as this embodiment is concerned, the core of this module is the format conversion algorithm, which uses the analysis text to convert the original text into annotated text. The BIOES tag system is used in this method. The following is the meaning of each BIOES tag:
[0076] B, which stands for Begin, means start
[0077] I, which stands for Intermediate, means intermediate
[0078] E, which stands for End, indicates the end
[0079] S, which stands for Single, indicates a single character
[0080] 0, which means Other, is used to mark irrelevant characters
[0081] The set T = {B, I, E, S, O} represents the label set, and the set M = {m0, m1, ..., m n}, where m i = (b, e, k) indicates that the i-th keyword in the text has a starting position of b, an ending position of e, and a category of k, representing a text analysis set; set C = {c0, c1, ..., c n} represents the original text, where c i Represents the characters in the text; set Q = {q0, q1, ..., q n} represents the conversion result set, where qi =(c, k, t) represents the conversion information corresponding to the i-th character c in the text. The category to which the character belongs is k and the corresponding label is t.
[0082] Algorithm 4 is a format conversion algorithm.
[0083] Input: label set T, text analysis set M, original text set C
[0084] Output: conversion result Q
[0085]
[0086]
[0087] The flowchart of Algorithm 4 is as follows Figure 6 As shown in Figure 4, the initial input of Algorithm 4 is the tag set T, the text analysis set M, the original text character set C, and the output is the format conversion result Q.
[0088] First, each character c in the character set C i Constructed into a triplet q i =(c i , k, O), used to initialize W. The initial k represents the empty type, and O represents all characters c in the initial state i All are irrelevant characters. Loop through the keyword triple information m from the text analysis set M i =(b, e, k). If b and e are equal, it means m i The keyword represented is composed of a single character, update If b and e are not equal, it means m i The keyword represented by it is composed of multiple characters. The first character of the keyword corresponds to The last character corresponds to q corresponding to characters other than the first and last characters j =(c j , m i (k), I), where m i (b)<j<m i (e).
[0089] In another embodiment, an automatic annotation system for cultural resource entity recognition training data includes a prefix set construction module, a keyword matching module, an interval merging module, a tool calling module and a format conversion module, wherein:
[0090] The prefix set construction module converts the keyword dictionary into a prefix set by using the prefix set construction algorithm through the keyword dictionary read in;
[0091] The keyword matching module receives the prefix set and the original text, identifies the domain keywords in the original text through the keyword matching algorithm, and records the position interval of the keywords in the original text into the information set;
[0092] The interval merging module receives the information set, solves the problem of keyword interval inclusion and intersection through the interval merging algorithm, finally generates a new information set, and saves the elements in the new information set into the analysis text;
[0093] The tool calling module is used to identify common proper nouns in the cultural field, add them to the new information set, and save the elements in the new information set into the analysis text;
[0094] The format conversion module converts the analysis text and the original text into mature annotated text through a format conversion algorithm.
[0095] Although the embodiments of the present invention are described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields, and the above specific embodiments are only illustrative and instructive, rather than restrictive. A person of ordinary skill in the art can also make many forms under the guidance of this specification and without departing from the scope of protection of the claims of the present invention, all of which belong to the protection of the present invention.
Claims
1. An automatic annotation system for cultural resource entity recognition training data, which includes a prefix set construction module, a keyword matching module, an interval merging module, a tool calling module and a format conversion module, wherein: The prefix set construction module converts the keyword dictionary into a prefix set by using the prefix set construction algorithm through the keyword dictionary read in; The keyword matching module receives the prefix set and the original text, identifies the domain keywords in the original text through the keyword matching algorithm, and records the position interval of the keywords in the original text into the information set; The interval merging module receives the information set, solves the problem of keyword interval inclusion and intersection through the interval merging algorithm, finally generates a new information set, and saves the elements in the new information set into the analysis text; The tool calling module is used to identify common proper nouns in the cultural field, add them to the new information set, and save the elements in the new information set into the analysis text; The format conversion module converts the analysis text and the original text into mature annotated text through a format conversion algorithm.
2. The system according to claim 1, wherein the prefix set construction algorithm is to put keywords with the same prefix into the same group.
3. The system according to claim 1, wherein the elements in the information set are in the form of tuples. is expressed in the form of Indicates the starting index of the domain keyword in the original text. Indicates the ending index of the domain keyword in the original text.
4. The system according to claim 1, wherein the elements in the new information set are triples. Indicates that Indicates the starting index of the domain keyword in the original text. Indicates the end index of the domain keyword in the original text. Indicates the category information of the field keyword.
5. The system according to claim 1, wherein the format conversion algorithm uses a BIOES tag system.
6. A method for automatically labeling cultural resource entity recognition training data, comprising the following steps: S100: converting the keyword dictionary into a prefix set by using a prefix set construction algorithm through the read keyword dictionary; S200: receiving the prefix set and the original text, identifying the domain keywords in the original text through a keyword matching algorithm and recording their position intervals in the original text into an information set; S300: receiving the information set, solving the problem of keyword interval inclusion and intersection by using interval merging algorithm, finally generating a new information set, and saving the elements in the new information set into the analysis text; S400: Identify common proper nouns in the cultural field, add them to the new information set, and save the elements in the new information set into the analysis text; S500: Convert the analyzed text and the original text into mature annotated text through a format conversion algorithm.
7. The method according to claim 6, wherein the prefix set construction algorithm is to put keywords with the same prefix into the same group.
8. The method according to claim 6, wherein the elements in the information set are in the form of binary tuples. is expressed in the form of Indicates the starting index of the domain keyword in the original text. Indicates the ending index of the domain keyword in the original text.
9. The method according to claim 6, wherein the elements in the new information set are triples. Indicates that Indicates the starting index of the domain keyword in the original text. Indicates the end index of the domain keyword in the original text. Indicates the category information of the field keyword.
10. The method according to claim 6, wherein the format conversion algorithm uses the BIOES tag system.
Citation Information
Patent Citations
Data stream predicting method based on rule antecedent generation tree matching
CN104809184A
Labeling method and device of electronic map
CN106484847A