A Rapid-Startup Interactive Relationship Annotation and Extraction Framework

Through the technology of few-sample learning and active learning, combined with manual proofreading information, an interactive relationship annotation and extraction framework with fast start is built, which solves the high cost problem of the supervised learning method during cold start, and realizes a relationship extraction system with low labor costs and high performance.

CN114118092BActive Publication Date: 2025-07-08SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111474423.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-07-08
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The existing supervised learning relationship extraction method is expensive to start cold and has poor transferability during actual business implementation, making it difficult to popularize, especially when applied in specific fields, requiring a large amount of labeled data and manpower investment.

Method used

Using a small sample learning technology combined with active learning technology, through manual proofreading of information and meta-learning methods, a small amount of labeled data is used for model pre-training and fine-tuning, to build a fast-start interactive relationship labeling and extraction framework, reduce the need for labeled data and improve model performance.

Benefits of technology

The fast start and low labor cost characteristics of the relationship extraction system are realized, which effectively solves the cold start problem, improves the performance of the model and reduces the need for fine-tuning time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114118092B_ABST
    Figure CN114118092B_ABST
Patent Text Reader

Abstract

The present invention relates to a fast-start interactive relationship annotation and extraction framework, which includes the following steps: S1: Pre-train a named entity recognition model using a general named entity recognition data set; S2: Pre-train a few-shot relationship extraction model using a general relationship extraction data set; S3: Set the relationships to be extracted and a small amount of annotated data; S4: Perform data preprocessing on the text to be extracted; S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted; S6: Manually pair the entities; S7: Perform preliminary relationship extraction on the pairing results; S8: Manually proofread the relationship extraction results; S9: Fine-tune the few-shot relationship extraction model; S10: Repeat S4 to S9 until all the texts to be extracted are processed. This solution overcomes the disadvantages of high start-up costs and heavy human cost investment in the prior art, and realizes relationship annotation and extraction with the characteristics of fast start-up and low labor costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interactive relationship annotation and extraction framework based on human-computer interaction and with fast startup, belonging to the technical fields of computer artificial intelligence and natural language processing. Background Art

[0002] Relation extraction is an important subtask in the field of information extraction, playing a key role in the construction of multiple application scenarios such as knowledge graphs, dialogue systems, and knowledge question-answering systems, and also having extensive application value in fields such as medical, military, and finance. The main goal of relation extraction is to extract the triple structure of <subject, predicate, object> or <head, relation, tail> from text. The common form of relation extraction is to input a text and two entities involved in it, and judge whether the text content describes the relationship existing between the two entities and infer what kind of relationship exists.

[0003] In past research, supervised learning relation extraction methods have achieved good results. However, the supervised learning method itself depends on a large amount of labeled data, and the acquisition of these labeled data often requires extremely high human and material resources, which makes the cold start cost of the supervised learning method very high in actual business implementation and difficult to popularize. In addition, the transferability of the supervised learning method is also poor. For example, a supervised learning relation extraction model trained with general domain corpus is difficult to be applied to a specific domain. Therefore, there are many problems in the actual application of the supervised learning relation extraction method.

[0004] Few-shot learning technology is an effective method to solve the cold start data requirement problem. Meta-learning technology is an important type of few-shot learning technology. Using meta-learning, the relation extraction task can be pre-trained to obtain a set of initial parameters of the relation extraction model. This set of initial parameters can converge quickly using a small amount of training data, thus solving the cold start data requirement problem in the relation extraction task.

[0005] Active learning technology is widely used to reduce the annotation cost and has achieved good results in the field of computer vision. Active learning technology obtains data samples that are difficult to classify by calculating metrics in the machine learning process. Then, these samples are manually proofread and reviewed, and the proofread data is reused for the training of the machine learning model, thereby improving the performance of the machine learning model and reducing the amount of labeled data. Summary of the Invention

[0006] The present invention precisely addresses the problems existing in the prior art and provides a fast-start interactive relationship annotation and extraction framework. This technical solution proposes an active learning technique that uses manually verified information to reduce annotation data and improve model performance, and combines few-shot relationship extraction techniques to enhance the cold-start performance of the model. Based on the framework disclosed in the present invention, the disadvantages of the high cold-start cost and heavy human cost investment of the existing relationship extraction system can be effectively overcome, and a relationship annotation and extraction system with the characteristics of fast start and low labor cost can be realized.

[0007] To achieve the above object, the technical solution of the present invention is as follows. A fast-start interactive relationship annotation and extraction framework includes the following steps:

[0008] S1: Pre-train a named entity recognition model using a general named entity recognition dataset;

[0009] S2: Pre-train a few-shot relationship extraction model using a general relationship extraction dataset;

[0010] S3: Set the relationships to be extracted and a small amount of annotated data;

[0011] S4: Perform data preprocessing on the text to be extracted;

[0012] S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted;

[0013] S6: Manually pair the entities;

[0014] S7: Perform preliminary relationship extraction on the pairing results;

[0015] S8: Manually proofread the relationship extraction results;

[0016] S9: Fine-tune the few-shot relationship extraction model;

[0017] S10: Repeat S4 to S9 until all the text to be extracted has been processed. This framework proposes an active learning technique that uses manually verified information to reduce annotation data and improve model performance, and combines few-shot relationship extraction techniques to enhance the cold-start performance of the model. The model is fine-tuned using the verified data to improve the extraction effect of the model. Based on the framework disclosed in the present invention, the disadvantages of the high cold-start cost and heavy human cost investment of the existing relationship extraction system can be effectively overcome, and a relationship annotation and extraction system with the characteristics of fast start and low labor cost can be realized.

[0018] As an improvement of the present invention, step S1: Use a general named entity recognition dataset to pre-train a named entity recognition model, and construct a fast-start interactive relation annotation and extraction framework, which includes: a named entity recognition model, a few-shot relation extraction model, a text repository to be processed, a general named entity recognition dataset, a general relation extraction dataset, and a dedicated relation extraction data repository. In addition, the framework also includes an artificial proofreading interaction method, a meta-learning training method, a parameter update method, and active learning.

[0019] As an improvement of the present invention, in step S2, use the general relation extraction dataset to pre-train the few-shot relation extraction model. Specifically, construct the named entity recognition model Net in the framework ner , and pre-train it using a general domain named entity class recognition dataset; construct the few-shot relation extraction model Net in the framework re , first train it in a meta-learning manner using the general domain relation extraction dataset to obtain the initial parameters θ0, and then fine-tune the parameters θ0 of Net re to obtain the parameters θ1.

[0020] As an improvement of the present invention, step S3: Set the relations to be extracted and a small amount of labeled data; select a text S to be extracted from the text repository to be processed.

[0021] As an improvement of the present invention, step S4: Perform data preprocessing on the text to be extracted; use the pre-trained named entity class recognition model to perform named entity recognition on the text to be extracted. For the convenience of annotators, mark the named entity recognition results {e1, e2,... e n} in the text to be processed.

[0022] As an improvement of the present invention, step S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted. Specifically, the annotator manually pairs the named entities recognized in S4, that is, selects the head and tail entity pairs {e h , e t} that need to perform relation extraction. The selected entity pairs {e h , e t}, the sentence S containing the entity pairs, the entity types {C h , C t}, and the relative positions of the entities in the sentence {Pos h , Pos t} are used as the input for the next step of relation extraction.

[0023] As an improvement of the present invention, step S6: perform manual pairing on entities, and the annotator manually pairs named entities: click on two entities in the text in sequence, and the entity clicked first is the head entity e h , with the corresponding type being C h , and the entity clicked later is e t , with the corresponding type being C t ; The entity calculates its relative position in the sentence according to the relationship between the clicked entity and the sentence where it is located. The specific method is as follows:

[0024] 1) If both e h and e t are included in sentence S, then mark the serial number of the first character of sentence S as 0, the serial number of the second character as 1, and mark the entire sentence S in sequence. Then Pos h = {h start , h end}. Pos t = {t start , t end}. Where h start is the serial number of the starting character of e h , h end is the serial number of the ending character of e h , t start is the serial number of the starting character of e t , and t end is the serial number of the ending character of e t ;

[0025] 2) If e h and e t are included in two consecutive sentences S1 and S2, then connect S1 and S2, denoted as S. If the length of S is less than or equal to the preset threshold L, process it according to the method described in 1); if the length of S is greater than the preset threshold L, no pairing is formed, and the annotator is prompted;

[0026] 3) If e h and e t are included in two non - consecutive sentences S1 and S2, then connect S1, the intermediate sentence, and S2, denoted as S. If the length of S is less than or equal to the preset threshold L, process it according to the method described in 1); if the length of S is greater than the preset threshold L, no pairing is formed, and the annotator is prompted.

[0027] As an improvement of the present invention, step S7: perform preliminary relationship extraction on the pairing results, specifically as follows,

[0028] S7: The annotator manually proofreads the extraction results in S6 to confirm whether the predicted relationship is correct. If the prediction is correct, record the result relationship as And record it together with the input of step S6 as a set of correct relation extraction results If the prediction is incorrect, the annotator needs to manually select the correct result relation from the candidate relation set R And record And record it together with the input of step S6 as a set of proofread relation extraction results

[0029] As an improvement of the present invention, step S8: Manually proofread the relation extraction results; specifically as follows:

[0030] The annotator manually proofreads the extraction results in S7 to confirm the predicted relation Whether it is correct, and the specific method is as follows:

[0031] 1) If the prediction is correct, record the result relation as And record it together with the input of S5 as a set of correct relation extraction results

[0032] 2) If the prediction is incorrect, the annotator needs to manually select the correct result relation from the candidate relation set R And record And record it together with the input of S5 as a set of proofread relation extraction results

[0033] As an improvement of the present invention, step S9: Fine-tune the few-shot relation extraction model; specifically as follows:

[0034] 1) When the number of stored correct relation extraction results is less than K + And the number of proofread relation extraction results is less than K - Use all the data in the dedicated relation extraction data warehouse to fine-tune Net re The parameter update formula is as follows:

[0035]

[0036] Among them, θ i-1 Is the parameter before update, θ i Is the parameter after update, D is all the data in the dedicated relation extraction data warehouse;

[0037] After the parameters are updated, take out all the proofread relation extraction results D in D - Use θ i To initialize the relation extraction model, and for D -Make a prediction, increment the error count for the results that are still predicted incorrectly by 1; and use the proofreading results that are still predicted incorrectly to fine-tune the parameters once.

[0038] 2) When the number of correct relation extraction results stored is greater than or equal to K + or the number of proofread relation extraction results is greater than or equal to K - at that time, randomly select K + correct relation extraction results from the correct relation extraction results D + Calculate the selection probability of each result from the proofread relation extraction results according to the following formula:

[0039]

[0040] where P i represents the probability that the i-th proofread relation extraction result is selected, and EC i represents the cumulative number of errors in the prediction after fine-tuning described in 1); calculate the probabilities of all proofread relation extraction results, and non-repeatedly select K - proofread relation extraction results according to the probability to form a fine-tuning data set, and fine-tune the parameters once.

[0041] After the parameters are updated, make a prediction for all proofread relation extraction results again as described in 1), increment the error count for the results that are still predicted incorrectly by 1; and use the proofreading results that are still predicted incorrectly to fine-tune the parameters once.

[0042] As an improvement of the present invention, step S10: The fine-tuned few-shot relation extraction model Net re is used for the extraction task of the subsequent text to be extracted.

[0043] Compared with the prior art, the present invention has the following advantages. The few-shot relation extraction described in the present invention is a relation extraction method for use in fields such as cold-start knowledge graph construction. In few-shot relation extraction, at system startup, the relations to be extracted need to be preset, which requires the system user to complete the entry of the relation names and relation descriptions of the relations to be extracted, and add a small amount of labeled extraction data to each relation to be extracted. Different from the supervised learning relation extraction method, the pre-training of few-shot relation extraction is carried out on datasets of different relations. This dataset needs to be organized by task. The basic unit of data organization is the task, and the composition of each task follows an experimental setting called N-way-K-shot. N-way means that the number of categories to be classified in each task is N, and K-shot means that in the support set of these tasks, each category contains K training samples. The support set refers to the data used to train the relation extraction model in a task, and the query set refers to the data used to test the relation extraction model. The tasks used in the meta-training stage are called meta-training tasks, and the tasks used in the meta-test stage are called meta-test tasks, and they contain different classification categories. Generally, several meta-training tasks form a round, and the meta-training stage contains one or more rounds. In the meta-test stage, the meta-learner quickly applies the meta-knowledge obtained in the meta-training stage to the training process of the support set, quickly fine-tunes the relation extraction model by using an extremely small amount of the support set, and uses the fine-tuning result to perform relation classification prediction on the dedicated relation extraction dataset. The framework has the characteristics of fast startup and low labor cost. The realization of these two characteristics depends on: 1) Using the few-shot relation extraction algorithm to achieve the characteristic of fast startup. The initialization parameters of the relation extraction model are pre-trained by the meta-learning method, and the initialization parameters can be used to complete fast startup with an extremely small amount of data; 2) Using human-computer interaction to provide supervision information, and using active learning technology to screen data to fine-tune the model, improving the model performance and reducing the time required for fine-tuning. Compared with the prior art, this framework considers the supervision information brought by human-computer interaction, uses the few-shot relation extraction technology to solve the cold-start problem in the startup stage of the relation extraction system, uses active learning technology to reduce the amount of data required for the model to be fine-tuned, improves the model performance and reduces the time required for fine-tuning. It effectively solves the cold-start problem of the relation extraction system and proposes a relation annotation and extraction system framework with the characteristics of fast startup and low labor cost. Based on this framework, a relation annotation and extraction system with the above characteristics can be developed, and such systems have broad application value and significance in knowledge graph construction, dialogue system construction, and question-answer system construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of a common form of relation extraction;

[0045] Figure 2It is the overall schematic diagram of the framework proposed in the present invention;

[0046] Figure 3 It is the structure of the relationship extraction neural network classification model in the embodiment of the present invention. Specific implementation manners

[0047] To deepen the understanding of the present invention, the following will make a detailed description of this embodiment with reference to the accompanying drawings.

[0048] Embodiment 1: Refer to Figures 1-3 , a fast-start interactive relationship annotation and extraction framework, including the following steps:

[0049] S1: Use a general named entity recognition dataset to pre-train the named entity recognition model;

[0050] S2: Use a general relationship extraction dataset to pre-train the few-shot relationship extraction model;

[0051] S3: Set the relationship to be extracted and a small amount of labeled data;

[0052] S4: Perform data preprocessing on the text to be extracted;

[0053] S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted;

[0054] S6: Manually pair the entities;

[0055] S7: Perform preliminary relationship extraction on the pairing results;

[0056] S8: Manually proofread the relationship extraction results;

[0057] S9: Fine-tune the few-shot relationship extraction model;

[0058] S10: Repeat S4 to S9 until all the texts to be extracted are processed.

[0059] Among them, the fast-start interactive relationship annotation and extraction framework in step S1 is as Figure 2 shown;

[0060] Construct a fast-start interactive relationship annotation and extraction framework, where the framework includes: a named entity recognition model, a few-shot relationship extraction model, a text warehouse to be processed, a general named entity recognition dataset, a general relationship extraction dataset, and a dedicated relationship extraction data warehouse. In addition, the framework also includes an artificial proofreading interaction method, a meta-learning training method, a parameter update method, and active learning.

[0061] Step S2 pre-trains the few-shot relation extraction model using a general relation extraction dataset as follows. Build the named entity recognition model Net in the framework ner , and pre-train it using a general domain named entity class recognition dataset; build the few-shot relation extraction model Net in the framework re , first train it in a meta-learning manner using a general domain relation extraction dataset to obtain the initial parameters θ0, and then fine-tune the parameters θ0 of Net re using a dedicated relation extraction data warehouse to obtain the parameters θ1.

[0062] The dedicated relation extraction data warehouse mentioned in step S2 only contains relation types formulated by the annotator himself / herself and a small amount of corresponding annotated data.

[0063] In the step of building the few-shot relation extraction model Net mentioned in step S2 re , the meta-learning training method is described in detail as follows

[0064]

[0065] Furthermore, in the step of building the few-shot relation extraction model Net mentioned in step S2 re , the meta-learning fine-tuning method is described in detail as follows

[0066]

[0067]

[0068] Step S3: Set the relation to be extracted and a small amount of annotated data; select a text S to be extracted from the text warehouse to be processed

[0069] Step S4: Perform data preprocessing on the text to be extracted; use the pre-trained named entity class recognition model to perform named entity recognition on the text to be extracted. For the convenience of the annotator, mark the named entity recognition results {e1, e2,... e n} in the text to be processed. The specific method is: mark the recognized entities in the text with different colors according to different types, where the entity types are predefined

[0070] Step S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted. Specifically, the annotator manually pairs the named entities recognized in S4, that is, selects the head and tail entity pairs {e h , e t} that need to perform relation extraction. The entity pairs {e h , e t} selected by the annotator, as well as the sentence S containing the entity pair and the entity types {Cj , C t {Relative position of the entity in the sentence} Pos h , Pos t} is used as the input for the next step of relation extraction.

[0071] Step S6: Manually pair the entities. The annotator manually pairs the named entities: Click on two entities in the text in sequence. The entity clicked first is the head entity e h , with the corresponding type C h , and the entity clicked later is e t , with the corresponding type C t ; Calculate the relative position of the entity in the sentence according to the relationship between the clicked entity and the sentence it is in. The specific method is as follows:

[0072] 1) If both e h and e t are included in sentence S, then mark the serial number of the first character of sentence S as 0, the serial number of the second character as 1, and mark the entire sentence S in sequence. Then Pos h = {h start , h end}, Pos t = {t start , t end}. Where h start is the serial number of the starting character of e h , h end is the serial number of the ending character of e h , t start is the serial number of the starting character of e t , t end is the serial number of the ending character of e t ;

[0073] 2) If e h and e t are included in two consecutive sentences S1 and S2, then connect S1 and S2, denoted as S. If the length of S is less than or equal to the preset threshold L, process it according to the method described in 1); if the length of S is greater than the preset threshold L, no pairing is formed and the annotator is prompted;

[0074] 3) If e h and e t are included in two non - consecutive sentences S1 and S2, then connect S1, the intermediate sentences, and S2, denoted as S. If the length of S is less than or equal to the preset threshold L, process it according to the method described in 1); if the length of S is greater than the preset threshold L, no pairing is formed and the annotator is prompted.

[0075] Step S7: Conduct preliminary relation extraction on the pairing results, specifically as follows,

[0076] S7: The annotator manually proofreads the extraction results in S6 to confirm the predicted relationship to be correct. If the prediction is correct, record the resulting relationship as and record it together with the input in step S6 as a set of correct relationship extraction results If the prediction is incorrect, the annotator needs to manually select the correct resulting relationship from the candidate relationship set R and record and record it together with the input in step S6 as a set of proofread relationship extraction results

[0077] Step S8: Manually proofread the relationship extraction results; specifically as follows:

[0078] The annotator manually proofreads the extraction results in S7 to confirm the predicted relationship to be correct. The specific approach is as follows:

[0079] 1) If the prediction is correct, record the resulting relationship as and record it together with the input in S5 as a set of correct relationship extraction results

[0080] 2) If the prediction is incorrect, the annotator needs to manually select the correct resulting relationship from the candidate relationship set R and record and record it together with the input in S5 as a set of proofread relationship extraction results

[0081] Step S9: Fine-tune the few-shot relationship extraction model; specifically as follows:

[0082] 1) When the number of stored correct relationship extraction results is less than K + and the number of proofread relationship extraction results is less than K - then use all the data in the dedicated relationship extraction data warehouse to fine-tune Net re The parameter update formula is as follows:

[0083]

[0084] where θ i-1 is the parameter before update, θ i is the parameter after update, and D is all the data in the dedicated relationship extraction data warehouse;

[0085] After the parameters are updated, take out all the proofread relationship extraction results D in D - and use θ iInitialize the relation extraction model and perform predictions on D - Make predictions, increment the error count for the results that are still predicted incorrectly by 1; and use the proofreading results that are still predicted incorrectly to fine-tune the parameters once;

[0086] 2) When the number of correct relation extraction results stored is greater than or equal to K + or the number of proofread relation extraction results is greater than or equal to K - At this time, randomly select K + correct relation extraction results from the correct relation extraction results D + Calculate the selection probability of each result from the proofread relation extraction results according to the following formula:

[0087]

[0088] where P i represents the probability that the i-th proofread relation extraction result is selected, and EC i represents the cumulative number of errors in the fine-tuned prediction described in 1); calculate the probabilities of all proofread relation extraction results and select K - non-repeated proofread relation extraction results according to the probability to form a fine-tuning dataset and fine-tune the parameters once;

[0089] After the parameters are updated, make another prediction on all proofread relation extraction results as described in 1), increment the error count for the results that are still predicted incorrectly by 1; and use the proofreading results that are still predicted incorrectly to fine-tune the parameters once.

[0090] Step S10: The few-shot relation extraction model Net after fine-tuning re is used for the subsequent extraction task of the text to be extracted.

[0091] Specific embodiment: Refer to Figure 1 — Figure 3 , in this embodiment, the general named entity recognition datasets are MUC-6 and MUC-7 datasets, the general relation extraction dataset is FewRel, and the text segments in the text repository to be processed are from Wikipedia. The named entity recognition model is a sequence labeling model based on conditional random fields, and the relation extraction model adopts the Prototypical Network structure. Among them, the PCNN model is used to encode the text sentences and entities, and its structure is as Figure 3 shown, and the GloVe word embedding vector is used as the pre-trained word vector to encode the words in the sentence.

[0092] In this embodiment, a fast-start interactive relation annotation and extraction framework provided by the present invention is applied, and its overall framework is as Figure 2 shown, and it specifically includes the following steps:

[0093] Step 1) Pre-train the named entity recognition model using a general named entity recognition dataset;

[0094] Set 7 types of named entities: Location, Person, Organization, Money, Percent, Date, Time; Use the training sets of MUC 6 and MUC 7 datasets as training data; In the training data file, the first in each line is a character, and the second is its label, separated by a space, and the label type is the combined label of "BIO-type"; Use the BERT pre-trained model to perform embedded encoding on the training data, and use the conditional random field as the decoder; Use the negative logarithm of the path score commonly used in the named entity class recognition task as the loss function;

[0095] Step 2) Pre-train the few-shot relation extraction model using a general relation extraction dataset;

[0096] Pre-train the few-shot relation extraction model using the FewRel dataset. In the training file, each training instance consists of three parts: the natural language sentence S, the head entity e h and the tail entity e t . Among them, each entity contains three parts: the entity mention, the relative position of the first character of the entity in S, and the relative position of the last character of the entity in S; Use GloVe to perform embedded vector encoding on the words in the natural language sentence S, and use PCNN to encode the embedded encoding of the sentence and the entity relative position information to obtain the entity vector; Use BiLSTM to encode the relation label to obtain the relation vector; Calculate the similarity score by using the cosine similarity to calculate the entity vector and the relation vector; Use the margin loss function as the loss function for training;

[0097] Step 3) Set the relations to be extracted and a small amount of labeled data;

[0098] The annotator manually adds several relations to be extracted, and each relation needs to add a relation name and a relation description; In addition, for each relation to be extracted, 10 pieces of labeled data are added, and the data structure of the labeled data is the same as that in the FewRel dataset;

[0099] Step 4) Perform data preprocessing on the text to be extracted;

[0100] The text to be extracted needs to be cleaned and preprocessed first to facilitate the subsequent extraction process. The specific operations are as follows:

[0101] 1) Filter out special symbols and punctuation. Use regular expression matching to filter out all other special symbols in the text except commas, periods, question marks, quotation marks, and book titles;

[0102] 2) To ensure the extraction speed, it will be stored in units of paragraphs. For long texts, in this embodiment, with a limit of 300 words, the long text is divided into paragraphs by sentences, and the end of the last natural sentence within the limit is used as the segmentation condition;

[0103] Step 5) Use a named entity recognition model to perform named entity recognition on the text to be extracted;

[0104] For the results of named entity recognition, according to the "BIO-Type" tags, different entities are divided, and the types of each entity are marked; in this embodiment, the entities are marked with special colors. Among the pre-defined seven types of entities, each type of entity is defined with a different color and is displayed together with the text;

[0105] Step 6) Manually pair the entities;

[0106] In this embodiment, the annotator clicks on two entities in the text in sequence for pairing. The entity clicked first is the head entity e h , with the corresponding type C h , and the entity clicked later is e t , with the corresponding type c t ; The entity calculates its relative position in the sentence according to the relationship between the clicked entity and the sentence it is in. The specific method is as follows:

[0107] 1) If both e h and e t are included in sentence S, then the serial number of the first character of sentence S is marked as 0, the serial number of the second character is marked as 1, and the entire sentence S is marked in sequence. Then Pos h = {h start , h end}, Pos t = {t start , t end}. Where h start is the serial number of the starting character of e h , h end is the serial number of the ending character of e h , t start is the serial number of the starting character of e t , and t end is the serial number of the ending character of e t ;

[0108] 2) If e h and e tIf it is included in two connected sentences S1 and S2, then S1 and S2 are connected and denoted as S. If the length of S is less than or equal to the preset threshold L, it is processed according to the method described in 1); if the length of S is greater than the preset threshold L, it does not form a pair and the annotator is prompted.

[0109] 3) If e h and e t are included in two non - connected sentences S1 and S2, then S1, the intermediate sentence, and S2 are connected and denoted as S. If the length of S is less than or equal to the preset threshold L, it is processed according to the method described in 1); if the length of S is greater than the preset threshold L, it does not form a pair and the annotator is prompted.

[0110] Step 7) Perform preliminary relationship extraction on the pairing results.

[0111] Input the pairing results into the relationship extraction model. The relationship extraction model extracts the sentence S, the entities e h and e t , the relative positions of the entities Pos h and Pos t , and encodes them into an instance vector; extracts the candidate relationships from the dedicated relationship extraction data warehouse and encodes them into relationship vectors; calculates the cosine similarity between the instance vector and the relationship vectors, and selects the relationship represented by the relationship vector with the highest similarity as the predicted relationship

[0112] Step 8) Manually proofread the relationship extraction results.

[0113] In this embodiment, the annotator needs to manually proofread the extraction results in step 7) to confirm the predicted relationship whether it is correct. The specific method is as follows:

[0114] 1) If the prediction is correct, record the result relationship as and record the input of step 7) together as a set of correct relationship extraction results

[0115] 2) If the prediction is incorrect, the annotator needs to manually select the correct result relationship from the candidate relationship set R and record and record the input of step 7) together as a set of proofread relationship extraction results

[0116] Step 9) Fine - tune the few - shot relationship extraction model.

[0117] In this embodiment, when the number of relationship extraction results stored in the dedicated relationship extraction data warehouse reaches a certain amount, the data therein is used to fine-tune the few-shot relationship extraction model Net re The specific method is as follows:

[0118] 1) When the number of correct relationship extraction results stored is less than K + and the number of verified relationship extraction results is less than K - use all the data in the dedicated relationship extraction data warehouse to fine-tune Net re The parameter update formula is as follows:

[0119]

[0120] where θ i-1 is the parameter before update, θ i is the parameter after update, and D is all the data in the dedicated relationship extraction data warehouse;

[0121] After the parameters are updated, take out all the verified relationship extraction results D - in D, initialize the relationship extraction model with θ i and make predictions on D - Increase the error count of the still wrongly predicted results by 1; and use the still wrongly predicted verified results to fine-tune the parameters once;

[0122] 2) When the number of correct relationship extraction results stored is greater than or equal to K + or the number of verified relationship extraction results is greater than or equal to K - randomly select K + correct relationship extraction results from the correct relationship extraction results D + and calculate the selection probability of each result from the verified relationship extraction results according to the following formula:

[0123]

[0124] where P i represents the selection probability of the i-th verified relationship extraction result, and EC i represents the cumulative number of errors in the prediction after the fine-tuning described in 1); calculate the probabilities of all the verified relationship extraction results and non-repeatedly select K - verified relationship extraction results according to the probabilities to form a fine-tuning data set and fine-tune the parameters once;

[0125] After the parameters are updated, make predictions on all the verified relationship extraction results again as described in 1), increase the error count of the still wrongly predicted results by 1; and use the still wrongly predicted verified results to fine-tune the parameters once;

[0126] Step 10) Repeat Step 4) to Step 9) until all the texts to be extracted are processed completely.

[0127] In summary, based on the supervision information brought by human-computer interaction, the method of the present invention proposes an interactive relationship annotation and extraction framework that combines few-shot relation extraction technology and active learning technology. This method uses few-shot relation extraction technology to solve the cold start problem in the startup phase of the relation extraction system, and uses active learning technology to reduce the amount of data required for model fine-tuning, improving the performance of the model and reducing the time required for fine-tuning. Based on this method, the disadvantages of high cold start cost and high human cost input in the existing relation extraction system can be effectively overcome, and a relation annotation and extraction system with the characteristics of fast startup and low labor cost can be realized. Such systems have broad application value and application prospects in natural language processing fields such as knowledge graph construction, dialogue system construction, and question answering system construction, as well as in the field of information extraction.

[0128] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention, and equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A fast-start interactive relationship annotation and extraction framework, characterized in that, Including the following steps: S1: Pre-train the named entity recognition model using a general named entity recognition dataset; S2: Pre-train the few-shot relation extraction model using a general relation extraction dataset; S3: Set the relations to be extracted and a small amount of labeled data; S4: Perform data preprocessing on the text to be extracted; S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted; S6: Manually pair the entities; S7: Perform preliminary relation extraction on the pairing results; S8: Manually check the relation extraction results; S9: Fine-tune the few-shot relation extraction model; S10: Repeat S4 to S9 until all texts to be extracted are processed; Among them, step S9: Fine-tune the few-shot relation extraction model; specifically as follows: 1) When the number of correctly extracted relationship results stored is less than K + and the number of relationship results extracted for proofreading is less than K - at this time, use the data in all dedicated relationship extraction data warehouses to fine-tune Net re The parameter update formula is as follows: Among them, θ i-1 is the parameter before update, θ i is the parameter after update, and D is the data in all dedicated relation extraction data warehouses; After the parameters are updated, extract all the proofreading relationship extraction results D in D - Take them out and use θ i Initialize the relationship extraction model and perform prediction on D - Increment the error count of the results that are still predicted incorrectly by 1; and use the proofreading results that are still predicted incorrectly to fine-tune the parameters once; 2) When the number of correct relation extraction results stored is greater than or equal to K + or the number of proofread relation extraction results is greater than or equal to K - At this time, randomly select K + correct relation extraction results from the correct relation extraction results D + Calculate the selection probability of each result from the proofread relation extraction results according to the following formula: Among them, P i represents the probability that the i-th proofreading relation extraction result is selected, and EC i represents the cumulative number of errors in the fine-tuned prediction described in 1); calculate the probabilities of all proofreading relation extraction results, and non-repeatedly select K - proofreading relation extraction results according to the probabilities to form a fine-tuning data set, and perform one fine-tuning on the parameters; After the parameters are updated, make a prediction on all the checked relation extraction results again as described in 1), add 1 to the number of errors of the results that are still predicted incorrectly; and use the checked results that are still predicted incorrectly to fine-tune the parameters once.

2. The quickly-started interactive relationship annotation and extraction framework according to claim 1, characterized in that Step S1: Pre-train the named entity recognition model using a general named entity recognition dataset, and construct a fast-start interactive relation annotation and extraction framework, which includes: a named entity recognition model, a few-shot relation extraction model, a text warehouse to be processed, a general named entity recognition dataset, a general relation extraction dataset, and a dedicated relation extraction data warehouse.

3. The fast-start interactive relationship annotation and extraction framework according to claim 2, characterized in that Step S2 pre-trains the few-shot relation extraction model using a general relation extraction dataset, specifically as follows. Build the named entity recognition model Net in the framework ner , and pre-train it using a general domain named entity class recognition dataset; build the few-shot relation extraction model Net in the framework re . First, train it in a meta-learning manner using a general domain relation extraction dataset to obtain the initial parameters θ0, and then fine-tune the parameters θ0 of Net re using a dedicated relation extraction data repository to obtain the parameters θ1.

4. The quickly-started interactive relationship annotation and extraction framework according to claim 3, wherein Step S3: Set the relations to be extracted and a small amount of labeled data; select a text S to be extracted from the text warehouse to be processed.

5. The quickly-started interactive relationship annotation and extraction framework according to claim 4, characterized in that Step S4: Perform data preprocessing on the text to be extracted; use a pre-trained named entity recognition model to perform named entity recognition on the text to be extracted, and mark the results {e1, e2,... e n} of the named entity recognition in the text to be processed.

6. The quickly-started interactive relationship annotation and extraction framework according to claim 5, characterized in that Step S5: Use the named entity recognition model to perform named entity recognition on the text to be extracted. Specifically, the annotator manually pairs the named entities recognized in S4, that is, selects the head and tail entity pairs {e h , e t} that need to perform relationship extraction. The entity pairs {e h , e t} selected by the annotator, the sentence S containing the entity pairs, the entity types {C h , C t}, and the relative positions of the entities in the sentence {Pos h , Pos t} are used as the input for the next step of relationship extraction.

7. The quickly-started interactive relationship annotation and extraction framework according to claim 6, characterized in that Step S6: Manually pair the entities. The annotator manually pairs the named entities by clicking on two entities in sequence in the text. The entity clicked first is the head entity e h , with the corresponding type C h , and the entity clicked later is e t , with the corresponding type C t ; The entity calculates its relative position in the sentence according to the relationship between the clicked entity and the sentence it is in. The specific method is as follows: 1) If e h and e t are both included in sentence S, then mark the serial number of the first character of sentence S as 0, the serial number of the second character as 1, and mark the entire sentence S in sequence. Then Pos h ={h start , h end}, Pos t ={t start , t end}, where h start is the serial number of the starting character of e h , h end is the serial number of the ending character of e h , t start is the serial number of the starting character of e t , and t end is the serial number of the ending character of e t ; 2) If e h and e t are included in two consecutive sentences S1 and S2, then connect S1 and S2, denoted as S. If the length of S is less than or equal to the preset threshold L, process it according to the method described in 1); if the length of S is greater than the preset threshold L, no pairing is formed, and the annotator is prompted; 3) If e h and e t are included in two non - connected sentences S1 and S2, then connect S1, the intermediate sentence, and S2, denoted as S. If the length of S is less than or equal to the preset threshold L, then process it according to the method described in 1); if the length of S is greater than the preset threshold L, then no pairing is formed and the annotator is prompted.

8. The quickly-started interactive relationship annotation and extraction framework according to claim 7, characterized in that Step S7: Perform preliminary relation extraction on the pairing results, specifically as follows S7: The annotator manually proofreads the extraction results in S6 to confirm the predicted relationship whether it is correct. If the prediction is correct, record the resulting relationship as and record it together with the input in step S6 as a set of correct relationship extraction results If the prediction is incorrect, the annotator needs to manually select the correct resulting relationship from the candidate relationship set R and record and record it together with the input in step S6 as a set of proofread relationship extraction results 9. The fast-start interactive relationship annotation and extraction framework according to claim 8, characterized in that Step S8: Manually check the relation extraction results; specifically as follows: The annotator manually proofreads the extraction results in S7 to confirm the predicted relationships Whether it is correct, and the specific method is as follows: 1) If the prediction is correct, record the result relationship as and record it together with the input of S5 as a set of correct relationship extraction results 2) If the prediction is incorrect, the annotator needs to manually select the correct result relationship from the candidate relationship set R and record and record it together with the input of S5 as a set of proofread relationship extraction results

Citation Information

Patent Citations

  • Remote supervision entity relationship extraction method based on man-machine interaction

    CN109635108A

  • Entity extraction method and device

    CN113128227A