A Chinese Discourse-Level Event Extraction Method and Device Incorporating Lexical Knowledge
By integrating vocabulary knowledge, the BMES set is constructed and the TF-IDF value is calculated, which solves the problem of low accuracy caused by ignoring word-level information in the prior art, and achieves higher precision Chinese-level event extraction.
Patent Information
- Application Number
- CN202111355005.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-11-16
AI Technical Summary
The existing Chinese-level event extraction method ignores word-level information, resulting in low accuracy.
By establishing an event template, collecting and annotating text, constructing a BMES set, calculating the TF-IDF value as word weight, and splicing and fusing the word list features with character-level features, and using the stochastic gradient descent method to train the neural network for event extraction.
It improves the accuracy of event extraction, avoids the influence of word segmentation errors, and improves the accuracy of Chinese-page event extraction.
Smart Images

Figure CN114036908B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and more particularly to a Chinese document-level event extraction method and device incorporating vocabulary knowledge. Background Art
[0002] Event extraction is an important subfield of information extraction, and its main goal is to study how to present the event information contained in unstructured text in a structured form. Event extraction is a basic task in natural language processing, and the valuable structured event information extracted is an important supplement to existing knowledge resources and is widely used as an upstream task in downstream tasks such as knowledge graph construction, text summarization, and information retrieval.
[0003] In recent years, with the booming development of text digitization, a large amount of unstructured text information has accumulated on the Internet. How to extract structured information from these unstructured texts has become a hot research direction. Event extraction mainly studies how to extract structured information from unstructured text, so the event extraction task has also become a hot research topic. With the continuous development of deep neural networks, significant progress has been made in the document-level event extraction task. However, due to some challenging factors in Chinese itself, including text ambiguity and word segmentation errors, this task is far from being solved.
[0004] As the name implies, the text granularity processed by document-level event extraction is at the document level. The general processing process can be summarized into the following parts: 1) Text mapping: Map the text into a text vector through some existing text-to-vector methods to serve subsequent related tasks; 2) Entity recognition: Extract event trigger word entities and event element entities through named entity recognition methods; 3) Event extraction: Identify events by using information related to event trigger word entities to determine the event type, and then classify event element entities for event elements to correspond event elements with corresponding event element roles one by one.
[0005] In the process of document-level event extraction, entity recognition plays an important role in event extraction, and the results of the extracted event trigger word entities and event element entities have an important impact on subsequent tasks. The existing document-level event extraction entity recognition subtask uses a character-level entity recognition model, which regards the entity recognition task as a character-level sequence labeling task. For example, Chinese Patent Publication No. CN112231447A discloses a method and system for event extraction of Chinese documents, which is based on character-level feature extraction. Although it has achieved good performance, these algorithms ignore word-level information, which has a very important auxiliary role in text understanding. Ignoring word-level information will result in low accuracy of Chinese document-level event extraction. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that the existing Chinese text-level event extraction method ignores word-level information, resulting in low accuracy.
[0007] The present invention solves the above technical problems through the following technical means: A Chinese text-level event extraction method incorporating vocabulary knowledge, the method comprising the following steps:
[0008] Step 1: Establish an event template;
[0009] Step 2: Collect texts and annotate the texts according to the event template, determine event trigger word entities and event element entities, and save the annotated texts in json format;
[0010] Step 3: Read the annotated texts, preprocess the texts and label the preprocessed texts according to the annotated entities;
[0011] Step 4: Convert the labeled texts into corresponding vectors to obtain character-level features;
[0012] Step 5: Construct a corresponding BMES set for each character in the labeled text. B represents all matching words starting with the corresponding character, M represents all matching words with the corresponding character in the middle, E represents all matching words ending with the corresponding character, and S is all matching words consisting of the corresponding character alone. If a word set is empty, use NONE to represent it; Encode the words in each set, convert them into corresponding text vectors, calculate the TF-IDF value of each word in each set as the weight of each word, and perform weighted summation on the words in each set to obtain vocabulary features; Concatenate and fuse the vocabulary features and character-level features to achieve feature extraction;
[0013] Step 6: Extract event trigger word entities and event element entities corresponding to events in the text;
[0014] Step 7: Event extraction;
[0015] Step 8: Input the extracted features, event trigger word entities, event element entities, and events into a neural network to train the network by the stochastic gradient descent method. The trained network is used as an optimized model, and the optimized model is used for Chinese text-level event extraction.
[0016] The present invention extracts character-level features, constructs a corresponding BMES set for each character, calculates the TF-IDF value of each word in each set as the weight of each word, performs weighted summation on the words in each set to obtain the vocabulary feature, splices and fuses the vocabulary feature with the character-level feature to achieve feature extraction. The extracted features incorporate vocabulary information, avoid the impact of word segmentation errors on event extraction, and improve the extraction accuracy.
[0017] Further, the event template is an asset freezing event, and the event elements included in the asset freezing event are the party whose assets are frozen, the frozen shares, the freezing start time, and the freezing end time.
[0018] Even further, the event trigger words include asset freezing.
[0019] Further, the preprocessing of the text includes: dividing the text into a test set, a training set, and a validation set, and the ratio of the test set, the training set, and the validation set is 7:2:1. Each text in the test set, the training set, and the validation set is defined to contain 64 sentences, and each sentence contains 128 words. Thus, each file is converted into a 64×128-dimensional vector. If it exceeds, it is truncated, and if it is insufficient, it is padded with 0s.
[0020] Even further, the tagging of the preprocessed text according to the labeled entities includes: adding the label B at the beginning of the labeled entity in the preprocessed text, adding the label E at the end of the entity, and adding the label O in front of the content that does not belong to the labeled entity.
[0021] Further, the initial learning rate in the stochastic gradient descent method is 0.0001, and the number of training times is 100 times. After reaching the number of training times, the final model converges to the optimal.
[0022] The present invention also provides a Chinese discourse-level event extraction device incorporating vocabulary knowledge. The device includes:
[0023] A template establishment module for establishing an event template;
[0024] A labeling module for collecting text and labeling the text according to the event template, determining the event trigger word entity and the event element entity, and saving the labeled text in json format;
[0025] A tagging module for reading the labeled text, preprocessing the text, and tagging the preprocessed text according to the labeled entities;
[0026] A character-level feature extraction module for converting the tagged text into a corresponding vector to obtain character-level features;
[0027] A feature fusion module, which is used to construct a corresponding BMES set for each character in the tagged text. B represents all matching words starting with the corresponding character, M represents all matching words with the corresponding character in the middle, E represents all matching words ending with the corresponding character, and S is all matching words consisting of the corresponding character alone. If a word set is empty, it is represented by the word "NONE". Encode the words in each set, convert them into corresponding text vectors, calculate the TF-IDF value of each word in each set as the weight of each word, and perform weighted summation on the words in each set to obtain the vocabulary feature. Concatenate and fuse the vocabulary feature with the character-level feature to achieve feature extraction.
[0028] An entity extraction module, which is used to extract the event trigger word entity and event element entity corresponding to the event in the text.
[0029] An event extraction module, which is used for event extraction.
[0030] A training optimization module, which is used to input the extracted features, event trigger word entities, event element entities, and events into a neural network and train the network by the stochastic gradient descent method. The trained network is used as an optimized model, and the optimized model is used for Chinese discourse-level event extraction.
[0031] Further, the event template is an asset freezing event, and the event elements included in the asset freezing event are the party to be frozen, frozen shares, freezing start time, and freezing end time.
[0032] Furthermore, the event trigger words include asset freezing.
[0033] Further, the preprocessing of the text includes: dividing the text into a test set, a training set, and a validation set, and the ratio of the test set, the training set, and the validation set is 7:2:1. Each text in the test set, the training set, and the validation set is defined to contain 64 sentences, and each sentence contains 128 words. Thus, each file is converted into a 64×128-dimensional vector. If it exceeds, it is truncated, and if it is not enough, it is padded with 0s.
[0034] Furthermore, the tagging of the preprocessed text according to the labeled entities includes: adding the label B at the beginning of the labeled entity in the preprocessed text, adding the label E at the end of the entity, and adding the label O in front of the content that does not belong to the labeled entity.
[0035] Further, the initial learning rate in the stochastic gradient descent method is 0.0001, and the number of training times is 100 times. After reaching the number of training times, the final model converges to the optimal.
[0036] The advantages of the present invention are as follows: The present invention extracts character-level features and constructs a corresponding BMES set for each character, calculates the TF-IDF value of each word in each set as the weight of each word, performs weighted summation on the words in each set to obtain the vocabulary feature, splices and fuses the vocabulary feature with the character-level feature to achieve feature extraction. The extracted features incorporate vocabulary information, avoiding the impact of word segmentation errors on event extraction and improving the extraction accuracy. Description of the Drawings
[0037] Figure 1 It is a network structure diagram of the event extraction algorithm in a Chinese discourse-level event extraction method incorporating vocabulary knowledge disclosed in an embodiment of the present invention;
[0038] Figure 2 It is a vocabulary structure diagram in a Chinese discourse-level event extraction method incorporating vocabulary knowledge disclosed in an embodiment of the present invention;
[0039] Figure 3 It is a schematic diagram of constructing a trie tree in a Chinese discourse-level event extraction method incorporating vocabulary knowledge disclosed in an embodiment of the present invention. Detailed Embodiments
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Embodiment 1
[0042] A Chinese discourse-level event extraction method incorporating vocabulary knowledge, the method comprising the following steps:
[0043] Step 1: Establish an event template; the event template is to determine how many event elements a certain event contains. Taking the asset freezing event as an example, it includes the following event elements: the party frozen, the frozen shares, the start time of freezing, and the end time of freezing.
[0044] Step 2: Collect texts and annotate the texts according to the event template to determine the event trigger word entity and the event element entity, and save the annotated texts in json format; the event trigger word entity is the text indicating the occurrence of a certain event, which plays an important role in event type classification. The event element entity is the text corresponding to the specific elements involved in the event, which plays an important role in event element classification. Taking the asset freezing event as an example, the event trigger words include asset freezing.
[0045] Step 3: Read the labeled text, preprocess the text, and tag the preprocessed text according to the labeled entities. Since the processed data is at the passage level, the number of sentences in each document and the number of characters in each sentence are generally not fixed. Without any processing, it is not conducive to parallel processing. To solve the above problems, the text is divided into a test set, a training set, and a validation set, and the ratio of the test set, the training set, and the validation set is 7:2:1. Each text in the test set, the training set, and the validation set is defined as containing 64 sentences, and each sentence contains 128 words. Thus, each file is converted into a 64×128-dimensional vector. If it exceeds, it is truncated, and if it is not enough, it is padded with 0. For the preprocessed text, add the label B at the beginning of the labeled entity and the label E at the end of the entity, and add the label O in front of the content that does not belong to the labeled entity. The above Step 2 belongs to manual annotation, and Step 3 is to tag the results of manual annotation, aiming to facilitate computer recognition.
[0046] Step 4: Convert the tagged text into a corresponding vector to obtain character-level features. Since the text cannot be directly understood by the computer, to solve the above problem, this embodiment uses the word2vec method to convert the text into a corresponding vector, thereby extracting the feature vector. It is mainly to map the corresponding characters to corresponding vectors using the trained word vectors. Since the dimension of the word vectors used in this method is 300, a document will obtain character-level features of 64×128×300 after being processed by word2vec.
[0047] Step 5: As Figure 2 shown, the relationship between texts is not a linear relationship but a graph structure. Words play an important role in this graph structure. An important task of integrating vocabulary knowledge is to learn this implicit information. The following will combine Figure 1 to illustrate the algorithm for integrating vocabulary information. A key step in integrating vocabulary information is how to construct the corresponding BMES set for each character in the tagged text. B represents all matching words starting with the corresponding character, M represents all matching words with the corresponding character in the middle, E represents all matching words ending with the corresponding character, and S is all matching words consisting of the corresponding character alone. If a word set is empty, it is represented by the word NONE. This embodiment obtains the vocabulary corresponding to each character mainly through the following steps: First, construct a trie tree for the pre-trained dictionary, and then use the trie tree to traverse and query to construct the corresponding BMES set for each character. A trie tree is an ordered tree used to store associated characters. The specific construction process can be seen in Figure 3, the trie tree regards each word as a character sequence and constructs a top-down tree structure according to the order of the character sequences in the string. Each edge in the tree structure corresponds to a character. It uses the common prefix of the string to reduce the query time overhead to achieve the purpose of improving efficiency; then it traverses each character of the input text with the constructed trie tree to construct the BMES set corresponding to each character.
[0048] After obtaining the BMES set corresponding to each character, the next step is to perform relevant encoding on the words in each set and convert them into corresponding text vectors. The specific process is as follows. First, each corresponding word is represented by a word-level vector containing text semantics through a word-level encoding model; then, the vectors in the four word sets are feature fused to be converted into feature vectors of a fixed dimension, and these features are concatenated as feature vectors that fuse semantic information and boundary information; in order to map these single-word sets into word vector representations of a fixed dimension, in this embodiment, weighted summation is used for the representation of each word in the word set to obtain the pooling representation of the word set. Considering computational efficiency, the dynamic weighting algorithm is not used to obtain the weight of each word, and the TF-IDF value of the word is used as its weight TF-IDF for weighting, and the words in each set are weighted and summed to obtain the vocabulary feature; the vocabulary feature is concatenated and fused with the character-level feature to achieve feature extraction.
[0049] Step Six: Extract the event trigger word entity and event element entity corresponding to the event in the text; identify the event trigger word entity and event element entity through the two modules of the Transformer feature extraction module and the CRF sequence annotation module.
[0050] Step Seven: Event extraction; there are mainly two tasks in event extraction: the first is event classification, which is a typical classification task and mainly classifies events according to the information provided by the event trigger word entity; the next is event element classification, which is also a typical classification task and mainly assigns the event element entity to the corresponding event and determines the type of the event. Entity extraction and event extraction belong to existing conventional technologies and will not be elaborated here.
[0051] Step Eight: Input the extracted features, event trigger word entity, event element entity, and event into the neural network to train the network by the stochastic gradient descent method. The trained network is used as an optimized model, and the optimized model is used for Chinese discourse-level event extraction. The initial learning rate in the stochastic gradient descent method is 0.0001, and the number of training times is 100 times. After reaching the number of training times, the final model converges to the optimal.
[0052] Through the above technical solutions, the present invention extracts character-level features, constructs a corresponding BMES set for each character, calculates the TF-IDF value of each word in each set as the weight of each word, performs weighted summation on the words in each set to obtain the vocabulary feature, splices and fuses the vocabulary feature with the character-level feature to achieve feature extraction. The extracted features incorporate vocabulary information, avoid the influence of word segmentation errors on event extraction, and improve the extraction accuracy.
[0053] Embodiment 2
[0054] Based on Embodiment 1, Embodiment 2 of the present invention further provides a Chinese discourse-level event extraction device incorporating vocabulary knowledge. The device includes:
[0055] A template establishment module for establishing event templates;
[0056] A labeling module for collecting texts and labeling the texts according to the event templates to determine event trigger word entities and event element entities, and saving the labeled texts in json format;
[0057] A tagging module for reading the labeled texts, preprocessing the texts and tagging the preprocessed texts according to the labeled entities;
[0058] A character-level feature extraction module for converting the tagged texts into corresponding vectors to obtain character-level features;
[0059] A feature fusion module for constructing a corresponding BMES set for each character in the tagged texts. B represents all matching words starting with the corresponding character, M represents all matching words with the corresponding character in the middle, E represents all matching words ending with the corresponding character, and S represents all matching words consisting of the corresponding character alone. If a word set is empty, it is represented by NONE; performing relevant encoding on each word in each set, converting it into a corresponding text vector, calculating the TF-IDF value of each word in each set as the weight of each word, performing weighted summation on each word in each set to obtain the vocabulary feature; splicing and fusing the vocabulary feature with the character-level feature to achieve feature extraction;
[0060] An entity extraction module for extracting event trigger word entities and event element entities corresponding to events in the texts;
[0061] An event extraction module for event extraction;
[0062] A training and optimization module for inputting the extracted features, event trigger word entities, event element entities and events into a neural network to train the network by the stochastic gradient descent method. The trained network is used as an optimized model, and the optimized model is used for Chinese discourse-level event extraction.
[0063] Specifically, the event template is an asset freeze event, and the event elements included in the asset freeze event are the party whose assets are frozen, the frozen shares, the start time of the freeze, and the end time of the freeze.
[0064] More specifically, the event trigger words include asset freeze.
[0065] Specifically, the preprocessing of the text includes: dividing the text into a test set, a training set, and a validation set, and the ratio of the test set, the training set, and the validation set is 7:2:1. Each text in the test set, the training set, and the validation set is defined to contain 64 sentences, and each sentence contains 128 words. Thus, each file is converted into a vector of 64×128 dimensions. If it exceeds, it is truncated, and if it is insufficient, it is padded with 0s.
[0066] More specifically, the tagging of the preprocessed text according to the labeled entities includes: adding the label B at the beginning of the labeled entities in the preprocessed text, adding the label E at the end of the entities, and adding the label O in front of the content that does not belong to the labeled entities.
[0067] Specifically, the initial learning rate in the stochastic gradient descent method is 0.0001, and the number of training times is 100 times. After reaching the number of training times, the final model converges to the optimal.
[0068] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Chinese discourse-level event extraction method incorporating vocabulary knowledge, characterized in that, The method includes the following steps: Step 1: Establish an event template, which is an asset freezing event. The event elements included in the asset freezing event are the party to be frozen, frozen shares, freezing start time, and freezing end time; Step 2: Collect texts and annotate the texts according to the event template to determine the event trigger word entity and event element entities, and save the annotated texts in json format; Step 3: Read the annotated texts, preprocess the texts, and tag the preprocessed texts according to the annotated entities; Step 4: Convert the tagged texts into corresponding vectors to obtain character-level features; Step 5: Construct a corresponding BMES set for each character in the tagged text. B represents all matching words starting with the corresponding character, M represents all matching words with the corresponding character in the middle, E represents all matching words ending with the corresponding character, and S is all matching words consisting of the corresponding character alone. If a word set is empty, it is represented by the word NONE; perform relevant encoding on the words in each set, convert them into corresponding text vectors, calculate the TF-IDF value of each word in each set as the weight of each word, and perform weighted summation on the words in each set to obtain the vocabulary feature; splice and fuse the vocabulary feature and the character-level feature to achieve feature extraction; Step 6: Extract the event trigger word entity and event element entities corresponding to the event in the text; Step 7: Event extraction; Step 8: Input the extracted features, event trigger word entity, event element entities, and event into a neural network to train the network by the stochastic gradient descent method. The trained network is used as an optimized model, and the optimized model is used for Chinese discourse-level event extraction.
2. The Chinese discourse-level event extraction method incorporating vocabulary knowledge according to claim 1, wherein The event trigger word includes asset freezing.
3. A Chinese discourse-level event extraction method incorporating vocabulary knowledge according to claim 1, characterized in that, The preprocessing of the text includes: dividing the text into a test set, a training set, and a validation set, and the ratio of the test set, the training set, and the validation set is 7:2:
1. Each text in the test set, the training set, and the validation set is defined to contain 64 sentences, and each sentence contains 128 words. Thus, each file is converted into a 64×128-dimensional vector. If it exceeds, it is truncated, and if it is not enough, it is filled with 0.
4. A Chinese discourse-level event extraction method incorporating vocabulary knowledge according to claim 3, characterized in that The tagging of the preprocessed text according to the annotated entities includes: adding the label B at the beginning of the annotated entity in the preprocessed text and adding the label E at the end of the entity, and adding the label O in front of the content that does not belong to the annotated entity.
5. A Chinese discourse-level event extraction method incorporating vocabulary knowledge according to claim 1, characterized in that, In the stochastic gradient descent method, the initial learning rate is 0.0001, and the number of training times is 100 times. After reaching the number of training times, the final model converges to the optimal.
6. A Chinese discourse-level event extraction device incorporating vocabulary knowledge, characterized in that, The device includes: A template establishment module for establishing an event template, which is an asset freezing event. The event elements included in the asset freezing event are the party to be frozen, frozen shares, freezing start time, and freezing end time; An annotation module for collecting texts and annotating the texts according to the event template to determine the event trigger word entity and event element entities, and saving the annotated texts in json format; The tagging module is used to read the annotated text, preprocess the text, and tag the preprocessed text according to the annotated entities; The character-level feature extraction module is used to convert the tagged text into corresponding vectors to obtain character-level features; The feature fusion module is used to construct a corresponding BMES set for each character in the tagged text. B represents all matching words starting with the corresponding character, M represents all matching words with the corresponding character in the middle, E represents all matching words ending with the corresponding character, and S is all matching words consisting of the corresponding character alone. If a word set is empty, it is represented by the word "NONE"; perform relevant encoding on the words in each set, convert them into corresponding text vectors, calculate the TF-IDF value of each word in each set as the weight of each word, and perform weighted summation on the words in each set to obtain the vocabulary feature; splice and fuse the vocabulary feature and the character-level feature to achieve feature extraction; The entity extraction module is used to extract the event trigger word entities and event element entities corresponding to the events in the text; The event extraction module is used for event extraction; The training optimization module is used to input the extracted features, event trigger word entities, event element entities, and events into the neural network to train the network by the stochastic gradient descent method. The trained network is used as the optimized model, and the optimized model is used for Chinese discourse-level event extraction.
7. An apparatus for Chinese discourse-level event extraction incorporating vocabulary knowledge according to claim 6, characterized in that, The event trigger words include asset freezing.
8. An apparatus for Chinese discourse-level event extraction incorporating vocabulary knowledge according to claim 6, characterized in that The preprocessing of the text includes: dividing the text into a test set, a training set, and a validation set, and the ratio of the test set, the training set, and the validation set is 7:2:
1. Each text in the test set, the training set, and the validation set is defined to contain 64 sentences, and each sentence contains 128 words. Thus, each file is converted into a 64×128-dimensional vector. If it exceeds, it is truncated, and if it is not enough, it is padded with 0s.
Citation Information
Patent Citations
Method and system for extracting Chinese document event
CN112231447A
Event trigger word extraction method based on document level attention mechanism
CN108829801A
Event element extraction method, computer equipment and storage medium
CN113434697A