Knowledge and text enhancement-based event causal relationship extraction method and system
By adopting knowledge-based and text enhancement methods in event causal relationship extraction, a causal association link network is constructed and iterative sparse convolution is used to solve the problems of poor model performance and weak generalization ability in the existing technology, and a more accurate and comprehensive event causal relationship extraction is achieved.
Patent Information
- Application Number
- CN202510144965.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
In the current technology, there are problems of poor model performance and weak generalization ability in event causality extraction. Especially when dealing with Chinese causality, long-distance dependencies cannot be effectively captured, resulting in missing parameters and impact on model effects.
Using a knowledge-based and text enhancement method, a causal association link network is constructed, domain knowledge and text features are integrated, and multi-scale semantic features are generated using iterative sparse convolution, and a two-way long and short-term memory network is used to mark event causality.
The performance and generalization capabilities of the model are improved, multi-scale semantic dependency features are mined, context information is fully considered, and complete event causal relationships are extracted.
Smart Images

Figure CN120069075A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language technology, and more specifically, to a method and system for extracting event causal relationships based on knowledge and text enhancement. Background Art
[0002] In natural language text, events are important carriers of semantic information. Identifying events and the relationships between events can help people better understand the meaning of the text. Event relationship extraction is an important topic in the field of information extraction in natural language processing. It refers to judging the logical relationships between events, including co-reference relationships, subordinate relationships, temporal relationships, and causal relationships. The causal relationship reflects the relationship between objective phenomena of cause and effect. Event causal relationship extraction is an important research topic in the field of natural language processing, aiming to identify the causal relationships between text events from unstructured text using a computer. It is also an important basis for multiple downstream tasks, such as information extraction, intelligent question answering, and event prediction. The event causal relationship extraction task starts from determining whether there is a causal relationship in the semantics, then extracts the cause (C) and effect (E) from the sentences containing the causal relationship, and finally forms a <cause, causal relationship, effect> triple. It can be seen from the essence of the causal relationship extraction process that it is essentially a causal extraction of two arguments. Therefore, it can be converted into a sequence labeling task. However, due to the ambiguity of event descriptions, the finiteness of the dataset, and the long-distance dependence of event causal relationships, designing an effective method for extracting event causal relationships remains a long-term research topic worthy of attention.
[0003] The prior art proposed a CSNN model for event causal relationship extraction. This model uses a CNN to capture important features in the causal relationship window and uses a self-attention mechanism to mine the semantics and related features between different features. Secondly, a BiLSTM is used to establish long-term dependence relationships between causal relationships to obtain deeper context semantic information. The CNN is used to extract text features, and the self-attention mechanism is used to establish associations between semantic features. However, the result parameter extraction of this prior art is incomplete, which may be because the CSNN does not effectively solve the long-distance dependence in Chinese causal relationships, resulting in missing parameters, which in turn affects the performance of the model.
[0004] The prior art proposed a neural network-based causal relationship extractor SCITE. Through BERT+SCITE, the context string embedding is transferred to a large corpus for training. Multiple prior knowledge is embedded into the embedding representation to generate a mixed embedding representation, and a multi-head attention mechanism is used to learn the correlation between causal words. The multi-head self-attention mechanism is introduced into SCITE to enable the model to capture the long-term dependence relationship between cause and effect. However, this prior art only considers the prior knowledge of events and ignores the influence of a large amount of causal knowledge existing in the text on the model. Summary of the Invention
[0005] The present invention aims to provide a method and system for event causality extraction based on knowledge and text enhancement to solve the problems of poor model performance and weak generalization ability.
[0006] The technical solution adopted by the present invention to solve its technical problems is: A method for event causality extraction based on knowledge and text enhancement, including:
[0007] Obtain the context vector of the input sentence to get a word embedding sequence, select event pairs with causal relationships from the labeled data, and obtain an n-gram set through Top-k causal n-gram recognition;
[0008] Use domain text as a sample, use the words in the sample as semantic nodes, and the relationships between words as association links to construct a causal association link network; input the input sentence into the causal association link network to obtain text features integrating domain knowledge;
[0009] Group the n-gram set to form a high-order n-gram semantic representation, and input the high-order n-gram semantic representation into iterative sparse convolution to generate multi-scale semantic features;
[0010] Concatenate the text features and multi-scale semantic features to form a fused feature; input the fused feature into a bidirectional long short-term memory network to obtain the annotation result of event causality.
[0011] Preferably, the step of obtaining the context vector of the input sentence to get a word embedding sequence, selecting event pairs with causal relationships from the labeled data, and obtaining an n-gram set through Top-k causal n-gram recognition includes:
[0012] Obtain the word sequence of the input text to get the word vector set of each word in the word sequence;
[0013] Convert the word vector set into a context vector set containing position information in the input sentence;
[0014] Select event pairs with causal relationships from the labeled data, and obtain an n-gram set through Top-k causal n-gram recognition.
[0015] Preferably, the step of using domain text as a sample, using the words in the sample as semantic nodes, and the relationships between words as association links to construct a causal association link network includes:
[0016] Apply a domain causal knowledge corpus to obtain a text data set;
[0017] Using the Pyltp library, for word segmentation, mining part-of-speech tagging, and identifying named entities, removing stop words and entities, selecting verbs and nouns, and using verbs, nouns, and phrases as nodes of the CALN to obtain the edge weights between nodes, specifically:
[0018]
[0019] Among them, Co(w i , w j ) represents the co-occurrence frequency of words wi and w j in the text dataset, that is, the number of times they appear in the same text; DF(w i ) is the frequency of the word w i appearing in the text dataset;
[0020] The input vector matrix of knowledge attention is obtained by concatenating the word embeddings and associated word embeddings of each k-gram. Use linear projection to project the input vector matrix, and calculate the attention scores through the attention function, specifically:
[0021]
[0022] The parameter matrix of the q-th linear projection is specifically:
[0023]
[0024] Among them, h is the number of heads;
[0025] The attention function is executed in parallel to obtain the final value of the k-gram, specifically:
[0026] a j = MultiHead(X i , X i , X i ) = Concat(head 1 , …, head h )
[0027] The semantic dependency feature representation of the j-th scale of the input sentence is:
[0028]
[0029] The feature representation of the input sentence learned from domain prior knowledge is:
[0030] A = Concat[A 1 , A 2 , … A l
[0031] Thus, a causal association link network is constructed.
[0032] Preferably, the input vector matrix is obtained through the input embedding of knowledge attention, and the input embedding of knowledge attention is specifically:
[0033]
[0034] where i is the i-th sliding window, k is the same as the size of the sliding window of the convolution, r represents the number of related words from CALN, and X i represents the input vector matrix.
[0035] Preferably, the related words are obtained by searching CALN, including:
[0036] Finding all adjacent nodes in CALN through the local text of the k-grams size of each word;
[0037] Selecting several adjacent nodes most relevant to the current word and determining the relevance between nodes through edge weights;
[0038] Intersecting the other words in the text except the current k-grams with the adjacent nodes to obtain the related words.
[0039] Preferably, the grouping of the n-gram set to form a high-order n-gram semantic representation includes:
[0040] Grouping the n-gram set according to the n-gram semantic features, and using the center of each cluster as the abstract feature of the n-gram to form a high-order n-gram semantic representation.
[0041] Preferably, the input of the high-order n-gram semantic representation into the iterative sparse convolution to generate multi-scale semantic features includes:
[0042] Taking the high-order n-gram semantic representation as the prior knowledge for convolution initialization, using the Naive Bayes method to select n-grams from the labeled data, obtaining the result n-grams through the following formula, and performing BERT encoding on the result n-grams to generate embedding representations;
[0043]
[0044] where c and e are the cause and the result respectively. Where n c i is the number of sentences containing the n-gram i in the cause c; ‖n c ‖ 1 is the number of sentences containing the n-gram in the cause c, and b is the smoothing parameter;
[0045] Select the event length with the largest proportion in the labeled data as the parameter of n-gram, and select the first 25% of n-gram as the semantic features for convolutional initialization;
[0046] In convolutional initialization, expand the input width of the convolution by skipping the dilation width once; among them, the enlarged convolution is expressed as:
[0047]
[0048] where ∪ is vector concatenation; W c is the filter width of r tokens; δ is the dilation width;
[0049] Generate several feature maps using convolution, and the feature generated by each i-th window is expressed as:
[0050] D i =[c 1 ,c 2 ,…,c m
[0051] where D i is the feature representation generated by the window vector at position i by m filters in the convolution, and the stride is 1;
[0052] The semantic feature of the j-th scale is expressed as:
[0053] S j ={D 1 ,D 2 ,…,D n}
[0054] The multi-scale semantic features learned from the data are expressed as:
[0055] S = Concat[S 1 ,S 2 ,…,S l .
[0056] Preferably, the text features and the multi-scale semantic features are concatenated to form a fusion feature; the fusion feature is input into a bidirectional long short-term memory network, and the hidden state sequence is output and labeled to obtain the labeled result of the event causality, including:
[0057] Model the input sentence from different perspectives with the text features and the multi-scale semantic features to form a fusion feature;
[0058] Input the fusion feature into the bidirectional long short-term memory to learn the semantic features of the sentence from two directions, specifically:
[0059]
[0060] Connect the forward and backward outputs of the bidirectional long short-term memory to obtain the bidirectional long short-term memory output depth context features, specifically:
[0061]
[0062] The bidirectional long short-term memory output generates the final representation. Use the conditional random field to decode the label sequence of the bidirectional long short-term memory output. Given the input sentence and its predicted label sequence y = (y 1 , y 2 , …, y n ), the conditional random field score is specifically:
[0063]
[0064] where represents the probability that the i-th word in the sentence is the y i label. The likelihood of converting the label y i-1 to the label y i is expressed as
[0065] The final convergence condition is to minimize the loss function, and the loss function is represented by the following formula:
[0066]
[0067] Obtain the annotation result of the event causal relationship according to the loss function.
[0068] Preferably, an event causal relationship extraction system based on knowledge and text enhancement applies the described event causal relationship extraction method based on knowledge and text enhancement, including:
[0069] A set acquisition module for obtaining the context vector of the input sentence to obtain the word embedding sequence, selecting event pairs with causal relationships from the labeled data, and obtaining the n-gram set through Top-k causal n-gram recognition;
[0070] A knowledge enhancement module for using the domain text as a sample, using the words in the sample as semantic nodes, and the relationships between the words as association links to construct a causal association link network; inputting the input sentence into the causal association link network to obtain the text features integrating domain knowledge;
[0071] A text enhancement module for grouping the n-gram set to form a high-order n-gram semantic representation, and inputting the high-order n-gram semantic representation into the iterative sparse convolution to generate multi-scale semantic features;
[0072] A relation extraction module, which is used to splice text features and multi-scale semantic features to form fused features; and input the fused features into a bidirectional long short-term memory network to obtain the annotation result of event causal relations.
[0073] The beneficial effects of the present invention are as follows:
[0074] Compared with the prior art, the present application provides a method and system for event causal relation extraction based on knowledge and text enhancement. By constructing a text-knowledge enhancement module, a CALN is established to capture event causal relations, and a parallel iterative extended convolutional neural network is used to encode global text information; the performance and generalization ability of the model are improved; multi-scale semantic dependency features are mined, and context information is fully considered to extract complete event mentions. Description of the Drawings
[0075] Figure 1 is a schematic flowchart of a method for event causal relation extraction based on knowledge and text enhancement according to the present invention;
[0076] Figure 2 is a schematic diagram of the module of a system for event causal relation extraction based on knowledge and text enhancement according to the present invention;
[0077] Figure 3 is a logical algorithm flowchart of an embodiment of the present invention.
[0078] The drawings are only for illustrative purposes and should not be construed as a limitation to the present invention; for better illustration of the embodiments, some components in the drawings may be omitted, enlarged or reduced, and do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted. Detailed Embodiments
[0079] The following will further describe the present invention in detail with reference to the drawings and specific embodiments.
[0080] As Figure 1 shown, a method for event causal relation extraction based on knowledge and text enhancement according to the present invention includes:
[0081] S1. Obtain the context vector of the input sentence to get the word embedding sequence, select event pairs with causal relations from the labeled data, and obtain the n-gram set through Top-k causal n-gram recognition;
[0082] S2. Use the domain text as a sample, use the words in the sample as semantic nodes, and the relationships between the words as association links to construct a causal association link network; input the input sentence into the causal association link network to obtain text features integrating domain knowledge;
[0083] S3. Group the n-gram set to form a high-order n-gram semantic representation, and input the high-order n-gram semantic representation into iterative sparse convolution to generate multi-scale semantic features;
[0084] S4. Concatenate the text features and the multi-scale semantic features to form a fused feature; input the fused feature into a bidirectional long short-term memory network to obtain the annotation result of the event causal relationship.
[0085] In the above solution, event causal relationship extraction is a challenging task in information extraction, which can automatically extract event descriptions and identify the causal relationships between events, and plays an important role in event prediction, scenario generation, question answering, and text entailment. Previous methods tend to use one-sided models to extract events and their related causal relationships. One-sided models usually focus on text content while ignoring the internal element transformation within events and the causal relationship transformation association between events. Moreover, most existing methods focus on the extraction of single-scale event causal relationships and fail to extract multi-scale (such as words, phrases, sentences) event causal relationships. Event causal relationship extraction should condense the complex relationships within events and the causal transition associations between events and consider multi-scale event causal relationship semantic information.
[0086] Therefore, this application constructs text enhancement and knowledge enhancement, and learns causal features of different scales from training data through parallel convolutional layers, fully considering global event mentions and causal transfer associations as well as multi-scale event causal relationship semantic information, thereby improving the accuracy of event causal relationship extraction.
[0087] Preferably, in step S1, to obtain the context vector of the input sentence, get the word embedding sequence, select event pairs with causal relationships from the labeled data, and obtain the n-gram set through Top-k causal n-gram recognition, including:
[0088] Obtain the word sequence of the input text to get the set of word vectors for each word in the word sequence;
[0089] Convert the set of word vectors into a set of context vectors containing position information in the input sentence;
[0090] Select event pairs with causal relationships from the labeled data and obtain the n-gram set through Top-k causal n-gram recognition.
[0091] In the above solution, W = {w 1 , w 2 , …, w n} is the word sequence of the input text, and X = {x 1 , x 2 , …, x n} is a set of word vectors for each word in the word sequence. Using a pre-trained model, each word in the word sequence is mapped to a position space to obtain the word vector representation of each word in the sentence; using the pre-trained language model Bert, the set of word vectors X is transformed into a set of sentence context vectors H = {h 1 , h 2 , …, h n}, where h i is the context vector of each word at position i, which contains its position information; select event pairs with causal relationships from the labeled data, and through Top-k causal n-gram recognition, obtain a set containing the most important causal n-grams.
[0092] Preferably, in step S2, constructing the causal association link network with the domain text as a sample, using the words in the sample as semantic nodes and the relationships between the words as association links includes:
[0093] Applying the domain causal knowledge corpus to obtain a text data set; specifically:
[0094] Applying a large-scale domain causal knowledge corpus to obtain a large number of explicit causal text data sets S c = {s 1 , s 2 , …, s n}, where s i is a sentence.
[0095] Using the Pyltp library for word segmentation, mining part-of-speech tagging and identifying named entities, deleting stop words and entities, selecting verbs and nouns, and using verbs, nouns and phrases as nodes of the CALN. The same word of different parts of speech is regarded as different nodes; obtaining the edge weights between the nodes, specifically:
[0096]
[0097] where Co(w i , w j ) represents the co-occurrence frequency of words w i and w j in the text data set, that is, the number of times they appear in the same text; DF(w i ) is the frequency of word w i in the text data set;
[0098] Concatenate the word embeddings and related word embeddings of each word in the k-gram to obtain the input vector matrix of knowledge attention. Use linear projection to project the input vector matrix and calculate the attention score through the attention function, specifically:
[0099]
[0100] The parameter matrix of the q-th linear projection is specifically:
[0101]
[0102] where h is the number of heads;
[0103] The attention functions are executed in parallel to obtain the final value of the k-gram, specifically:
[0104] a j = MultiHead(X i , X i , X i ) = Concat(head 1 , …, head h )
[0105] Represent the semantic dependency feature of the j-th scale of the input sentence as:
[0106]
[0107] The feature representation of the input sentence learned from domain prior knowledge is:
[0108] A = Concat[A 1 , A 2 , … A l
[0109] Thus, a causal association link network is constructed.
[0110] In the above solution, a causal association link network is constructed. Using the words in the domain text as semantic nodes and the relationships between words as association links, a weighted network is established. It shows whether there is semantic relevance between words or phrases and provides a weight for the degree of this relevance.
[0111] Preferably, the input vector matrix is obtained through the input embedding of knowledge attention, and the input embedding of knowledge attention is specifically:
[0112]
[0113] where i is the i-th sliding window, k is the same as the sliding window size of the convolution, r represents the number of associated words from the CALN, and X i represents the input vector matrix.
[0114] Preferably, the associated words are obtained by searching the CALN, including:
[0115] Finding all adjacent nodes in the CALN through the local text of the k-grams size of each word;
[0116] Select several adjacent nodes that are most relevant to the current word, and determine the relevance between nodes through edge weights;
[0117] Intersect other words in the text except for the current k-grams with adjacent nodes to obtain associated words.
[0118] Preferably, in step S3, the grouping of the n-gram set to form a high-order n-gram semantic representation includes:
[0119] Group the n-gram set according to the n-gram semantic features, and use the center of each cluster as the abstract feature of the n-gram to form a high-order n-gram semantic representation.
[0120] Preferably, in step S3, the input of the high-order n-gram semantic representation into the iterative sparse convolution to generate multi-scale semantic features includes:
[0121] Use the high-order n-gram semantic representation as the prior knowledge for convolution initialization, select n-grams from the labeled data using the Naive Bayes method, obtain the result n-grams through the following formula, and perform BERT encoding on the result n-grams to generate embedding representations;
[0122]
[0123] where c and e are the cause and the result respectively. Among them is the number of sentences containing n-gram i in cause c; ||n c || 1 is the number of sentences containing n-grams in cause c, and b is the smoothing parameter;
[0124] In the above solution, extract important causal n-gram semantic features from the labeled data; use the clustering method to group semantically similar n-gram semantic features together to generate a high-level n-gram semantic representation; then input the high-level n-gram representation into the initialization process of the iterative dilated convolution. Use high-level n-gram features for some filters, and randomly initialize the remaining positions to let the model learn more useful features by itself. Use the n-gram semantic features as the prior knowledge for convolution initialization, select effective n-grams from the labeled data using the Naive Bayes method, obtain the top k cause or result n-grams, and perform BERT encoding on each n-gram to generate embedding representations.
[0125] Select the event length with the largest proportion in the labeled data as the parameter of the n-gram, and select the first 25% of the n-gram as the semantic features for convolution initialization;
[0126] In convolution initialization, the input width of the convolution is extended by skipping the dilation width once; where the enlarged convolution is expressed as:
[0127]
[0128] where ∪ is vector concatenation; W c is the filter width of the r tokens; δ is the dilation width;
[0129] Using the convolution to generate a number of feature maps, the feature generated by each i-th window is expressed as:
[0130] D i = [c 1 , c 2 , …, c m
[0131] where D i is the feature representation generated by the m filters in the convolution for the window vector at position i, with a stride of 1;
[0132] The semantic feature representation of the j-th scale is expressed as:
[0133] S j = {D 1 , D 2 , …, D n}
[0134] The multi-scale semantic features learned from the data are expressed as:
[0135] S = Concat[S 1 , S 2 , …, S l .
[0136] Preferably, in step S4, the splicing of the text features and the multi-scale semantic features to form a fusion feature; inputting the fusion feature into a bidirectional long short-term memory network, outputting a hidden state sequence and annotating to obtain the annotation result of the event causal relationship includes:
[0137] Modeling the input sentence from different perspectives with the text features and the multi-scale semantic features to form a fusion feature;
[0138] Inputting the fusion feature into the bidirectional long short-term memory to learn the semantic features of the sentence from two directions, specifically:
[0139]
[0140] Connecting the forward and backward outputs of the bidirectional long short-term memory to obtain the bidirectional long short-term memory output depth context feature, specifically:
[0141]
[0142] The bidirectional long short-term memory outputs the final representation, and the conditional random field is used to decode the label sequence output by the bidirectional long short-term memory. Given the input sentence and its predicted label sequence y = (y 1 , y 2 , …, y n ), the conditional random field score is specifically:
[0143]
[0144] where represents the probability that the i-th word in the sentence is the y i label, and the likelihood of converting the label y i-1 to the label y i is expressed as
[0145] The final convergence condition is to minimize the loss function, and the loss function is represented by the following formula:
[0146]
[0147] The annotation result of the event causal relationship is obtained according to the loss function.
[0148] In the above solution, in order to make the most of the domain prior knowledge and the multi-scale semantic features of the event causal relationship in the data, the output of the text enhancement channel and the output of the knowledge enhancement channel are used as the input of the BiLSTM-CRF, and two independent features are used to model the sentence from different perspectives.
[0149] Preferably, as Figure 2 shown, an event causal relationship extraction system based on knowledge and text enhancement applies the above-mentioned event causal relationship extraction method based on knowledge and text enhancement, including:
[0150] A set acquisition module, used to obtain the context vector of the input sentence, obtain the word embedding sequence, select the event pairs with causal relationships from the labeled data, and obtain the n-gram set through Top-k causal n-gram recognition;
[0151] A knowledge enhancement module, used to use the domain text as a sample, use the words in the sample as semantic nodes, and the relationships between the words as association links to construct a causal association link network; input the input sentence into the causal association link network to obtain the text features integrating domain knowledge;
[0152] A text enhancement module, which is used to group the n-gram set to form a high-order n-gram semantic representation, and input the high-order n-gram semantic representation into an iterative sparse convolution to generate multi-scale semantic features;
[0153] A relation extraction module, which is used to splice the text features and the multi-scale semantic features to form a fused feature; and input the fused feature into a bidirectional long short-term memory network to obtain the annotation result of the event causal relationship.
[0154] In the above solution, as Figure 3 shown, the present application proposes a text-knowledge dual-mode enhanced neural network. By constructing a text-knowledge enhancement module, it establishes CALN to mine multi-scale semantic dependency features and event causal relationships, and uses a parallel iterative extended convolutional neural network to encode the global text information, considering multi-scale information to further mine the causal knowledge of the text; it can capture the element transformation within the event and the causal relationship between events, capture multi-scale semantic information from the data using parallel convolutions of different scales, and improve the performance and generalization ability of the model; by constructing CALN, the relevant information between words or phrases is discovered. Then the relevant information is introduced into the model to mine multi-scale semantic dependency features and event causal relationships, and it can fully consider the context information to extract complete event mentions.
[0155] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for extracting event causal relationships based on knowledge and text enhancement, characterized in that: include: Get the context vector of the input sentence to get the word embedding sequence, select event pairs with causal relationships from the labeled data, and obtain the n-gram set through Top-k causal n-gram recognition; Take the domain text as a sample, use the words in the sample as semantic nodes, and the relationship between words as associative links to build a causal associative link network; input the input sentence into the causal associative link network to obtain text features that integrate domain knowledge; The n-gram sets are grouped to form high-order n-gram semantic representations, and the high-order n-gram semantic representations are input into iterative sparse convolution to generate multi-scale semantic features; Combine text features and multi-scale semantic features to form fusion features; The fused features are input into the bidirectional long short-term memory network to obtain the labeling results of the event causal relationship.
2. According to the method for extracting event causal relationships based on knowledge and text enhancement according to claim 1, it is characterized in that: The context vector of the input sentence is obtained to obtain a word embedding sequence, and event pairs with causal relationships are selected from the labeled data. Through Top-k causal n-gram recognition, the n-gram set obtained includes: Get the word sequence of the input text and obtain the word vector set of each word in the word sequence; Convert the word vector set into a context vector set containing position information in the input sentence; Event pairs with causal relationships are selected from the labeled data, and n-gram sets are obtained through Top-k causal n-gram recognition.
3. The event causal relationship extraction method based on knowledge and text enhancement according to claim 2 is characterized in that: The method of using the domain text as a sample, using the words in the sample as semantic nodes, and the relationship between the words as an association link to construct a causal association link network includes: Apply domain causal knowledge corpus to obtain text datasets; The Pyltp library is used for word segmentation, part-of-speech tagging and named entity identification, stop words and entities are deleted, verbs and nouns are selected, and verbs, nouns and phrases are used as nodes of CALN to obtain the edge weights between nodes, specifically: Where Co(w i ,w j ) represents word w i and w j The co-occurrence frequency in a text dataset is the number of times it appears in the same text; DF(w i ) is the word w i Frequency of occurrence in a text dataset; The input vector matrix of knowledge attention is obtained by concatenating each word embedding and the related word embedding of the k-gram. The input vector matrix is projected using linear projection, and the attention score is calculated through the attention function, which is: The parameter matrix of the qth linear projection is specifically: Among them, h is the number of heads; The attention functions are executed in parallel to obtain the final value of k-gram, which is: a j =MultiHead(X i ,X i ,X i )=Concat(head1,…,head h ) The j-th scale semantic dependency feature of the input sentence is expressed as: The input sentence features learned from domain prior knowledge are expressed as: A=Concat[A 1 ,A 2 ,…A l ] Thus, a causal associative link network is constructed.
4. The event causal relationship extraction method based on knowledge and text enhancement according to claim 3 is characterized in that: The input vector matrix is obtained by input embedding of knowledge attention, and the input embedding of knowledge attention is specifically: Where i is the i-th sliding window, k is the same as the sliding window size of the convolution, r represents the number of associated words from CALN, and X i Represents the input vector matrix.
5. The event causal relationship extraction method based on knowledge and text enhancement according to claim 3 is characterized in that: The associated words are obtained by searching CALN, including: In CALN, all neighboring nodes are found through the local context of k-grams size of each word; Select several adjacent nodes that are most relevant to the current word, and determine the correlation between the nodes through edge weights; Intersect other words in the text except the current k-grams with adjacent nodes to get related words.
6. The event causal relationship extraction method based on knowledge and text enhancement according to claim 3 is characterized in that: The grouping of n-gram sets to form high-order n-gram semantic representations includes: The n-gram sets are grouped according to the n-gram semantic features, and the center of each cluster is used as the abstract feature of the n-gram to form a high-order n-gram semantic representation.
7. The method for extracting event causal relationships based on knowledge and text enhancement according to claim 6 is characterized in that: The step of inputting the high-order n-gram semantic representation into iterative sparse convolution to generate multi-scale semantic features includes: The high-order n-gram semantic representation is used as the prior knowledge for convolution initialization, and the naive Bayes method is used to select n-grams from the labeled data. The resulting n-grams are obtained by the following formula, and the resulting n-grams are BERT encoded to generate embedded representations; Where c and e are the cause and result respectively. is the number of sentences containing n-gram i in reason c; ||n c ||1 is the number of sentences containing n-grams in reason c, and b is the smoothing parameter; The event length that accounts for the largest proportion in the labeled data is selected as the parameter of n-gram, and the first 25% of n-gram is selected as the semantic features for convolution initialization; In the convolution initialization, the input width of the convolution is expanded by skipping the dilation width once; the enlarged convolution is expressed as: Where ∪ is the vector connection; W c is the filtering width of r tokens; δ is the dilation width; Several feature maps are generated using convolution, and the features generated by each i-th window are expressed as: D i =[c1,c2,…,c m ] Where D i It is the feature representation generated by the window vector at position i by m filters in the convolution, with a stride of 1; The semantic feature of the jth scale is expressed as: S j ={D1,D2,…,D n } The multi-scale semantic features learned from the data are expressed as: S=Concat[S 1 ,S 2 ,…,S l ]。 8. The event causal relationship extraction method based on knowledge and text enhancement according to claim 7 is characterized in that: The text features and multi-scale semantic features are spliced to form fusion features; The fused features are input into the bidirectional long short-term memory network, and the hidden state sequence is output and labeled. The labeling results of the event causal relationship include: Model the input sentence from different perspectives using text features and multi-scale semantic features to form fusion features; The fused features are input into the bidirectional long short-term memory to learn the semantic features of the sentence from two directions, specifically: The forward and backward outputs of the bidirectional long short-term memory are connected to obtain the bidirectional long short-term memory output deep context feature, specifically: The bidirectional long short-term memory output generates the final representation, and the conditional random field is used to decode the label sequence of the bidirectional long short-term memory output. Given the input sentence and its predicted label sequence y = (y1, y2, ..., y n ), the conditional random field score is specifically: in Indicates that the i-th word in the sentence is y i The probability of the label, label y i-1 Convert to label y i The possibility is expressed as The final convergence condition is to minimize the loss function, which is expressed by the following formula: The labeling results of event causal relationships are obtained based on the loss function.
9. A system for extracting event causal relationships based on knowledge and text enhancement, applying the method for extracting event causal relationships based on knowledge and text enhancement as claimed in claim 1, characterized in that: include: The set acquisition module is used to obtain the context vector of the input sentence, obtain the word embedding sequence, select event pairs with causal relationships from the labeled data, and obtain the n-gram set through Top-k causal n-gram recognition; The knowledge enhancement module is used to construct a causal associative link network using domain text as samples, words in the samples as semantic nodes, and the relationships between words as associative links; Input the input sentence into the causal association link network to obtain text features that integrate domain knowledge; The text enhancement module is used to group n-gram sets to form high-order n-gram semantic representations, and input the high-order n-gram semantic representations into iterative sparse convolution to generate multi-scale semantic features; The relation extraction module is used to combine text features and multi-scale semantic features to form fusion features; The fused features are input into the bidirectional long short-term memory network to obtain the labeling results of the event causal relationship.
Citation Information
Patent Citations
Event causal relationship extraction method and device based on background knowledge and storage medium
CN116341519A
Cited By
Text semantic visualization presentation method and system based on multi-modal large model
CN121188194A