A Predicate Extraction Method Based on Graph Isomorphism Network

Through the predicate extraction method based on graph isomorphic networks, using technologies such as DDParser, Bert model and GIN network, the shortcomings in applicability and efficiency of existing methods are solved, and higher extraction accuracy and cross-domain adaptability are achieved.

CN114330293BActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111638017.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-27
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The existing predicate extraction methods have limited applicability, difficulty in transplantation, time-consuming and difficult to maintain, especially when dealing with large-scale and cross-domain data.

Method used

The predicate extraction method based on graph isomorphic network is adopted, sentences are analyzed through the DDParser tool, word embedding is used using the Bert model, and node embedding vectors are obtained in combination with the GIN network and attention mechanism, and finally input into the binary classifier for predicate recognition.

Benefits of technology

This method learns the structural template characteristics of sentences through deep learning, reducing human workload and improving cross-domain adaptability and extraction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330293B_ABST
    Figure CN114330293B_ABST
Patent Text Reader

Abstract

The present invention discloses a predicate extraction method based on a graph isomorphism network. The present invention uses the DDParser tool to parse text sentences, and generalizes proper nouns in the word segmentation sequence by using the part-of-speech results obtained after sentence parsing. Adjust the embedding part of Bert, add the encoding of part-of-speech information, and input the generalized word sequence into the fine-tuned Bert model for encoding. Use the GIN network to obtain the embedding vector of each node in the dependency tree and the representation vector of the dependency subtree. After that, through a layer of attention mechanism, the semantic information and the dependency structure information are fused to obtain the final node embedding vector. Finally, the present invention inputs the set of the final word embedding vectors into a binary classifier to obtain the predicate result. The present invention uses deep learning to learn the structural template features of sentences, greatly reducing the workload of people, having strong cross-domain and adaptability, and effectively improving the accuracy of the predicate extraction method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information extraction, and more specifically, to a predicate extraction method based on graph isomorphism network. Background Art

[0002] Information extraction, that is, extracting specific facts or factual information from natural language texts to help us automatically classify, extract, and reconstruct massive content. These information usually include entities, relationships, events, etc. For example, extracting information such as time, place, and person from news information, and extracting information such as patient symptoms, medication conditions, and diseases from case data. Compared with other natural language tasks, the information extraction task is more purposeful and can present the extracted information in a specified structure, so as to achieve the purpose of extracting factual information of interest to users from natural language, and has a wide range of applications in the field of knowledge graphs.

[0003] Triple extraction is a classic information extraction task. Common triple extraction results can be represented by triples in the SPO triple structure, that is, Subject, Predication, and Object.

[0004] In triple extraction, how to extract predicates is a very important issue. Commonly used predicate extraction methods in the past include artificial template methods, statistical generation methods, and dependency-based methods. Among them, both artificial template and statistical generation methods regard triple extraction as a whole task and match the triples existing in the text by formulating templates. The basic starting point of the artificial template method is to statistically summarize a large amount of pattern information manually, and domain experts define the characters, grammatical features, etc. expressed by the predicate in the context, and use it as a pattern to match the text, and finally obtain the desired triple results. To reduce people's workload, the statistical generation method is proposed. This method mainly generates templates based on search engines. Specifically, this method uses the known triple facts as query statements, returns the first n result documents through the search engine and retains the set of sentences containing the triple, and finally uses the longest string containing the triple as the statistical template and retains the templates with higher confidence for triple extraction. These two methods have high accuracy, but their applicability is limited and it is difficult to transplant. The dependency-based method divides triple extraction into two steps. First, it extracts predicates through information such as the part of speech and dependency structure of the text, and then uses this predicate as a starting point to construct rules to extract the subject and object by using the connections and relationships between the various components in the sentence. This method has higher accuracy than the artificial template and statistical template methods and is applicable to small-scale data sets, but it also has problems such as time-consuming, laborious, and difficult to maintain. Summary of the Invention

[0005] After comprehensively considering the above problems, in view of the problems existing in the prior art, the present invention proposes a predicate extraction method based on a graph isomorphism network.

[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0007] Step (1): Use the DDParser tool to parse the input sentence to obtain the word segmentation result, part-of-speech, and syntactic dependency tree information;

[0008] Step (2): According to the part-of-speech and word segmentation result, generalize the proper nouns in the word segmentation. Fine-tune the word embedding part of Bert and add the encoding of part-of-speech information. Use the generalized word sequence and the part-of-speech data in step (1) as the input of the fine-tuned Bert model, and output a set of hidden vectors;

[0009] Step (3): Traverse the subtrees formed by any two nodes in the syntactic dependency tree in step (1), convert the information of each edge in this subtree into an edge vector, and then input the information of this subtree and the hidden vector in step (2) into the GIN network to obtain node embedding vectors, and perform pooling processing on the node embedding vectors to obtain the representation vector of the subtree;

[0010] Step (4): Calculate the attention weights using the subtree representation vector in step (3) and each hidden vector in step (2), and then multiply this weight by the embedding vector of each node in step (3) to obtain the final set of node embedding vectors;

[0011] Step (5): Input the set of node embedding vectors with semantic information obtained in step (4) into a binary classifier to obtain a binary sequence, and each binary in the sequence indicates whether the corresponding word is a predicate.

[0012] The beneficial effects of the present invention are as follows:

[0013] The present invention proposes a predicate extraction method based on a graph isomorphism network. First, the present invention uses the DDParser tool to parse text sentences, and generalizes proper nouns in the word segmentation sequence using the part-of-speech results obtained after sentence parsing to weaken the influence of some meaningless semantic information on the results. At the same time, the present invention adjusts the embedding part of Bert, adds encoding of part-of-speech information, and inputs the generalized word sequence into the fine-tuned Bert model for encoding. In addition, in order to emphasize the dependency structure information in the original sentence, the present invention uses a GIN network to obtain the embedding vector of each node in the dependency tree and the representation vector of the dependency subtree. After that, through an attention mechanism, the semantic information and the dependency structure information are fused to obtain the final node embedding vector. Finally, the present invention inputs the set of final word embedding vectors into a binary classifier to obtain the predicate result. Compared with existing technologies, the present invention uses deep learning to learn the structural template features of sentences, greatly reducing the workload of people, having strong cross-domain and adaptability, and effectively improving the accuracy of the predicate extraction method. Description of the Drawings

[0014] Figure 1 Flowchart of the overall implementation scheme of the present invention

[0015] Figure 2 Overall architecture diagram of the model of the present invention

[0016] Figure 3 Word embedding construction diagram of the present invention

[0017] Figure 4 Attention mechanism enhanced information diagram of the present invention Detailed Implementation Modes

[0018] The present invention will be further described below with reference to the drawings.

[0019] As Figure 1 and 2 shown, a predicate extraction method based on a graph isomorphism network includes the following steps:

[0020] A predicate extraction method based on a graph isomorphism network includes the following steps:

[0021] Step (1) uses the DDParser tool to parse the input sentence to obtain the word segmentation result, part of speech, and syntactic dependency tree information;

[0022] Step (2) generalizes the proper nouns in the word segmentation according to their parts of speech, obtaining a generalized word sequence corresponding to the input sentence after generalization; fine-tunes the word embedding part of the Bert model, adding the encoding of the part-of-speech information to the word embedding part; uses the generalized word sequence and the part-of-speech information in step (1) as the input of the fine-tuned Bert model, and outputs a set of hidden vectors;

[0023] Step (3) traverses any subtree in the syntactic dependency tree information in step (1), converts the information of each edge in this subtree into an edge vector, and then inputs the information of this subtree and the set of hidden vectors in step (2) into the GIN network to obtain node embedding vectors, and performs pooling processing on the node embedding vectors to obtain the representation vector of the subtree;

[0024] Step (4) calculates the attention weights using the representation vector of the subtree in step (3) and each hidden vector in step (2), and then multiplies this attention weight by each node embedding vector in step (3) to obtain the final set of node embedding vectors;

[0025] Step (5) inputs the final set of node embedding vectors with semantic information obtained in step (4) into a binary classifier to obtain a binary sequence, where each binary indication in the sequence corresponds to whether the corresponding word is a predicate.

[0026] Further, the specific implementation process of step (1) is as follows:

[0027] Use DDParser to parse the text sentence, and obtain the results:

[0028] X = (x 1 , x 2 ,..., x n ) (1)

[0029] T(X) = (t 1 , t 2 ,..., t n ) (2)

[0030] D(X) = Dependency_Parser(X) (3)

[0031] Among them, X represents the segmented sequence, x 1 , x 2 ,..., x n in formula (1) represents the word segmentation result, and t 1 , t 2 ,..., t n in formula (2) corresponds to x 1 , x 2 ,..., x nThe part-of-speech tagging result, where D(X) is the syntactic dependency tree.

[0032] Further, the specific implementation process of the step (2) is as follows:

[0033] 2-1 Perform generalization processing on the original sequence X according to the part-of-speech tagging result T(X). The specific rule content is as follows: For the part-of-speech tagging results of "LOC", "f", "s", " <time>”、" <loc> ”、" <per> ”、" <org>Replace the words "nw", "nz" with the "PN" tag to obtain the generalized word sequence X':

[0034] X' = (x' 1 , x' 2 ,..., x' n ) (4)

[0035] where x' 1 , x' 2 ,..., x' n represent the generalized vocabulary;

[0036] 2-2 As shown in Figure 3 , fine-tune the embedding structure of the Bert model, add a Postag Embedding layer to the original embedding structure to add part-of-speech information; perform word embedding processing on the generalized word sequence X', send the generalized word sequence X' into the Token Embedding layer to convert each word into a vector form, send the generalized word sequence X' into the Position Embedding layer to obtain the sequential features of each word, send the part-of-speech tagging result T(X) into the PostagEmbedding layer to obtain the part-of-speech features of each word, and finally concatenate these three results and input them into the Bert model to obtain the final word embedding and get the output hidden vector set;

[0037] The word embedding process can be expressed as the following formula:

[0038] H = BERT(X', T(X)) = {h 1 , h 2 ,..., h n} (5)

[0039] where H is the set of output hidden vectors I, and h 1 , h 2 ,..., h n are hidden vectors.

[0040] Furthermore, the specific implementation process of step (3) is as follows:

[0041] 3-1 Traverse any two nodes in the dependency tree, calculate the lowest common ancestor node of these two nodes, and obtain the subtree d(X) with the common ancestor node as the root and the two nodes as the leaves; convert all edge information in the subtree d(X) into edge vectors to obtain the result:

[0042] E = {e 1 , e 2 ,..., e q} (6)

[0043] where q represents the total number of edges in the current subtree;

[0044] 3-2 Input the hidden vector set H and the subtree d(X) into the GIN network to obtain node embedding information. The GIN network consists of m layers of graph isomorphism convolutional layer groups, and the calculation process of each layer is as follows:

[0045]

[0046] where, represents the hidden vector output by node i in the k-th layer of the graph isomorphism convolutional layer. In the first layer of the graph isomorphism convolutional layer, is the hidden vector output by Bert in step (2). ε is a learnable parameter. N(i) represents the set of all adjacent nodes of node i, E(i) represents the set of all adjacent edges of node i, and e p is the edge embedding corresponding to the edge. MLP is the multi-layer perceptron algorithm;

[0047] 3-3 Perform max pooling on the final node embedding vectors obtained in step 3-2 to obtain the representation vector of the subtree:

[0048]

[0049] where, h child-tree represents the representation vector of the subtree, represents the node embedding vector.

[0050] Furthermore, as Figure 4 shown, the specific implementation process of step (4) is as follows:

[0051] 4-1 Use the attention mechanism to enhance the effective information in the representation vector. Take the subtree representation vector h child-tree as Q, the set of hidden vectors {h 1 , h 2 ,..., h n} output by the Bert model as K, and the node embedding vector output by the GIN network as V. First, calculate the attention weight w i using Q and K. The detailed calculation process is as follows:

[0052]

[0053] Next, the model applies the attention weight p i to the corresponding V to obtain the final node embedding vector o i . The detailed calculation process is as follows:

[0054]

[0055] Further, the specific implementation process of step (5) is as follows:

[0056] The final hidden vector is input into a binary classifier to assign a binary label to each word, which indicates whether the current word is a predicate. The detailed calculation process is as follows:

[0057] p i = σ(Wo i + b) (11)

[0058] where both W and b are learnable parameters, and σ is the sigmoid function;

[0059] During the training process, the loss function is defined as:

[0060] Loss = CE(P, Y) (12)

[0061] where P represents the prediction result of the label, Y represents the true label, and CE represents the cross-entropy loss function.< / org> < / per> < / loc> < / time>

Claims

1. A predicate extraction method based on graph isomorphism network, characterized in that it includes the following steps: Step (1) Use the DDParser tool to parse the input sentence to obtain the word segmentation result, part-of-speech, and syntactic dependency tree information; Step (2) Generalize the proper nouns in the word segmentation according to the part-of-speech to obtain the generalized word sequence corresponding to the input sentence after generalization; Fine-tune the word embedding part of the Bert model and add the encoding of part-of-speech information to the word embedding part; Take the generalized word sequence and the part-of-speech information in step (1) as the input of the fine-tuned Bert model and output a set of hidden vectors; Step (3) Traverse any subtree in the syntactic dependency tree information in step (1), convert the information of each edge in this subtree into an edge vector, and then input the information of this subtree and the set of hidden vectors in step (2) into the GIN network to obtain node embedding vectors, and perform pooling on the node embedding vectors to obtain the representation vector of the subtree; Step (4) Calculate the attention weights using the representation vector of the subtree in step (3) and each hidden vector in step (2), and then multiply this attention weight by each node embedding vector in step (3) to obtain the final set of node embedding vectors; Step (5) Input the final set of node embedding vectors with semantic information obtained in step (4) into a binary classifier to obtain a binary sequence, and each binary indicator in the sequence indicates whether the corresponding word is a predicate; The specific implementation process of step (2) is as follows: 2-1 Perform generalization processing on the original sequence X according to the part-of-speech tagging result T(X), and the specific rule content is as follows: The part-of-speech tagging results are "LOC", "f", "s", " <time>”、" <loc> ”、" <per> ”、" <org>Replace the words of "、", "nw", "nz" with the "PN" label to obtain the generalized word sequence X′: < / org> < / per> < / loc> < / time> X′ = (x′ 1 , x′ 2 , …, x′ n )(4) Among them, x' 1 , x' 2 , …, x' n represent the generalized vocabulary; 2-2 Fine-tune the embedding structure of the Bert model, add a PostagEmbedding layer to the original embedding structure to add part-of-speech information; Perform word embedding processing on the generalized word sequence X′, send the generalized word sequence X′ into the Token Embedding layer to convert each word into a vector form, send the generalized word sequence X′ into the PositionEmbedding layer to obtain the sequential features of each word, send the part-of-speech tagging result T(X) into the Postag Embedding layer to obtain the part-of-speech features of each word, and finally splice these three results and input them into the Bert model to obtain the final word embedding, and obtain the output set of hidden vectors; The word embedding process can be expressed as the following formula: H = BERT(X′, T(X)) = {h 1 , h 2 , …, h n} (5) Among them, H is the set of output hidden vectors Ⅰ, h 1 , h 2 , …, h n are hidden vectors; The specific implementation process of step (3) is as follows: 3-1 Traverse any two nodes in the dependency tree, calculate the lowest common ancestor node of these two nodes, and obtain the subtree d(X) with the common ancestor node as the root and the two nodes as the leaves; Convert all edge information in the subtree d(X) into edge vectors to get the result: E = {e 1 , e 2 , …, e q}(6) where q represents the total number of edges in the current subtree; 3-2 Input the set of hidden vectors H and the subtree d(X) into the GIN network to obtain node embedding information, where the GIN network consists of m layers of graph isomorphism convolutional layer groups, and the calculation process of each layer is as follows: Among them, represents the hidden vector output by node i in the k-th layer of the graph isomorphism convolutional layer. In the first layer of the graph isomorphism convolutional layer, is the hidden vector output by Bert in step (2). ε is a learnable parameter. N(i) represents the set of all adjacent nodes of node i, and E(i) represents the set of all adjacent edges of node i. e p is the edge embedding corresponding to the edge. MLP is the multi-layer perceptron algorithm; 3-3 Perform max pooling on the final node embedding vectors obtained in step 3-2 to obtain the representation vector of the subtree: Among them, h child-tree represents the representation vector of the subtree, and represents the node embedding vector.

2. A predicate extraction method based on graph isomorphism network according to claim 1, characterized in that the specific implementation process of step (1) is as follows: Use DDParser to parse the text sentence to obtain the result: X = (x 1 , x 2 , …, x n )(1) T(X) = (t 1 , t 2 , …, t n )(2) D(X) = Dependency_Parser(X) (3) Among them, X represents the segmented sequence, and \(x\) in formula (1), 1 , \(x\) 2 , …, \(x\) n represent the word segmentation results, and \(t\) in formula (2), 1 , \(t\) 2 , …, \(t\) n correspond to the part-of-speech tagging results of \(x\) in formula (1), 1 , \(x\) 2 , …, \(x\) n and \(D(X)\) is the syntactic dependency tree.

3. A predicate extraction method based on graph isomorphism network according to claim 2, the specific implementation process of step (4) is as follows: 4-1 uses the attention mechanism to enhance the effective information in the representation vector, and the subtree representation vector h chile-tree is used as Q, and the set of hidden vectors {h 1 , h 2 , …, h n} output by the Bert model is used as K, and the node embedding vector output by the GIN network is used as V. First, the attention weight w i is calculated using Q and K. The detailed calculation process is as follows: Next, the model will apply the attention weight p i to the corresponding V to obtain the final node embedding vector o i , and the detailed calculation process is as follows:

4. A predicate extraction method based on graph isomorphism network according to claim 3, characterized in that the specific implementation process of step (5) is as follows: Input the final hidden vector into a binary classifier, and assign a binary label to each word, which indicates whether the current word is a predicate. The detailed calculation process is as follows: p i = σ(Wo i + b) (11) where W and b are both learnable parameters, and σ is the sigmoid function; During the training process, the loss function is defined as: Loss = CE(P, Y) (12) where P represents the prediction result of the label, Y represents the true label, and CE represents the cross-entropy loss function.

Citation Information

Patent Citations

  • Predicate extraction method

    CN111581365A

  • Graph convolutional network relationship extraction method based on multi-dependency relationship representation mechanism

    CN113239186A