A Remote Supervision Relationship Extraction Method Integrating Knowledge and Constraint Graphs

Through the method of fusion knowledge and constraint graph, combined with the multi-source fusion attention mechanism, the data noise and long-tail problems in remote supervision relationship extraction are solved, and the effect of relationship extraction and the ability to mine deep semantic features are improved.

CN115545005BActive Publication Date: 2025-06-27BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211185558.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-06-27
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

The existing relationship extraction method based on remote supervision has difficulties in dealing with data noise and long-tail relationships, resulting in poor model training results and the inability to effectively mine deep semantic features of text.

Method used

By integrating knowledge and constraint graphs, information is transmitted using entity knowledge context and entity relationship constraint graph, and a multi-source fusion attention mechanism is adopted to integrate sentence semantic information, entity context information and entity relationship constraint information to help express sentences and entity relationships.

Benefits of technology

It effectively reduces data noise in remote supervision relationship extraction, solves the long-tail problem of relationship extraction, improves the effect of relationship extraction, and can better extract structured fact information from unstructured text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115545005B_ABST
    Figure CN115545005B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote supervision relation extraction method integrating knowledge and constraint graphs, belonging to the technical field of text data relation extraction in computer natural language processing. In this method, additional information is supplemented by using entity knowledge context. Information is transmitted between relations through entity types and relation constraint graphs, and a multi-source fusion attention mechanism is used to fuse sentence semantic information, entity context information, and entity relation constraint information to assist in the representation learning of sentences and entity relations, improving the effect of relation extraction. This method also solves the data noise problem and relation long-tail problem in remote supervision relation extraction, and is particularly suitable for relation extraction under large-scale text data and complex text environments, and is very effective for extracting structured factual information from unstructured text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote supervision relation extraction method integrating knowledge and constraint graphs, belonging to the technical field of text data relation extraction in computer natural language processing. Background Art

[0002] Natural Language Processing (NLP) is an important direction in the fields of computer and artificial intelligence, mainly realizing various theories and methods for effective communication between humans and computers in natural language.

[0003] With the development of artificial intelligence technology, in computer natural language processing, relation extraction based on machine learning and deep learning has become a hot topic in this field. Currently, by using large-scale text corpora to train relation extraction to automatically mine the deep semantic features of language texts and achieve intelligent semantic analysis and relation extraction, it breaks the limitations of the traditional method of designing rules manually for relation extraction and has become the main means of relation extraction technology. Automatically performing relation extraction under large-scale text data and complex text environments through deep learning to extract structured factual information from unstructured texts has broad application prospects in tasks such as question answering systems, knowledge graphs, and search engines in the field of natural language processing.

[0004] In recent years, relation extraction based on remote supervision is a popular direction of relation extraction technology based on deep learning. This method aligns large knowledge bases and corpora to automatically obtain a large amount of labeled text data for training relation extraction models. Specifically, assuming that there is a "<head entity, tail entity, relation>" triple in the knowledge base, the remote supervision hypothesis is that sentences containing the head entity and the tail entity in the corpus all express this triple relation. Compared with other relation extraction methods, the relation extraction method based on remote supervision uses a large amount of automatically labeled data for training, which can solve the problem of difficult acquisition of training data for relation extraction, and the trained relation model has strong generalization ability.

[0005] However, the data set automatically labeled by remote supervision has two serious problems: one is the noise problem, that is, most of the sentences in the labeled data set are mislabeled, which affects the training effect of the relation extraction model. The other is the relation long-tail problem, that is, a small number of relations occupy the vast majority of the data, resulting in the relation extraction model being unable to fully learn the long-tail relations. Existing relation extraction methods based on remote supervision are difficult to solve the above two problems of the remote supervision data set, or can only focus on solving one of them, resulting in the inability to effectively train the model, making the actual effect of relation extraction poor and seriously affecting the mining of the deep semantic features of texts. Summary of the Invention

[0006] The purpose of the present invention is to address the deficiencies of the prior art. In order to effectively solve the data noise problem and long-tail problem faced in remote supervision relation extraction, a remote supervision relation extraction method integrating knowledge and constraint graph is creatively proposed.

[0007] The innovation of this method lies in: supplementing additional information through the use of entity knowledge context, transmitting information between relations through entity types and relation constraint graphs, and adopting a multi-source fusion attention mechanism to fuse sentence semantic information, entity context information, and entity relation constraint information to assist in the representation learning of sentences and entity relations, improving the effect of relation extraction. This method can effectively identify the relations between entity pairs for unstructured texts annotated with entity information, especially for relation extraction under large-scale text data and complex text environments, and is very effective in extracting structured factual information from unstructured texts.

[0008] The present invention is implemented by the following technical solutions.

[0009] A remote supervision relation extraction method integrating knowledge and constraint graph includes the following steps:

[0010] Step 1: Collect the neighbor entities of the entities in the remote supervision dataset in the knowledge base, including one-hop and two-hop neighbor entities. An entity set is composed of the entities in the remote supervision dataset and their neighbor entities. An entity neighbor graph is constructed using this entity set, and a constraint graph is constructed in combination with the set of relations between entities. For the remote supervision dataset, sentences with the same entity pair are combined into a sentence bag.

[0011] Among them, for the entities in the remote supervision dataset, the Flair named entity recognition tool can be used to identify their entity types to form an entity type set.

[0012] Step 2: Obtain the word embedding vectors of each word in the sentences in the bag, and the feature vector representation of the sentences.

[0013] Specifically, for the sentence bag obtained in Step 1, the sentences in the bag obtain the word embedding vectors of each word in the sentence through the word2vec tool, and the feature vector representation of the sentence is obtained through a piecewise convolutional neural network (PCNN) as the sentence encoder.

[0014] Step 3: Using the attribute encoder, for each entity in the entity set, collect its entity attribute information in the knowledge base, including entity name, entity alias, entity type, and entity description. Each entity splices this attribute information and inputs it into the attribute encoder, and then the output matrix is taken and column vector averaging is performed to obtain the attribute vector corresponding to the entity.

[0015] Specifically, BERT can be used as the attribute encoder.

[0016] Step 4: Construct an adjacency matrix using the entity neighbor graph, and use the adjacency matrix and the entity attribute vector as inputs to the neighbor graph encoder constructed by the graph convolutional neural network to obtain the knowledge context vector representation of the target entity.

[0017] Step 5: Construct an adjacency matrix using the constraint graph, and use the adjacency matrix and the vector representations of entity types and relationships as inputs to the constraint graph encoder constructed by the graph convolutional neural network to obtain the vector representations of entity types and relationships.

[0018] Step 6: Use the sentence feature vector representation, the entity context vector representation, and the vector representations of entity types and relationships as inputs, and through the multi-source fusion attention mechanism, calculate to obtain the feature vector representation of the sentence pack.

[0019] Step 7: For the feature vector representation of the sentence pack, through the relation classifier, predict the relation labels of the sentence pack.

[0020] Furthermore, in Step 1, the knowledge base contains entity pairs and the relationships corresponding to the entity pairs, as well as attribute information such as the entity name, entity alias, entity type, and entity description for each entity. The remote supervision dataset is the training corpus annotated by the remote supervision method, and the entity pairs and corresponding relationships in the knowledge base are used to annotate the natural language text. Specifically, assuming that there is a "<head entity, tail entity, relationship>" triple in the knowledge base, any sentence containing the head entity and the tail entity is considered to express the triple relationship, and the annotated data is obtained accordingly.

[0021] Furthermore, in Step 1, the entity type set can use 18 entity types defined in OntoNotes 5.0 and a special entity type "Other", for a total of 19 entity types, and the Flair named entity recognition tool is used to identify the types of entities in the dataset.

[0022] Furthermore, in Step 2, the piecewise convolutional neural network (PCNN) is a neural network model that takes the sequence of feature vectors of the words in a sentence as input and generates the sentence feature vector representation through convolution and piecewise pooling based on the positions of the two entities in the sentence.

[0023] Furthermore, in Step 3, BERT is a deep neural network built based on the multi-head self-attention mechanism and pre-trained using a large amount of corpus, and can output a feature vector representation with semantic information for the input text.

[0024] Further, in steps 4 and 5, the graph convolutional neural network is a neural network that realizes information propagation and aggregation through convolutional operations on graph data to extract feature information.

[0025] Further, in step 6, the multi-source fusion attention mechanism is a technical solution for information fusion based on the attention mechanism proposed for the first time in the present invention. It fuses sentence semantic information, entity knowledge context, and constraint information of entity relationships to obtain a feature vector representation of the sentence pack.

[0026] Further, in step 7, the specific implementation of the relation classifier is as follows: Input the feature vector representation of the sentence pack and the feature representation of each relation in the relation set, calculate the prediction score through dot product operation, and then calculate the probability that the sentence pack is classified as a certain relation through softmax operation.

[0027] Beneficial effects

[0028] The present invention has the following advantages compared with the prior art:

[0029] First, the method of the present invention incorporates the information of entities in the knowledge base into the relation extraction model as a supplement to the entity knowledge context, helping the relation extraction model to judge whether the remotely supervised annotation is correct, and effectively reducing the data noise in remotely supervised relation extraction.

[0030] Second, the method of the present invention uses the constraint graph of entity types and inter-entity relationships to enable different relationships to have an indirect connection through entity types, helping information to be transmitted between relationships, and at the same time solving the problem of relation long tail in remotely supervised relation extraction.

[0031] Third, the method of the present invention uses the multi-source fusion attention method to fuse the entity knowledge context, constraint graph information, and in-pack sentence information. The obtained feature vector representation of the pack can better reflect the relationship of the target entity pair, and at the same time solves the data noise problem and relation long tail problem in remotely supervised relation extraction. Brief description of the drawings

[0032] Figure 1 It is a schematic diagram of the overall framework of the method of the present invention. Detailed implementation manners

[0033] The method of the present invention will be further described in detail below with reference to the drawings.

[0034] A remotely supervised relation extraction method that fuses knowledge and constraint graphs, as Figure 1 shown, includes the following steps:

[0035] Step 1: Collect the one-hop or two-hop neighbor entities of the entities in the remotely supervised dataset in the knowledge base. The entities in the remotely supervised dataset and their neighbor entities form an entity set, and use this entity set to construct an entity neighbor graph. According to the entities in the remotely supervised dataset, use the Flair named entity recognition tool to identify their entity types to form an entity type set, and construct a constraint graph in combination with the set of entity relationships. For the remotely supervised dataset, combine the sentences with the same entity pair into sentence bags.

[0036] Specifically, define the entity neighbor graph as graph K = {E, N}, where E represents the set of entity nodes, that is, the entity set; N represents the set of edges. If two entities e1, e2 in the set E appear in a triple in the knowledge base at the same time, then there is an edge (e1, e2) ∈ N.

[0037] Define the constraint graph as graph G = {T, R, C}, where T is the set of entity type nodes. The 18 rough entity types defined in OntoNotes5.0 can be used. For entity types that do not belong to these 18 types, use a special node "Others" to represent them. Therefore, the set T includes 19 entity type nodes. Use the Flair named entity recognition tool to identify the types of entities in the dataset.

[0038] Let R be the set of relationship nodes composed of all relationships, and C be the set of constraint edges. If the entity types of entities e1, e2 are and entities e1, e2 have a relationship r, then there is a constraint Each constraint corresponds to and two edges.

[0039] Step 2: For the sentence bags obtained in Step 1, the sentences in the bag can use the word2vec tool to obtain the word embedding vectors of each word in the sentence, and a piecewise convolutional neural network (PCNN) can be used as the sentence encoder to obtain the feature vector representation of the sentence.

[0040] Specifically, for the sentences in the bag n s is the length of sentence s, and the input of each word w i ∈ s consists of its own word embedding vector and position feature vector. Among them, the word embedding vector v i can be pre-trained through the word2vec tool, and let the vector dimension be d w ; the position feature vector is the embedding vector representation of the two relative distances of word w i from the target entity pair (e h , e t ) in the sentence, and the vector dimension is dp Among them, the relative distance takes the position where the entity pair first appears in the sentence as the reference position to calculate the relative distances of other words.

[0041] The word w is obtained by concatenation i The input representation of w i , d = d w + 2d p , denotes a vector. The input representation of w i As shown in Equation 1:

[0042]

[0043] Among them, ";" denotes the vector concatenation operation. Then the input representation of the sentence s is a matrix

[0044] Use a piecewise convolutional neural network (PCNN) to encode the input representation matrix X of the sentence s to obtain a sentence feature vector with a fixed dimension size.

[0045] Among them, the piecewise convolutional neural network includes a convolutional layer and a piecewise max-pooling layer. Among them, the parameter matrix W of the convolutional layer is expressed as: w represents the length of the convolutional sliding window, and the sub-matrix q of the matrix X under the m-th sliding window m As shown in Equation 2:

[0046] q m = X m-w+1:m (1 ≤ m ≤ l s + w - 1) (2)

[0047] Among them, l s represents the length of the sentence s, and m - w + 1:m represents the index range of the word sequence under the sliding window in all word sequences of the original sentence.

[0048] Then the relationship between the sub-matrix q m and the parameter matrix c of the convolutional kernel m is shown in Equation 3:

[0049]

[0050] Among them, denotes the convolution operation.

[0051] Specifically, during the convolution process, sliding convolution is performed with a stride of 1, and the part of the convolution window that exceeds the sentence boundary is filled with zero vectors. Finally, the feature vector c representing the matrix X is obtained.

[0052] In practical use, in order to capture different features of a sentence, multiple convolutional kernels are usually used to perform convolution on the representation matrix of the sentence. In the present invention, d c convolutional kernels are adopted, and the set of convolutional kernels is denoted as After convolution calculation, the representation matrix X corresponds to d c feature vectors

[0053] For the piecewise convolutional neural network piecewise max-pooling layer, taking the positions of entity pairs in the sentence s as segmentation points, the feature vectors can be split into three parts, and then the max-pooling operation is applied to each part respectively. Specifically, for any one feature vector Three feature sub-vectors are generated after segmentation: {c i,1 ; c i,2 ; c i,3}. Max-pooling is performed on each feature sub-vector to obtain a pooled feature vector f i , as shown in Equation 4:

[0054] f i = [max(c i,1 ) ; max(c i,2 ) ; max(c i,3 )] (4)

[0055] where max(·) represents the operation of taking the maximum value.

[0056] The matrix X corresponds to d c feature vectors After passing through the piecewise max-pooling layer respectively, the obtained pooled feature vectors are concatenated, and after passing through the activation function tanh(·), the feature vector representation of X is obtained As shown in Equation 5:

[0057]

[0058] where, represents d c pooled feature vectors. The obtained is the feature vector representation of the sentence s.

[0059] Step 3: Use BERT as the attribute encoder. For each entity in the entity set in Step 1, collect four entity attribute information of its entity name, entity alias, entity type, and entity description in the knowledge base. Each entity inputs these attribute information through concatenation into BERT, then outputs a matrix, and column vector averaging is taken to obtain the attribute vector corresponding to the entity.

[0060] Specifically, for each entity e in the entity set i∈ E, the corresponding attribute vector obtained d a represents the dimension of the entity attribute vector. For entity e i 's attribute vector is calculated as follows:

[0061]

[0062] where, Mean(·) represents the operation of taking the column mean; BERT(·) represents the output matrix obtained by calculation using the BERT pre-trained model; name i represents the entity name of entity e i ; alias i represents the entity alias of entity e i ; type i represents the entity type of entity e i in the knowledge base; description i represents the entity description of entity e i . The attribute information is separated by the symbol [SEP]; [CLS] is the starting identification symbol that needs to be added when inputting to BERT.

[0063] Step 4: Use the entity neighbor graph in Step 1 to construct an adjacency matrix, and use the adjacency matrix and the entity attribute vector obtained in Step 3 as inputs to obtain the knowledge context vector representation of the target entity through the neighbor graph encoder constructed by the graph convolutional neural network.

[0064] Among them, the target entity can be any entity in the entity set in Step 1. For the entity neighbor graph K = {E, N}, its entity node set is E, and the set of edges is N. Its adjacency matrix is calculated as shown in Equation 7:

[0065]

[0066] where, |E| represents the number of entity nodes; v i , v j ∈ E represents any two entity nodes in the entity node set; (v i , v j ) ∈ N means that there is an edge between the entity nodes v i , v j in the entity neighbor graph.

[0067] Select a two-layer graph convolutional neural network as the neighbor graph encoder, and use the entity attribute vector obtained in Step 3 as the initial vector of entity e i ∈ E in the entity neighbor graph. For the output of node v i ∈ E in the k-th graph convolutional layer Its calculation method is as shown in Equation 8:

[0068]

[0069] Among them, Relu(·) is a non-linear activation function, |E| represents the number of entity nodes, and W (k) represents the weight matrix of the k-th layer of the neighbor graph encoder, and b (k) represents the bias term of the k-th layer. represents the feature vector output by the (k - 1)-th layer of the j-th entity node.

[0070] Select the output of the last layer of the neighbor graph encoder as the knowledge context representation vector of entity e i ∈E. d n represents the output vector dimension of the last layer of the neighbor graph encoder.

[0071] For the set E of entity nodes of the entity neighbor graph K, a knowledge context matrix is obtained through the neighbor graph encoder. Each row vector in the matrix K corresponds to the knowledge context vector representation of an entity, and this knowledge context vector representation integrates the information of the entity's neighbors and their attributes.

[0072] Step 5: Use the constraint graph constructed in Step 1 to construct an adjacency matrix, and use the adjacency matrix and the vector representations of entity types and relationships as inputs to the constraint graph encoder constructed by the graph convolutional neural network to obtain the vector representations of entity types and relationships.

[0073] Specifically, for the constraint graph G = {T, R, C}, the nodes in the graph include entity type nodes and relationship nodes, and the node set V G of the constraint graph = T ∪ R. Define the edge set of the constraint graph G as D G . For each constraint D G corresponds to and two edges. The calculation formula of the adjacency matrix of the constraint graph G is as shown in Equation 9:

[0074]

[0075] Among them, |V G | represents the number of nodes in the constraint graph; v i , v j ∈V G are any two nodes in the node set of the constraint graph; (v i , v j ) ∈ DG Indicates that there is an edge between nodes v i and v j in the constraint graph.

[0076] Select a two-layer graph convolutional neural network as the constraint graph encoder. Randomly initialize the initial input vectors of each node in the constraint graph node set V G . For node v i ∈V G , the output at the k-th graph convolutional layer is calculated as shown in Equation 10:

[0077]

[0078] where Relu(·) is a non-linear activation function, |V G | represents the number of nodes in the constraint graph, M (k) represents the weight matrix of the k-th layer of the constraint graph encoder, q (k) represents the bias term of the k-th layer, represents the feature vector of the output of the (k - 1)-th layer of the j-th node.

[0079] Select the output of the last layer of the constraint graph encoder as the feature vector representation of node v i ∈V G . For the node set V G of the constraint graph G, after passing through the constraint graph encoder, a feature matrix d g represents the output vector dimension of the last layer of the constraint graph encoder. Each row vector in the feature matrix V G corresponds to the vector representation of a node in the node set V G of the constraint graph G.

[0080] The node set V G includes entity type nodes and relationship nodes. By splitting the feature matrix V G , the feature matrix of the entity type node set T and the feature matrix of the relationship node set R |n t | represents the number of entity types in the entity type set T, and |n r | represents the number of relationships in the relationship node set R.

[0081] Step 6: Use the sentence feature vector representation obtained in Step 2, the entity context vector representation obtained in Step 4, and the vector representations of entity types and relationships obtained in Step 5 as inputs, and calculate the feature vector representation of the sentence pack through a multi-source fusion attention mechanism.

[0082] Specifically, for the sentence pack n b is the number of sentences in sentence pack B, and its corresponding target entity pair is (e h , e t ), and the corresponding relation label is r ∈ R.

[0083] For each sentence s i ∈ B in sentence pack B, its feature vector s i is obtained through the sentence encoder; for the target entity pair (e h , e t ) corresponding to the sentence pack, the knowledge context vector representations corresponding to entities e h and e t are obtained through the entity attribute encoder and the neighbor graph encoder. For the target entity pair (e h , e t ) and relation label r of the sentence pack, the corresponding entity type vector and relation vector r ∈ R′ are obtained through the constraint graph encoder.

[0084] The multi-source fusion attention mechanism provides more knowledge background information and constraint information for relation extraction by adding the knowledge context vector and entity type vector of the entity to the feature vectors of the sentence and the relation. Specifically, the final vector representation i of sentence s ∈ B is obtained by concatenating the sentence feature vector s i , the entity type vector of the target entity pair and the knowledge context vector of the target entity pair , as shown in Equation 11:

[0085]

[0086] where d s = 3d c + 2d g + 2d n , the dimension of vector s i is 3d c , the dimension of vector is d g , and the dimension of vector is d n ; ";" represents the vector concatenation operation.

[0087] The calculation method of the final vector representation r f of relation r ∈ R is:

[0088]

[0089] Among them, Linear(·) represents a linear fully connected layer, aiming to align the vector r f and the vector dimensionality of the vector

[0090] ";" represents the vector concatenation operation

[0091] After obtaining the final vector representation of the sentence s i ∈ B and the final vector representation r of the relation label r f After that, the feature vector representation of the sentence bag B is obtained through the attention mechanism The calculation method is shown in Equation 13

[0092]

[0093] Among them, n b is the number of sentences in the sentence bag B, and α i represents the weight of the final vector representation of the sentence s i ∈ B and its calculation method is shown in Equation 14

[0094]

[0095] Among them, e i represents the matching score between the sentence vector and the relation label r of the sentence bag, and its calculation method is shown in Equation 15

[0096]

[0097] Among them, "·" represents the vector dot product operation

[0098] Step 7: For the feature vector representation of the sentence bag obtained in Step 6, the relation classifier predicts the relation label of the sentence bag

[0099] After obtaining the feature vector B of the bag B through the multi-source fusion attention mechanism, the relation classifier predicts the relation label of the sentence bag B. The calculation formula of the prediction score for each relation r i in the sentence bag B and the relation set R is shown in Equation 17

[0100]

[0101] Among them, represents the final vector representation of the relation r i and represents the bias term

[0102] When obtaining each relation r iAfter the predicted score, the probability P that the sentence pack B is classified as the relationship r is calculated through the softmax function, and the calculation method is shown in Equation 17: i

[0103]

[0104] where |n r | represents the number of relationships in the relationship set R, represents the predicted score of the sentence pack B and the j-th relationship r in the relationship set R, and θ is the parameter of the relationship extraction model. j

[0105] Furthermore, in the relationship extraction model of the present invention, the cross-entropy loss function is used as the loss function for training, and its calculation formula is shown in Equation 18:

[0106]

[0107] where n B represents the number of sentence packs in the training set used, represents the relationship label of the sentence pack B i The parameters of the relationship extraction model used in the relationship extraction method of the present invention include all the parameters in the above steps: the parameters of the sentence encoder in step 2, the parameters of the attribute encoder in step 3, the parameters of the entity neighbor graph encoder in step 4, the parameters of the constraint graph encoder in step 5, the parameters of the multi-source fusion attention mechanism in step 6, and the parameters of the relationship classifier in step 7. The parameters include the weight matrix and the bias term.

[0108] In the present invention, the relationship extraction model used by the relationship extraction method can use the mini-batch stochastic gradient descent optimization algorithm SGD to minimize the loss function J(θ), so as to better optimize the parameters.​​

Claims

1. A remote supervision relation extraction method integrating knowledge and constraint graphs, characterized in that Including the following steps: Step 1: Collect the neighbor entities of the entities in the remotely supervised dataset in the knowledge base, including one-hop and two-hop neighbor entities; Form an entity set from the entities in the remotely supervised dataset and their neighbor entities, use this entity set to construct an entity neighbor graph, and construct a constraint graph in combination with the set of relationships between entities; For the remotely supervised dataset, combine the sentences with the same entity pair into a sentence bag; Step 2: Obtain the word embedding vectors of each word in the sentences in the bag, and the feature vector representation of the sentences; Step 3: Use the attribute encoder to collect the entity attribute information of each entity in the entity set in the knowledge base, including entity name, entity alias, entity type, and entity description; each entity inputs these attribute information through concatenation and then outputs a matrix and takes the column vector mean to obtain the attribute vector of the corresponding entity; Step 4: Use the entity neighbor graph to construct an adjacency matrix, and use the adjacency matrix and the entity attribute vector as inputs to the neighbor graph encoder constructed by the graph convolutional neural network to obtain the knowledge context vector representation of the target entity; Step 5: Use the constraint graph to construct an adjacency matrix, and use the adjacency matrix and the vector representations of entity types and relationships as inputs to the constraint graph encoder constructed by the graph convolutional neural network to obtain the vector representations of entity types and relationships; Step 6: Use the feature vector representation of the sentence, the context vector representation of the target entity, and the vector representations of entity types and relationships as inputs, and calculate the feature vector representation of the sentence bag through the multi-source fusion attention mechanism; Step 7: For the feature vector representation of the sentence bag, predict the relationship label of the sentence bag through a relationship classifier.

2. The remote supervision relation extraction method integrating knowledge and constraint graph according to claim 1, characterized in that, In Step 1, the knowledge base contains entity pairs and the relationships corresponding to the entity pairs, as well as the attribute information of each entity: entity name, entity alias, entity type, and entity description; The remotely supervised dataset is a training corpus annotated by the remote supervision method. Use the entity pairs and corresponding relationships in the knowledge base to annotate the natural language text. Assume that there is a "<head entity, tail entity, relationship>" triple in the knowledge base, and any sentence containing the head entity and the tail entity is considered to express the relationship of this triple, and thus the annotated data is obtained.

3. The remote supervision relation extraction method integrating knowledge and constraint graph according to claim 1, characterized in that In Step 2, the sentences in the bag use the word2vec tool to obtain the word embedding vectors of each word in the sentence, and use the segmented convolutional neural network as the sentence encoder to obtain the feature vector representation of the sentence. The segmented convolutional neural network is a neural network model that takes the sequence of feature vectors of the words in the sentence as input and generates the feature vector representation of the sentence through convolution and segmented pooling based on the positions of the two entities in the sentence.

4. A remote supervision relation extraction method integrating knowledge and constraint graphs according to claim 1, characterized in that In Step 1, define the entity neighbor graph as graph K = {E, N}, where E represents the set of entity nodes, that is, the entity set; N represents the set of edges; if two entities e1, e2 in the set E appear in a triple in the knowledge base at the same time, then there is an edge (e1, e2) ∈ N; Define the constraint graph as graph G = {T, R, C}, where T is the set of entity type nodes, and use the Flair named entity recognition tool to identify the types of entities in the dataset; Let R be the set of relationship nodes consisting of all relationships, and C be the set of constraint edges. If the entity types of entities e1 and e2 are and entities e1 and e2 have a relationship r, then there exists a constraint Each constraint corresponds to and two edges; In step 2, for the sentence pack obtained in step 1, for the sentences within the pack n s is the length of sentence s, and the input of each word w i ∈s consists of its own word embedding vector and position feature vector. Among them, the word embedding vector v i , let the vector dimension be d w ; the position feature vector is the embedding vector representation of the two relative distances between word w i and the target entity pair (e h , e t ) in the sentence, and the vector dimension is d p ; among them, the relative distance takes the position where the entity pair first appears in the sentence as the reference position to calculate the relative distances of other words; The word w is obtained by splicing i The input representation of w i , d = d w + 2d p , represents a vector; the input representation of w i As shown in Equation 1: where ";" represents the vector concatenation operation; then the input representation of sentence s is a matrix Encode the input representation matrix X of sentence s using a segmented convolutional neural network to obtain a sentence feature vector with a fixed dimensionality; wherein, the segmented convolutional neural network includes a convolutional layer and a segmented max pooling layer; wherein, the parameter matrix W of the convolutional layer is expressed as: w represents the length of the convolutional sliding window, and the sub-matrix q of matrix X under the m-th sliding window m As shown in Equation 2: q m = X m-w+1:m (1 ≤ m ≤ l s + w - 1) (2) where l s represents the length of sentence s, and m - w + 1:m represents the index range of the word sequence under the sliding window in all word sequences of the original sentence; Then the sub-matrix q m and the parameter matrix c of the convolution kernel m are related as shown in Equation 3: Among them, represents a convolution operation; During the convolution process, sliding convolution is performed with a stride of 1, and the part of the convolution window that exceeds the sentence boundary is filled with zero vectors. Finally, the feature vector c representing the matrix X is obtained. Using d c convolution kernels, and the set of convolution kernels is denoted as After convolution calculation, it represents that matrix X corresponds to d c eigenvectors The segmented max - pooling layer of the segmented convolutional neural network uses the positions of entity pairs in sentence s as segmentation points to split the feature vectors into three parts, and then applies the max - pooling operation to each part separately; for any feature vector After segmentation, three feature sub - vectors are generated: {c i,1 ; c i,2 ; c i,3}; Max - pooling is performed on each feature sub - vector to obtain a pooled feature vector f i , as shown in Equation 4: f i = [max(c i,1 ); max(c i,2 ); max(c i,3 )] (4) Among them, max(·) represents the maximum value operation; Matrix X corresponds to d c eigenvectors After passing through the segmented max-pooling layer respectively, the obtained pooled feature vectors are concatenated, and after passing through the activation function tanh(·), the feature vector representation of X is obtained As shown in Equation 5: Among them, represents d c pooling feature vectors; The resulting is the feature vector representation of sentence s; In step 3, BERT is used as the attribute encoder. For each entity in the entity set of step 1, four types of entity attribute information, namely the entity name, entity alias, entity type, and entity description in the knowledge base, are collected. Each entity inputs these attribute information after splicing into BERT, then outputs a matrix, and the column vector averaging is taken to obtain the attribute vector corresponding to the entity; For each entity e in the entity set i ∈ E, the corresponding attribute vector d a represents the dimension of the entity attribute vector; the attribute vector of entity e i is calculated as follows: Mean(BERT([CLS]+name i +[SEP]+alias i +[SEP]+type i +[SEP]+description i +[SEP])) (6) Among them, Mean(·) represents the operation of taking the column mean; BERT(·) represents the output matrix obtained by calculation using the BERT pre-trained model; name i represents the entity name of entity e i ; alias i represents the alias of entity e i ; type i represents the entity type of entity e i in the knowledge base; description i represents the entity description of entity e i ; The attribute information is separated by the symbol [SEP]; [CLS] is the starting identification symbol that needs to be added when inputting to BERT; In step 4, the adjacency matrix is constructed using the entity neighbor graph in step 1. Using the adjacency matrix and the entity attribute vectors obtained in step 3 as inputs, the knowledge context vector representation of the target entity is obtained through the neighbor graph encoder constructed by the graph convolutional neural network; Among them, the target entity is any entity in the entity set in Step 1; for the entity neighbor graph K = {E, N}, its entity node set is E, the edge set is N, and its adjacency matrix is calculated as shown in Equation 7: Among them, |E| represents the number of entity nodes; v i , v j ∈E represents any two entity nodes in the set of entity nodes; (v i , v j ) ∈ N means that there is an edge between the entity nodes v i , v j in the entity neighbor graph; Select a two-layer graph convolutional neural network as the neighbor graph encoder, and use the entity attribute vector obtained in step 3 as the initial vector of entity e i ∈ E in the entity neighbor graph. For the output of node v i ∈ E in the k-th graph convolutional layer its calculation method is shown in Equation 8: Among them, Relu(·) is a non-linear activation function, |E| represents the number of entity nodes, and W (k) represents the weight matrix of the k-th layer of the neighbor graph encoder, and b (k) represents the bias term of the k-th layer, represents the feature vector output by the (k-1)-th layer of the j-th entity node, Select the output of the last layer of the neighbor graph encoder as the entity e i The knowledge context representation vector of ∈ E d n Denote the output vector dimension of the last layer of the neighbor graph encoder; For the set of entity nodes E of the entity neighbor graph K, a knowledge context matrix is obtained after passing through the neighbor graph encoder. Each row vector in the matrix K corresponds to the knowledge context vector representation of an entity, and this knowledge context vector representation integrates the entity's neighbors and their attribute information. In step 5, the adjacency matrix is constructed using the constraint graph in step 1. Using the adjacency matrix and the vector representations of entity types and relationships as inputs, through the constraint graph encoder constructed by the graph convolutional neural network, the vector representations of entity types and relationships are obtained; For a constraint graph G = {T, R, C}, the nodes in the graph include entity type nodes and relationship nodes, and the node set V of the constraint graph G = T ∪ R; the edge set of the constraint graph G is defined as D G , for each constraint D G corresponding in and two edges; the adjacency matrix of the constraint graph G is calculated by the formula shown in Equation 9: Among them, |V G | represents the number of nodes in the constraint graph; v i , v j ∈ V G are any two nodes in the set of nodes of the constraint graph; (v i , v j ) ∈ D G means that there is an edge between nodes v i , v j in the constraint graph; Select a two-layer graph convolutional neural network as the constrained graph encoder, and randomly initialize the initial input vectors of each node in the constrained graph node set V G For each node, perform a random initialization operation on its initial input vector. For node v i ∈V G The output at the k-th graph convolutional layer Its calculation method is shown in Equation 10: Among them, Relu(·) is a non-linear activation function, |V G | represents the number of nodes in the constraint graph, M (k) represents the weight matrix of the k-th layer of the constraint graph encoder, q (k) represents the bias term of the k-th layer, represents the feature vector of the output of the (k-1)-th layer of the j-th node; Select the output of the last layer of the constrained graph encoder as the feature vector representation of node v i ∈ V G ; for the node set V of the constrained graph G G , a feature matrix is obtained after passing through the constrained graph encoder d g represents the output vector dimension of the last layer of the constrained graph encoder; each row vector in the feature matrix V G corresponds to the vector representation of a node in the node set V of the constrained graph G G ; Node set V G includes entity type nodes and relationship nodes; by splitting the feature matrix V G we obtain the feature matrix of the entity type node set T and the feature matrix of the relationship node set R |n t | represents the number of entity types in the entity type set T, and |n r | represents the number of relationships in the relationship node set R; In step 6, taking the sentence feature vector representation obtained in step 2, the entity context vector representation obtained in step 4, and the vector representations of entity types and relationships obtained in step 5 as inputs, through the multi-source fusion attention mechanism, the feature vector representation of the sentence packet is calculated; For the sentence pack n b is the number of sentences in sentence pack B, and the corresponding target entity pair is (e h , e t ), and the corresponding relation label is r ∈ R; For each sentence s in sentence bag B i ∈ B, obtain its feature vector s through the sentence encoder i ; for the target entity pair (e h , e t ) corresponding to the sentence bag, obtain the knowledge context vector representations corresponding to entities e h and e t through the entity attribute encoder and the neighbor graph encoder For the target entity pair (e h , e t ) and relationship label r of the sentence bag, obtain the corresponding entity type vector as well as relationship vector r ∈ R'; The multi-source fusion attention mechanism provides more knowledge background information and constraint information for relation extraction by adding the knowledge context vector and entity type vector of the entity to the feature vectors of the sentence and relation; sentence s i The final vector representation of By concatenating the sentence feature vector s i , the entity type vector of the target entity pair and the knowledge context vector of the target entity pair is obtained, as shown in Equation 11: Among them, d s = 3d c + 2d g + 2d n , the dimension of vector s i is 3d c , vector has a dimension of d g , vector has a dimension of d n ; ";" represents the vector concatenation operation; The final vector representation r of the relationship r ∈ R f is calculated as follows: Among them, Linear(·) represents a linear fully connected layer, aiming to align the vector r f and the vector vector dimensions; "; " represents the vector concatenation operation; After obtaining the final vector representation of sentence s i ∈ B and the final vector representation r of relation label r f After that, the feature vector representation of sentence bag B is obtained through the attention mechanism The calculation method is shown in Equation 13: where n b is the number of sentences in sentence pack B, and α i represents the final vector representation of sentence s i ∈ B, and its weight is calculated as shown in Equation 14: ​ where m i represents the matching score between the sentence vector and the relationship label r of the sentence pack, and the calculation method is shown in Equation 15: Among them, "·" represents the vector dot product operation; In step 7, for the feature vector representation of the sentence packet obtained in step 6, the relationship label of the sentence packet is predicted through the relationship classifier; After obtaining the feature vector B of bag B through the multi-source fusion attention mechanism, the relation classifier predicts the relation label of sentence bag B; for each relation r in the relation set R of sentence bag B i The calculation formula of the prediction score is shown in Equation 16: Among them, represents the final vector representation of relationship r i , and represents the bias term; After obtaining the prediction scores of the sentence pack B for each relationship r in the relationship set R i the probability P that the pack B is classified as the relationship r i is calculated by the softmax function, and the calculation method is shown in Equation 17: where |n r | represents the number of relationships in the relationship set R, represents the predicted score of the sentence pack B and the j-th relationship r in the relationship set R j , and θ is the parameter of the relationship extraction model.

5. The remote supervision relation extraction method integrating knowledge and constraint graph according to claim 4, characterized in that The relationship extraction model is trained using the cross-entropy loss function as the loss function, and its calculation formula is shown in Equation 18: where n B represents the number of sentence packs in the training set used, represents the relationship label of sentence pack B i θ are the parameters of the relationship extraction model, including all the parameters in the above steps: the parameters of the sentence encoder in step 2, the parameters of the attribute encoder in step 3, the parameters of the entity neighbor graph encoder in step 4, the parameters of the constraint graph encoder in step 5, the parameters of the multi-source fusion attention mechanism in step 6, and the parameters of the relationship classifier in step 7; the parameters include weight matrices and bias terms.