Method and apparatus for entity and relationship knowledge extraction with statement-oriented feature dimension enhancement
By conducting vectorization and joint detection of entity relationships on input statements, the problems of overlapping triplets and error propagation are solved, and effective extraction of entity and relational features and diversity and reliability of triplets are realized.
Patent Information
- Application Number
- CN202211150030.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-21
AI Technical Summary
The prior art has problems of overlapping triplets and error propagation in joint extraction of entity relationships, and the interaction between entities and relationships is not effectively considered.
The statement-oriented feature dimension enhancement method is adopted, and the input statement is vectorized, entities and relationships are jointly detected and characterized, and information strengthened processing is carried out to finally form a triple.
It effectively avoids overlapping triplets and error propagation, ensures the diversity and reliability of triplet information, and improves the extraction effect of entity and relationship characteristics.
Smart Images

Figure CN115510239B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a method and device for entity and relationship knowledge extraction with enhanced feature dimensions for statements. Background Art
[0002] Extracting entities and relationships from unstructured text provides the main knowledge source for subsequent automatic construction of a knowledge graph and is a necessary step in knowledge graph construction. The extracted knowledge generally exists in the form of triples such as (subject, relationship, object) or (s, r, o). Among them, the subject and object in the triple are two entities connected by a certain relationship in the knowledge graph.
[0003] Traditional triple extraction methods use a pipeline approach, which divides the extraction task into two steps. First, named entity prediction (NER) is performed on the input statement, and then relationship classification (RC) is performed on the predicted entity pairs. However, due to the strict order requirements of this method, the obvious problem is that it will lead to error propagation. To solve this problem, researchers have proposed methods for joint entity-relationship extraction. Recent research results show that the joint extraction method can better integrate entity and relationship information, and the overall extraction effect is indeed better than that of the pipeline approach. Recently, deep learning-based joint extraction methods have become very popular due to their outstanding effects. However, there are still the following challenging problems in relationship entity triple extraction:
[0004] 1) Overlapping triples, which include two types of overlaps: EntityPairOverlap (EPO) and SingleEntityOverlap (SEO), as Figure 1 shown. To solve this problem, many people have adopted a method based on the decomposition of the subject and object for extraction, but such a method is prone to error propagation problems.
[0005] 2) Error propagation, which results from the strict prediction order process. For example, in the pipeline approach, to solve this problem, some researchers have proposed a method based on a decoder to extract triples, but such methods still do not consider the interaction between entities and relationships, and between sentences and relationships.
[0006] 3) Ignoring the interaction between entities and relationships, and between relationships and statements. For the triple information extracted from a sentence, we believe that there is a certain association between the sentence and the relationship, and between the entity and the relationship. Extracting them separately will not be able to learn the deep relationship between them, and such triple information will be very difficult to extract. Summary of the Invention
[0007] The main objective of the present invention is to overcome the drawbacks and deficiencies of the prior art, and provide a method and apparatus for entity and relationship knowledge extraction with enhanced feature dimensions oriented to sentences. By means of the method of joint entity-relationship extraction, the problems of overlapping triples and error propagation are solved.
[0008] To achieve the above objective, the present invention adopts the following technical solutions:
[0009] On the one hand, the present invention provides a method for entity and relationship knowledge extraction with enhanced feature dimensions oriented to sentences, including the following steps:
[0010] Vectorize the input sentence to obtain a vectorized sentence with context semantic features;
[0011] Perform entity detection and characterization as well as relationship detection and characterization on the vectorized sentence to respectively obtain entity feature information and relationship feature information; the entity feature information refers to the subject information and object information extracted from the vectorized sentence; the relationship feature information refers to the association features existing between the subject and the object extracted from the vectorized sentence;
[0012] Perform joint prediction of entities and relationships on the vectorized sentence, and use the entity feature information and relationship feature information as auxiliary dimension feature information for information enhancement processing to obtain the feature information of joint prediction of entities and relationships;
[0013] Perform splicing or link prediction on the feature information of joint prediction of entities and relationships, and finally form triples.
[0014] Preferably, the vectorization of the input sentence is specifically as follows:
[0015] Extract the hidden features of each word in the input sentence through the encoder in the Bert model, and convert the input sentence into a vectorized sentence with context semantic features. The expression of the vectorized sentence H is as follows:
[0016] H = Bert[{x1, x2,..., x n ,..., x m} * mask]
[0017] H = [h1, h2,.., h n ,..., h m
[0018] Wherein, x1, x2,..., x n ,..., x m are the IDs in the Bert model's corresponding dictionary to which each word in the input statement is mapped. n represents the length of the input statement sequence, m is the total length of the statement after vectorization and padding, mask is the actual valid statement information in the input statement, and h1, h2, .., h n , ..., h m are word vectors incorporating context information.
[0019] Preferably, the entity refers to the subject and the object;
[0020] The entity detection and characterization are specifically as follows:
[0021] Input the vectorized statement H into a fully connected layer, calculate the start position probability and end position probability of the entity. If the probability of the start position is greater than a preset first threshold, then determine this start position as the start position of the entity in the vectorized statement; similarly, if the probability of the end position is greater than a preset second threshold, then determine this end position as the end position of the entity in the vectorized statement; meanwhile, the neural network of the fully connected layer will be trained according to the label information of the training set, and continuously adjust the trainable weight values W and b;
[0022] The calculation formulas for the start position probability and end position probability of the entity are as follows:
[0023] p i start_sub(obj) = sigmoid(W start h i + b start )
[0024] p i end_sub(obj) = sigmoid(W end h i + b end )
[0025] Among them, p i start_sub(obj) is the probability that the i-th position in the input statement is marked as the start position of the entity, and p i end _sub(obj) start start end end is the probability that the i-th position in the input statement is marked as the end position of the entity; hi is the output result of the encoder layer, W end and b end are the trainable weight values for calculating the start position probability of the entity, and W
[0026] After determining the probabilities of the start position and end position of the entity, the main body information T is extracted. i sub and the object information T i obj , the formula is:
[0027] T i sub =(p i start_sub , p i end_sub )
[0028] T i obj =(p i start_obj , p i end_obj )
[0029] Among them, p i start_sub is the probability that the i-th position is marked as the start position of the main body, and p i end_sub is the probability that the i-th position is marked as the end position of the main body; p i start_obj is the probability that the i-th position is marked as the start position of the object, and p i end_obj is the probability that the i-th position is marked as the end position of the object.
[0030] Preferably, the relationship detection and characterization are specifically:
[0031] Embed all the preset relationship tags into a high-dimensional vector, and then through a linear mapping layer, represent the final result as the most relevant initial relationship node embedding. The calculation formula of the initial relationship node embedding is:
[0032] R m =W r *E([r1, r2,..., r m )+b r
[0033]
[0034] Among them, r i is the one-hot vector of the relationship index in the predefined relationship, m is the number of predefined relationships, E is the relationship embedding matrix, W r and b r are the trainable parameters of the predefined process of the relationship node, and R m is the initial relationship node, is a high-dimensional relationship vector;
[0035] Predict the initial relational node information contained in the vectorized input statement. First, add the obtained initial relational node information to the initial statement, and then add the initial statement with the initial relational node information to a fully connected layer for neural network calculation. Finally, obtain the relational information feature through the sigmoid function. At the same time, the high-dimensional feature vector changes the weights of W r and b r during continuous training, and then determine the feature of the relational information. The calculation formula of the relational information feature is as follows:
[0036]
[0037] where is the high-dimensional relational vector obtained in the previous step, h i is the output result of the encoder layer, W r and b r are the trainable weights in the relational detection process, and sigmoid is the activation function.
[0038] Preferably, perform joint prediction of entities and relations on the vectorized statement, and use the entity feature information and relational feature information as entity auxiliary dimension features for information enhancement processing, specifically:
[0039] Add the entity head information feature and entity tail feature to the statement feature respectively, and then multiply by the relational feature information. Use two fully connected layer networks, one network for predicting the subject-relation and the other network for predicting the object-relation. After self-adjustment and training of the network, obtain the feature information of joint prediction of entities and relations.
[0040] Preferably, splice or perform link prediction on the predicted entity and relational feature information to finally form a triple, specifically:
[0041] Perform category judgment on the feature information of joint prediction of entities and relations. The judgment method is to construct two one-dimensional matrices with the same length as the number of relational databases. By traversing the results of the joint prediction of both parties, map the IDs corresponding to the relational values predicted by both parties to the array subscript positions, so as to register the number of relations. Finally, obtain two categories: unique relation matching and multi-relation matching. The unique relation matching means that the prediction numbers of the subject-relation and object-relation under the same relation are both not greater than 1. The multi-relation matching means that the prediction numbers of the subject-relation and object-relation under the same relation are both greater than 1.
[0042] For the unique relation matching category, adopt the direct splicing principle, and perform matching splicing on the data under the same relation to obtain the triple.
[0043] For multi-relation matching classes, the start position information of the subject and the two parts of the start position information of the relation and the object are concatenated by relation to form Tr = [Sub start , Obj start , rel], where Sub start is the start position information of the subject, Obj start is the start position information of the object, and rel represents the relation information corresponding to the start position information of the subject and the start position information of the object; then, probability prediction is performed again on the concatenated Tr, and the calculation formula is as follows:
[0044] P i Tr = sigmoid(W i Tr i + b i )
[0045] where Tr is formed by concatenating the main start relation matrix and the target start relation matrix through the relation, W i and b i are the trainable weight values for re-prediction, and sigmoid is the activation function; through the obtained T i sub = (p i start_sub , p i end_sub ) and T i obj = (p i start_obj , p i end_obj ) information, the head information is completely expanded into triple information; where T i sub is the subject information, T i obj is the object information; p i start_sub is the probability that the i-th position is marked as the start position of the subject, p i end_sub is the probability that the i-th position is marked as the end position of the subject; p i start_obj is the probability that the i-th position is marked as the start position of the object, p i end_obj is the probability that the i-th position is marked as the end position of the object.
[0046] Another aspect of the present invention provides a method and system for entity and relationship knowledge extraction with enhanced feature dimensions for statements. Applied to the method for entity and relationship knowledge extraction with enhanced feature dimensions for statements, it includes a vectorization module, a detection and characterization module, a joint prediction module, and a triple output module;
[0047] The vectorization module is used to vectorize the input statement to obtain a vectorized statement with context semantic features;
[0048] The detection and characterization module is used to perform entity detection and characterization and relationship detection and characterization on the vectorized statement, and respectively obtain entity feature information and relationship feature information;
[0049] The joint prediction module is used to perform joint prediction of entities and relationships on the vectorized statement, and perform information enhancement processing on the entity feature information and relationship feature information as auxiliary dimension feature information to obtain entity and relationship joint prediction feature information;
[0050] The triple output module is used to splice or perform link prediction on the feature information of the joint prediction of entities and relationships, and finally form a triple.
[0051] Another aspect of the present invention provides an electronic device, characterized in that the electronic device includes:
[0052] At least one processor; and,
[0053] A memory communicatively connected to the at least one processor; wherein,
[0054] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the entity and relationship knowledge extraction method based on feature dimension enhancement of statement orientation.
[0055] Another aspect of the present invention provides a computer-readable storage medium storing a program, and when the program is executed by a processor, it implements the entity and relationship knowledge extraction method based on feature dimension enhancement of statement orientation provided in another aspect of the present invention.
[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0057] 1. The present invention deeply analyzes the features of text statements, first detects and analyzes the two most important feature dimensions in the statements: entity and relationship features, and characterizesthe detection results again to facilitate subsequent statement feature enhancement, and also strengthens the internal connection between statements and entities and between statements and relationships.
[0058] 2. In the process of model construction, to avoid possible overlapping triples and propagation errors, the present invention adopts a method of joint extraction of entities and relationships. At the same time, when predicting the subject, object, and relationship respectively, the entity and relationship feature information analyzed from the sentence is added for feature enhancement. Thus, the feature information of the subject-relationship and object-relationship is directly obtained, and the obtained information is divided into two categories: 1) unique relationship matching: the prediction numbers of the subject / object-relationship under the same relationship are both not greater than 1; 2) multiple relationship matching: the prediction numbers of the subject / object-relationship under the same relationship are both greater than 1. And the second type of information obtained, which is less in number, is separately re-predicted for the subject and object under the same relationship. Finally, both types of information form triples, ensuring the diversity and reliability of the triple information. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0060] Figure 1 It is a schematic diagram of the types of overlapping triples of the present invention;
[0061] Figure 2 It is a framework flowchart of the entity and relationship knowledge extraction method with enhanced feature dimensions for sentences according to an embodiment of the present invention;
[0062] Figure 3 It is a structural diagram of the joint extraction model according to an embodiment of the present invention;
[0063] Figure 4 It is a structural diagram of the entity and relationship knowledge extraction system with enhanced feature dimensions for sentences according to an embodiment of the present invention;
[0064] Figure 5 It is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0066] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.
[0067] The Bert model involved in this application is an autoencoding language model that trains the language model in the way of MaskLM. Generally speaking, when inputting a sentence, some words to be predicted are randomly selected and then replaced with a special symbol (MASK), and then the model is allowed to learn the words to be filled in these places according to the given labels; finally, the vector representation of each word in the text after integrating the semantic information of the full text is output.
[0068] In addition, the sigmoid function mentioned in the text is a non-linear activation function of neurons and is widely used in neural networks. The learning of neural networks is based on a set of samples, including inputs and outputs. The number of input and output neurons corresponds to the number of components of the inputs and outputs; initially, the weights (Weight) and thresholds (Threshold) of the neural network are arbitrarily given, and the training is to gradually adjust the weights and thresholds to make the actual output of the network consistent with the expected output.
[0069] Please refer to Figure 2 , in an embodiment of this application, a method for extracting entity and relationship knowledge with enhanced feature dimensions for statements is provided, including the following steps:
[0070] S1. Vectorize the input statement to obtain a vectorized statement with context semantic features.
[0071] Further, the vectorization of the input statement is specifically as follows:
[0072] Extract the hidden features of each word in the input statement through the encoder in the Bert model, and convert the input statement into a vectorized statement with context semantic features. The expression of the vectorized statement is as follows:
[0073] H = Bert[{x1, x2,..., x n ,..., x m} * mask]
[0074] H = [h1, h2,.., h n ,..., h m
[0075] Where x1, x2,..., xn , ..., x m are the IDs of each word in the input sentence mapped to the corresponding dictionary of the Bert model. n represents the length of the input sentence sequence, m is the total length of the sentence after vectorization and padding, mask is the actual valid sentence information in the input sentence, h1, h2, .., h n , ..., h m are word vectors incorporating context information.
[0076] S2. Perform entity detection and characterization as well as relationship detection and characterization on the vectorized sentence to obtain entity feature information and relationship feature information respectively.
[0077] Furthermore, the entity refers to the subject and the object;
[0078] S21. The entity detection and characterization are specifically as follows:
[0079] Input the vectorized sentence H into a fully connected layer, calculate the start position probability and end position probability of the entity. If the probability of the start position is greater than a preset first threshold, then determine this start position as the start position of the entity in the vectorized sentence; similarly, if the probability of the end position is greater than a preset second threshold, then determine this end position as the end position of the entity in the vectorized sentence; meanwhile, the neural network of the fully connected layer will be trained according to the label information of the training set, and continuously adjust the trainable weight values W and b;
[0080] The calculation formulas for the start position probability and end position probability of the entity are as follows:
[0081] p i start_sub(obj) = sigmoid(W start h i + b start )
[0082] p i end_sub(obj) = sigmoid(W end h i + b end )
[0083] where p i start_sub(obj) is the probability that the i-th position in the input sentence is marked as the start position of the entity, p i end _sub(obj) i is the probability that the i-th position in the input sentence is marked as the end position of the entity; h i start is the output result of the encoder layer, W startThe trainable weight values W for calculating the probability of the start position of the entity end and b end are the trainable weight values for calculating the probability of the end position of the entity, and sigmoid is the activation function;
[0084] After determining the probability of the start position of the entity and the probability of the end position of the entity, the subject information T is extracted i sub and the object information T i obj , and the formula is:
[0085] T i sub =(p i start_sub , p i end_sub )
[0086] T i obj =(p i start_obj , p i end_obj )
[0087] Wherein, p i start_sub is the probability that the i-th position is marked as the start position of the subject, and p i end_sub is the probability that the i-th position is marked as the end position of the subject; p i start_obj is the probability that the i-th position is marked as the start position of the object, and p i end_obj is the probability that the i-th position is marked as the end position of the object.
[0088] S22. The relationship detection and characterization are specifically as follows:
[0089] Embed all the preset relationship labels into a high-dimensional vector, and then through a linear mapping layer, represent the final result as the most relevant node embedding. The embedding formula of the relationship node is:
[0090] R m =W r *E([r1, r2,..., r m )+b r
[0091]
[0092] Wherein, r i is the one-hot vector of the relationship index in the predefined relationship, m is the number of predefined relationships, E is the relationship embedding matrix, W r and br Trainable parameters for the predefined process of the relationship node, R m Is the initial relationship node, Is a high-dimensional relationship vector;
[0093] Predict the potential relationship information contained in the feature-vectorized input statement. The specific steps are as follows: First, add the obtained initial relationship node information to the initial statement, and add them together to a fully connected layer for neural network calculation. Then, finally obtain the relationship information feature through the sigmoid function; At the same time, the high-dimensional feature vector W r , b r The weights change, and then determine the feature of the relationship information. The calculation formula of the relationship information feature is as follows:
[0094]
[0095] Among them, Is the high-dimensional relationship vector obtained in the previous step, h i Is the statement information after Bert output, W r And b r Are the trainable weights for the relationship detection process, and sigmoid is the activation function.
[0096] S3. Perform entity and relationship joint prediction on the vectorized statement, and use the entity feature information and relationship feature information as auxiliary dimension feature information for information enhancement processing. The specific steps are as follows:.
[0097] Add the entity head information feature and the entity tail feature to the statement feature respectively, and then multiply by the relationship feature information; Two fully connected layer networks are used. One network is used to predict the subject-relationship, and the other network is used to predict the object-relationship. After the self-adjustment and training of the network, the joint feature information of the entity and the relationship is obtained;
[0098] The calculation formula of the joint feature information of the entity and the relationship is as follows:
[0099]
[0100]
[0101] Among them, sigmoid is the activation function, T i start And T i end Are the subject feature information and the object feature information h i,relation Is the vectorized statement feature, Is the relationship feature result calculated by prediction, W start , bstart , W end , and b end are trainable weight parameters.
[0102] S4. Concatenate or perform link prediction on the predicted entity and relationship feature information, and finally form triples.
[0103] Furthermore, the concatenating or performing link prediction on the predicted entity and relationship feature information to finally form triples is specifically as follows:
[0104] S41. Perform category judgment on the entity-relationship jointly predicted feature information. The judgment method is to construct two one-dimensional matrices with the same length as the number of relationship libraries. By traversing the results of the two-party joint prediction, map the IDs corresponding to the predicted relationship values of the two parties to the array subscript positions, so as to register the number of relationships. Finally, two categories of unique relationship matching and multi-relationship matching are obtained; the unique relationship matching means that the prediction numbers of the subject-relationship and object-relationship under the same relationship are both not greater than 1; the multi-relationship matching means that the prediction numbers of the subject-relationship and object-relationship under the same relationship are both greater than 1.
[0105] S42. For the unique relationship matching category, adopt the direct concatenation principle, and perform matching and concatenation on the data under the same relationship to obtain triples;
[0106] For the multi-relationship matching category, concatenate the two parts of the matrix of the start position information of the subject and the relationship and the start position information of the object and the relationship according to the relationship to form Tr = [Sub start , Obj start , rel], where Sub start is the start position information of the subject, Obj start is the start position information of the object, and then re-perform probability prediction on the concatenated Tr. The calculation formula is as follows:
[0107] P i Tr = sigmoid(W i Tr i + b i )
[0108] where Tr is formed by concatenating the main start relationship matrix and the target start relationship matrix through the relationship. W i and b i are re-prediction trainable weight values, and sigmoid is an activation function;
[0109] S43. Through the obtained T i sub =(p i start_sub , pi end_sub ) and T i obj = (p i start_obj , p i end_obj ) information, expand the header information into complete triple information; among them, T i sub is the subject information, T i obj is the object information; p i start_sub is the probability that the i-th position is marked as the start position of the subject, p i end_sub is the probability that the i-th position is marked as the end position of the subject; p i start _obj The probability that the i-th position is marked as the start position of the object, p i end_obj The probability that the i-th position is marked as the end position of the object.
[0110] Please refer to Figure 3 , in another embodiment of the present application, a method for extracting entity and relationship knowledge with enhanced feature dimensions for statements is provided, including the following steps:
[0111] Step 1: Input "Tom was born in 1942 and has lived in New York ever since." into the encoder in the Bert model to extract its hidden features, and obtain a vectorized statement H = [0.89902386 0.09758244 -0.06996521 0.20864412 0.03722338 -1.117653 0.3860746 -0.08808775 0.5787261 -0.2631619......] with context semantic features;
[0112] Step 2: Perform entity detection and characterization and relationship detection and characterization on the vectorized statement H; first, input the vectorized statement H in Step 1 into the fully connected layer, and calculate the start position probability and end position probability of the entity. The calculation formulas for the start position probability and end position probability of the entity are as follows:
[0113] p i start_sub(obj) = sigmoid(W start h i + b start )
[0114] p iend_sub(obj) = sigmoid(W end h i + b end )
[0115] After determining the start position probability and end position probability of the entity, the main body information T is extracted using the following formula i sub and the object information T i obj
[0116] T i sub = (p i start_sub , p i end_sub )
[0117] T i obj = (p i start_obj , p i end_obj )
[0118] The extracted main body information includes: Tom; The object information includes: New York, 1942
[0119] Secondly, for the vectorized statement, relationship detection and characterization are performed. All preset relationship labels are embedded into a high-dimensional vector, and then through a linear mapping layer, the final result is represented as the most relevant initialized relationship node embedding. The embedding formula of the initialized relationship node is as follows:
[0120] R m = W r * E([r1, r2,..., r m ) + b r
[0121]
[0122] Then, predict the potential relationship information contained in the feature vectorized input statement, and further determine the characteristics of the relationship information. The calculation formula of the relationship information characteristics is as follows:
[0123]
[0124] The obtained relationship information includes: Live in; Birth data; Birth place;
[0125] Step 3: Add the entity head information feature and the entity tail feature to the statement feature respectively, and then multiply by the relationship feature information. Two fully connected layer networks are used. One network is used to predict the subject-relationship, and the other network is used to predict the object-relationship. After the self-adjustment and training of the network, the combined feature information of the entity and the relationship is obtained;
[0126] The calculation formula of the combined feature information of the entity and the relationship is as follows:
[0127]
[0128]
[0129] Step 4: The obtained combined prediction feature information of the entity and the relationship is divided into two categories: unique relationship matching and multiple relationship matching.
[0130] The obtained subject-relationship and object-relationship information are respectively: (Tom, live in), (live in, NewYork); (Tom, Birth data), (Birth data, 1942); (Tom, Birth place), (Birth place, NewYork);
[0131] For the multiple relationship matching category, splice the two parts of the matrix of the start position information of the subject and the relationship and the start position information of the object and the relationship according to the relationship to form Tr, re-predict the probability of Tr, and then through the subject information and object information obtained in Step 2, expand the head information into triple information;
[0132] For the unique relationship matching category, use the direct splicing principle to obtain triples (Tom, Live in, New York), (Tom, Birth data, 1942), (Tom, Birth place, New York).
[0133] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0134] Based on the same idea as the method for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement in the above embodiments, the present invention also provides a system for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement, which can be used to execute the above method for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement. For the sake of illustration, in the structural schematic diagram of the embodiment of the system for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement, only the parts related to the embodiments of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than those illustrated, or combine certain components, or arrange different components.
[0135] Please refer to Figure 4 , in another embodiment of the present application, a system 100 for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement is provided. The system includes a vectorization module 101, a detection and characterization module 102, a joint prediction module 103, and a triple output module 104. The vectorization module 101 is configured to input a statement and perform vectorization using the BERT model to obtain a vectorized statement; the detection and characterization module 102 is configured to perform entity detection and characterization and relationship detection and characterization on the vectorized statement to obtain entity feature information and relationship feature information;
[0136] The joint prediction module 103 is configured to perform joint prediction of entities and relationships on the vectorized statement, and perform information enhancement processing on the entity features and relationship features as auxiliary dimension features to obtain entity and relationship joint prediction feature information;
[0137] The triple module 104 is configured to splice or perform link prediction on the predicted entity and relationship feature information to finally form a triple.
[0138] It should be noted that the system for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement of the present invention corresponds one-to-one with the method for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement of the present invention. The technical features and their beneficial effects described in the embodiments of the above method for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement are applicable to the embodiments of the system for extracting entity and relationship knowledge with statement-oriented feature dimension enhancement. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be repeated here. This is hereby declared.
[0139] In addition, in the implementation of the entity and relationship knowledge extraction system with statement-oriented feature dimension enhancement in the above embodiments, the logical division of each program module is only for illustration. In practical applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the entity and relationship knowledge extraction system with statement-oriented feature dimension enhancement is divided into different program modules to complete all or part of the functions described above.
[0140] Please refer to Figure 5 , in one embodiment, an electronic device for implementing a method for entity and relationship knowledge extraction with statement-oriented feature dimension enhancement is provided. The electronic device 200 may include a first processor 201, a first memory 202, and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as an entity and relationship knowledge extraction program 203 with feature dimension enhancement.
[0141] Among them, the first memory 202 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the first memory 202 may be an internal storage unit of the electronic device 200, such as the mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 may also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 200. Further, the first memory 202 may also include both the internal storage unit and the external storage device of the electronic device 200. The first memory 202 can be used not only to store application software installed in the electronic device 200 and various types of data, such as the code of the entity and relationship knowledge extraction program 203 with feature dimension enhancement, but also to temporarily store data that has been output or will be output.
[0142] In some embodiments, the first processor 201 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The first processor 201 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing programs or modules stored in the first memory 202, and calling data stored in the first memory 202, to execute various functions of the electronic device 200 and process data.
[0143] Figure 5 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 5 the shown structure does not constitute a limitation on the electronic device 200, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0144] The entity and relationship knowledge extraction program 203 with enhanced feature dimensions of the statements stored in the first memory 202 in the electronic device 200 is a combination of multiple instructions. When running in the first processor 201, it can achieve:
[0145] Vectorize the input statement to obtain a vectorized statement with context semantic features;
[0146] Perform entity detection and characterization and relationship detection and characterization on the vectorized statement to obtain entity feature information and relationship feature information respectively;
[0147] Perform entity and relationship joint prediction on the vectorized statement, and use the entity feature information and relationship feature information as auxiliary dimension feature information for information enhancement processing to obtain entity and relationship joint prediction feature information;
[0148] Stitch or perform link prediction on the entity and relationship joint prediction feature information to finally form a triple.
[0149] Furthermore, if the modules / units integrated in the electronic device 200 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory).
[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0151] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0152] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention should be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for entity and relationship knowledge extraction with enhanced feature dimensions oriented to statements, characterized in that, Including the following steps: Vectorize the input statement to obtain a vectorized statement with context semantic features; Perform entity detection and characterization and relationship detection and characterization on the vectorized statement to obtain entity feature information and relationship feature information respectively; The entity feature information refers to the subject information and object information extracted from the vectorized statement; the relationship feature information refers to the association features existing between the subject and the object extracted from the vectorized statement; Perform entity-relationship joint prediction on the vectorized statement, and use the entity feature information and relationship feature information as auxiliary dimension feature information for information enhancement processing to obtain entity-relationship joint prediction feature information; Concatenate or perform link prediction on the feature information of the entity-relationship joint prediction to finally form a triple; The entity refers to the subject and the object; The entity detection and characterization specifically are: Input the vectorized statement H into a fully connected layer, calculate the start position probability and end position probability of the entity. If the probability of the start position is greater than a preset first threshold, then determine this start position as the start position of the entity in the vectorized statement; similarly, if the probability of the end position is greater than a preset second threshold, then determine this end position as the end position of the entity in the vectorized statement; meanwhile, the neural network of the fully connected layer will be trained according to the label information of the training set, and continuously adjust the trainable weight values W and b; The calculation formulas for the start position probability and end position probability of the entity are as follows: p i start_sub(obj) = sigmoid(W start h i + b start ) p i end_sub(obj) = sigmoid(W end h i + b end ) where p i start_sub(obj) is the probability that the i-th position in the input sentence is marked as the start position of an entity, and p i end_sub(obj) is the probability that the i-th position in the input sentence is marked as the end position of an entity; h i is the output result of the encoder layer, W start and b start are the trainable weight values for calculating the start position probability of an entity, and W end and b end are the trainable weight values for calculating the end position probability of an entity, and sigmoid is the activation function; After determining the probabilities of the start position and end position of the entity, the subject information T is extracted i sub and the object information T i obj , and the formula is: T i sub = (p i start_sub , p i end_sub ) T i obj = (p i start_obj , p i end_obj ) Among them, p i start_sub is the probability that the i-th position is marked as the start position of the subject, p i end_sub is the probability that the i-th position is marked as the end position of the subject; p i start_obj The probability that the i-th position is marked as the start position of the object, p i end_obj is the probability that the i-th position is marked as the end position of the object; The calculation formula for obtaining the entity-relationship joint prediction feature information is as follows: Among them, sigmoid is the activation function, T i start and T i end are the main feature information and the feature information of the object respectively, h i,relation is the vectorized statement feature, is the predicted relationship feature result, W start and b start and W end and b end are trainable weight parameters.
2. The entity and relationship knowledge extraction method for statement-oriented feature dimension enhancement according to claim 1, wherein The vectorization of the input statement specifically is: Extract the hidden features of each word in the input statement through the encoder in the Bert model, and convert the input statement into a vectorized statement with context semantic features. The expression of the vectorized statement H is as follows: H = Bert[{x1, x2,..., x n ,..., x m} * mask] H = [h1, h2,.., h n ,..., h m where x1, x2, ..., x n , ..., x m are the IDs of each word in the input sentence mapped to the corresponding dictionary of the Bert model, n represents the length of the input sentence sequence, m is the total length of the sentence after vectorization and padding, mask is the actual valid sentence information in the input sentence, h1, h2, .., h n , ..., h m are the word vectors incorporating context information.
3. The method for extracting entity and relationship knowledge with enhanced statement-oriented feature dimensions according to claim 1, wherein The relationship detection and characterization specifically are: Embed all preset relationship labels into a high-dimensional vector, and then through a linear mapping layer, represent the final result as the initial relationship node embedding with the most relationship. The calculation formula for the initial relationship node embedding is: R m = W r * E([r1, r2,..., r m ) + b r Among them, r i is the one-hot vector of the relationship index in the predefined relationship, m is the number of predefined relationships, E is the relationship embedding matrix, W r and b r are the trainable parameters of the predefined process of the relationship node, R m is the initial relationship node, which is a high-dimensional relationship vector; Predict the initial relational node information contained in the feature vectorized input statement. First, add the obtained initial relational node information to the initial statement, add the initial statement with the initial relational node information to a fully connected layer for neural network calculation, and finally obtain the relational information feature through the sigmoid function. At the same time, the weights of the high-dimensional feature vectors W r , b r change during continuous training, and then determine the features of the relational information. The calculation formula for the relational information feature is as follows: Among them, is the high-dimensional relationship vector obtained in the previous step, h i is the output result of the encoder layer, W r and b r are the trainable weights for the relationship detection process, and sigmoid is the activation function.
4. The method for entity and relationship knowledge extraction with enhanced statement-oriented feature dimensions according to claim 1, wherein Perform entity-relationship joint prediction on the vectorized statement, and use the entity feature information and relationship feature information as entity auxiliary dimension features for information enhancement processing specifically: Add the entity head information feature and entity tail feature to the statement feature respectively, and then multiply by the relationship feature information. Use two fully connected layer networks, one network for predicting the subject-relationship, and the other network for predicting the object-relationship; after the self-adjustment and training of the network, obtain the entity-relationship joint prediction feature information.
5. The method for extracting entity and relationship knowledge with enhanced statement-oriented feature dimensions according to claim 1, wherein Concatenate or perform link prediction on the predicted entity and relationship feature information to finally form a triple specifically: Perform category judgment on the feature information of the joint prediction of entities and relationships. The judgment method is to construct two one-dimensional matrices with the same length as the number of relationship databases. By traversing the results of the joint prediction output by both parties, map the IDs corresponding to the relationship values predicted by both parties to the array subscript positions, so as to register the number of relationships. Finally, two categories of unique relationship matching and multiple relationship matching are obtained; the unique relationship matching means that the prediction numbers of the subject-relationship and object-relationship under the same relationship are both not greater than 1; the multiple relationship matching means that the prediction numbers of the subject-relationship and object-relationship under the same relationship are both greater than 1. Adopt the direct splicing principle for the unique relationship matching category, and perform matching splicing on the data under the same relationship to obtain triples. For the multi-relation matching class, the start position information of the subject and the two matrices of the start position information of the relation and the object and the relation are concatenated by relation to form Tr = [Sub start , Obj start , rel], where Sub start is the start position information of the subject, Obj start is the start position information of the object, and rel represents the relation information corresponding to the start position information of the subject and the start position information of the object; then, probability prediction is performed again on the concatenated Tr, and the calculation formula is as follows: Among them, Tr is formed by splicing the main start relationship matrix and the target start relationship matrix through relationships, W i and b i are trainable weight values for re-prediction, and sigmoid is an activation function; by obtaining T i sub =(p i start_sub , p i end_sub ) and T i obj =(p i start_obj , p i end_obj ) information, the head information is completely extended into triple information; among them, T i sub is the subject information, and T i obj is the object information; p i start_sub is the probability that the i-th position is marked as the start position of the subject, and p i end_sub is the probability that the i-th position is marked as the end position of the subject; p i start_obj is the probability that the i-th position is marked as the start position of the object, and p i end_obj is the probability that the i-th position is marked as the end position of the object.
6. A statement-oriented entity and relationship knowledge extraction system with enhanced feature dimensions, characterized in that Applied to the entity and relationship knowledge extraction method with enhanced feature dimensions for statements described in any one of claims 1-5, including a vectorization module, a detection and characterization module, a joint prediction module, and a triple output module. The vectorization module is used to vectorize the input statement to obtain a vectorized statement with context semantic features. The detection and characterization module is used to perform entity detection and characterization and relationship detection and characterization on the vectorized statement, and respectively obtain entity feature information and relationship feature information. The joint prediction module is used to perform joint prediction of entities and relationships on the vectorized statement, and use the entity feature information and relationship feature information as auxiliary dimension feature information for information enhancement processing to obtain the feature information of the joint prediction of entities and relationships. The triple output module is used to splice or perform link prediction on the feature information of the joint prediction of entities and relationships, and finally form triples.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the entity and relationship knowledge extraction method with enhanced feature dimensions for statements described in any one of claims 1-5.
8. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the entity and relationship knowledge extraction method with enhanced feature dimensions for statements described in any one of claims 1-5.
Citation Information
Patent Citations
Multi-triad joint extraction method based on knowledge graph embedding
CN111444305A
Entity relation joint extraction method based on global pointer network
CN114417839A