A triple entity alignment method, system, computer device and storage medium

By constructing positive and negative samples triplets and using the Bert-EA model for entity alignment, the problem of traditional algorithms ignoring semantic information is solved, the accuracy of entity alignment is improved, and the simplification of the knowledge graph and structural rationality are ensured.

CN117807249BActive Publication Date: 2025-05-06AIR FORCE UNIV PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311855702.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-05-06
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

The discrete attribute information obtained by traditional entity alignment algorithms ignores many implicit semantic information, resulting in too low alignment accuracy.

Method used

The entity relationship triplets are extracted through knowledge extraction technology, and positive sample triplets and negative sample triplets are constructed. The triplets are encoded and fused using the Bert-EA model to consider the potential relationship between entities and improve the accuracy of alignment.

Benefits of technology

By considering multiple semantic information, the accuracy of entity alignment is significantly improved, ensuring the simplification of the knowledge graph and the rationality of the structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117807249B_ABST
    Figure CN117807249B_ABST
Patent Text Reader

Abstract

The present invention discloses a triple entity alignment method, system, computer equipment and storage medium, which relates to the technical field of entity alignment, and includes the following steps: extracting entity relationship triples; obtaining entity pairs with the same attributes and different attributes in the entity relationship triples, wherein each pair of entity pairs with the same attributes introduces an "equal to" relationship and constitutes a positive sample triple, and each pair of entity pairs with different attributes introduces a "not equal to" relationship and constitutes a negative sample triple; training a Bert‑EA model through positive sample triples and negative sample triples; and aligning and judging the entity pairs of entity relationship triples through the trained Bert‑EA model. The positive and negative sample triples of the present invention not only take into account the attributes between any two entities, but also take into account the potential relationships that may exist between entities. A Bert‑EA model is also proposed, which integrates various aspects of semantic information and greatly improves the alignment accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of entity alignment, and in particular to a triple entity alignment method, system, computer equipment and storage medium. Background Art

[0002] With the continuous advancement of military construction, more and more novel technologies are gradually being applied to the daily construction of the military. The widespread promotion of new technologies has led to an increasingly rich case reserve of fault diagnosis and fault repair data. However, due to the immaturity of data management and data reading technology in the military, a large amount of fault diagnosis and fault repair data cannot be applied to daily diagnosis and maintenance work through intelligent reading technology. In order to solve this problem, it is necessary to use the currently widely used knowledge graph as a tool to visualize these fault data, so as to facilitate the review and reading during the equipment inspection and maintenance process.

[0003] The current research focus is on how to construct a knowledge graph of fault data. By consulting a large number of reference materials, it is concluded that the core tasks of constructing a knowledge graph are named entity recognition and relationship extraction. After extracting triple data through knowledge extraction technology, these triple data are initially represented in a symbolic form as a knowledge graph. In order to further improve the knowledge graph, it is necessary to further process the previously extracted data, such as performing entity alignment on the extracted triples and matching triplets with similar semantics together, thereby reducing the number of unnecessary nodes in the constructed knowledge graph and simplifying the structure of the graph.

[0004] Most traditional entity alignment algorithms are based on calculating the similarity between word vectors to achieve entity alignment. The reason for calculating the distance between word vectors is that it is necessary to determine whether two entities are "synonymous isomers" during entity alignment. The so-called "synonymous isomer" refers to the same entity having several different names. For example, the common fuel tank component "fuel supply pipe pressure sensor" is also called "oil pressure sensor". When drawing a knowledge graph, if there are a large number of similar "synonymous isomers", the graph will be too bloated, which is not convenient for the implementation of knowledge reasoning based on entity relationship triples. Therefore, it is necessary to use an entity alignment algorithm to alleviate the problem of too bloated nodes in the knowledge graph. The idea of ​​the distance-based entity alignment algorithm is to convert the original entity into the corresponding word vector form through the Word2Vector method, and then determine whether the two entities can be merged by calculating the distance between the two entity word vectors.

[0005] However, the discrete attribute information obtained by traditional entity alignment algorithms ignores many implicit semantic information, such as the association semantics between attributes and the semantic information between triple structures, which makes the alignment accuracy too low. Summary of the invention

[0006] The present invention provides a triple entity alignment method, system, computer equipment and storage medium, which solves the problem that discrete attribute information obtained by traditional entity alignment algorithms ignores various implicit semantic information.

[0007] The present invention provides a triple entity alignment method, comprising the following steps:

[0008] Extract entity relationship triples through knowledge extraction technology;

[0009] Obtain entity pairs with the same attributes and different attributes in the entity relationship triplet, where each pair of entity pairs with the same attributes introduces an "equal to" relationship and constitutes a positive sample triplet, and each pair of entity pairs with different attributes introduces a "not equal to" relationship and constitutes a negative sample triplet;

[0010] The Bert-EA model is trained using positive sample triplets and negative sample triplets;

[0011] The trained Bert-EA model is used to align entity pairs of entity relationship triples.

[0012] Preferably, the Bert-EA model includes:

[0013] The knowledge encoding layer encodes the triples to obtain the vector representation of each character and the overall semantic representation of the triples;

[0014] The information fusion layer fuses the vector representation of each character and the overall semantic representation of the triple to obtain the semantic information of the triple;

[0015] The scoring layer calculates the scores of the semantic information of the triples.

[0016] Preferably, the Bert-EA model is trained by using positive sample triplets and negative sample triplets, comprising the following steps:

[0017] Input each triple into the Bert model to obtain the vector representation of each character and the overall semantic representation of the triple;

[0018] The pooling layer is used to obtain a combined character representation with the same dimension as the overall semantic representation of the triple;

[0019] The overall semantic representation of the triple is fused with the character merged representation through a nonlinear layer to obtain the semantic information representation of the triple;

[0020] The semantic information representation of the triple is calculated through the scoring layer to obtain the corresponding score;

[0021] If the score exceeds the threshold, the two entities are judged to point to the same object and aligned, otherwise no alignment is performed.

[0022] Preferably, the semantic information of the triple is represented as follows:

[0023]

[0024] In the formula, σ is the nonlinear activation function, W is the connection weight parameter, b is the bias parameter, It is the character combination representation of the triple, h cls It is the overall semantic representation of the triple.

[0025] Preferably, the semantic information representation of the triple is calculated by the following formula:

[0026]

[0027] In the formula, is the sigmoid activation function, W o is the connection weight of the scoring layer, b o is the bias of the scoring layer, For score.

[0028] Preferably, the Bert-EA model is trained by minimizing the cross entropy loss function, which is as follows:

[0029]

[0030] Where, θ=(W,b,W o ,b o ), ((h,r,t),y) is a triplet, y is the label, and L is the loss function.

[0031] Preferably, after the trained Bert-EA model performs alignment judgment on the entity pairs of the entity relationship triples, additional judgments are required, including:

[0032] Additionally, it is determined whether the two entities in the fused entity pair have the same surrounding relationship;

[0033] If the same surrounding relationship exists, the entity pair is considered to be correctly fused; if the same surrounding relationship does not exist, it is considered that the potential connection between the pair of entities needs to be further judged manually.

[0034] A triple entity alignment system in the process of fault knowledge graph construction, comprising:

[0035] The extraction module is used to extract entity relationship triples through knowledge extraction technology;

[0036] A sample construction module is used to obtain entity pairs with the same attributes and different attributes in the entity relationship triplet, where each pair of entity pairs with the same attributes introduces an "equal" relationship and constitutes a positive sample triplet, and each pair of entity pairs with different attributes introduces a "not equal" relationship and constitutes a negative sample triplet;

[0037] Model training module, used to train the Bert-EA model through positive sample triplets and negative sample triplets;

[0038] The judgment module is used to align the entity pairs of entity relationship triples through the trained Bert-EA model.

[0039] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the triple entity alignment method when executing the program.

[0040] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the triple entity alignment method is implemented.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The present invention first determines whether the attributes of any two entities are the same. If they are the same, the entity pair and the "equal" relationship are combined to form a positive sample triple. If the attributes of any two entities are not the same, a "not equal" relationship is introduced to form a negative sample triple. The positive sample triples and negative sample triples constructed by the present invention not only take into account the attributes between any two entities, but also take into account the potential relationship that may exist between the entities. At the same time, a Bert-EA model is proposed, which first encodes the triples to obtain the vector representation of each character and the overall semantic representation of the triple, and then fuses the two to obtain the final semantic information of the triple. The present invention integrates various aspects of semantic information and greatly improves the accuracy of entity alignment. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 A schematic diagram of the knowledge representation process based on a deep neural network of the present invention;

[0045] Figure 2It is a schematic diagram of the Bert-EA model structure of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Reference Figure 1 The present invention provides a triple entity alignment method, which specifically includes the following steps:

[0048] Step 1: Extract entity relationship triples through knowledge extraction technology.

[0049] A large amount of triple data is obtained through knowledge extraction technology. A large number of entities in these triples point to the same object in the real world. Therefore, these entities have similar information. This causes the extracted knowledge to contain a lot of redundant information. Entity alignment technology is needed to eliminate this redundant information. Entity alignment aims to determine whether two or more entities from different information sources point to the same object in the real world. If multiple entities represent the same object, an alignment relationship is constructed between these entities, and the information contained in the entities is fused and aggregated. In recent years, with the continuous development of knowledge representation learning technology, researchers have proposed many entity alignment methods based on knowledge representation learning, which are significantly better than traditional methods.

[0050] Step 2: Obtain entity pairs with the same attributes and different attributes in the entity relationship triplet, where each pair of entity pairs with the same attributes introduces an "equal" relationship and constitutes a positive sample triplet, and each pair of entity pairs with different attributes introduces a "not equal" relationship and constitutes a negative sample triplet.

[0051] Considering that deep learning has good performance in knowledge representation, the present invention proposes an entity alignment method for deep knowledge representation to solve the entity alignment problem. Since this model uses the pre-trained Bert model to encode triples, the model is called Bert based Entity Alignment, or Bert-EA model for short.

[0052] The core of this model is: for the alignment operation, a new relationship called "equal" is introduced, that is, the relationship "equal" is introduced for two aligned entities, so that the entity alignment problem is transformed into a link prediction problem of discovering the "equal" relationship, which not only takes into account the attributes between entities, but also takes into account the potential relationship between entities. With the advantage of pre-trained deep network models that can well encode text information, the semantic information of triples is encoded using the pre-trained Bert model that is widely used today. The basic process of the entire algorithm: introduce the "equal" relationship, add the "equal" relationship to the seed entity pair, and construct a new triple. All triples are input into the pre-trained Bert model for encoding. A two-layer fully connected network is constructed on the Bert model to fuse the triple information. Finally, an output layer for calculating fact scores is constructed to predict the score value of the triple.

[0053] Step 3: Train the Bert-EA model using positive sample triplets and negative sample triplets.

[0054] Reference Figure 1 ,The Bert-EA model includes knowledge encoding layer, information fusion layer and scoring layer.

[0055] Knowledge encoding layer: Through the pre-trained Bert model, each pair of triples is converted into corresponding character vectors in the form of each character in turn, so as to obtain the semantic information of each pair of triples after encoding.

[0056] Information fusion layer: The obtained word fragment matrix is ​​input into the non-linear layer, and the encoding of all characters is fused into a vector representation using the pooling operation. Then, the overall encoding information of the triple is fused with the triple vector representation through the non-linear layer to obtain the semantic information of the final triple.

[0057] Scoring layer: Through nonlinear operations, it outputs the score of whether the triple is a fact. The higher the score, the more likely it is a fact triple.

[0058] By analyzing the ideas of the algorithm, the specific algorithm process is reproduced as follows:

[0059] (1) Before selecting any two entities to form an entity pair, first determine whether the attribute values ​​of the two entities are the same. If the attributes of the two entities are the same, the entity pair and the "equal" relationship form a triple. If the attributes of any two entities are different, the pair of entities is used as a negative sample to optimize the model parameters. Finally, the set of entity pairs with the same attribute values ​​D is obtained:

[0060] (h i ,t i )∈D'

[0061] (2)(hi ,t i )∈D' For entity pairs in set D, introduce label 1 to the factual triple (positive sample), denoted as ((h,r,t),y),y=1. Correspondingly, obtain some non-factual triples (negative samples) and introduce label 0, denoted as ((h,r,t),y),y=0. Assuming that there are k kinds of relations, introduce the k+1th relation "equal", indicating that the head entity and the tail entity point to the same real target, that is, the aligned entity pair. Construct the known aligned entity pairs into triples, denoted as ((h,r k ,t),1), merged into the original triple.

[0062] (3) Randomly select a sample ((h, r, t), y) and input (h, r, t) into the Bert model to obtain the vector representation F of each character i , and the overall semantic representation of the triple h cls , merging the character representations into And through pooling, we get cls Representation of the same dimension

[0063] (4) The vector representation h of the triplet is fused through a nonlinear layer cls and The final triplet fusion semantic information representation h is obtained:

[0064]

[0065] (5) After the scoring layer, predict the score of the input triple

[0066]

[0067] in is the sigmoid activation function, W o With b o Represents the connection weights and biases of the scoring layer.

[0068] The above is the entire data encoding process of the Bert-EA model. The model is trained by minimizing the cross entropy loss function. The specific loss function is:

[0069]

[0070] where θ=(W,b,W o ,b o) are the parameters of the model, and the model is trained by the gradient descent method. When testing, set the threshold in advance. First, determine whether the attributes of any two entities are the same. If they are the same, the entity pair and the "equal" relationship form a triple. If the attributes of any two entities are not the same, the pair of entities are used as negative samples to optimize the model parameters. In this way, not only the attributes between any two entities are taken into account, but also the potential relationship that may exist between the entities. Input the trained Bert-EA model to get the score value. If the score value exceeds the threshold, it is considered that the two entities point to the same object and are aligned, otherwise no alignment is performed.

[0071] For the loss function, the following algorithm is used to solve it:

[0072] because:

[0073]

[0074] Then we can get:

[0075]

[0076] So we can deduce:

[0077]

[0078] Thus, Where α is the learning rate.

[0079] Set the gradient descent method for the learning phase of the model to update the iterative parameter w j , so that the loss function L(θ) gradually decreases, so that the model training is carried out by minimizing the loss function loss.

[0080] Step 4: Use the trained Bert-EA model to align entity pairs of entity relationship triples.

[0081] Finally, when testing the performance of the model, it is necessary to determine whether the two entities in the fused entity pair have the same surrounding relationship. If the same surrounding relationship exists, the entity pair is considered to be correctly fused. If the same surrounding relationship does not exist, it is considered that the potential connection between the pair of entities needs to be further manually determined. Finally, a trained model is obtained.

[0082] Based on the same concept, the present invention also provides a triple entity alignment system in the process of constructing a fault knowledge graph, including an extraction module, a sample construction module, a model training module and a judgment module.

[0083] The extraction module is used to extract entity relationship triples through knowledge extraction technology.

[0084] The sample construction module is used to obtain entity pairs with the same attributes and different attributes in the entity relationship triplet, where each pair of entity pairs with the same attributes introduces an "equal" relationship and constitutes a positive sample triplet, and each pair of entity pairs with different attributes introduces a "not equal" relationship and constitutes a negative sample triplet.

[0085] The model training module is used to train the Bert-EA model using positive sample triplets and negative sample triplets.

[0086] The judgment module is used to perform alignment judgment on entity pairs of entity relationship triples through the trained Bert-EA model.

[0087] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned triple entity alignment method when executing the program.

[0088] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned triple entity alignment method is implemented.

[0089] Example 1

[0090] The data used in the experiment is a set of triples extracted from 100 text data about drone maintenance using the Hor-Ver-Casrel model, and some aligned entity pairs are manually annotated. There are 800 extracted triples and 50 aligned entity pairs. Model training and testing are based on these data.

[0091] The data set used in this experiment is manually constructed, and its specific format is shown in Table 1:

[0092] Table 1. Artificially constructed data format

[0093] 1 equal <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Fuel pipe oil pressure sensor< / e2> . 2 No wait <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Pressure servo valve< / e2> . 3 No wait <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Fuel tank sensor< / e2> . 4 equal <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Oil pressure sensor< / e2> . 5 No wait <e1> Fuel supply pipe oil pressure sensor< / e1> and <e2> Pressure servo valve< / e2> . 6 No wait <e1> Fuel pipe oil pressure sensor< / e1> and <e2> Fuel tank sensor< / e2> . 7 equal <e1> Fuel supply pipe oil pressure sensor< / e1> and <e2> Oil pressure sensor< / e2> . 8 No wait <e1> Pressure servo valve< / e1> and <e2> Fuel tank sensor< / e2> . 9 No wait <e1> Pressure servo valve< / e1> and <e2> Oil pressure sensor< / e2> . 10 No wait <e1> Fuel tank sensor< / e1> and <e2> Oil pressure sensor< / e2> .

[0094] The idea of ​​building the dataset is mainly based on the extracted named entities. First, 50 pairs of entities that have been aligned are extracted as positive samples and labeled "equal". Then, 50 pairs of entities that do not have an alignment relationship are randomly selected as negative samples and labeled "unequal" to construct the dataset. Then, special symbols are added before and after the first entity of each entity pair. <e1> and< / e1> , and add special symbols before and after the second entity <e2> and< / e2> Finally, the label of each entity pair is placed at the beginning to complete the construction of the dataset. For the training set and test set, 90% of the constructed dataset is used as the training set, and the remaining part is used as the test set to train the model.

[0095] The final model is on the test set:

[0096] The traditional distance-based entity alignment model has an accuracy of 73% and a recall of 67%;

[0097] The Bert-EA model achieved a precision of 85% and a recall of 75%.

[0098] Through experimental verification, it is found that the effects of the Bert-EA model meet the requirements of the technical indicator list, that is, "the accuracy of the algorithm is not less than 80%, and the recall rate is not less than 70%." Because the Bert-EA model not only considers the attribute information of the entity, but also the surrounding relationship information of the entity, the model has better effects.

[0099] Example 2

[0100] Taking the abnormal oil supply pipe pressure failure as an example, the model introduced above is applied to a specific failure case for entity alignment.

[0101] The format of the training set for the abnormal oil supply pipe pressure failure case is as follows:

[0102] equal <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Fuel supply pipe oil pressure sensor< / e2> .

[0103] No wait <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Pressure servo valve< / e2> .

[0104] No wait <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Fuel tank sensor< / e2> .

[0105] equal <e1> Fuel supply pipe pressure sensor< / e1> and <e2> Oil pressure sensor< / e2> .

[0106] No wait <e1> Fuel supply pipe oil pressure sensor< / e1> and <e2> Pressure servo valve< / e2> .

[0107] No wait <e1> Fuel pipe oil pressure sensor< / e1> and <e2> Fuel tank sensor< / e2> .

[0108] equal <e1> Fuel supply pipe oil pressure sensor< / e1> and <e2> Oil pressure sensor< / e2> .

[0109] No wait <e1> Pressure servo valve< / e1> and <e2> Fuel tank sensor< / e2> .

[0110] No wait <e1> Pressure servo valve< / e1> and <e2> Oil pressure sensor< / e2> .

[0111] No wait <e1> Fuel tank sensor< / e1> and <e2> Oil pressure sensor< / e2> .

[0112] The final entity alignment results are as follows:

[0113] The fuel supply pipe is equal to the engine fuel supply pipe

[0114] Cable plug equals to connection plug

[0115] The oil supply pressure is equal to the oil supply pipe pressure

[0116] Cable plug is equivalent to cable connection plug

[0117] Connection plug equivalent to cable connection plug

[0118] The fuel supply pressure sensor is equal to the fuel supply pipe pressure sensor

[0119] By comparing the above entity alignment results with the real dictionary of entity alignment, it is found that the final entity alignment results are exactly the same as the real results. Therefore, it is believed that the performance of the model meets the standard and the model has learned deep semantic information between entities during the training process.

[0120] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0121] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A triple entity alignment method, characterized in that: The following steps are involved: Extract entity relationship triples through knowledge extraction technology; Obtain entity pairs with the same attributes and different attributes in the entity relationship triples, where each pair of entity pairs with the same attributes introduces an "equal" relationship and constitutes a positive sample triple, and each pair of entity pairs with different attributes introduces a "not equal" relationship and constitutes a negative sample triple; The Bert-EA model is trained using positive sample triplets and negative sample triplets; The trained Bert-EA model is used to align entity pairs of entity relationship triples; The Bert-EA model includes: The knowledge encoding layer encodes the triples to obtain the vector representation of each character and the overall semantic representation of the triples; The information fusion layer fuses the vector representation of each character and the overall semantic representation of the triple to obtain the semantic information of the triple; The scoring layer calculates the scores of the semantic information of the triples; The semantic information of the triple is represented as follows: ; In the formula, h is the semantic information representation of the triple, is a nonlinear activation function, W is the connection weight parameter, b is the bias parameter, It is the character combination representation of the triple. is the overall semantic representation of the triple; The semantic information representation of the triple is calculated by the following formula: ; In the formula, is the sigmoid activation function, W o is the connection weight of the scoring layer, b o is the bias of the scoring layer, For fractions; The Bert-EA model is trained by minimizing the cross entropy loss function, which is as follows: ; In the formula, ,(( h , r , t ), y ) is a triple, y For labels, L is the loss function.

2. A triple entity alignment method as claimed in claim 1, characterized in that: The Bert-EA model is trained using positive and negative sample triplets, including the following steps: Input each triple into the Bert model to obtain the vector representation of each character and the overall semantic representation of the triple; The pooling layer is used to obtain a combined character representation with the same dimension as the overall semantic representation of the triple; The overall semantic representation of the triple is fused with the character merged representation through a nonlinear layer to obtain the semantic information representation of the triple; The semantic information representation of the triple is calculated through the scoring layer to obtain the corresponding score; If the score exceeds the threshold, the two entities are judged to point to the same object and aligned, otherwise no alignment is performed.

3. A triple entity alignment method as claimed in claim 1, characterized in that: After the trained Bert-EA model is used to align the entity pairs of entity relationship triples, additional judgments are required, including: Additionally, it is determined whether the two entities in the fused entity pair have the same surrounding relationship; If the same surrounding relationship exists, the entity pair is considered to be correctly fused; if the same surrounding relationship does not exist, it is considered that the potential connection between the pair of entities needs to be further judged manually.

4. A triple entity alignment system, characterized in that: include: The extraction module is used to extract entity relationship triples through knowledge extraction technology; A sample construction module is used to obtain entity pairs with the same attributes and different attributes in the entity relationship triplet, where each pair of entity pairs with the same attributes introduces an "equal" relationship and constitutes a positive sample triplet, and each pair of entity pairs with different attributes introduces a "not equal" relationship and constitutes a negative sample triplet; Model training module, used to train the Bert-EA model through positive sample triplets and negative sample triplets; The judgment module is used to align the entity pairs of entity relationship triples through the trained Bert-EA model; The Bert-EA model includes: The knowledge encoding layer encodes the triples to obtain the vector representation of each character and the overall semantic representation of the triples; The information fusion layer fuses the vector representation of each character and the overall semantic representation of the triple to obtain the semantic information of the triple; The scoring layer calculates the scores of the semantic information of the triples; The semantic information of the triple is represented as follows: ; In the formula, h is the semantic information representation of the triple, is a nonlinear activation function, W is the connection weight parameter, b is the bias parameter, It is the character combination representation of the triple. is the overall semantic representation of the triple; The semantic information representation of the triple is calculated by the following formula: ; In the formula, is the sigmoid activation function, W o is the connection weight of the scoring layer, b o is the bias of the scoring layer, For fractions; The Bert-EA model is trained by minimizing the cross entropy loss function, which is as follows: ; In the formula, ,(( h , r , t ), y ) is a triple, y For labels, L is the loss function.

5. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the triple entity alignment method described in any one of claims 1 to 3 is implemented.

6. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the triple entity alignment method described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Mapping knowledge domain relation inference method and device, computer equipment, and storage medium

    CN108446769A

  • Neuromorphic hardware and method for storing and / or processing a knowledge graph

    EP4030350A1