A Knowledge Graph Fusion Method Based on Entity Merging
By using recurrent neural networks in knowledge graph fusion to embed entity attributes into the same feature space, the problems of structural differences and attribute inconsistencies between different graphs are solved, and efficient and accurate knowledge graph fusion is achieved.
Patent Information
- Application Number
- CN202210964519.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-08-11
AI Technical Summary
In the process of knowledge graph fusion, there are problems such as large structural differences between different graphs, inconsistent text descriptions, and inconsistent number of entity attributes and arrangement order, which makes it difficult to calculate similarity.
Using a recurrent neural network-based method, treating the properties of an entity as a context, mapping the entities embedding vectors in the two graphs to the same feature space, so that the embedding vectors contain all the attribute information of the entity and the same dimensions.
It effectively solves the problem that entities are difficult to calculate similarity between different maps, and achieves fast and accurate knowledge graph fusion.
Smart Images

Figure CN115438188B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the fields of computer artificial intelligence and knowledge graphs, and particularly relates to a knowledge graph fusion method based on entity merging. Background Art
[0002] With the booming development of the Internet, how to effectively construct, utilize, and mine the knowledge covered by Internet data has become a new challenge. A knowledge graph describes the entities and semantic associations existing in the objective world and provides users with structured knowledge. As an excellent knowledge organization method, knowledge graphs have gradually received widespread attention in the academic and industrial communities. To build a high-quality knowledge graph, an efficient approach is to fuse the knowledge graphs generated from multiple knowledge bases. Therefore, graph fusion is an essential link in knowledge graph construction. During the process of knowledge graph fusion, problems such as duplication among multi-source knowledge, diverse semantic ambiguities, and uneven quality often occur. Therefore, how to identify the same entities in multiple graphs and merge them into one entity is a key link in knowledge graph fusion.
[0003] An entity alignment method for a GCN twin network (CN110472065B). When processing attributes, this method uses the frequency of the attributes appearing in the graph as the embedding vector of the attributes, lacking semantic features. At the same time, this method needs to count all the attributes in the entire knowledge graph and construct an attribute vector with the same dimension as the number of knowledge graph attributes and input it into the convolutional network to generate the attribute embedding vector.
[0004] During the process of identifying the same entities between graphs, entity alignment technology is required. A relatively common method in entity alignment is to use deep learning methods to generate the embedding vectors of each attribute of the entities in the graph, and find the same entities between the two graphs by calculating the similarity of the embedding vectors. However, in the knowledge graph fusion task of specific domains, there are problems such as large structural differences between different graphs, inconsistent text descriptions, inconsistent numbers and arrangement orders of entity attributes, etc., resulting in difficulties in directly calculating the similarity of entities between different graphs. The present invention borrows the idea of using a recurrent neural network to process text, takes the attributes of the entities as context features, and maps the embedding vectors of the entities in different graphs to the same feature space, so that the embedding vectors contain all the attribute information of the entities and have the same dimension. Summary of the Invention
[0005] In view of the deficiencies in the existing knowledge graph fusion, the present invention provides a knowledge graph fusion method based on entity merging. Based on a recurrent neural network, the attributes of an entity are regarded as context, and the entity embedding vectors in two graphs are mapped to the same feature space. Moreover, the embedding vectors can contain all the attribute information of the entity and have the same dimension, solving the problem that it is difficult to calculate similarity due to large structural differences, inconsistent text descriptions, inconsistent numbers and arrangement orders of entity attributes, etc. among different graphs in the knowledge graph fusion of specific fields, with high speed and accuracy.
[0006] The present invention is implemented by at least one of the following technical solutions.
[0007] A knowledge graph fusion method based on entity merging includes the following steps:
[0008] (a) For several knowledge graphs to be merged, obtain their structured entities and attributes, and calculate the word embedding vectors of each attribute;
[0009] (b) Concatenate all the attribute word embedding vectors generated by a single entity into a sentence embedding vector and input it into a recurrent neural network. Use the output of the last hidden layer of the recurrent neural network as the attribute embedding vector of the entity, so as to map the attribute embedding vectors in two graphs to the same feature space. The attribute embedding vector includes all the attribute information of the entity and has the same dimension;
[0010] (c) For the entities between graphs, use the cosine similarity algorithm to calculate the similarity of the attribute embedding vectors, and merge the two entities with a similarity exceeding the set threshold and the highest similarity to obtain a fused knowledge graph.
[0011] Further, for several knowledge graphs to be fused, define the entity set as where represents the T i th entity in the knowledge graph, T i represents the number of entities in the i-th knowledge graph, and I represents the number of knowledge graphs to be fused;
[0012] Define the entity attribute set as where is the N t th attribute value of the t-th entity in the i-th graph, N t is the number of attributes of the entity, and the numbers and orders of entity attributes are different.
[0013] Further, for numerical attributes, extract the numerical value and unit of the numerical attribute through regular expressions. For the numerical value, construct a zero vector v with the same dimension as the word vector output by the used word embedding generation algorithm 0, and add this value to the last dimension of the vector to obtain the word embedding vector of this value; for the unit name, use the word embedding generation algorithm to generate the word embedding vector.
[0014] Further, for the entity e t in step (b), for the m-th attribute a m,t of it, several word embedding vectors of the attribute are generated through step (a) where is the word embedding vector generated by the N m -th word of the m-th attribute; N m is the number of words of the m-th attribute; M is the number of attributes, and they are concatenated into a sentence embedding vector:
[0015]
[0016] Among them, a zero vector v 0 with the same dimension is added between the attributes. Split the several word embedding vectors of different attributes, and input the sentence embedding vector into the recurrent neural network to obtain the output of its last hidden layer as the entity embedding vector ev t .
[0017] Further, the training method of the recurrent neural network is to obtain the entities in multiple graphs in the training dataset through manual annotation, use the entity pairs belonging to the same thing as positive samples, and the entity pairs belonging to different things as negative samples; use the positive and negative samples as the training set to train the recurrent neural network.
[0018] Further, the training set uses the knowledge graph training set, or is generated by crawling data and through entity extraction and relationship extraction methods.
[0019] Further, for the recurrent neural network that needs to be trained, the training loss function is:
[0020]
[0021] p k = cos_sim(f(e i ), f(e j ))
[0022] where, N represents the total number of samples; y k represents the label of the k-th sample in the training dataset, 1 for positive samples and 0 for negative samples; p k represents the similarity of the entity embedding vectors of the k-th sample entity pair (e i , e j ); cos_sim represents the cosine similarity calculation function; f(e i ) represents the entity e iThe entity embedding vector output after passing through the recurrent neural network f, where f(e j ) represents the entity e j The vector output after passing through the recurrent neural network f.
[0023] Furthermore, for two graphs to be fused, select an entity from graph A and calculate the similarity of the entity attribute embedding vectors with all entities in another graph B. When the similarity exceeds the set threshold Q and is the highest among all entities, merge these two entities into one entity in the new graph C; complete the merging of the two graphs by looping through all entities to obtain the new graph C.
[0024] Furthermore, for multiple knowledge graphs, adopt a pairwise merging method. First, randomly select two graphs, perform steps (a)-(c) to merge them into one graph, and then merge the resulting graph with the unmerged graphs, thereby fusing multiple knowledge graphs into one knowledge graph.
[0025] Furthermore, for the numerical attributes of entities, obtain their numerical magnitudes and units through regular expressions. Then, for the numerical values, add them to the zero vector as the word embedding vector; for the units, use the general word embedding generation algorithm to generate the word embedding vector; for the text attributes of entities, use the general word embedding generation algorithm to generate the word embedding vector.
[0026] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0027] This method effectively maps the embedding vectors of entity attributes to the same feature space. The embedding vectors contain all the attribute information of the entities and have the same dimension, enabling the direct calculation of the similarity between entities in different graphs for merging; when processing the attributes of entities, the word embedding generation algorithm is used to generate attribute word vectors with semantic features, making the attribute embedding vectors carry more information and being more accurate when calculating entity similarity. At the same time, this method only needs to construct attribute vectors consistent with the dimension of the number of entity attributes and input them into the recurrent neural network to generate attribute embedding vectors. Since the dimension of the input vector is reduced, the corresponding model structure and computational complexity will be reduced, making this method computationally efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a flowchart of the knowledge graph fusion method based on entity merging in the implementation manner;
[0029] Figure 2 It is a sample structure diagram of the entity embedding vector constructed for the knowledge graph entities;
[0030] Figure 3 It is a flowchart of training the recurrent neural network in the implementation method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following further illustrates the implementation manners of the present invention in conjunction with embodiments, but the implementation of the present invention is not limited thereto.
[0032] Embodiment 1
[0033] A knowledge graph fusion method based on entity merging includes the following steps:
[0034] S1. For several knowledge graphs to be merged, obtain their structured entities and attributes, and calculate the word embedding vectors of each attribute;
[0035] For the numerical attributes of entities, obtain their numerical magnitudes and units through regular expressions. Then, for the numerical values, add them to the zero vector as the word embedding vectors; for the units, use the general word embedding generation algorithm to generate the word embedding vectors;
[0036] For the text attributes of entities, use the general word embedding generation algorithm to generate the word embedding vectors;
[0037] For several knowledge graphs to be fused, define the entity set as where represents the Tth entity in the knowledge graph, T i represents the number of entities in the ith knowledge graph, and I represents the number of knowledge graphs to be fused; i
[0038] Define the entity attribute set as where is the Nth attribute value of the tth entity in the ith graph, N t represents the number of attributes of the entity, and the number of attributes and the order of attributes of different entities are different. t
[0039] For the mth attribute a t of the entity e m,t , generate several word embedding vectors of the attribute where is the word embedding vector generated by the Nth word of the mth attribute; N m represents the number of words of the mth attribute; M represents the number of attributes; m
[0040] Concatenate the word embedding vectors of all attributes into a sentence embedding vector:
[0041]
[0042] where, where is the Nth M The word embedding vectors generated by a word; add a zero vector v of the same dimension between attributes 0 Separate several word embedding vectors of different attributes. By inputting the sentence embedding vector into a recurrent neural network, obtain the output of its last hidden layer as the entity embedding vector ev t .
[0043] S2. Concatenate the generated word embedding vectors into a sentence embedding vector and input it into a recurrent neural network. Use the output of the last hidden layer of the recurrent neural network as the attribute embedding vector of the entity, so as to map the attribute embedding vectors in the two knowledge graphs to the same feature space. The attribute embedding vector includes all attribute information of the entity and the same dimension;
[0044] The training method of the recurrent neural network is to obtain entities in multiple knowledge graphs in the training dataset through manual annotation. Entities pairs that belong to the same thing in the real world are used as positive samples, and entities pairs that belong to different things in the real world are used as negative samples. For example, tomatoes and tomatoes belong to the same thing in the real world and have the same attributes, such as being red, herbaceous plants, belonging to the Solanaceae family and Solanum genus in biology; apples can represent fruits or mobile phones, belonging to different things in the real world, one with the attribute of a plant and the other with the attribute of a mobile phone. Use the positive and negative samples as the training set to train the recurrent neural network. The loss function for training is:
[0045]
[0046] p k = cos_sim(f(e i ), f(e j ))
[0047] where N represents the total number of samples; y k represents the label of the k-th sample in the training dataset, 1 for positive samples and 0 for negative samples; p k represents the similarity of the entity embedding vectors of the k-th sample entity pair (e i , e j ); cos_sim represents the cosine similarity calculation function; f(e i ) represents the vector output after the entity e i passes through the recurrent neural network f, and f(e j ) represents the vector output after the entity e j passes through the recurrent neural network f.
[0048] S3. For entities between the two knowledge graphs, use the cosine similarity algorithm to calculate the similarity of the attribute embedding vectors, and merge the two entities with the highest similarity that exceed the set threshold to obtain a fused knowledge graph.
[0049] Example 2
[0050] As Figure 1 shown, the knowledge graph fusion method based on entity merging includes the following steps:
[0051] First step, obtain all entities and attributes of the two graphs to be merged.
[0052] Put the entities of the two graphs into the sets and where represents the T i th entity in the knowledge graph, T i represents the number of entities in the i-th knowledge graph, and each entity contains N t attributes is the nN t th attribute value of the t-th entity in the i-th graph, and N t is the number of attributes of the entity. The number of entity attributes and the attribute order are different. For example, in one knowledge graph, there is an entity "tomato" with attributes (red, Solanaceae, Solanum, herbaceous plant), and in another knowledge graph, there is an entity "tomato" with attributes (Solanum, Solanaceae, herbaceous plant). At the same time, without loss of generality, this embodiment assumes that two knowledge graphs need to be merged.
[0053] Second step, convert numerical attributes into word embedding vectors.
[0054] According to the actual application scenario, count all the units of the attributes in the graph, such as length unit (meter), weight unit (kilogram), area unit (square meter), etc., and define a standard unit for each unit type and convert all non-standard units to the standard unit. Extract the numerical value and unit of the numerical attribute through regular expressions.
[0055] Taking the length unit as an example, meter (m) can be set as the standard unit, and the rest of the non-standard length units are converted to meters. Generally speaking, the regular expression can be expressed as:
[0056] ”\d+\.\d*[mM]”
[0057] where, "\d" represents any one of the digits 0-9, "+" represents matching the previous "\d" once or more, "\." represents the decimal point symbol, and "*" represents matching the previous "\d" zero times or more. "[mM]"
[0058] represents matching one of the letters m or M.
[0059] This expression is used to match a string with a floating-point value and the last character being'm' (or 'M'). By splitting the'm' (or 'M') and the floating-point value, the numerical value and the unit name of this attribute can be obtained.
[0060] For the numerical value, construct a zero vector with a dimension of 768, add the numerical value to the last element of the vector, and after normalization, the word embedding vector of the numerical value can be obtained.
[0061] For the unit name, use the Bert word embedding generation algorithm to generate a word embedding vector with a dimension of 768. Finally, concatenate the numerical word embedding vector and the unit word embedding vector to obtain the embedding vector of this numerical attribute. The generation method of the above numerical attribute can ensure that there are significant differences between the numerical attribute and the text attribute, enabling the text attribute and the numerical attribute to share the same recurrent neural network model without calculating the text attribute of one entity and the numerical attribute of another entity as the same attribute in subsequent processing.
[0062] Step 3: Convert the text attribute into a word embedding vector.
[0063] For the text attribute, directly input the attribute string into the Bert model to generate several word embedding vectors where n m is the m-th attribute a m,t The number of generated word embedding vectors. Among them is the n-th m word-generated word embedding vector of the m-th attribute.
[0064] Step 4: Concatenate the word embedding vectors of all attributes of the entity.
[0065] Generate several word embedding vectors of the attribute through the second and third steps Concatenate the word embedding vectors of all attributes into a sentence embedding vector:
[0066]
[0067] Among them, a zero vector v with the same dimension needs to be added between attributes 0 Separate several word embedding vectors of different attributes. The specific operation is as Figure 2 shown.
[0068] Step 5: Input into the recurrent neural network to obtain the attribute embedding vector.
[0069] By inputting the sentence embedding vector into the recurrent neural network as shown in Figure 3 obtain the output of the last hidden layer as the entity embedding vector ev tFor a recurrent neural network, a unidirectional long short-term memory artificial neural network (LSTM) can be used, and it needs to be trained before use. The training method is as follows: Obtain the same entity pairs (e i , e j ) in multiple graphs in the training dataset through manual annotation as positive samples. The training set can use a publicly available knowledge graph training set, or crawl data from websites such as Baidu Encyclopedia and generate a knowledge graph through entity extraction and relationship extraction methods. By selecting other entities e' j in the same graph from the positive samples to replace e j , negative samples (e i , e' j ) are constructed. Use the positive and negative samples as the training set to train the recurrent neural network. The loss function for training is:
[0070]
[0071] p k = cos_sim(f(e i ), f(e j ))
[0072] where N represents the total number of samples; y k represents the label of the k-th sample in the training set, 1 for positive samples and 0 for negative samples; p k represents the similarity of the entity embedding vectors of the k-th sample entity pair (e i , e j ); cos_sim represents the cosine similarity calculation function; f(e i ) represents the entity embedding vector output after the entity e i passes through the recurrent neural network f, and f(e j ) represents the vector output after the entity e j passes through the recurrent neural network f.
[0073] Step 6: Calculate the entity similarity between graphs and merge entities.
[0074] For two graphs to be fused, select one entity from one graph and all entities from the other graph, and calculate the similarity of the two entity attribute embedding vectors through the second to fifth steps. When the similarity exceeds the set threshold Q = 0.9 and is the highest similarity among all entities, the two entities from the two graphs can be merged into one entity. Complete the merging of the two graphs by looping through all entities.
[0075] For multiple knowledge graphs, a pairwise merging method can be adopted. First, randomly select two graphs, merge them into one graph through the first to fifth steps, and then merge the resulting graph with the unmerged graphs, thus fusing multiple knowledge graphs into one knowledge graph.
[0076] Example 3
[0077] A knowledge graph fusion method based on entity merging includes the following steps:
[0078] First step: Obtain all entities and attributes of the two graphs to be merged. For example, the Chinese and English versions of the introduction to Chinese cities can be obtained from Baidu Encyclopedia through web crawlers, and the Chinese and English knowledge graphs of Chinese cities can be formed through entity extraction and relationship extraction algorithms. For example, for the introduction of Guangdong Province, the entity Guangzhou can be extracted, and the attributes include an area of 74,344 square kilometers, provincial capital city, etc.
[0079] Put the entities of the two graphs into a set and Each entity contains N t attributes The number and order of attributes of different entities are different.
[0080] Second step: Convert numerical attributes into word embedding vectors.
[0081] According to the actual application scenario, count all the units of attributes in the graph, such as length unit (meter), weight unit (kilogram), area unit (square meter), etc., and define a standard unit for each unit type and convert all non-standard units into standard units. Extract the numerical value and unit of the numerical attribute through regular expressions.
[0082] For the numerical value, construct a zero vector with a dimension of 768, add the numerical value to the last digit of the vector and normalize it to obtain the word embedding vector of the numerical value.
[0083] For the unit name, use the Bert word embedding generation algorithm to generate a word embedding vector with a dimension of 768. Finally, concatenate the numerical word embedding vector and the unit word embedding vector to obtain the embedding vector of the numerical attribute.
[0084] Third step: Convert text attributes into word embedding vectors.
[0085] For text attributes, directly input the attribute string into the Bert model to generate several word embedding vectors where n m is the m-th attribute a m,t The number of generated word embedding vectors. Among them is the n-th m word of the m-th attribute generated word embedding vector.
[0086] Fourth step: Concatenate the word embedding vectors of all attributes of the entity.
[0087] Generate several word embedding vectors of attributes through the second and third steps Concatenate the word embedding vectors of all attributes into a sentence embedding vector:
[0088]
[0089] Among them, a zero vector v with the same dimension needs to be added between attributes 0 Separate several word embedding vectors of different attributes. The specific operation is as Figure 2 shown
[0090] Step 5: Construct a dataset
[0091] The dataset can use the publicly available knowledge graph training set, or crawl data from websites such as Baidu Encyclopedia and generate a knowledge graph through entity extraction and relationship extraction methods. Obtain the same entity pair (e i , e j ) between the two graphs in the training dataset through manual annotation as positive samples, and label the sample label as 1. Among them, e i is an entity in the first graph E 1 , and e i is an entity in the second graph E 2 , and they belong to the same thing in the real world. For example, Guangzhou and GuangZhou are two entities in the Chinese knowledge graph and the English knowledge graph, but they belong to the same thing in the real world. By selecting other entities e' 2 in the first graph E j from the positive samples to replace e j , construct negative samples (e i , e' j ), and its sample label is 0. For example, Guangzhou and ShangHai are two entities in the Chinese knowledge graph and the English knowledge graph, and they belong to different things in the real world
[0092] Step 6: Train a recurrent neural network model
[0093] For the dataset constructed in the fifth step, divide it into a training set and a validation set according to a ratio of 9:1. As Figure 3 shown, the training set is used for model training. In each stage of training, obtain an entity pair and its sample label, input them into the recurrent neural network, and calculate the loss at this stage. The loss function for training is as follows:
[0094]
[0095] p k = cos_sim(f(e i ), f(ej ))
[0096] Among them, N represents the total number of samples, that is, the total number of positive and negative samples in the training set constructed in the fifth and sixth steps; y k represents the label of the k-th sample in the training set, where the positive sample is 1 and the negative sample is 0; p k represents the similarity of the entity embedding vectors of the k-th sample entity pair (e i , e j ); cos_sim represents the cosine similarity calculation function; f(e i ) represents the entity embedding vector output after the entity e i passes through the recurrent neural network f, and f(e j ) represents the vector output after the entity e j passes through the recurrent neural network f.
[0097] At the beginning of training, initialize a global model and set the global accuracy to 0. Traverse all training set samples in each round of training and update the model parameters. In each iteration of a round of training, update the parameters of the model through the gradient descent algorithm, and the learning rate R of the gradient descent is 0.01. After completing a round of training, obtain a temporary model and calculate its accuracy on the validation set. If the accuracy of the temporary model in this round is higher than the global accuracy, use the temporary model as the global model and update the global accuracy to the accuracy of this round. Training ends after 20 rounds.
[0098] Step 7: Input the recurrent neural network to obtain the attribute embedding vector.
[0099] By inputting the sentence embedding vector into the recurrent neural network as shown in Figure 3 , obtain the output of its last hidden layer as the entity embedding vector ev t .
[0100] Step 8: Calculate the entity similarity between graphs and merge entities.
[0101] For two graphs to be fused, select one entity from graph E 1 and all entities of graph E 2 , such as selecting the entity Guangzhou in the Chinese knowledge graph and all entities GuangZhou, ShangHai, Beijing, etc. in the English knowledge graph. Calculate the similarity of the two entity attribute embedding vectors through the second to fourth steps. When the similarity exceeds the set threshold Q = 0.9 and is the highest similarity among all entities, the two entities of the two graphs can be merged into one entity. Complete the merging of the two graphs by looping through all entities of graph E 1 .
[0102] For multiple knowledge graphs, a pairwise merging method can be adopted. First, randomly select two graphs, merge them from the first step to the fifth step into one graph, and then merge the resulting graph with the unmerged graphs, so as to fuse multiple knowledge graphs into one knowledge graph.
[0103] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A knowledge graph fusion method based on entity merging, characterized in that, it includes the following steps: (a) For several knowledge graphs to be merged, obtain their structured entities and attributes, and calculate the word embedding vectors of each attribute; (b) Concatenate all the attribute word embedding vectors generated by a single entity into a sentence embedding vector and input it into a recurrent neural network. Use the output of the last hidden layer of the recurrent neural network as the attribute embedding vector of the entity, so as to map the attribute embedding vectors in the two graphs to the same feature space. The attribute embedding vector includes all the attribute information of the entity and has the same dimension; For entity e in step (b) t the m-th attribute a m,t , a number of word embedding vectors of the attribute are generated through step (a) where is the word embedding vector generated by the N m -th word of the m-th attribute; N m is the number of words of the m-th attribute; M is the number of attributes, and they are concatenated into a sentence embedding vector: Among them, a zero vector v of the same dimension is added between the attributes 0 Separate the embedding vectors of several words with different attributes. By inputting the sentence embedding vector into a recurrent neural network, obtain the output of its last hidden layer as the entity embedding vector ev t ; for the numerical attributes of the entity, obtain its numerical magnitude and unit through regular expressions. Then, for the numerical value, add it to the zero vector as the word embedding vector; for the unit, use a general word embedding generation algorithm to generate the word embedding vector; for the text attributes of the entity, use a general word embedding generation algorithm to generate the word embedding vector; (c) For the entities between graphs, use the cosine similarity algorithm to calculate the similarity of the attribute embedding vectors, and merge the two entities with the similarity exceeding the set threshold and the highest similarity to obtain a fused knowledge graph.
2. The knowledge graph fusion method based on entity merging according to claim 1, characterized in that: For several knowledge graphs to be integrated, define the entity set as 1 ≤ i ≤ I, where represents the T i -th entity in the knowledge graph, T i represents the number of entities in the i-th knowledge graph, and I represents the number of knowledge graphs to be integrated; Define the set of entity attributes as where is the N t th attribute value of the t t th entity in the i th atlas, N t is the number of attributes of the entity, and the number of attributes and the order of attributes of different entities are different.
3. The knowledge graph fusion method based on entity merging according to claim 1, characterized in that: For numerical attributes, the numerical value and unit of the numerical attribute are extracted through regular expressions. For a numerical value, a zero vector v is constructed with the same dimension as the word vector output by the word embedding generation algorithm used, 0 and the numerical value is added to the last dimension of the vector to obtain the word embedding vector of the numerical value; For the unit name, use the word embedding generation algorithm to generate word embedding vectors.
4. The knowledge graph fusion method based on entity merging according to claim 1, characterized in that: The training method of the recurrent neural network is to obtain the entities in multiple graphs in the training dataset through manual annotation. Take the entity pairs belonging to the same thing as positive samples and the entity pairs belonging to different things as negative samples; use the positive and negative samples as the training set to train the recurrent neural network.
5. The knowledge graph fusion method based on entity merging according to claim 4, characterized in that: The training set uses the knowledge graph training set, or is generated by crawling data and through entity extraction and relationship extraction methods.
6. The knowledge graph fusion method based on entity merging according to claim 1, characterized in that: For the recurrent neural network that needs to be trained, the training loss function is: p k = cos_sim(f(e i ), f(e j )) Among them, N represents the total number of samples; y k represents the label of the k-th sample in the training set, where the positive sample is 1 and the negative sample is 0; p k represents the similarity of the entity embedding vectors of the k-th sample entity pair (e i , e j ); cos_sim represents the cosine similarity calculation function; f(e i ) represents the entity embedding vector output after the entity e i passes through the recurrent neural network f, and f(e j ) represents the vector output after the entity e j passes through the recurrent neural network f.
7. The knowledge graph fusion method based on entity merging according to claim 1, characterized in that: For the two graphs to be fused, select an entity of graph A and calculate the similarity of the entity attribute embedding vectors with all the entities of another graph B. When the similarity exceeds the set threshold Q and is the highest similarity among all entities, merge these two entities into an entity of the new graph C; complete the merging of the two graphs by looping through all entities to obtain the new graph C.
8. The knowledge graph fusion method based on entity merging according to claim 1, characterized in that: For multiple knowledge graphs, adopt a pairwise merging method. First, randomly select two graphs, perform steps (a) - (c) to merge them into one graph, and then merge the resulting graph with the unmerged graphs, so as to fuse multiple knowledge graphs into one knowledge graph.
Citation Information
Patent Citations
A Cross-Language Knowledge Graph Entity Alignment Method Based on GCN Siamese Network
CN110472065B
Cross-language knowledge graph alignment and fusion method and device and storage medium
CN113111657A