A cross-language knowledge graph alignment method based on graph representation learning
Through a graph representation learning method, multilingual knowledge graphs are embedded in a unified vector space and aligned, which solves the problem that the existing technology cannot fully utilize graph structure information and fused cross-language data, and achieves more accurate and efficient knowledge graph fusion.
Patent Information
- Application Number
- CN202210020693.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-01-10
AI Technical Summary
The existing knowledge graph cross-language alignment technology cannot fully utilize graph structure information and cannot accurately and efficiently integrate rich cross-language data, resulting in poor comprehensive integration of knowledge graphs at the semantic level.
Using a graph representation learning method, triples are constructed by crawling multilingual data, filtering and extracting structured data in the knowledge graph construction stage, and using graph embedding to embed knowledge graphs from different sources into a unified vector space and aligned according to the distance of entities.
By making full use of the structural information of the knowledge graph, the entities of different languages are merged into a unified space, and more accurate and comprehensive fact fusion is achieved, improving the efficiency and accuracy of knowledge graph fusion across language fields.
Smart Images

Figure CN114443855B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a knowledge graph cross-language alignment method, and in particular to a knowledge graph cross-language alignment method based on graph representation learning, which belongs to the technical field of natural language processing. Background Art
[0002] Knowledge graph, as a knowledge base that represents concepts, entities and the relationships between entities in the objective world in the form of a graph, is essentially a large-scale semantic network that can organize massive amounts of data into an interconnected network graph. Since the rise of mobile Internet, information has exploded, and large-scale knowledge graphs have emerged in an endless stream, resulting in problems such as knowledge duplication and unclear relationships between knowledge graphs, which has affected the comprehensive integration of knowledge graphs at the semantic level. Typical multilingual knowledge graphs include: DBpedia, YAGO, and Freebase. Each knowledge graph contains a large amount of knowledge description, but due to differences in data sources and different data languages, it is actually difficult to construct a knowledge graph that contains comprehensive facts.
[0003] Entity alignment is also described as entity matching or entity resolution in fields such as machine translation, question-answering systems, and information retrieval. The goal of the entity alignment task is to identify whether the objects referred to by different knowledge graphs are entity pairs of the same thing in the real world. The entity alignment technology of knowledge graphs can connect knowledge, merge similar knowledge graphs into larger-scale, more authoritative domain knowledge graphs, and provide knowledge guarantee for downstream applications.
[0004] The cross-language alignment task of knowledge graphs usually requires complex calculations. Traditional cross-language entity alignment methods usually use methods based on manually defined features, which not only consumes a lot of manpower but is also difficult to migrate to actual application scenarios. Most of the cross-language alignment methods of knowledge graphs that have emerged in recent years focus on encoding triple information, but do not fully utilize the structural information of knowledge graphs. In addition, labeled data for cross-language entity alignment is difficult to obtain. Therefore, how to pre-train on a large amount of unlabeled text and maximize the value of a small amount of labeled data is of great significance to the development and integration of large-scale knowledge graphs.
[0005] In terms of cross-language alignment of knowledge graphs, many current methods focus on text data, calculate the similarity between texts, or embed knowledge graphs based on the idea of translation models. These methods do not fully utilize the structural information of knowledge graphs and cannot achieve good results in cross-language alignment of knowledge graphs. Summary of the invention
[0006] The purpose of the present invention is to creatively propose a knowledge graph cross-language alignment method based on graph representation learning in response to the technical problems that the current knowledge graph cross-language data information sources are numerous and the content is complex, while the existing knowledge graph cross-language alignment technology cannot fully utilize the graph structure information and cannot accurately and efficiently integrate sufficiently rich cross-language data.
[0007] The innovation of the present invention is that in the knowledge graph construction stage, website data is crawled as a source. Then, multilingual entities are filtered and screened and their structured data are extracted to form triples to construct a knowledge graph. In the alignment stage, knowledge graphs from different sources are generated into corresponding embedding matrices through graph representation learning. On the basis of graph embedding, entities in knowledge graphs of different languages are merged into a unified space based on aligned entities, and aligned according to the distance between entities in the joint semantic space.
[0008] The present invention is achieved through the following technical solutions.
[0009] A knowledge graph cross-language alignment method based on graph representation learning includes the following steps:
[0010] Step 1: Get multilingual data.
[0011] Among them, obtaining multilingual data includes data from various encyclopedia websites;
[0012] Specifically, step 1 includes the following steps:
[0013] Step 1.1: Crawl encyclopedia multilingual website data and save it locally in HTML format;
[0014] Step 1.2: Classify the data crawled in step 1.1 and remove dirty data (Dirty Read refers to data in the source system that is not within the given range or is meaningless to the actual business, or the data format is illegal, and there is non-standard coding and ambiguous business logic in the source system).
[0015] The reason for classifying the data is that the crawled data usually contains some non-entity data, which will affect the subsequent construction of the knowledge graph.
[0016] Specifically, the following methods can be used to classify data:
[0017] Step 1: Traverse the locally stored data obtained in step 1.1 and obtain a list of entity names containing all data.
[0018] Step 2: Based on the list of data entity names obtained in the first step, randomly select M data, manually label these M data, and divide them into training set and validation set.
[0019] Step 3: Use the Bert model to pre-train and fine-tune the training set obtained in the second step, and perform cross-validation on the validation set. When the accuracy reaches more than 90%, all M data obtained in the second step are input into the Bert model for training to obtain a complete pre-trained model.
[0020] Step 4: Use the pre-trained model obtained in step 3 to classify the list of all data entity names obtained in step 1, remove the dirty data in the crawling results, and obtain the final list containing data entity names.
[0021] Step 5: Based on the final list containing data entity names, filter and save the local HTML data obtained in step 1.1.
[0022] Step 2: Parse the multilingual data in HTML format obtained in step 1 and process it into JSON format data of triple type.
[0023] Since the original HTML data has a large difference in form, if it is not converted into a unified format, it will be difficult to store and will not be suitable for the subsequent construction of the knowledge graph.
[0024] Specifically, step 2 includes the following steps:
[0025] First, use the bs4 library to traverse the multilingual data in HTML format obtained in step 1 and find the table information;
[0026] Then, based on the above table information, extract the text content and establish entity-relationship-entity triples according to the data entity names;
[0027] Finally, the triples obtained above are stored as JSON format data files, saved locally, and a part of the triples are marked to obtain seed alignment entities.
[0028] Step 3: Create a multilingual knowledge graph based on the JSON format data obtained in step 2.
[0029] Specifically, step 3 includes the following steps:
[0030] Step 3.1: Create an index for the crawled data from different sources;
[0031] Step 3.2: Based on the index established in step 3.1, construct knowledge graphs for data from different sources;
[0032] Specifically, the following methods can be used to build a knowledge graph:
[0033] Step 1: Based on the JSON format data file obtained in step 2, traverse the triples of each language data to obtain its head node, relationship, and tail node.
[0034] Step 2: Based on the head node, relationship and tail node obtained in the first step, create fields for the data entity names to obtain all attribute information of each language data.
[0035] Step 3: Find data from different sources based on the index established in step 3.1. For data from the same source, use the py2neo library to mark them. Import the data obtained in the second step into the relational database Neo4j, and establish knowledge graphs based on different data sources and languages.
[0036] Step 4: Embed the multilingual knowledge graphs from different sources obtained in step 3 into a unified vector space.
[0037] The reason for embedding into a unified vector space is that entities, relationships and other components in the knowledge graph are converted into a continuous vector space and represented as dense low-dimensional vectors. Compared with simple one-hot encoding, the graph representation learning dimension is lower and is not easily affected by sparse data. It can improve computing efficiency and better express the semantic information between knowledge graph objects. The closer the distance between two objects in the space, the greater their similarity.
[0038] Specifically, step 4 includes the following steps:
[0039] Step 4.1: Relation embedding;
[0040] Among them, for each knowledge graph from different sources obtained in step 3, relationship embedding is performed respectively;
[0041] Specifically, the steps of relation embedding are as follows:
[0042] Step 1: Based on the knowledge graphs from different sources obtained in step 3, establish the adjacency matrix A of the knowledge graph according to its entity-relationship-entity structure.
[0043] Step 2: Add self-loop I to the adjacency matrix obtained in the first step, where I is the unit matrix, and get the matrix
[0044] Step 3: Calculate the matrix obtained in step 2 The diagonal matrix
[0045] Step 4: Randomly initialize the network's weight matrix W.
[0046] Step 5: Calculate the matrix obtained in step 2 The characteristic matrix H (i).
[0047] Step 6: Based on formula (1), the feature matrix H of the current layer obtained in step 5 is (i) , calculate the output H of this layer (i +1) , H (i+1) That is the relational embedding expression of the knowledge graph.
[0048]
[0049] Among them, σ represents the activation function.
[0050] Step 4.2: Embedding space transformation;
[0051] The purpose of embedding space transformation is to embed knowledge graphs from different sources into a unified vector space to improve the evaluation of entity similarity in graph representation learning;
[0052] Specifically, the steps of embedding space transformation are as follows:
[0053] Step 1: Randomly initialize the network's weight matrix M.
[0054] Step 2: Input the seed alignment entities obtained in step 2 and the relational embedding expressions of the knowledge graphs from various sources obtained in step 4.1 into the fully connected layer to train the matrix M.
[0055] Step 3: Based on the matrix M obtained in the second step, encode the knowledge graphs from different sources into a unified embedding space.
[0056] Step 5: Calculate the distance between entities in vector space and align them.
[0057] Specifically, step 5 includes the following steps:
[0058] Step 1: Based on the multilingual knowledge graph obtained in step 3, traverse the entities in the knowledge graph of one of the data sources.
[0059] Step 2: Map each of the above entities according to the vector space obtained in step 4 to obtain the vector expression of each entity.
[0060] Step 3: Traverse the vector expressions of entities in the knowledge graphs of all other data sources, calculate the cosine similarity with the vector expressions of each entity obtained in the second step, and store the calculation results in the result table.
[0061] Step 4: Sort the above result table in descending order. The entity with the highest score is the aligned entity of each entity in the knowledge graph selected in the first step.
[0062] Step 5: Add the aligned entities obtained in step 4 to the knowledge graph selected in step 1 to obtain the final knowledge graph cross-language alignment result.
[0063] Beneficial Effects
[0064] Compared with the prior art, the method of the present invention has the following advantages:
[0065] 1. This method makes full use of the structural information of the knowledge graph, merges the entities in the knowledge graphs of different languages into a unified space through a graph representation learning method, and aligns the entities according to their distance in the joint semantic space, ensuring that the fused data is more accurate and comprehensive.
[0066] 2. This method provides a means to extract structured knowledge from massive text data, and further integrates and analyzes multilingual data, standardizes the unified description and organizational association of entity data in each language, displays the structured relationship between data, and improves the efficiency of rapid analysis and intelligent search in cross-language fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is the overall process of the method of the present invention;
[0068] Figure 2 is a data acquisition flow chart of the method of the present invention;
[0069] Figure 3 is a flow chart of data processing and establishing a multilingual knowledge graph according to the method of the present invention;
[0070] Figure 4 It is the detailed architecture of the graph representation learning model on which the method of the present invention relies.
[0071] Figure 5 It is the system architecture of the method of the present invention. DETAILED DESCRIPTION
[0072] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0073] Example
[0074] This embodiment describes a specific embodiment of the method of the present invention.
[0075] Implementation diagram Figure 1 The overall process is shown in Figure 4It is a detailed architecture of a graph representation learning model based on a graph representation learning-based knowledge graph cross-language alignment method of the present invention. In the specific implementation of the present invention, the data set obtained in step 1 is multilingual data collected from various encyclopedia websites, which is cleaned and saved in a Neo4j graph database.
[0076] Using the method proposed in the present invention, a knowledge graph is constructed for multilingual data in a graph database, and the constructed knowledge graph is embedded into a vector space through graph representation learning. The multi-source knowledge graph is then processed into a unified vector space through pre-labeled seed alignment entities, and entity similarity is calculated and aligned in this space, which is then saved in the graph database and can be viewed by users through the database's built-in display interface.
[0077] Figure 2 This is the data acquisition process of the knowledge graph cross-language alignment method based on graph representation learning in the present invention.
[0078] According to step 1 introduced in the present invention, data is crawled from various encyclopedia websites, all the crawled HTML data is stored locally, the data is classified and cleaned, and dirty data is removed.
[0079] Figure 3 This is the data processing flow of a knowledge graph cross-language alignment method based on graph representation learning in the present invention.
[0080] According to step 2 introduced in the present invention, all HTML files in the local folder are read, the HTML data is parsed, the index is updated to Table 1, the relationship triples therein are extracted, converted into JSON format, and updated to Table 2.
[0081] In order to use the graph representation learning method for entity alignment, it is necessary to first build a knowledge graph. According to step 3 introduced in the present invention, multi-source json format data is imported into the graph database neo4j, and attributes are marked for each source of data in the graph database, and different knowledge graphs are built according to different sources, and the relevant information is synchronized to neo4j and input into the graph convolutional neural network model used for graph representation learning.
[0082] Table 1 Index table
[0083]
[0084]
[0085] Table 2 json data table
[0086]
[0087] Figure 4It is a detailed architecture of a graph representation learning model based on which a knowledge graph cross-language alignment method based on graph representation learning is based on the present invention.
[0088] In order to better utilize the graph structure information of the knowledge graph, according to step 4.1 introduced in the present invention, when performing knowledge representation learning, it is necessary to first extract the attribute information in the data, add the extracted entity-attribute-attribute value triples to the vector space matrix, and input the vector space matrices from different knowledge graphs into the graph convolutional neural network respectively to obtain embedded data from different vector space matrices. According to step 4.2 introduced in the present invention, the pre-aligned seed alignment entities obtained in step 2 introduced in the present invention are used to embed knowledge graphs from different sources into a unified vector space to improve the accuracy of entity alignment after graph representation learning.
[0089] Figure 5 This is the system architecture of the knowledge graph cross-language alignment method based on graph representation learning described in the present invention.
[0090] First, data is acquired according to step 1 introduced in the present invention, and data is preprocessed according to step 2 introduced in the present invention. Then, a multilingual knowledge graph is constructed according to step 3 introduced in the present invention and added to the neo4j graph database.
[0091] Then, all multilingual knowledge graphs in the graph database are read, and according to step 4 introduced in the present invention, the knowledge graphs of different languages are embedded into different vector spaces, and the seed alignment entities obtained in step 2 are used to unify the vector space.
[0092] Finally, according to step 5 introduced in the present invention, the similarity between entities is calculated in a unified vector space, and the knowledge graphs of different languages are automatically fused. At the same time, due to the effectiveness of entity alignment based on graph representation learning, it can ensure that the fused cross-language knowledge graph is more accurate and rich in information.
[0093] The above is only a preferred embodiment of the present invention, and the present invention should not be limited to the contents disclosed in the embodiment and the drawings. Any equivalent or modification completed without departing from the spirit disclosed in the present invention shall fall within the scope of protection of the present invention.
Claims
1. A knowledge graph cross-language alignment method based on graph representation learning, characterized in that: The following steps are involved: Step 1: Obtain multilingual data, including the following steps: First, crawl the data of multilingual websites of encyclopedias and save them locally in HTML format; Then, the crawled data is classified and dirty data is removed; Step 2: Parse the multilingual data in HTML format obtained in step 1 and process it into JSON format data of triple type; Step 3: Create a multilingual knowledge graph based on the JSON format data obtained in step 2, including the following steps: Step 3.1: Create an index for the crawled data from different sources; Step 3.2: Based on the index established in step 3.1, construct knowledge graphs for data from different sources; Step 1: Based on the JSON format data file obtained in step 2, traverse the triples of each language data to obtain its head node, relationship, and tail node; Step 2: Based on the head node, relationship and tail node obtained in the first step, create fields for the data entity names to obtain all attribute information of each language data; Step 3: Find data from different sources based on the index created in step 3.
1. For data from the same source, use the py2neo library to mark them. Import the data obtained in the second step into the relational database Neo4j, and create knowledge graphs based on different data sources and languages. Step 4: Embed the multilingual knowledge graphs from different sources obtained in step 3 into a unified vector space, including the following steps: Step 4.1: Relation embedding, where for each knowledge graph from different sources obtained in step 3, relationship embedding is performed separately; Step 4.2: Transform the embedding space as follows: Step 1: Randomly initialize the network's weight matrix M; Step 2: Input the seed alignment entities obtained in step 2 and the relational embedding expressions of the knowledge graphs from various sources obtained in step 4.1 into the fully connected layer to train the matrix M; Step 3: According to the matrix M obtained in the second step, the knowledge graphs from different sources are encoded into a unified embedding space; Step 5: Calculate the distance between entities in vector space and align them.
2. The method for cross-language alignment of knowledge graphs based on graph representation learning according to claim 1, characterized in that: In step 1, the data is classified using the following method: Step 1: Traverse the data stored locally and get a list of entity names containing all data; Step 2: According to the list of data entity names obtained in the first step, randomly select M data, manually annotate these M data, and divide them into training set and validation set; Step 3: Use the Bert model to pre-train and fine-tune the training set obtained in the second step, and perform cross-validation on the validation set. When the accuracy reaches more than 90%, all the M data obtained in the second step are input into the Bert model for training to obtain a complete pre-trained model. Step 4: Use the pre-trained model obtained in step 3 to classify the list of all data entity names obtained in step 1, remove the dirty data in the crawled results, and obtain the final list of data entity names; Step 5: Based on the final list containing data entity names, filter and save the locally stored HTML data.
3. The knowledge graph cross-language alignment method based on graph representation learning according to claim 1, characterized in that: Step 2 includes the following steps: First, traverse the multilingual data in HTML format obtained in step 1 to find the table information; Then, based on the above table information, extract the text content and establish entity-relationship-entity triples according to the data entity names; Finally, the triples obtained above are stored as JSON format data files, saved locally, and a part of the triples are marked to obtain seed alignment entities.
4. The knowledge graph cross-language alignment method based on graph representation learning according to claim 1, characterized in that: In step 4.1, the steps of relation embedding are as follows: Step 1: Based on the knowledge graphs from different sources obtained in step 3, establish the adjacency matrix A of the knowledge graph according to its entity-relationship-entity structure; Step 2: Add self-loop I to the adjacency matrix obtained in the first step, where I is the unit matrix, and get the matrix Step 3: Calculate the matrix obtained in step 2 The angle matrix Step 4: Randomly initialize the network's weight matrix W; Step 5: Calculate the matrix obtained in step 2 The characteristic matrix H (i) ; Step 6: Based on formula (1), the feature matrix H of the current layer obtained in step 5 is (i) , calculate the output H of this layer (i+1) , H (i+1) That is, the relational embedding expression of the knowledge graph; Among them, σ represents the activation function.
5. The knowledge graph cross-language alignment method based on graph representation learning according to claim 1, characterized in that: Step 5 includes the following steps: Step 1: Based on the multilingual knowledge graph obtained in step 3, traverse the entities in the knowledge graph of one of the data sources; Step 2: Map each of the above entities according to the vector space obtained in step 4 to obtain the vector expression of each entity; Step 3: Traverse the vector expressions of entities in the knowledge graphs of all other data sources, calculate the cosine similarity between the vector expressions of each entity obtained in the second step and the vector expressions of each entity, and store the calculation results in the result table; Step 4: Sort the above result table in descending order, and the one with the highest score is the aligned entity of each entity in the knowledge graph selected in the first step; Step 5: Add the aligned entities obtained in step 4 to the knowledge graph selected in step 1 to obtain the final knowledge graph cross-language alignment result.
Citation Information
Patent Citations
Cross-language entity alignment method based on subgraph embedding
CN111931505A
Method, apparatus and system for monitoring internet media events based on industry knowledge mapping database
WO2018036239A1