Knowledge graph alignment method and device, equipment, storage medium and program product
By converting entities, relationships and attributes in the knowledge graph into text sequences, and using preset language models to generate multi-dimensional entity encodings, calculating the cosine similarity of entity encodings to achieve alignment, the problem of sharp increase in computing resources and storage space requirements during large-scale knowledge graph alignment in the prior art is solved, and efficient and accurate alignment effects are achieved.
Patent Information
- Application Number
- CN202411955094.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
Existing embedded-based knowledge graph alignment methods have increased dramatically when processing large-scale knowledge graphs, making it difficult to complete the alignment tasks efficiently and accurately.
The cosine similarity of entity encodings is calculated for alignment by editing entity names, relational triples and attribute triples in the knowledge graph into text sequences and generating multi-dimensional entity encodings using preset language models.
While fully representing entity features, this method reduces the impact of graph scale on computing resources and storage space, improves the accuracy and efficiency of entity encoding, and can effectively process large-scale knowledge graphs.
Smart Images

Figure CN119990286A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a knowledge graph alignment method, device, equipment, storage medium and program product. Background Art
[0002] Currently, embedding-based knowledge graph alignment has become a hot research area that has attracted much attention. Its core operating mechanism is to first map the entities in the knowledge graph into a low-dimensional vector space through a specific mapping (embedding) method, and then accurately explore the alignment relationship between entities based on the distance measurement between the entity embedding vectors. This process can be likened to placing entities in a specific "vector space context", in which entities with relatively close distances are most likely to correspond to each other and need to be aligned urgently, so as to achieve the association integration and collaborative matching between knowledge graphs.
[0003] However, the knowledge graphs in the real world are usually very large in scale, covering a very large number of entity elements. Most existing embedding-based methods usually treat the embedding representation of entities as learnable parameters in their technical implementation. This approach directly leads to a significant problem, that is, as the number of entities continues to increase, the computational resource overhead required to learn and process these parameters and the amount of storage space occupied by storing these parameters will show a sharp increase.
[0004] In view of this, most of these methods can currently only be effectively applied in small and medium-sized knowledge graph scenarios. When faced with the alignment task of large-scale knowledge graphs, they often appear to be stretched to their limits and find it difficult to properly handle such a vast amount of data and efficiently achieve the goal of accurate alignment. Summary of the invention
[0005] The present invention provides a knowledge graph alignment method, device, equipment, storage medium and program product to solve the problem that the methods in the prior art are difficult to efficiently and accurately handle alignment work with huge amounts of data.
[0006] The present invention provides a knowledge graph alignment method, comprising: obtaining a first graph, wherein the first graph comprises an entity name, a relationship triple and an attribute triple; editing the entity name, the relationship triple and the attribute triple into a text sequence, and determining the sorting method of the relationship and the attribute according to the uniqueness of the relationship pointed to by the head entity and the tail entity, so as to obtain a target text sequence; inputting the target text sequence into a preset language model, and generating a multi-dimensional entity encoding according to a CLS vector of the starting position of the output hidden layer of the preset language model; and calculating the cosine similarity of the multi-dimensional entity encoding with the entity encoding of the second graph according to the multi-dimensional entity encoding, so as to obtain an alignment result.
[0007] According to a knowledge graph alignment method provided by the present invention, before inputting the target text sequence into a preset language model, the method also includes: obtaining a training set, wherein the training set includes aligned seed entities; and fine-tuning the feedforward neural network of the preset language model based on the training set so that the distance between the aligned entity encodings is closer and the distance between the non-aligned entity encodings is farther.
[0008] According to a knowledge graph alignment method provided by the present invention, the order of determining relations and attributes based on the uniqueness of the pointing relationship between the head entity and the tail entity includes: determining the function value of the relation triple in the graph and the function value of the attribute triple in the graph, wherein the function value is used to indicate the uniqueness of the pointing relationship between the head entity and the tail entity; and determining the order of sorting relations and attributes according to the order of the function values from large to small.
[0009] According to a knowledge graph alignment method provided by the present invention, the cosine similarity between the entity encoding of the multidimensional entity encoding and the entity encoding of the second graph is calculated to obtain an alignment result, including: calculating the cosine similarity based on the multidimensional entity encoding to obtain a similarity matrix of the entity name, a similarity matrix of the relationship, and a similarity matrix of the attribute; performing weighted summation on the similarity matrix of the entity name, the similarity matrix of the relationship, and the similarity matrix of the attribute to obtain an aggregated similarity matrix; performing row and column normalization processing on the aggregated similarity matrix to obtain a doubly random matrix; converting the doubly random matrix into a cost matrix in an assignment problem, and discretizing the cost matrix into a permutation matrix to obtain an alignment result.
[0010] According to a knowledge graph alignment method provided by the present invention, the aggregated similarity matrix is subjected to row-column normalization processing to obtain a double random matrix, including: using a Sinkhorn algorithm to perform row-column normalization processing on the aggregated similarity matrix, and when the Sinkhorn algorithm converges, a double random matrix is obtained.
[0011] According to a knowledge graph alignment method provided by the present invention, the cost matrix is discretized into a permutation matrix to obtain an alignment result, including: using the Hungarian algorithm to discretize the cost matrix into a permutation matrix to obtain an alignment result.
[0012] The present invention also provides a knowledge graph alignment device, comprising the following modules: an acquisition module and a processing module; the acquisition module is used to acquire a first graph, the first graph comprising an entity name, a relationship triple and an attribute triple; the processing module is used to edit the entity name, the relationship triple and the attribute triple into a text sequence, and determine the sorting method of the relationship and the attribute according to the uniqueness of the relationship pointed to by the head entity and the tail entity, so as to obtain a target text sequence; the target text sequence is input into a preset language model, and a multi-dimensional entity encoding is generated according to the CLS vector of the starting position of the output hidden layer of the preset language model; the cosine similarity of the multi-dimensional entity encoding with the entity encoding of the second graph is calculated according to the multi-dimensional entity encoding to obtain an alignment result.
[0013] According to a knowledge graph alignment device provided by the present invention, the acquisition module is used to acquire a training set, and the training set includes aligned seed entities; the processing module is used to fine-tune the feedforward neural network of the preset language model based on the training set, so that the distance between the aligned entity encodings is closer and the distance between the non-aligned entity encodings is farther.
[0014] According to a knowledge graph alignment device provided by the present invention, the processing module is used to determine the function value of the relationship triple in the graph and the function value of the attribute triple in the graph, and the function value is used to indicate the uniqueness of the relationship pointed to by the head entity and the tail entity; the sorting order of the relationship and the attribute is determined according to the order of the function value from large to small.
[0015] According to a knowledge graph alignment device provided by the present invention, the processing module is used to calculate the cosine similarity based on the multi-dimensional entity encoding to obtain the similarity matrix of the entity name, the similarity matrix of the relationship and the similarity matrix of the attribute; perform weighted summation on the similarity matrix of the entity name, the similarity matrix of the relationship and the similarity matrix of the attribute to obtain an aggregated similarity matrix; perform row and column normalization processing on the aggregated similarity matrix to obtain a doubly random matrix; convert the doubly random matrix into a cost matrix in the assignment problem, and discretize the cost matrix into a permutation matrix to obtain an alignment result.
[0016] According to a knowledge graph alignment device provided by the present invention, the processing module is used to use the Sinkhorn algorithm to perform row and column normalization processing on the aggregated similarity matrix, and when the Sinkhorn algorithm converges, a double random matrix is obtained.
[0017] According to a knowledge graph alignment device provided by the present invention, the processing module is used to discretize the cost matrix into a permutation matrix using the Hungarian algorithm to obtain an alignment result.
[0018] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the knowledge graph alignment method as described above is implemented.
[0019] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the knowledge graph alignment methods described above.
[0020] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the knowledge graph alignment methods described above.
[0021] The knowledge graph alignment method, apparatus, device, storage medium and program product provided by the present invention, on the one hand, can use a preset language model to perform multi-dimensional encoding on the entity names, relationship triples and attribute triples in the first graph, thereby reducing the impact of the graph scale while fully representing the entity characteristics and improving the accuracy and efficiency of entity encoding; on the other hand, since the target text sequence can be input into the preset language model, unified and effective encoding of entity names, relationships and attributes can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1 It is a flowchart of the knowledge graph alignment method provided by the present invention; Figure 2 It is a structural schematic diagram of the knowledge graph alignment device provided by the present invention; Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0025] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0026] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0027] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and order of execution.
[0028] The embodiments of the present application describe some exemplary embodiments for the purpose of explanation. It should be understood that the present application can be implemented in other ways that are not specifically shown in the drawings.
[0029] like Figure 1 As shown, the embodiment of the present application provides a knowledge graph alignment method, which can be applied to a knowledge graph alignment device. The knowledge graph alignment method may include S101-S104: S101. The knowledge graph alignment device obtains a first graph.
[0030] The first graph includes entity names, relationship triples and attribute triples.
[0031] Specifically, the triples in the first graph can be composed of a head entity, a relationship, and a tail entity to describe a certain relationship between two entities; or they can be composed of an entity, an attribute, and an attribute value to describe the attributes of an entity itself.
[0032] S102. The knowledge graph alignment device edits the entity name, the relationship triple and the attribute triple into a text sequence, and determines the sorting method of the relationship and the attribute according to the uniqueness of the relationship pointed by the head entity and the tail entity to obtain the target text sequence.
[0033] Optionally, the knowledge graph alignment device may edit the entity name, the relationship triple and the attribute triple into text sequences respectively, wherein the feature sequence of the entity name, the feature sequence of the relationship triple and the feature sequence of the attribute triple may be respectively expressed as: (1) Entity name: The name is an important clue to determine whether two entities are equivalent. Let e be the entity name. The feature sequence of the entity name can be expressed as: (2) Entity relationship: is the set of neighbors of entity e and its associated relationships, where It's a neighbor. is the relationship between entities and neighbors. The feature sequence of the relationship triple can be expressed as: (3) Entity attributes: Let is the set of attributes and attribute values of entity e, where is an attribute, is the attribute value corresponding to the attribute, and the characteristic sequence of the attribute triple can be expressed as: Optionally, the knowledge graph alignment device determines the sorting method of the relationship triples and the attribute triples according to the uniqueness of the pointing relationship between the head entity and the tail entity, including: determining the function value of the relationship triples in the graph and the function value of the attribute triples in the graph, the function values are used to indicate the uniqueness of the pointing relationship between the head entity and the tail entity; determining the sorting order of the relationships and attributes in descending order of the function values.
[0034] Specifically, when linearizing the relationship triples and attribute triples, this application adopts a sorting method based on function values: For a relation triple, it can be regarded as a function under certain conditions, that is, when a head entity is given, the relation points to a unique tail entity. The function value of the relation triple is used to measure the uniqueness of its pointing, and the specific calculation formula is: ; T refers to the set of relation triples in the graph. The numerator of the formula counts the number of head entities s that meet specific conditions in the set of relation triples with relation r as the intermediate element; the denominator counts the number of all different combinations of head entities and tail entities (s, o) in the same triples with relation r as the intermediate element. The ratio of the two reflects the uniqueness of the relationship r pointing to. The higher the function value, the more unique the pointing is.
[0035] Similarly, the function value of the attribute triple can be defined in a similar way. The specific calculation formula is: ; T refers to the set of attribute triples in the graph. The numerator counts the number of head entities s that meet the relevant conditions in the triples with attribute a as the middle element in the attribute triple set T; the denominator counts the number of all different head entities and attribute value combinations (s, v) in these triples with attribute a as the middle element. The ratio of the two reflects the uniqueness of the attribute a pointing to. The higher the function value, the more unique the pointing is.
[0036] Finally, the relations and attributes can be sorted from high to low according to the function value to achieve linear arrangement of triples and obtain the target text sequence.
[0037] S103. The knowledge graph alignment device inputs the target text sequence into a preset language model, and generates a multi-dimensional entity code according to the CLS vector of the starting position of the output hidden layer of the preset language model.
[0038] Specifically, the knowledge graph alignment device can input the generated target text sequence as the initial code into the single-layer feedforward neural network of the preset language model, and obtain the multi-dimensional entity code by extracting the CLS vector at the starting position of the hidden layer output by the preset language model. .
[0039] Optionally, before inputting the target text sequence into a preset language model, the knowledge graph alignment device can obtain a training set, which includes aligned seed entities; and fine-tune the feedforward neural network of the preset language model based on the training set so that the distance between the aligned entity encodings is closer and the distance between the non-aligned entity encodings is farther.
[0040] Specifically, in order to obtain the optimal encoding of entities, the knowledge graph alignment device can use the known aligned seed entities as the training set and fine-tune the preset language model to adjust the encoding distribution to minimize the distance between the aligned entity encodings and maximize the distance between the non-aligned entity encodings. That is, given a known aligned seed entity pair S, , train the feedforward neural network to reduce the error L, ; where γ is a hyperparameter, , is the negative sampling set.
[0041] Based on the above scheme, the feedforward neural network of the preset language model can be fine-tuned based on the training set. Since the present application can be trained with only a small number of seed entities, the problem that seed entity annotation is expensive and difficult to obtain in the real world can be solved.
[0042] In the present application, the preset language model can be a BERT model. The knowledge graph alignment device can fine-tune the three BERT-based encoders respectively to obtain the encoding of names, relationships, and attributes. In the negative sampling part, relevant research in the prior art shows that whether the negative sampling method is effective will greatly affect the accuracy of the results. A simpler method is to randomly sample from all possible negative samples, and a better method is to select more difficult negative samples to improve the encoder's ability to distinguish between equivalent and non-equivalent entities. This application refers to the latter and selects negative samples with higher similarity to entities from all possible negative samples to participate in training.
[0043] S104. The knowledge graph alignment device calculates the cosine similarity between the multi-dimensional entity coding and the entity coding of the second graph to obtain an alignment result.
[0044] Optionally, the knowledge graph alignment device calculates the cosine similarity of the entity coding with the entity coding of the second graph according to the multidimensional entity coding to obtain an alignment result, including: calculating the cosine similarity according to the multidimensional entity coding to obtain a similarity matrix of the entity name, a similarity matrix of the relationship and a similarity matrix of the attribute; performing weighted summation on the similarity matrix of the entity name, the similarity matrix of the relationship and the similarity matrix of the attribute to obtain an aggregated similarity matrix; performing row and column normalization processing on the aggregated similarity matrix to obtain a double random matrix; converting the double random matrix into a cost matrix in the assignment problem, and discretizing the cost matrix into a permutation matrix to obtain an alignment result.
[0045] Optionally, the knowledge graph alignment device may use a Sinkhorn algorithm to perform row and column normalization processing on the aggregated similarity matrix, and when the Sinkhorn algorithm converges, a double random matrix is obtained.
[0046] Optionally, the cost matrix is discretized into a permutation matrix using a Hungarian algorithm to obtain an alignment result.
[0047] Specifically, the knowledge graph alignment device can calculate the cosine similarity of the entity encodings in the first graph and the second graph to form a similarity matrix , ,in, is the dimension index.
[0048] In order to integrate the multi-dimensional similarity matrix, the knowledge graph alignment device can perform weighted summation on the name, relationship and attribute similarity matrices to obtain the aggregated similarity matrix S: ; Next, the Sinkhorn algorithm is applied to the aggregated similarity matrix to perform row and column normalization. When the Sinkhorn algorithm converges, the doubly random matrix M is obtained: M=Sinkhorn(S) When predicting the final alignment relationship, the double random matrix M is first converted into the cost matrix in the assignment problem, and then the Hungarian algorithm is applied to the cost matrix to discretize the cost matrix into the permutation matrix P. Finally, the alignment result is obtained: , source entity i is aligned with target entity j.
[0049] In the embodiments of the present application, on the one hand, since the preset language model can be used to perform multi-dimensional encoding on the entity names, relationship triples and attribute triples in the first graph, it is possible to fully represent the entity characteristics while reducing the impact of the graph scale and improving the accuracy and efficiency of entity encoding; on the other hand, since the target text sequence can be input into the preset language model, unified and effective encoding of entity names, relationships and attributes can be achieved.
[0050] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0051] The knowledge graph alignment method provided in the embodiment of the present application can be executed by a knowledge graph alignment device or a control module for knowledge graph alignment in the knowledge graph alignment device. In the embodiment of the present application, the knowledge graph alignment device provided in the embodiment of the present application is described by taking the execution of the knowledge graph alignment method by the knowledge graph alignment device as an example.
[0052] It should be noted that the embodiment of the present application can divide the functional modules of the knowledge graph alignment device according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. Optionally, the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0053] like Figure 2 As shown, an embodiment of the present application provides a knowledge graph alignment device 200. The knowledge graph alignment device 200 includes: an acquisition module 201 and a processing module 202. The acquisition module 201 can be used to acquire a first graph, wherein the first graph includes an entity name, a relationship triple and an attribute triple; the processing module 202 is used to edit the entity name, the relationship triple and the attribute triple into a text sequence, and determine the sorting method of the relationship and the attribute according to the uniqueness of the relationship pointed to by the head entity and the tail entity, so as to obtain a target text sequence; the target text sequence is input into a preset language model, and a multi-dimensional entity encoding is generated according to the CLS vector of the starting position of the output hidden layer of the preset language model; the cosine similarity of the multi-dimensional entity encoding with the entity encoding of the second graph is calculated according to the multi-dimensional entity encoding to obtain an alignment result.
[0054] Optionally, the acquisition module 201 is used to acquire a training set, which includes aligned seed entities; the processing module 202 is used to fine-tune the feedforward neural network of the preset language model based on the training set, so that the distance between the aligned entity encodings is closer and the distance between the non-aligned entity encodings is farther.
[0055] Optionally, the processing module 202 is used to determine the function value of the relationship triple in the graph and the function value of the attribute triple in the graph, and the function value is used to indicate the uniqueness of the relationship pointed to by the head entity and the tail entity; and to determine the sorting order of the relationship and the attribute according to the order of the function value from large to small.
[0056] Optionally, the processing module 202 is used to calculate the cosine similarity based on the multi-dimensional entity encoding to obtain the similarity matrix of the entity name, the similarity matrix of the relationship and the similarity matrix of the attribute; perform weighted summation on the similarity matrix of the entity name, the similarity matrix of the relationship and the similarity matrix of the attribute to obtain an aggregated similarity matrix; perform row and column normalization processing on the aggregated similarity matrix to obtain a double random matrix; convert the double random matrix into a cost matrix in the assignment problem, and discretize the cost matrix into a permutation matrix to obtain an alignment result.
[0057] Optionally, the processing module 202 is used to perform row and column normalization processing on the aggregated similarity matrix using a Sinkhorn algorithm, and when the Sinkhorn algorithm converges, a double random matrix is obtained.
[0058] Optionally, the processing module 202 is used to discretize the cost matrix into a permutation matrix using a Hungarian algorithm to obtain an alignment result.
[0059] In the embodiments of the present application, on the one hand, since the preset language model can be used to perform multi-dimensional encoding on the entity names, relationship triples and attribute triples in the first graph, it is possible to fully represent the entity characteristics while reducing the impact of the graph scale and improving the accuracy and efficiency of entity encoding; on the other hand, since the target text sequence can be input into the preset language model, unified and effective encoding of entity names, relationships and attributes can be achieved.
[0060] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the knowledge graph alignment method, which includes: obtaining a first graph, wherein the first graph includes an entity name, a relationship triple and an attribute triple; editing the entity name, the relationship triple and the attribute triple into a text sequence, and determining the ordering method of the relationship and the attribute according to the uniqueness of the pointing relationship between the head entity and the tail entity to obtain a target text sequence; inputting the target text sequence into a preset language model, and generating a multi-dimensional entity encoding according to the CLS vector of the output hidden layer start position of the preset language model; and calculating the cosine similarity of the multi-dimensional entity encoding with the entity encoding of the second graph according to the multi-dimensional entity encoding to obtain an alignment result.
[0061] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0062] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the knowledge graph alignment method provided by the above methods, which method includes: obtaining a first graph, wherein the first graph includes an entity name, a relationship triple and an attribute triple; editing the entity name, the relationship triple and the attribute triple into a text sequence, and determining the sorting method of the relationship and the attribute based on the uniqueness of the relationship pointed to by the head entity and the tail entity to obtain a target text sequence; inputting the target text sequence into a preset language model, and generating a multidimensional entity encoding based on the CLS vector of the starting position of the output hidden layer of the preset language model; calculating the cosine similarity of the multidimensional entity encoding with the entity encoding of the second graph to obtain an alignment result.
[0063] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the knowledge graph alignment method provided by the above-mentioned methods, the method comprising: obtaining a first graph, the first graph comprising an entity name, a relationship triple and an attribute triple; editing the entity name, the relationship triple and the attribute triple into a text sequence, and determining the sorting method of the relationship and the attribute according to the uniqueness of the relationship pointed to by the head entity and the tail entity, to obtain a target text sequence; inputting the target text sequence into a preset language model, and generating a multidimensional entity encoding according to the CLS vector of the starting position of the output hidden layer of the preset language model; calculating the cosine similarity of the multidimensional entity encoding with the entity encoding of the second graph to obtain an alignment result.
[0064] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0065] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A knowledge graph alignment method, characterized in that: include: Obtain a first graph, wherein the first graph includes entity names, relationship triples, and attribute triples; The entity name, the relationship triple and the attribute triple are edited into a text sequence, and the order of the relationship and the attribute is determined according to the uniqueness of the relationship pointed by the head entity and the tail entity to obtain a target text sequence; Inputting the target text sequence into a preset language model, and generating a multi-dimensional entity code according to a CLS vector of a starting position of an output hidden layer of the preset language model; The cosine similarity between the multi-dimensional entity encoding and the entity encoding of the second graph is calculated to obtain an alignment result.
2. The knowledge graph alignment method according to claim 1, characterized in that: Before inputting the target text sequence into a preset language model, the method further includes: Obtain a training set, wherein the training set includes aligned seed entities; The feedforward neural network of the preset language model is fine-tuned based on the training set so that the distance between the aligned entity codes is closer and the distance between the unaligned entity codes is farther.
3. The knowledge graph alignment method according to claim 1, characterized in that: The method of determining the order of relations and attributes according to the uniqueness of the pointing relationship between the head entity and the tail entity includes: Determine the function value of the relationship triple in the graph and the function value of the attribute triple in the graph, wherein the function value is used to indicate the uniqueness of the pointing relationship between the head entity and the tail entity; The order of relations and attributes is determined according to the descending order of the function values.
4. The knowledge graph alignment method according to claim 1, characterized in that: The calculating the cosine similarity between the multi-dimensional entity code and the entity code of the second graph to obtain an alignment result includes: Calculating cosine similarity according to the multi-dimensional entity encoding to obtain a similarity matrix of the entity name, a similarity matrix of the relationship, and a similarity matrix of the attribute; Performing weighted summation on the similarity matrix of the entity name, the similarity matrix of the relationship, and the similarity matrix of the attribute to obtain an aggregated similarity matrix; Performing row and column normalization processing on the aggregated similarity matrix to obtain a double random matrix; The double random matrix is converted into a cost matrix in the assignment problem, and the cost matrix is discretized into a permutation matrix to obtain an alignment result.
5. The knowledge graph alignment method according to claim 4, characterized in that: The step of performing row and column normalization processing on the aggregated similarity matrix to obtain a double random matrix includes: The Sinkhorn algorithm is used to perform row and column normalization processing on the aggregated similarity matrix, and when the Sinkhorn algorithm converges, a double random matrix is obtained.
6. The knowledge graph alignment method according to claim 4, characterized in that: The step of discretizing the cost matrix into a permutation matrix to obtain an alignment result includes: The cost matrix is discretized into a permutation matrix using the Hungarian algorithm to obtain an alignment result.
7. A knowledge graph alignment device, characterized in that: include: Acquisition module and processing module; The acquisition module is used to acquire a first graph, wherein the first graph includes entity names, relationship triples, and attribute triples; The processing module is used to edit the entity name, the relationship triple and the attribute triple into a text sequence, and determine the sorting method of the relationship and the attribute according to the uniqueness of the relationship pointed to by the head entity and the tail entity to obtain a target text sequence; input the target text sequence into a preset language model, and generate a multidimensional entity encoding according to the CLS vector of the starting position of the output hidden layer of the preset language model; calculate the cosine similarity of the multidimensional entity encoding with the entity encoding of the second graph according to the multidimensional entity encoding to obtain an alignment result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the knowledge graph alignment method as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the knowledge graph alignment method as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the knowledge graph alignment method as described in any one of claims 1 to 6 is implemented.