Knowledge Graph Representation Learning Method, Device and Medium
By using preset stripping algorithm, graph attention network and clustering algorithm in knowledge graph representation learning, the shared feature vector is determined and nonlinear changes are performed, the problem of excessive parameter quantity in traditional methods is solved, and more efficient knowledge graph representation learning is achieved.
Patent Information
- Application Number
- CN202510153141.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Traditional knowledge graph representation learning methods are inefficient when processing large-scale knowledge graphs, making it difficult to perform complex reasoning, and the amount of parameters is proportional to the number of entities, resulting in excessive parameter volume.
The preset stripping algorithm is used to determine the core entity from the entity, and the initial feature vector is determined based on the distance between the entity and the core entity. The graph attention network is used to aggregate neighbor information, and the shared feature vector is determined through the clustering algorithm, and the shared feature correlation degree is nonlinear to obtain the target feature vector.
The decoupling of the parameter quantity and the number of entities is realized, which significantly reduces the parameter quantity of knowledge graph representation learning, and avoids the problem of excessive parameters caused by the individual definition of representation vectors for each entity in traditional methods.
Smart Images

Figure CN119622000B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network information processing, and particularly to a knowledge graph representation learning method, device and medium. Background Art
[0002] At present, with the explosive growth of data volume, the scale of knowledge graphs is constantly expanding. Traditional knowledge graph representation learning methods face problems such as low efficiency in processing large-scale knowledge graphs and difficulty in performing complex reasoning. Although existing lightweight knowledge graph representation learning methods reduce the number of parameters to a certain extent, they still cannot solve the problem that the number of parameters is proportional to the number of entities, resulting in an excessively high number of parameters when facing large-scale knowledge graphs. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to overcome the deficiencies in the prior art and provide a knowledge graph representation learning method, device and medium. The present invention provides the following technical solutions:
[0004] In a first aspect, the present invention provides a knowledge graph representation learning method, and the method includes:
[0005] Obtain a knowledge graph, where the knowledge graph includes N entities and the relationships between the entities;
[0006] Determine M core entities from the N entities by using a preset peeling algorithm, where M < N;
[0007] Determine the initial feature vector corresponding to each entity according to the distance between each entity and each core entity;
[0008] For the initial feature vector corresponding to the i-th entity, use a graph attention network to aggregate neighbor information to obtain the aggregated feature vector corresponding to the i-th entity, where 1 ≤ i ≤ N;
[0009] Use a preset clustering algorithm to divide the aggregated feature vectors into multiple clusters, and determine the clustering centers of the clusters as shared feature vectors respectively;
[0010] Determine the shared feature correlation degree of each aggregated feature vector according to the distance between each aggregated feature vector and each shared feature vector;
[0011] Perform non-linear transformation on each shared feature correlation degree respectively to obtain the target feature vector corresponding to each entity.
[0012] In one embodiment, determining M core entities from the N entities by using a preset peeling algorithm includes: traversing each of the entities in a loop, and sequentially adding the entities that meet the preset conditions to a preset entity list in the traversal order; determining the last M entities in the preset entity list as core entities respectively.
[0013] In one embodiment, determining the initial feature vector corresponding to each entity according to the distances between each entity and each core entity includes: for the i-th entity, determining the distances from the i-th entity to each of the core entities respectively; determining the initial feature vector corresponding to the i-th entity according to the distances from the i-th entity to each of the core entities respectively.
[0014] In one embodiment, aggregating neighbor information for the initial feature vector corresponding to the i-th entity by using a graph attention network to obtain the aggregated feature vector corresponding to the i-th entity includes:
[0015] Determining the entities directly connected to the i-th entity as the neighbor entities of the i-th entity respectively;
[0016] Determining the attention weights of the i-th entity for each of the neighbor entities respectively, and the calculation formula is:
[0017]
[0018]
[0019] Determining the aggregated feature vector corresponding to the i-th entity according to each of the attention weights, and the formula is:
[0020]
[0021] where, represents the attention weight of the i-th entity for its neighbor entity , represents the i-th entity, represents the -th neighbor entity, represents the relationship between the i-th entity and the neighbor entity, represents the set of neighbor entities, and are different trainable mapping matrices respectively, is a transformation vector related to the connection relationship and is used to map the relationship between entities into a vector space of a preset relationship, represents the initial feature vector corresponding to the i-th entity, represents the initial feature vector corresponding to the neighbor entity e Denote the aggregated feature vector corresponding to the \(i\)-th said entity.
[0022] In one embodiment, the preset clustering algorithm is used to divide each of the aggregated feature vectors into multiple clusters, and the clustering loss function is:
[0023] where denotes the entity to the shared feature vector distance, minpool represents the min pooling operation, \(V\) represents the entity set, and \(L\) represents the shared feature vector set.
[0024] In one embodiment, the determining the shared feature correlation degree of each of the aggregated feature vectors according to the distances between each of the aggregated feature vectors and each of the shared feature vectors includes: determining the distances from the aggregated feature vector to each of the shared feature vectors respectively, and determining the minimum distance among the distances as the shared feature correlation degree of the aggregated feature vector.
[0025] In one embodiment, the respectively performing non-linear transformation on each of the shared feature correlation degrees includes: using a multi-layer perceptron to respectively perform non-linear transformation on each of the shared feature correlation degrees.
[0026] In one embodiment, the knowledge graph further includes a positive triple set, the positive triple set includes a plurality of positive samples, and the method further includes:
[0027] Based on the positive triple set, constructing a plurality of negative samples through random negative sampling, and each of the negative samples constitutes a negative triple set;
[0028] Determining a triple loss function according to each of the positive samples in the positive triple set and each of the negative samples in the negative triple set, and the calculation formula is:
[0029]
[0030]
[0031]
[0032] where \(S\) represents the triple loss function, represents the loss of the positive sample represents the loss of the negative sample represents the positive triple set, represents the negative triple set, represents the hyperparameter of the margin ranking loss, represents the positive sample The head entity represents a positive sample The relationship represents a positive sample The tail entity represents a negative sample The head entity represents a negative sample The relationship represents a negative sample The tail entity represents a positive sample The shared feature correlation degree of the head entity represents a positive sample The shared feature correlation degree of the tail entity represents a negative sample The shared feature correlation degree of the head entity represents a negative sample The shared feature correlation degree of the tail entity.
[0033] In a second aspect, the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the computer program runs on the processor, it executes the knowledge graph representation learning method described in the first aspect.
[0034] In a third aspect, the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it performs the knowledge graph representation learning method described in the first aspect.
[0035] The knowledge graph representation learning method, device, and medium provided by the embodiments of the present application. The method includes: obtaining a knowledge graph, where the knowledge graph includes N entities and the relationships between the entities; determining M core entities from the N entities by using a preset peeling algorithm, where M < N; determining the initial feature vectors corresponding to the entities according to the distances between the entities and the core entities; for the initial feature vector corresponding to the i-th entity, aggregating neighbor information by using a graph attention network to obtain the aggregated feature vector corresponding to the i-th entity, where 1 ≤ i ≤ N; using a preset clustering algorithm to divide the aggregated feature vectors into multiple clusters, and determining the clustering centers of the clusters as the shared feature vectors respectively; determining the shared feature correlation degrees of the aggregated feature vectors according to the distances between the aggregated feature vectors and the shared feature vectors; respectively performing non-linear transformation on the shared feature correlation degrees to obtain the target feature vectors corresponding to the entities respectively. By constructing the shared feature vectors between entities and representing each entity in the knowledge graph based on the shared feature vectors, the present application realizes the decoupling of the number of parameters and the number of entities, significantly reduces the number of parameters in knowledge graph representation learning, and avoids the problem of excessive parameters caused by separately defining representation vectors for each entity in the traditional method.
[0036] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific embodiments are given, and detailed descriptions are made in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 FIG. 1 shows a flowchart of the knowledge graph representation learning method provided by the embodiments of the present application;
[0039] Figure 2 FIG. 2 shows another flowchart of the knowledge graph representation learning method provided by the embodiments of the present application;
[0040] Figure 3 FIG. 3 shows still another flowchart of the knowledge graph representation learning method provided by the embodiments of the present application;
[0041] Figure 4 FIG. 4 shows yet another flowchart of the knowledge graph representation learning method provided by the embodiments of the present application;
[0042] Figure 5It shows yet another schematic flow diagram of the knowledge graph representation learning method provided by the embodiments of the present application;
[0043] Figure 6 It shows a schematic structural diagram of an electronic device provided by the embodiments of the present application.
[0044] Main component symbol description:
[0045] 600 - Electronic device; 601 - Transceiver; 602 - Processor; 603 - Memory. Detailed implementation manners
[0046] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0047] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this template herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0049] Embodiment 1
[0050] As a structured knowledge representation method, a knowledge graph organizes entities and the relationships between entities in the form of a graph, where nodes represent entities and edges represent relationships between entities. Traditional knowledge graph representation methods have problems such as low efficiency and difficulty in complex reasoning when dealing with large-scale knowledge graphs. Thus, knowledge graph representation learning emerged. Knowledge graph representation learning aims to map entities and relationships in a knowledge graph to a vector space to facilitate efficient storage, retrieval, and reasoning of data. Existing knowledge graph representation learning usually assigns one or more vector representations to each entity and relationship in the knowledge graph. In some cases, even additional complex parameters are required to describe specific relationship types or entity interaction methods. As the scale of the knowledge graph expands, existing knowledge graph representation learning faces the problem of excessive parameter quantities. For this, please refer toFigure 1 , this application proposes a lightweight knowledge graph representation learning method based on shared features, including steps S110 to S170.
[0051] Step S110, obtain a knowledge graph, where the knowledge graph includes N entities and the relationships between each of the entities.
[0052] An entity is the basic unit in a knowledge graph and can be any object with independent significance, such as: person names, place names, organizations, concepts, time, etc. A relationship is the semantic link connecting two entities and is used to describe the semantic association between entities. Entities and the relationships between each entity form a knowledge graph.
[0053] In this embodiment, the knowledge graph is defined as , where represents the N entities included in the knowledge graph, represents the relationships between each entity, represents all triples. It should be noted that a triple is the smallest semantic unit in a knowledge graph and is used to represent a complete semantic fact in the knowledge graph. A triple is usually defined as , where represents the head entity, represents the relationship between the two entities, represents the tail entity.
[0054] It can be understood that the goal of traditional knowledge graph representation learning is to learn vector representations for entities and relationships based on triple information, that is , where is the dimension of the vector, and this vector can effectively reflect the structural information of the knowledge graph. Therefore, the number of parameters in the traditional knowledge graph representation is , which is linearly related to the number of entities .
[0055] In this embodiment, to achieve decoupling of the number of parameters from the number of entities in the knowledge graph, a lightweight representation method based on shared features is adopted, representing entities as a combination of shared features, thereby avoiding defining a separate representation vector for each entity in the traditional method. For details, please refer to the following text.
[0056] Step S120, use a preset stripping algorithm to determine M core entities from the N entities, where M < N.
[0057] It should be noted that in traditional knowledge graph representation learning, the way to represent the uniqueness of an entity is the entity index. However, the index cannot well reflect the structural features of the entity in the knowledge graph.
[0058] In this embodiment, entity location identifiers are used to identify the uniqueness of entities. Specifically, a preset peeling algorithm, such as the K-shell algorithm, is used to select core entities with important location structures from all entities in the knowledge graph, and then the uniqueness of each entity is characterized by the distance between each entity and each core entity.
[0059] In one embodiment, please refer to Figure 2 , step S120 includes steps S121 to S122.
[0060] Step S121, loop through each of the entities, and in the order of traversal, add the entities that meet the preset conditions to a preset entity list in sequence.
[0061] Specifically, define a preset entity list. It can be understood that the preset entity list is an empty list. Loop through each entity in the knowledge graph. In each loop, add the entities with a degree of 1 to the preset entity list in the order of traversal. This process continues until all entities with a degree of 1 in the knowledge graph are processed.
[0062] Step S122, determine the last M entities in the preset entity list as core entities respectively.
[0063] It can be understood that the multiple entities included in the preset entity list obtained after the traversal are arranged in ascending order of the importance of the location structure. The entities with a more complex connection structure are located further back. Select the last M entities in the location list as core entities respectively, and further characterize the uniqueness of the entities based on the distance between each entity and the core entities.
[0064] Step S130, determine the initial feature vector corresponding to each entity according to the distance between each entity and each core entity.
[0065] In this embodiment, the initial feature vector is a set of fixed low-dimensional vectors , where represents the distance from the corresponding entity to the th core entity. It can be understood that through the distance between each entity and each core entity, different entities can be effectively distinguished, and the location characteristics of the entities in the knowledge graph can be provided to a certain extent.
[0066] In one embodiment, please refer to Figure 3 , step S130 includes steps S131 to S132.
[0067] S131, for the i-th entity, determine the distances from the i-th entity to each of the core entities.
[0068] Taking each core entity as a reference point, calculate the shortest path distance from the i-th entity to each core entity.
[0069] S132. Determine the initial feature vector corresponding to the i-th entity according to the distances from the i-th entity to each of the core entities.
[0070] For example, the knowledge graph includes core entities: 、 and , and the distances from the i-th entity to each core entity are respectively: 、 、 , then the initial feature vector corresponding to the i-th entity is: .
[0071] Step S140. For the initial feature vector corresponding to the i-th entity, use a graph attention network to aggregate neighbor information to obtain the aggregated feature vector corresponding to the i-th entity, where 1 ≤ i ≤ N.
[0072] In this embodiment, after obtaining the initial feature vector of the entity, according to the connection situation of the entity in the knowledge graph, further use a graph attention network to aggregate the neighbor information of the entity to enrich the structural information of the entity.
[0073] In one implementation manner, please refer to Figure 4 , step S140 includes steps S141 to S143.
[0074] S141. Respectively determine the neighbor entities of the i-th entity as the entities directly connected to the i-th entity.
[0075] In this embodiment, use to represent the set of neighbor entities of the i-th entity, where e represents a neighbor entity.
[0076] S142. Respectively determine the attention weights of the i-th entity to each of the neighbor entities, and the calculation formula is:
[0077]
[0078]
[0079] where represents the attention weight of the i-th entity to its neighbor entity , represents the i-th entity, represents the -th neighbor entity, represents the relationship between the i-th entity and the neighbor entity Represents a set of neighbor entities, and are respectively different trainable mapping matrices, is a transformation vector related to the connection relationship, used to map the relationship between entities into the vector space of the preset relationship, represents the initial feature vector corresponding to the i-th entity, represents the initial feature vector corresponding to the neighbor entity e, represents the aggregated feature vector corresponding to the i-th entity.
[0080] It can be understood that is the attention score of the neighbor entity e, used to measure the importance of the neighbor entity to the central entity, that is, the i-th entity. The attention weight confirmation method provided in this application considers the relationship type between entities, enabling the attention mechanism to better capture the semantic information in the heterogeneous graph.
[0081] S143. According to each of the attention weights, determine the aggregated feature vector corresponding to the i-th entity. The formula is:
[0082]
[0083] By aggregating neighbor information, the aggregated feature vectors corresponding to each entity are obtained, and then each entity is characterized by richer information.
[0084] Step S150. Use a preset clustering algorithm to divide each of the aggregated feature vectors into multiple clusters, and determine the clustering centers of each of the clusters as the shared feature vectors respectively.
[0085] In this embodiment, based on the aggregated feature vectors corresponding to all entities respectively, a clustering method based on a neural network is used to discover at least one shared feature vector. Specifically, it includes: clustering each of the aggregated feature vectors into multiple clusters, and determining the centers of each of the clusters as the shared feature vectors respectively. Among them, the clustering loss function is:
[0086]
[0087] Among them, represents the entity to the shared feature vector distance, minipool represents the min-pooling operation, V represents the entity set, and L represents the shared feature vector set.
[0088] It should be noted that the number of clusters can be selected according to the data characteristics, calculation cost, and application scenario by appropriate methods, such as: elbow method, silhouette coefficient, gap statistic, or method based on neural network, which is not specifically limited.
[0089] Step S160: Determine the shared feature correlation degree of each aggregation feature vector according to the distance between each aggregation feature vector and each shared feature vector.
[0090] In this embodiment, a multi-layer perceptron (MLP) is used to calculate the distance from each aggregation feature vector to each shared feature vector, and the shared feature correlation degree of each entity is further determined according to the calculated distances. .
[0091] In one implementation manner, the determining the shared feature correlation degree of each aggregation feature vector according to the distance between each aggregation feature vector and each shared feature vector includes: determining the distance from each aggregation feature vector to each shared feature vector, and determining the minimum distance among the distances as the shared feature correlation degree of the aggregation feature vector.
[0092] It can be understood that taking one entity as an example, calculate the distance from the aggregation feature vector corresponding to the entity to each shared feature vector, and determine the shortest distance as the correlation degree between the entity and the shared feature vector, that is, the shared feature correlation degree. Among them, the shared feature correlation degree is a vector.
[0093] Step S170: Perform non-linear transformation on each shared feature correlation degree respectively to obtain the target feature vector corresponding to each entity respectively.
[0094] In this embodiment, after obtaining the shared feature correlation degree corresponding to each entity respectively, further perform non-linear transformation on each shared feature correlation degree to map it to the final representation vector, that is, the target feature vector. The calculation formula is: , where represents a multi-layer perceptron.
[0095] In one implementation manner, the performing non-linear transformation on each shared feature correlation degree respectively includes: using a multi-layer perceptron to perform non-linear transformation on each shared feature correlation degree respectively.
[0096] Perform non-linear transformation on the shared feature correlation degree corresponding to each entity to further optimize the vector representation of each entity, so that the final target feature vector of each entity can better capture complex semantic information.
[0097] In one implementation manner, the knowledge graph further includes a set of positive triples, and the triple set includes multiple positive samples. Please refer to Figure 5 , and the method further includes: steps S510 to S520.
[0098] S510. Based on the positive triple set, construct multiple negative samples through random negative sampling, and each of the negative samples constitutes a negative triple set.
[0099] In this embodiment, the positive triple set includes multiple original triples, that is, positive samples . Based on each positive sample in the original triples, construct multiple negative samples by means of random negative sampling , where , , are respectively , , the head entity, relation, and tail entity after random replacement.
[0100] S520. Determine the triple loss function according to each positive sample in the positive triple set and each negative sample in the negative triple set. The calculation formula is:
[0101]
[0102]
[0103]
[0104] where S represents the triple loss function, represents the loss of the positive sample , represents the loss of the negative sample , represents the positive triple set, represents the negative triple set, represents the hyperparameter of the margin ranking loss, represents the head entity of the positive sample , represents the relation of the positive sample , represents the tail entity of the positive sample , represents the head entity of the negative sample , represents the relation of the negative sample , represents the tail entity of the negative sample , represents the shared feature correlation degree of the head entity of the positive sample , represents the shared feature correlation degree of the tail entity of the positive sample , represents the shared feature correlation degree of the head entity of the negative sample Indicates a negative sample The shared feature correlation degree of the tail entity of
[0105] Furthermore, the target loss function can be determined according to the triplet loss function and the clustering loss function. Specifically, the target loss function is equal to the sum of the triplet function and the clustering loss function. The core purpose of the target loss function is to measure the difference between the predicted value and the true value, and to adjust the parameters by optimizing this difference, thereby improving the learning performance.
[0106] To verify the effectiveness of the knowledge graph representation learning method proposed in this application in lightweight knowledge graph representation learning, the industry-standard WN18RR and FB15K237 datasets are used for verification below. The verification metrics are MRR (Mean Reciprocal Rank) and Hits@10 for link prediction. MRR is obtained by calculating the average of the sum of the reciprocals of the ranks of all query results, and Hits@10 checks whether the correct answer is among the top 10 positions of these candidate answers. Please refer to Table 1 below for the experimental results. According to the experimental results, the method proposed in the present invention has higher accuracy in the link prediction task than many existing latest lightweight learning methods, including methods based on knowledge distillation (DualDE), lemmas (LightKG), and transposed convolution (LN-TransE), indicating the excellent characteristics of the shared features.
[0107] Table 1
[0108]
[0109] Meanwhile, the method proposed in the present invention requires fewer parameters in these two datasets, greatly improving the applicable range of knowledge graph representation learning. Specifically, the comparison results of the number of parameters are shown in Table 2 below.
[0110] Table 2
[0111]
[0112] The knowledge graph representation learning method provided by the embodiments of the present application includes: obtaining a knowledge graph, where the knowledge graph includes N entities and the relationships between the entities; determining M core entities from the N entities by using a preset stripping algorithm, where M < N; determining the initial feature vectors corresponding to the entities according to the distances between the entities and the core entities; for the initial feature vector corresponding to the i-th entity, aggregating neighbor information by using a graph attention network to obtain the aggregated feature vector corresponding to the i-th entity, where 1 ≤ i ≤ N; using a preset clustering algorithm to divide the aggregated feature vectors into multiple clusters, and determining the clustering centers of the clusters as the shared feature vectors respectively; determining the shared feature correlation degrees of the aggregated feature vectors according to the distances between the aggregated feature vectors and the shared feature vectors; and performing non-linear transformation on the shared feature correlation degrees respectively to obtain the target feature vectors corresponding to the entities respectively. By constructing the shared feature vectors between entities and representing each entity in the knowledge graph based on the shared feature vectors, the present application realizes the decoupling of the number of parameters and the number of entities, significantly reduces the number of parameters in knowledge graph representation learning, and avoids the problem of excessive parameters caused by separately defining representation vectors for each entity in the traditional method.
[0113] Embodiment 2
[0114] In addition, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the computer program runs on the processor, it executes the knowledge graph representation learning method provided in Embodiment 1.
[0115] Specifically, please refer to Figure 6 , the electronic device 600 includes: a transceiver 601, a bus interface, and a processor 602. The processor 602 is configured to obtain a knowledge graph, where the knowledge graph includes N entities and the relationships between the entities; determine M core entities from the N entities by using a preset stripping algorithm, where M < N; determine the initial feature vectors corresponding to the entities according to the distances between the entities and the core entities; for the initial feature vector corresponding to the i-th entity, aggregate neighbor information by using a graph attention network to obtain the aggregated feature vector corresponding to the i-th entity, where 1 ≤ i ≤ N; use a preset clustering algorithm to divide the aggregated feature vectors into multiple clusters, and determine the clustering centers of the clusters as the shared feature vectors respectively; determine the shared feature correlation degrees of the aggregated feature vectors according to the distances between the aggregated feature vectors and the shared feature vectors; and perform non-linear transformation on the shared feature correlation degrees respectively to obtain the target feature vectors corresponding to the entities respectively.
[0116] In one embodiment, determining M core entities from the N entities by using a preset peeling algorithm includes: traversing each of the entities in a loop, and sequentially adding the entities that meet the preset conditions to a preset entity list according to the traversal order; determining the last M entities in the preset entity list as core entities respectively.
[0117] In one embodiment, determining the initial feature vector corresponding to each entity according to the distances between each entity and each core entity includes: for the i-th entity, determining the distances from the i-th entity to each of the core entities respectively; and determining the initial feature vector corresponding to the i-th entity according to the distances from the i-th entity to each of the core entities.
[0118] In one embodiment, aggregating neighbor information for the initial feature vector corresponding to the i-th entity by using a graph attention network to obtain the aggregated feature vector corresponding to the i-th entity includes:
[0119] Determining the entities directly connected to the i-th entity as the neighbor entities of the i-th entity respectively;
[0120] Determining the attention weights of the i-th entity for each of the neighbor entities respectively, and the calculation formula is:
[0121]
[0122]
[0123] Determining the aggregated feature vector corresponding to the i-th entity according to each of the attention weights, and the formula is:
[0124]
[0125] Among them, represents the attention weight of the i-th entity for its neighbor entity , represents the i-th entity, represents the -th neighbor entity, represents the relationship between the i-th entity and the neighbor entity, represents the neighbor entity set, and are different trainable mapping matrices respectively, is a transformation vector related to the connection relationship, and is used to map the relationship between entities into a vector space of a preset relationship, represents the initial feature vector corresponding to the i-th entity, represents the initial feature vector corresponding to the neighbor entity e Represents the aggregated feature vector corresponding to the i-th entity.
[0126] In one embodiment, the preset clustering algorithm is used to divide each of the aggregated feature vectors into multiple clusters, and the clustering loss function is:
[0127] Wherein, Represents the entity To the shared feature vector The distance of, minpool represents the min pooling operation, V represents the entity set, and L represents the shared feature vector set.
[0128] In one embodiment, the method for determining the shared feature correlation degree of each of the aggregated feature vectors according to the distances between each of the aggregated feature vectors and each of the shared feature vectors includes: determining the distances from the aggregated feature vector to each of the shared feature vectors respectively, and determining the minimum distance among the distances as the shared feature correlation degree of the aggregated feature vector.
[0129] In one embodiment, the method for respectively performing non-linear transformation on each of the shared feature correlation degrees includes: using a multi-layer perceptron to respectively perform non-linear transformation on each of the shared feature correlation degrees.
[0130] In one embodiment, the knowledge graph further includes a positive triple set, the positive triple set includes a plurality of positive samples, and the method further includes:
[0131] Based on the positive triple set, a plurality of negative samples are constructed by random negative sampling, and each of the negative samples constitutes a negative triple set;
[0132] According to each of the positive samples in the positive triple set and each of the negative samples in the negative triple set, a triple loss function is determined, and the calculation formula is:
[0133]
[0134]
[0135]
[0136] Wherein, S represents the triple loss function, Represents the loss of the positive sample Of, Represents the negative sample Of the loss, Represents the positive triple set, Represents the negative triple set, Represents the hyperparameter of the margin ranking loss, Represents the positive sample The head entity, represents a positive sample The relationship, represents a positive sample The tail entity, represents a negative sample The head entity, represents a negative sample The relationship, represents a negative sample The tail entity, represents a positive sample The shared feature correlation degree of the head entity, represents a positive sample The shared feature correlation degree of the tail entity, represents a negative sample The shared feature correlation degree of the head entity, represents a negative sample The shared feature correlation degree of the tail entity.
[0137] In the embodiment of the present invention, the electronic device 600 further includes: a memory 603. In Figure 6 , the bus architecture may include any number of interconnected buses and bridges, specifically various circuits of one or more processors represented by the processor 602 and the memory represented by the memory 603 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, will not be further described herein. The bus interface provides an interface. The transceiver 601 may be a plurality of elements, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on the transmission medium. The processor 602 is responsible for managing the bus architecture and general processing, and the memory 603 may store data used by the processor 602 when executing operations.
[0138] The electronic device 600 provided by the embodiment of the present invention can execute the knowledge graph representation learning method provided in the above method embodiment 1. To avoid repetition, it will not be elaborated here.
[0139] Embodiment 3
[0140] In addition, the embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the knowledge graph representation learning method provided in Embodiment 1 is implemented.
[0141] In this embodiment, the computer-readable storage medium may be a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disk, or the like.
[0142] The computer-readable storage medium provided in this embodiment can implement the knowledge graph representation learning method provided in Embodiment 1. To avoid repetition, it will not be elaborated here.
[0143] In all the examples shown and described here, any specific value should be construed as merely exemplary, not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0144] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0145] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A knowledge graph representation learning method, characterized in that: The method comprises: Obtain a knowledge graph, wherein the knowledge graph includes N entities and relationships between the entities; Determine M core entities from the N entities using a preset stripping algorithm, where M<N; Determine the initial feature vector corresponding to each of the entities according to the distance between each of the entities and each of the core entities; For the initial feature vector corresponding to the i-th entity, the graph attention network is used to aggregate the neighbor information to obtain the aggregated feature vector corresponding to the i-th entity, where 1≤i≤N; Using a preset clustering algorithm to divide each of the aggregated feature vectors into a plurality of clusters, and determining the cluster centers of each of the clusters as shared feature vectors; Determining the shared feature correlation of each of the aggregated feature vectors according to the distance between each of the aggregated feature vectors and each of the shared feature vectors; Performing nonlinear changes on the correlation degrees of the shared features respectively to obtain target feature vectors corresponding to the entities respectively; The method of using a preset stripping algorithm to determine M core entities from the N entities includes: looping through the entities, and sequentially adding the entities that meet the preset conditions to a preset entity list in the order of traversal; and determining the last M entities in the preset entity list as core entities respectively; The preset clustering algorithm is used to divide each of the aggregated feature vectors into multiple clusters, and the clustering loss function is: in, Representing Entities to the shared feature vector distance, minpool represents the minimum pooling operation, V represents the entity set, and L represents the shared feature vector set; Determining the shared feature association degree of each aggregated feature vector based on the distance between each aggregated feature vector and each shared feature vector includes: determining the distance from the aggregated feature vector to each shared feature vector respectively, and determining the minimum distance among the distances as the shared feature association degree of the aggregated feature vector.
2. The knowledge graph representation learning method according to claim 1, characterized in that: The determining, according to the distance between each of the entities and each of the core entities, an initial feature vector corresponding to each of the entities comprises: For the i-th entity, determine the distances from the i-th entity to each of the core entities; An initial feature vector corresponding to the ith entity is determined according to the distances from the ith entity to each of the core entities.
3. The knowledge graph representation learning method according to claim 1, characterized in that: The initial feature vector corresponding to the i-th entity is obtained by aggregating neighbor information using a graph attention network to obtain an aggregated feature vector corresponding to the i-th entity, including: Determine a plurality of entities directly connected to the i-th entity as neighbor entities of the i-th entity; The attention weight of the i-th entity to each of the neighboring entities is determined respectively, and the calculation formula is: According to each of the attention weights, the aggregate feature vector corresponding to the i-th entity is determined, and the formula is: in, Represents the relationship between the i-th entity and its neighbor entity The attention weight, represents the i-th entity, Indicates said neighbor entities, represents the relationship between the i-th entity and the neighbor entity, represents the set of neighbor entities, and are different trainable mapping matrices, It is a transformation vector related to the connection relationship, which is used to map the relationship between entities into the vector space of the preset relationship. represents the initial feature vector corresponding to the i-th entity, represents the initial feature vector corresponding to the neighbor entity e, Represents the aggregated feature vector corresponding to the i-th entity.
4. The knowledge graph representation learning method according to any one of claims 1 to 3, characterized in that: The non-linearly changing the correlation degree of each shared feature respectively includes: A multi-layer perceptron is used to perform nonlinear changes on the correlation degree of each shared feature.
5. The knowledge graph representation learning method according to claim 1, characterized in that: The knowledge graph further includes a positive triple set, the positive triple set includes a plurality of positive samples, and the method further includes: Based on the positive triplet set, construct a plurality of negative samples by random negative sampling, each of the negative samples constituting a negative triplet set; A ternary loss function is determined according to each of the positive samples in the positive triplet set and each of the negative samples in the negative triplet set, and the calculation formula is: Among them, S represents the ternary loss function, Represents a positive sample of loss, Represents negative samples of loss, represents the set of positive triples, represents the negative triple set, represents the hyperparameter of the edge ranking loss, Represents a positive sample The head entity, Represents a positive sample relationship, Represents a positive sample The tail entity, Represents negative samples The head entity, Represents negative samples relationship, Represents negative samples The tail entity, Represents a positive sample The shared feature correlation of the head entity, Represents a positive sample The shared feature association of the tail entity, Represents negative samples The shared feature correlation of the head entity, Represents negative samples The shared feature association degree of the tail entity.
6. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and when the computer program runs on the processor, it executes the knowledge graph representation learning method described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the knowledge graph representation learning method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Knowledge graph entity alignment method and device
CN114036307A
Log anomaly detection method and system based on knowledge graph semantic embedding
CN118860714A