Multi-party knowledge graph embedding method based on federal learning

Through the multi-party knowledge graph embedding method based on federated learning, the improved graph attention model is used to solve the problem of effectively using other data for knowledge graph embedding training while protecting data privacy, achieving efficient completion of knowledge graphs and improving the accuracy of inter-entity prediction.

CN119940502APending Publication Date: 2025-05-06HUNAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510021516.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

While protecting the privacy of multi-party data, it is difficult for the prior art to effectively use other data to train local knowledge graph embedding, and traditional embedding models are difficult to completely capture the multi-relational information between two entities, resulting in the loss of relational information.

Method used

The multi-party knowledge graph embedding method based on federated learning is adopted, and the knowledge graph completion is completed by building a central entity server, multi-party client entity embedding data is aggregated, and the improved graph attention model is used to train, and the utility of other data is fused to complete the knowledge graph.

Benefits of technology

While protecting data privacy, it improves the accuracy of inter-entity predictions, completes the knowledge graph, and enhances the robustness and performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940502A_ABST
    Figure CN119940502A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-party knowledge graph embedding method based on federal learning, which comprises the following steps of: constructing an entity central server, and determining initialized global entity embedding E0; pre-training an entity embedding # imgabs0 # and a relation embedding # imgabs1 # initialized by K local knowledge maps by using global embedding E0, downloading the initialized E0 into a client, and performing e rounds of iterative training of local entity embedding and relation embedding by using a map attention network local model; the K locally-trained entity embedding units are uploaded to the central server side to be globally aggregated to obtain global entity embedding units; repeating n rounds of downloading from global embedding to local embedding for local training and aggregating from local embedding to global embedding; embedding parameters trained under federal learning are fused with embedding parameters only based on local training, and the K party obtains a robust local embedding model. The method can be used for multiple users, and under the condition that data privacy is protected, the data of other users are efficiently utilized for model parameter training, the accuracy of relation prediction between entities is improved, and knowledge graph completion is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a field of knowledge graph federation, and in particular to a multi-party knowledge graph embedding method based on federated learning. Background Art

[0002] Multi-party knowledge graphs are a common case in practical knowledge graph applications, which can be viewed as a set of related individual knowledge graphs, where different knowledge graphs contain different aspects of the relationship between entities. Intuitively, for each knowledge graph, its completion depends largely on the triples defined and labeled in other knowledge graphs. However, due to the privacy and sensitivity of the data, a set of related knowledge graphs cannot simply complement each other's knowledge graphs by collecting data from different knowledge graphs.

[0003] It is difficult to fully capture the multi-relation information between two entities using traditional embedding models in a single knowledge graph. This is because these methods mainly learn entity embeddings and focus on incorporating relation information into entity embeddings, which leads to the loss of relation information to a certain extent.

[0004] Therefore, the present invention provides a method for embedding a multi-party knowledge graph based on federated learning to solve the above problems. The method can solve the problem of using the data of other parties to train the local knowledge graph embedding while protecting the privacy of multi-party data. In addition, the improved graph attention model is used in the local knowledge graph embedding model to improve the model performance and complete the task of completing the knowledge graph. Summary of the invention

[0005] The purpose of the present invention is to provide a multi-party knowledge graph embedding method based on federated learning. While protecting the privacy of multi-party data, the present invention efficiently uses other-party data to perform local model training, improves the accuracy of inter-entity prediction, and completes the knowledge graph.

[0006] In order to achieve the above object, the present invention provides a multi-party knowledge graph embedding method based on federated learning, which is characterized by the following working steps:

[0007] S1. Build an entity central server to aggregate entity embedding data of multiple clients, aggregate the entities of K clients to the central server, randomly initialize the global entity embedding E0 of the central server, and K parties download E0 to their respective clients for pre-training to obtain the initialized local entity embedding and relation embedding parameters

[0008] S2, the client downloads the global embedding from the central server, uses the local graph attention network for training, uses the client's triples and relationship embeddings to update the entity embedding and relationship embedding, and updates the graph attention network model parameters. The central server aggregates the entity embeddings updated by each client and updates the global embedding;

[0009] S3. Repeat S2 for n rounds. K clients obtain the embedding parameters trained under federated learning and the local graph attention model parameters. The training of these parameters incorporates the utility of other parties’ data.

[0010] S4. By integrating the embedding parameters trained under federated learning with the embedding parameters based only on local training, K-squared obtains a robust local embedding model.

[0011] The steps of multi-party pre-training in S1 include:

[0012] (1) The global entity embedding E0 downloaded from the central server is converted into a local entity embedding through the mapping matrix In the mapping matrix, the corresponding entity in the local and global is 1, and the non-corresponding entity is 0. A set of mapping matrices for the K-party client is expressed as n is the number of entities corresponding to the global embedding, n k is the number of entities corresponding to the k-th client,

[0013] (2) Use distance-based loss function for training and initialize local relationships

[0014] (3) The learnable attention coefficients and matrices in local self-attention are assigned using random initialization.

[0015] The specific steps of global training and local training in S2 are:

[0016] (1) The client downloads the global entity embedding from the central server and calculates the local entity embedding through the mapping matrix;

[0017] (2) Using an improved graph attention network for encoding in the local network and integrating it into the full graph structure of a single knowledge graph;

[0018] (3) Use distance-based loss function for decoding training;

[0019] (4) Repeat steps (1) and (2) for e rounds to train and update local entity embedding, relationship embedding, and graph attention network model parameters;

[0020] (5) The trained local entity embeddings are uploaded to the central server for aggregation and update of the global entity embeddings.

[0021] The improved graph attention network mentioned in the specific steps of S2 is that in each attention layer, the graph structure information of the node is focused on the nodes connected to it and the relationship between them, and is represented by the distance vector of adjacent nodes and relationships. In the calculation of the attention coefficient, the distance vector of adjacent nodes and relationships is also paid attention to.

[0022] The distance-based loss function mentioned in the specific steps of S2 is to generate multiple groups of negative samples for each entity, and the triples of the knowledge graph are positive samples. The distances of the head entity, relationship, and tail entity in the positive and negative samples are calculated. Training reduces the distance of positive samples and increases the distance of negative samples.

[0023] The fusion steps in S4 are:

[0024] (1) After completing the model embedding training under federated learning and the local model embedding training, the client performs a connection operation on the triple-based distance formulas in the two cases;

[0025] (2) The concatenated formula is input as a feature vector into a linear classifier, which is trained using a marginal ranking loss so that positive triplets are ranked higher than negative triplets. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of a multi-party knowledge graph embedding method based on federated learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The detailed description of the drawings is intended as an illustration of some embodiments of the present application, and is not intended to represent the only form in which the present application can be implemented. It should be understood that the same or equivalent functions can be accomplished by different embodiments intended to be included in the spirit and scope of the present application.

[0028] See also Figure 1 In one embodiment, a multi-party knowledge graph embedding method based on federated learning has the following working steps:

[0029] S1. Build an entity central server to aggregate entity embedding data of multiple clients, aggregate the entities of K clients to the central server, randomly initialize the global entity embedding E0 of the central server, and K parties download E0 to their respective clients for pre-training to obtain the initialized local entity embedding and relation embedding parameters

[0030] The client pre-training is as follows: (1) the global entity embedding E0 downloaded from the central server is converted into a local entity embedding through a mapping matrix In the mapping matrix, the corresponding entity in the local and global is 1, and the non-corresponding entity is 0. A set of mapping matrices for the K-party client is expressed as n is the number of entities corresponding to the global embedding, n k is the number of entities corresponding to the k-th client, (2) Use distance-based loss function for training and initialize local relationships The loss function formula is:

[0031]

[0032] Among them, (h, r, t) is the triple of the local knowledge graph (head node, relationship, tail node), (h, r, t' i ) is the i-th negative sample triplet selected from n negative samples, γ is a hyperparameter, p(h,r,t' i ) is the weight value of the negative sample, and its formula is

[0033]

[0034] The training reduces the distance between positive samples and increases the distance between negative samples; (3) The learnable attention coefficients and matrices in the local self-attention are assigned using random initialization. This completes the initialization operation of the federated knowledge graph model.

[0035] S2. The client downloads the global embedding on the central server, trains it using the local graph attention network, updates the entity embedding and relationship embedding using the client's triples and relationship embedding, and updates the graph attention network model parameters. The central server aggregates the updated entity embeddings of each client and updates the global embedding.

[0036] In parallel, the models of K clients are trained. The global entity embedding is downloaded to each client. The mapping matrix is ​​used to obtain the local entity embedding of each client. E rounds of training are performed in the local local training. The length of each round of local training is (1) Encoding is performed using the improved graph attention network, which integrates the whole graph structure. In each layer of the graph attention network, the nodes can be represented as

[0037]

[0038] Where l represents the lth layer, N(i) is the adjacent set of nodes i, and denote the relation embedding and adjacent node entity embedding of the lth layer respectively, and W dir is a learnable matrix, which can be divided into three categories:

[0039]

[0040] Loopr The edge type is self-loop, In r For edge type inward, Out r For edge type outward, α ij is the attention coefficient of the jth neighboring node of node i

[0041]

[0042] in,

[0043]

[0044] α T is the transposed learnable vector, and the relationship can be expressed as

[0045]

[0046] is a learnable matrix of l-layer relations. After multi-layer graph attention training, local entity embedding and relationship embedding that integrate the whole graph structure are obtained and input into the decoder; (2) decoding training is performed using a distance-based loss function; (3) local entity embedding and relationship embedding as well as graph attention network model parameters are updated.

[0047] After the above local training is completed, K clients upload their trained entity embeddings to the central server for aggregation. The aggregation operation is

[0048]

[0049] in, represents a vector of all 1s, Θ represents the division of the corresponding elements of the vector one by one, represents vector product, V k Represents the existence vector of the entity in the kth client in the global entity, n is the number of global entities, and the value is 1 when the client exists at the position of the global vector, and 0 when it does not exist.

[0050] S3. Repeat S2 for n rounds. K clients obtain the embedding parameters trained under federated learning and the local graph attention model parameters. The training of these parameters incorporates the effectiveness of other party’s data.

[0051] S4. By integrating the embedding parameters trained under federated learning with the embedding parameters based only on local training, K-squared obtains a robust local embedding model.

[0052] The embedding model trained under federated learning is a global embedding. For the local model, it is necessary to fuse the embedding based on the local training. The specific fusion steps are as follows: (1) After the client completes the model embedding training under federated learning and the local model embedding training, it connects the distance formulas based on triples in the two cases respectively.

[0053] f(h * ,r * ,t * )=-||h * +r * -t * ||,

[0054] f(h * ,r * ,t * ) is close to 0, indicating a small distance. * 、r * and t * are respectively the embedded vector representations after the trained improved graph attention network encoding, and f overall (h * ,r * ,t * ) represents the embedding vector function obtained by training under federated learning after being encoded by the local graph attention model under the corresponding global state, f part (h * ,r * ,t * ) represents the embedding vector function obtained based only on local data training after being encoded by the local graph attention model under the corresponding local area, which is

[0055] x=[f overall (h * ,r * ,t * );f part (h * ,r * ,t * )]

[0056] Perform a one-dimensional vector up and down concatenation operation; (2) Input the concatenated equation as a feature vector into a linear classifier

[0057] f final (h * ,r * ,t * )=α T x+b,

[0058] Among them, α Trepresents the learnable weight vector, b is the bias value, and a linear classifier is used through marginal sorting loss

[0059] L(h*,r*,t*)=max(0,δ-f final (h * ,r * ,t * )+f final (h * ,r * ,(t') * )

[0060] Training is performed, where δ represents the marginal value, (h * ,r * ,(t') * ) represents a negative sample, so that the positive triples are ranked higher than the negative triples. After the above operation, the weight parameters in the classifier are obtained after training, and the embedding under federated learning is integrated with the embedding based only on local data to form the final robust local embedding model.

[0061] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention is described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A multi-party knowledge graph embedding method based on federated learning, characterized in that: The method includes the following: S1. Build an entity central server to aggregate entity embedding data of multiple clients, aggregate the entities of K clients to the central server, randomly initialize the global entity embedding E0 of the central server, and K parties download E0 to their respective clients for pre-training to obtain the initialized local entity embedding and relation embedding parameters S2, the client downloads the global embedding from the central server, uses the local graph attention network for training, uses the client's triples and relationship embeddings to update the entity embedding and relationship embedding, and updates the graph attention network model parameters. The central server aggregates the entity embeddings updated by each client and updates the global embedding; S3. Repeat S2 for n rounds. K clients obtain the embedding parameters trained under federated learning and the local graph attention model parameters. The training of these parameters incorporates the utility of other parties’ data. S4. By integrating the embedding parameters trained under federated learning with the embedding parameters based only on local training, K-squared obtains a robust local embedding model.

2. The multi-party knowledge graph embedding method based on federated learning as claimed in claim 1, characterized in that: The specific steps of the S1 multi-party pre-training are: S11, the global entity embedding E0 downloaded from the central server is converted into a local entity embedding through the mapping matrix In the mapping matrix, the corresponding entity in the local and global is 1, and the non-corresponding entity is 0. A set of mapping matrices for the K-party client is expressed as n is the number of entities corresponding to the global embedding, n k is the number of entities corresponding to the k-th client, S12. Use distance-based loss function for training and initialize local relationships S13. The learnable attention coefficients and matrices in local self-attention are assigned using random initialization.

3. The multi-party knowledge graph embedding method based on federated learning as claimed in claim 1, characterized in that: The steps of training the S2 local graph attention network are: S21. Use the improved graph attention network for encoding in the local network and integrate it into the full graph structure of a single knowledge graph; S22, use distance-based loss function for decoding training; S23. Repeat steps S21 and S22 for e rounds to train and update local entity embedding, relationship embedding and graph attention network model parameters.

4. The multi-party knowledge graph embedding method based on federated learning as claimed in claim 1, characterized in that: The S4 fusion steps are: S41, after completing the model embedding training under federated learning and the local model embedding training, the client performs a connection operation on the triple-based distance formulas in the two cases respectively; S42. The concatenated formula is input as a feature vector into a linear classifier. The linear classifier is trained by marginal ranking loss so that the positive triples are ranked higher than the negative triples.

5. The multi-party knowledge graph embedding method based on federated learning as described in claims 2 and 3, characterized in that: The distance loss function is based on generating multiple groups of negative samples for each entity, and the triples of the knowledge graph are positive samples. The distances of the head entity, relationship and tail entity in the positive and negative samples are calculated. The training reduces the distance of the positive samples and increases the distance of the negative samples.

6. The multi-party knowledge graph embedding method based on federated learning as claimed in claim 3, characterized in that: The improved graph attention network focuses on the graph structure information of the node on the connected nodes and the relationship between them in each attention layer, and uses the distance vector of adjacent nodes and relationships to represent it. In the calculation of the attention coefficient, the distance vector of adjacent nodes and relationships is also paid attention to.

Citation Information

Cited By

  • Knowledge graph completion method, device, equipment and medium

    CN120598014A

  • A knowledge graph completion method, device, equipment and medium

    CN120598014B