Multi-Task Knowledge Graph Completion Method Based on Federated Learning

Through the multi-task knowledge graph completion method based on federated learning, the entity and relationship embedding representation of the central server is used to aggregate the client's entity and relationships, and efficient and secure knowledge graph completion under privacy protection is achieved, model deviation and privacy protection problems in multi-task training are solved, and model accuracy is improved.

CN115687640BActive Publication Date: 2025-07-11NINGBO UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211271214.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2025-07-11
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

The existing knowledge graph embedding method is only applicable to a single knowledge graph, and cannot be effectively applied in the case of joint training of multiple knowledge graphs, and there are privacy protection issues.

Method used

Using a multi-task knowledge graph completion method based on federated learning, two central servers are constructed to aggregate entities and relationship embedding representations respectively. The client uploads parameters for model fusion after local training, ensuring privacy protection while improving training efficiency and model accuracy.

Benefits of technology

It realizes the safe use of multiple knowledge graph information to complete under privacy protection, improves model training efficiency and accuracy, solves the problem of model deviation under the federal framework, and provides a knowledge graph completion method suitable for multi-tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115687640B_ABST
    Figure CN115687640B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task knowledge graph completion method based on federated learning. By constructing two central servers, which are respectively used to obtain the entity and relation embedding representations from each client after secure aggregation, and aggregating the aggregated component embedding representations and sending them back to each client. One round of operation of the knowledge graph completion method is the backpropagation after the parameter aggregation of the central server. After the knowledge graph is locally trained by the client, the parameters are uploaded to the central server for aggregation operation and then sent back to each client for iterative training. After the training of the knowledge graph embedding representation under the federated framework converges to the parameters, a model fusion operation is performed with the local knowledge graph training, and finally the fused knowledge graph model is output. The present invention efficiently and securely jointly utilizes the knowledge graph information provided by multiple clients under the premise of privacy protection, completes the knowledge graphs of each client, and can improve the performance of its downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for knowledge graph completion, in particular to a multi-task knowledge graph completion method based on federated learning. Background Art

[0002] With the rapid development of knowledge graph technology, knowledge graphs are used more and more widely among major enterprise organizations. The information of traditional knowledge graphs is often incomplete, and knowledge completion for them is a very important task. However, due to factors such as commercial interests and privacy protection, major enterprise institutions cannot publicly disclose knowledge graphs explicitly. Existing knowledge graph embedding methods represent the components of a knowledge graph as vectors in a continuous vector space and are effectively applied to knowledge graph completion technology. However, such methods are only applicable to a single knowledge graph and are not suitable for the case of joint training of multiple knowledge graphs. Therefore, there is an urgent need to develop a multi-task knowledge graph completion method applicable to joint training of multiple knowledge graphs. Summary of the Invention

[0003] The purpose of the present invention is to provide a multi-task knowledge graph completion method based on federated learning. On the premise of privacy protection, the present invention efficiently and securely jointly utilizes the knowledge graph information provided by multiple clients to complete the knowledge graphs of each client and improve the performance of its downstream tasks.

[0004] The technical solution of the present invention: A multi-task knowledge graph completion method based on federated learning includes the following steps:

[0005] Step S1, respectively construct two central servers for aggregating entity embedding representations and relationship embedding representations. The two central servers respectively aggregate the initial entity embedding representations and initial relationship embedding representations from each client, respectively create an entity table and a relationship table, randomly initialize the entity table and the relationship table, and distribute the aggregated entity embedding representations and relationship embedding representations to the corresponding clients, and enter step S2;

[0006] Step S2, each client downloads only the entity embedding representations and relationship embedding representations contained locally from the two central servers respectively, and enters step S3;

[0007] Step S3, each client uses a unified knowledge graph embedding representation method to perform training for a set number of epochs, updates the knowledge graph model locally, and enters step S4;

[0008] Step S4: The two central servers respectively aggregate the entity embedding representations and relationship embedding representations randomly selected from different clients, update these entity embedding representations and relationship embedding representations from different clients according to their respective parameter update rules, and then send them back to the corresponding clients through the entity table, and enter Step S5;

[0009] Step S5: Each client downloads the updated entity embedding representations and relationship embedding representations from the two central servers respectively, and determines whether the updated parameters converge;

[0010] If the parameters converge, enter Step S6;

[0011] If the parameters do not converge, return to Step S3;

[0012] Step S6: After the parameters in Step S5 converge, a knowledge graph model updated based on the federated framework is obtained, and a model fusion operation is performed on it and the knowledge graph model updated locally, and finally the fused knowledge graph model is output.

[0013] Compared with the prior art, the beneficial effects of the present invention are reflected in: by constructing two trusted central servers, one for aggregating entity information and the other for aggregating relationship information, the present invention can effectively avoid the leakage of triples. At the same time, the local client downloads all the entity and relationship embedding vector information only contained locally after aggregation processing from the central server, which greatly improves the training efficiency of the local model. After the knowledge graph is locally trained by the client, the parameters are uploaded to the central server for aggregation operation and then sent back to each client for iterative training, achieving the safe utilization of information from other knowledge graphs for knowledge graph completion under the premise of protecting privacy. In addition, the model fusion step of the present invention solves the model deviation caused by training based on the global minimum loss under the federated framework, further improving the model accuracy. The present invention provides an idea for the knowledge graph completion technology under the federated framework, and has the characteristics of strong privacy protection and high availability of the model results.

[0014] The two central servers respectively obtain the entity and relationship embedding representations after secure aggregation from each client, and send the aggregated component embedding representations back to each client after aggregation. Such a round of operation of the knowledge graph completion method is the parameter aggregation and backpropagation of the central server once.

[0015] In the foregoing multi - task knowledge graph completion method based on federated learning, the central server for aggregating entity embedding representations is the entity server. When the entity server performs the aggregation operation on entity embedding representations from each client for the first time in step S1, it records the unique entities from all different clients, establishes a complete entity table, maps the entities from different clients to the entity table, and after establishing and randomly initializing the entity table, distributes it to the corresponding clients;

[0016] After each client has undergone training for a set period, the entity server aggregates the entity embedding representations from different clients according to the entity table, updates these entity embedding representations from each client according to the central server update rule, and then sends them back to the corresponding clients through the entity table. Thus, the entity embedding representations of each client are updated.

[0017] In the foregoing multi - task knowledge graph completion method based on federated learning, the central server for aggregating relationship embedding representations is the relationship server. When the relationship server performs the aggregation operation on relationship embedding representations from each client for the first time in step S1, it records the unique relationships from all different clients, establishes a complete relationship table, maps the relationships from different clients to the relationship table, and after establishing and randomly initializing the relationship table, distributes it to the corresponding clients;

[0018] After each client has undergone training for a set period, the relationship server aggregates the relationship embedding representations from different clients according to the entity table, updates these entity relationship representations from each client according to the central server update rule, and then sends them back to the corresponding clients through the relationship table. Thus, the relationship embedding representations of each client are updated.

[0019] In the foregoing multi - task knowledge graph completion method based on federated learning, for the knowledge graph completion technology under the federated framework, it is required that each client use a specific and unified knowledge graph embedding representation method locally. This is because different knowledge graph embedding representation methods will result in non - uniform continuous vector spaces mapped by each client, mismatched parameters uploaded to the server for aggregation, and ultimately affect the training results of the model. Therefore, in the present invention, local knowledge graph representation learning is performed on all clients, and the parameters obtained after learning are uploaded to two different central servers for aggregation operations. The specific steps for a client to locally train the knowledge graph embedding representation (i.e., step S3) include:

[0020] Step S3.1: The client obtains the randomly initialized entity vector matrix and relationship vector matrix from the entity server and the relationship server respectively;

[0021] Step S3.2: For each triple on the client side, use a unified knowledge graph embedding representation method to train for a set number of epochs. Currently, the mainstream methods include TransE, RotatE, etc.

[0022] Step S3.3: After training locally for a set number of epochs, when receiving upload instructions from two central servers, upload the updated entity vector matrix to the entity server and upload the updated relation vector matrix to the relation server.

[0023] In the aforementioned multi-task knowledge graph completion method based on federated learning, the update processes on the server side and the client side are continuously iterated until the parameters converge. However, the learned knowledge graph embedding representation is based on the global minimum of all clients, not the local minimum of each client. The finally converged knowledge graph embedding representation is inaccurate for a single independent client. Therefore, through the model fusion step, the federated knowledge graph and the local knowledge graph are personalized integrated to make the learned knowledge graph embedding representation more accurate. The specific steps of model fusion (i.e., Step S6) include:

[0024] Step S6.1: After the client completes the training of the federated knowledge graph and the local knowledge graph, respectively perform a concatenation operation on the scoring functions corresponding to the vector embeddings of the knowledge graph triples locally and federally.

[0025] Step S6.2: Use the concatenated scoring function as a feature vector and input it into a linear classifier, and then train through margin loss ranking to obtain a scoring function based on the local minimum loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of the present invention;

[0027] Figure 2 is a framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The present invention will be further described below in conjunction with the drawings and embodiments, but it shall not be used as a basis for limiting the present invention.

[0029] Embodiment: A multi-task knowledge graph completion method based on federated learning, the flowchart is as Figure 1 shown, including the following steps:

[0030] Step S1: Construct two central servers for aggregating entity embedding representations and relationship embedding representations respectively. The two central servers aggregate the initial entity embedding representations and initial relationship embedding representations from each client respectively, create an entity table and a relationship table, randomly initialize the entity table and the relationship table, and distribute the aggregated entity embedding representations and relationship embedding representations to the corresponding clients, then enter Step S2.

[0031] The central server for aggregating entity embedding representations is the entity server. When the entity server performs the aggregation operation on the entity embedding representations from each client for the first time in Step S1, record the unique entities from all different clients, establish a complete entity table, map the entities of different clients to the entity table, and after establishing and randomly initializing the entity table, distribute it to the corresponding clients.

[0032] The central server for aggregating relationship embedding representations is the relationship server. When the relationship server performs the aggregation operation on the relationship embedding representations from each client for the first time in Step S1, record the unique relationships from all different clients, establish a complete relationship table, map the relationships of different clients to the relationship table, and after establishing and randomly initializing the relationship table, distribute it to the corresponding clients.

[0033] Step S2: Each client downloads the entity embedding representations and relationship embedding representations that are only locally contained from the two central servers respectively, and then enters Step S3.

[0034] Step S3: Each client uses a unified knowledge graph embedding representation method to perform training for a set number of epochs, and updates the knowledge graph model locally, then enters Step S4.

[0035] Step S3 specifically includes:

[0036] Step S3.1: The client obtains the randomly initialized entity vector matrix and relationship vector matrix from the entity server and the relationship server respectively;

[0037] Step S3.2: For each triple in the client, use a unified knowledge graph embedding representation method to perform training for a set number of epochs;

[0038] Step S3.3: After training for a set number of epochs locally, when receiving the upload instructions from the two central servers, upload the updated entity vector matrix to the entity server and upload the updated relationship vector matrix to the relationship server.

[0039] Step S4: The two central servers respectively aggregate the entity embedding representations and relationship embedding representations randomly selected from different clients, update these entity embedding representations and relationship embedding representations from different clients according to their respective parameter update rules, and then send them back to the corresponding clients through the entity table, and enter Step S5.

[0040] Step S5: Each client downloads the updated entity embedding representations and relationship embedding representations from the two central servers respectively, and determines whether the updated parameters converge;

[0041] If the parameters converge, enter Step S6;

[0042] If the parameters do not converge, return to Step S3.

[0043] Step S6: After the parameters in Step S5 converge, a knowledge graph model updated based on the federated framework is obtained, and a model fusion operation is performed on it and the knowledge graph model updated locally, and finally the fused knowledge graph model is output.

[0044] Step S6 specifically includes:

[0045] Step S6.1: After the client completes the training of the federated knowledge graph and the local knowledge graph training, respectively connect the scoring functions corresponding to the vector embedding representations of the knowledge graph triples locally and the vector embedding representations federally;

[0046] Step S6.2: Use the connected scoring function as the feature vector and input it into a linear classifier, and then train through margin loss ranking to obtain the scoring function based on the local minimum loss.

[0047] The framework schematic diagram of the present invention is as Figure 2 shown, and the content in Steps S1 to S6 is further explained based on the framework.

[0048] The entity server is responsible for aggregating the entity embeddings from different clients and sending the aggregated entity embeddings back to each client. The model iterates continuously, and according to the convergence rule, stops updating after the parameters converge.

[0049] The specific operations of the entity server are as follows:

[0050] 1) The entity server aggregates all the entities of the clients into the same entity table T and constructs a corresponding set of existence matrices and existence vectors where C represents the total number of clients, n represents how many entities are included in the entity table T in total, and n c represents how many entities are included in client c in total. The permutation matrix Pc Map the entity of client c to the entity table T. For example, It means that the i-th entity in the entity table T corresponds to the j-th entity of client c. It means that client c has the i-th element in the entity table T.

[0051] 2) The entity server randomly initializes the entity vector matrix according to all the entities existing in the entity table T where d e refers to the dimension size of the entity vector embedding representation. represents the corresponding vector space. Then, the entity server distributes the entity vector matrix E0 to the corresponding clients according to the existence matrix respectively. Client c downloads the entity vector from the entity server and updates the entity vector matrix

[0052]

[0053] 3) In the t-th round, the entity server randomly selects F×C clients from them, where F ∈ [0, 1], to construct a new entity set C t , after the specified entity server locally trains for p epochs, where p is a hyperparameter, the entity server aggregates the entity vectors from the specified clients.

[0054]

[0055] where represents the all-ones vector. represents element-wise division. represents element-wise multiplication.

[0056] 4) After the t-th round ends, the entity matrix of client c will be updated in the same way through the existence matrix P c to get updated.

[0057]

[0058] The operation of the entity server for the corresponding algorithm is as follows:

[0059] Input: Global entity table T, number of clients C, client entity set E c , existence matrix P c , existence vector V c , entity matrix E0.

[0060] Output: Entity matrices of each client after federated knowledge graph training

[0061] 1. Randomly initialize the entity matrix E0

[0062]

[0063]

[0064] The relationship server is responsible for aggregating the relationship embeddings from different clients and sending the aggregated relationship embeddings back to each client. The model iterates continuously and stops updating according to the convergence rule when the parameters converge.

[0065] The specific operations of the relationship server are as follows:

[0066] 1) The relationship server aggregates all the relationships of the clients into the same relationship table T * and constructs a corresponding set of existence matrices and existence vectors where C represents the total number of clients, and n * represents how many relationships are included in the relationship table T * in total, and represents how many relationships client c contains. The permutation matrix maps the relationships of client c to the relationship table T * .

[0067] 2) The relationship server randomly initializes the relationship vector matrix * according to all the relationships existing in the relationship table T where d e refers to the dimension size of the relationship vector embedding representation, and represents the corresponding vector space. Then, the relationship server distributes the randomly initialized relationship vectors to the corresponding clients according to the existence matrix The relationship vector matrix R0. Client c downloads the relationship vectors from the relationship server and updates the relationship vector matrix

[0068]

[0069] 3) In the t-th round, the relationship server randomly selects F×C clients from them, where F ∈ [0, 1], to construct a new relationship set C t , and after the specified relationship server locally trains for p epochs, where p is a hyperparameter, the relationship server aggregates the relationship vectors from the specified clients,

[0070]

[0071] where, represents the all-ones vector, and represents element-wise division, Denotes element-wise multiplication.

[0072] 4) After the end of the t-th round, the relationship matrix of client c will be updated in the same way through the existence matrix to obtain an update

[0073]

[0074] The algorithm corresponding to the relationship server operation is as follows:

[0075] Input: Global relationship table T * , the number of clients C, the client relationship set R c , the existence matrix existence vector relationship matrix R0.

[0076] Output: The relationship matrices of each client after the federated knowledge graph training

[0077] 1. Randomly initialize the relationship matrix R0

[0078]

[0079] Client update steps:

[0080] Client c obtains the randomly initialized entity vector matrix and relationship vector matrix

[0081] For each triple in client c Use a specific and unified knowledge graph embedding representation method for training. Different knowledge graph embedding representation models have different scoring functions f r (h, t), and the loss function of each triple (h, r, t) in client c is as follows:

[0082]

[0083] Among them, γ represents the margin value, (h, r, t i ′) represents the negative sample corresponding to the triple (h, r, t), m represents the number of negative samples, p(h, r, t i ′) represents the weight of the self-adversarial negative sample (h, r, t i ′), and is defined as follows:

[0084]

[0085] Among them, α represents the temperature parameter.

[0086] After training for a set period on the client side, the updated entity vector matrix is obtained. and the relation vector matrix will be uploaded to the entity server and the relation server respectively, where t represents the communication round.

[0087] The algorithm for client update is as follows:

[0088] Input: Entity vector matrix E c , client vector matrix R c , local knowledge graph

[0089] Output: Entity vector matrix E after local training c′ and client vector matrix R c′ .

[0090]

[0091]

[0092] Model fusion steps:

[0093] The central server aggregation operation and the client update operation are iterated until the model converges. However, the vector embedding representation learned by the knowledge graph under the federated framework is based on the global minimum loss, not the local minimum loss of each client's knowledge graph. The goal of this process is to integrate the information learned from the federated knowledge graph and the information of the local knowledge graph to minimize the local loss of the client.

[0094] First, after completing the training based on the federated knowledge graph and the local knowledge graph, each triple obtains two different embedding representations, one representing the global model information and the other representing the local model information of the client. According to a specific knowledge graph embedding representation model, two scoring functions and are obtained. Connecting the two scoring functions results in a feature vector. Among them, the scoring function utilizes the knowledge graph entity and relation vector representations obtained from local training, and the scoring function represents the knowledge graph entity and relation vector representations obtained from training in the federated scenario.

[0095] Then, the connected feature vector is input into a linear classifier, and the result is output as a scoring function:

[0096] s (h,r,t) = Wx + b,

[0097]

[0098] Among them, W represents the weight matrix, b represents the bias, and [;] represents the connection operation between the scoring functions.

[0099] This linear classifier is trained by the margin ranking loss, making the ranking of positive triples higher than that of negative triples. The loss of model fusion is as follows:

[0100] L f (h, r, t) = max(0, β - s (h,r,t) + s (h,r,t′) ),

[0101] where β represents the margin value, and (h, r, t′) represents the negative sample triple. The training objective of the model fusion loss process is to minimize the model fusion loss of all triples.

[0102] The algorithm corresponding to model fusion is as follows:

[0103] Input: Locally trained scoring function Federated training scoring function

[0104] Output: Final scoring function s (h,r,t) .

[0105]

[0106] The above is only the preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A multi-task knowledge graph completion method based on federated learning, characterized in that, Including the following steps: Step S1: Construct two central servers for aggregating entity embedding representations and relationship embedding representations respectively. The two central servers respectively aggregate the initial entity embedding representations and initial relationship embedding representations from each client, create an entity table and a relationship table respectively, randomly initialize the entity table and the relationship table, and distribute the aggregated entity embedding representations and relationship embedding representations to the corresponding clients, then enter step S2; Step S2: Each client downloads the entity embedding representations and relationship embedding representations only contained locally from the two central servers respectively, and then enters step S3; Step S3: Each client uses a unified knowledge graph embedding representation method to perform training for a set number of cycles, and updates the knowledge graph model locally, then enters step S4; Step S4: The two central servers respectively aggregate the entity embedding representations and relationship embedding representations randomly selected from different clients, update these entity embedding representations and relationship embedding representations from different clients according to their respective parameter update rules of the two central servers, and then send them back to the corresponding clients through the entity table, and enter step S5; Step S5: Each client downloads the updated entity embedding representations and relationship embedding representations from the two central servers respectively, and judges whether the updated parameters converge; If the parameters converge, then enter step S6; If the parameters do not converge, then return to step S3; Step S6: After the parameters in step S5 converge, obtain a knowledge graph model updated based on the federated framework, perform a model fusion operation on it and the knowledge graph model updated locally, and finally output the fused knowledge graph model.

2. The multi-task knowledge graph completion method based on federated learning according to claim 1, wherein: The central server for aggregating entity embedding representations is the entity server. When the entity server performs the aggregation operation on the entity embedding representations from each client for the first time in step S1, record the unique entities from all different clients, establish a complete entity table, map the entities of different clients to the entity table, and after establishing and randomly initializing the entity table, distribute it to the corresponding clients.

3. The multi-task knowledge graph completion method based on federated learning according to claim 2, wherein: The central server for aggregating relationship embedding representations is the relationship server. When the relationship server performs the aggregation operation on the relationship embedding representations from each client for the first time in step S1, record the unique relationships from all different clients, establish a complete relationship table, map the relationships of different clients to the relationship table, and after establishing and randomly initializing the relationship table, distribute it to the corresponding clients.

4. The multi-task knowledge graph completion method based on federated learning according to claim 3, characterized in that: The specific steps of step S3 include: Step S3.1: The client respectively obtains the randomly initialized entity vector matrix and relationship vector matrix from the entity server and the relationship server; Step S3.2: For each triple in the client, use a unified knowledge graph embedding representation method to perform training for a set number of cycles; Step S3.3: After locally performing training for a set number of cycles, when receiving the upload instructions from the two central servers, upload the updated entity vector matrix to the entity server and upload the updated relationship vector matrix to the relationship server.

5. The multi-task knowledge graph completion method based on federated learning according to claim 4, wherein: The specific steps of step S6 include: Step S6.1: After the client completes the training of the federated knowledge graph and the local knowledge graph training, the client respectively performs a connection operation on the scoring functions corresponding to the vector embedding representations of the knowledge graph triples locally and federally. Step S6.2: The scoring function after the connection operation is used as a feature vector and input into a linear classifier, and then trained by margin loss ranking to obtain a scoring function based on local minimum loss.

Citation Information

Patent Citations

  • Knowledge graph representation method based on federal learning

    CN113886598A

  • Updating method and device of video processing model, electronic equipment and storage medium

    CN114844889A