Method and device for multi-party joint training of graph learning model
The full-scale label data and homoproton graph are constructed through the federated label propagation algorithm, which solves the problem of graph reconstruction attacks in federated incomplete graph learning, and realizes the protection of data privacy and the maintenance of model availability.
Patent Information
- Application Number
- CN202510359310.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing federal non-complete graph learning framework guarantees data privacy at the computing level, but there is a risk of privacy leakage caused by graph reconstruction attacks and cannot effectively resist the leakage of graph topology data.
Full-scale label data is constructed through the federated label propagation algorithm, completely homoproton graphs are generated, and their connected edges are added to the original edge set to form an obfuscated edge set, which is used to jointly train the graph to learn the model, resist graph reconstruction attacks, and maintain the usability of the model.
Effectively resist graph reconstruction attacks, significantly maintain the availability of graph learning models, and ensure data privacy is not leaked.
Smart Images

Figure CN120296782A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the technical field of data security processing, and in particular to a method and apparatus for jointly training a graph learning model by multiple parties, a computer-readable storage medium, and a computing device. Background Art
[0002] Graph data has shown significant application potential in many fields such as recommendation systems and financial fraud detection. However, small enterprises often face difficulties in data acquisition. Especially in the startup period or when the scale is limited, it is difficult to establish an effective association network among target users in a timely manner. In contrast, platform enterprises own and operate a large amount of graph topologies. However, due to strict privacy protection regulations, these large enterprises cannot easily provide or share their topology data with external demanders.
[0003] Federated Incomplete Graph Learning (FedIGL) provides a feasible solution for secure collaboration among institutions. Even if the graph topology and node features are scattered among different parties, technologies such as homomorphic encryption can be used to ensure data privacy and achieve secure collaboration of multi-party data. This gives small enterprises the opportunity to use graph topologies to improve their capabilities, thereby providing better services for target users.
[0004] However, current federated incomplete graph learning only guarantees data privacy at the computational level. Recent research has found that there is also a privacy risk of reconstructing the graph topology using node representations in federated graph learning. Therefore, the embodiments of this specification disclose an improved solution to further eliminate or reduce the privacy leakage risk caused by graph reconstruction attacks on the basis of the federated incomplete graph learning framework. Summary of the Invention
[0005] The embodiments of this specification describe a method and apparatus for jointly training a graph learning model by multiple parties. Based on homogeneity measurement, the graph data is obfuscated, and the graph learning model obtained by training through the obfuscated edge set can not only effectively resist graph reconstruction attacks but also significantly maintain the usability of the graph learning model.
[0006] According to a first aspect, there is provided a method for jointly training a graph learning model by multiple parties, where the multiple parties include a first business party that holds the original label data of some nodes in the full set of target nodes, and a second party that holds the original edge set of the full set of target nodes. The method is applied to the first party and includes:
[0007] Based on the original label data, jointly calculate the full - scale label data with the second party based on the original edge set, which includes the pseudo - label data of another part of the nodes in the full - scale target nodes. Based on the full - scale label data, construct multiple connecting edges, where any two nodes connected by a connecting edge have the same label. Send the multiple connecting edges to the second party for it to add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
[0008] In some embodiments, the first party also holds the feature data of the full - scale target nodes.
[0009] In some embodiments, the federated calculation includes several rounds of iterative updates for the plaintext label matrix, and the updated plaintext label matrix corresponds to the full - scale label data. Any round of iterative update includes:
[0010] Encrypt the current plaintext label matrix using the public key of the first party to obtain a ciphertext label matrix; any i - th row or column in the plaintext label matrix is the one - hot encoding corresponding to the label of the i - th node in the full - scale target nodes. Send the ciphertext label matrix to a trusted third party for it to perform a homomorphic ciphertext multiplication process on the ciphertext label matrix and the ciphertext adjacency matrix, where the ciphertext adjacency matrix is determined by the second party encrypting the adjacency matrix corresponding to the original edge set using the public key. Decrypt the result matrix ciphertext of the homomorphic ciphertext multiplication process received from the trusted third party using the private key corresponding to the public key to obtain a result matrix plaintext. Update the current plaintext label matrix using the result matrix plaintext.
[0011] In some embodiments, the federated calculation includes several rounds of iterative updates for the plaintext label matrix, and the updated plaintext label matrix corresponds to the full - scale label data. Any round of iterative update includes:
[0012] Encrypt the current plaintext label matrix using the public key of the first party to obtain a ciphertext label matrix; any i - th row or column in the plaintext label matrix is the one - hot encoding corresponding to the label of the i - th node in the full - scale target nodes. Send the ciphertext label matrix to the second party for it to perform a homomorphic plain - ciphertext multiplication process on the ciphertext label matrix and the adjacency matrix corresponding to the original edge set. Decrypt the result matrix ciphertext of the homomorphic plain - ciphertext multiplication process received from the second party using the private key corresponding to the public key to obtain a result matrix plaintext. Update the current plaintext label matrix using the result matrix plaintext.
[0013] Further, in some specific embodiments, when any round of iterative update is the first round, before encrypting the current plaintext label matrix with the public key of the first party, the method further includes: initializing the rows or columns in the plaintext label matrix corresponding to the partial nodes based on the original label data; initializing the rows or columns in the plaintext label matrix corresponding to the other partial nodes with all-zero vectors.
[0014] In some specific embodiments, updating the current plaintext label matrix with the result matrix plaintext includes: for any i-th row in the plaintext label matrix, determining the column number j where the largest element in this row is located, and updating the value at the j-th column in this row to a predetermined positive value, and updating the values at other columns to 0.
[0015] In some embodiments, the original label data involves K category labels; wherein, based on the full amount of label data, constructing multiple connection edges includes: based on the full amount of label data, dividing the full amount of target nodes into K groups corresponding to the K category labels; for each group of nodes, establishing connection edges between each pair of nodes therein to obtain corresponding subsets of connection edges; sampling based on the K subsets of connection edges to obtain the multiple connection edges.
[0016] Further, in some specific embodiments, sampling based on the K subsets of connection edges to obtain the multiple connection edges includes: determining the union of the K subsets of connection edges; calculating the sampling ratio, which is inversely proportional to the number of connection edges in the union and directly proportional to the number of nodes in the full amount of target nodes; sampling the ratio of connection edges from the union as the multiple connection edges.
[0017] According to a second aspect, a method for multi-party joint training of a graph learning model is provided, where the multi-party includes a first party that holds the original label data of some nodes in the full amount of target nodes, and further includes a second party that holds the original edge set of the full amount of target nodes. The method is applied to the second party and includes:
[0018] Based on the original edge set, performing federated calculation with the first party based on the original label data, so that the first party obtains the full amount of label data, which includes the pseudo-label data of the other part of the nodes in the full amount of target nodes. Receiving from the first party multiple connection edges constructed based on the full amount of label data, where any two nodes connected by a connection edge have the same label. Adding the multiple connection edges to the original edge set to obtain a confused edge set for jointly training the graph learning model with the first party.
[0019] According to a third aspect, there is provided a method for multi-party joint training of a graph learning model, where the multi-party includes a first party that holds the feature data of all target nodes and the original label data of some of the nodes, and also includes a second party that holds the original edge set of all the target nodes. The method applied to the first party includes:
[0020] Based on the feature data and the original label data of the partial nodes, train a node classifier, and use the node classifier to determine the pseudo-label data of another part of the nodes in all the target nodes. Based on the original label data and the pseudo-label data, construct multiple connecting edges, where any two nodes connected by a connecting edge have the same label. Send the multiple connecting edges to the second party for it to add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training the graph learning model with the first party.
[0021] According to a fourth aspect, there is provided a method for multi-party joint training of a graph learning model, where the multi-party includes a first party that holds the feature data of all target nodes and the original label data of some of the nodes, and also includes a second party that holds the original edge set of all the target nodes. The method applied to the second party includes:
[0022] Receive from the first party multiple connecting edges constructed based on all the label data, where any two nodes connected by a connecting edge have the same label; where the all label data includes the original label data and the pseudo-label data of another part of the nodes in all the target nodes determined by a node classifier, and the node classifier is trained based on the feature data and the original label data of the partial nodes. Add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training the graph learning model with the first party.
[0023] According to a fifth aspect, there is provided a device for multi-party joint training of a graph learning model, where the multi-party includes a first party that holds the original label data of some of the nodes in all the target nodes, and also includes a second party that holds the original edge set of all the target nodes. The device is integrated in the first party and includes:
[0024] A federated label propagation module configured to obtain all the label data including the pseudo-label data of another part of the nodes in all the target nodes through federated computing with the second party based on the original edge set based on the original label data. A connecting edge construction module configured to construct multiple connecting edges based on all the label data, where any two nodes connected by a connecting edge have the same label. A connecting edge sending module configured to send the multiple connecting edges to the second party for it to add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training the graph learning model with the first party.
[0025] According to a sixth aspect, there is provided an apparatus for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the original label data of some nodes in the full set of target nodes, and a second party that holds the original edge set of the full set of target nodes. The apparatus is integrated in the second party and includes:
[0026] A federated label propagation module, configured to perform federated computation with the first party based on the original edge set and the original label data, so that the first party obtains the full set of label data, including the pseudo-label data of another part of the nodes in the full set of target nodes. A connected edge receiving module, configured to receive from the first party a plurality of connected edges constructed based on the full set of label data, where the two nodes connected by any one of the connected edges have the same label. An edge set obfuscation module, configured to add the plurality of connected edges to the original edge set to obtain an obfuscated edge set for jointly training the graph learning model with the first party.
[0027] According to a seventh aspect, there is provided an apparatus for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the feature data of the full set of target nodes and the original label data of some of the nodes, and a second party that holds the original edge set of the full set of target nodes. The apparatus is integrated in the first party and includes:
[0028] A classifier training module, configured to train a node classifier based on the feature data and the original label data of some of the nodes, and use the node classifier to determine the pseudo-label data of another part of the nodes in the full set of target nodes. A connected edge construction module, configured to construct a plurality of connected edges based on the original label data and the pseudo-label data, where the two nodes connected by any one of the connected edges have the same label. A connected edge sending module, configured to send the plurality of connected edges to the second party for it to add the plurality of connected edges to the original edge set to obtain an obfuscated edge set for jointly training the graph learning model with the first party.
[0029] According to an eighth aspect, there is provided an apparatus for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the feature data of the full set of target nodes and the original label data of some of the nodes, and a second party that holds the original edge set of the full set of target nodes. The apparatus is integrated in the second party and includes:
[0030] A connection edge receiving module, configured to receive, from the first party, multiple connection edges constructed based on full-scale label data, where two nodes connected by any connection edge have the same label; wherein, the full-scale label data includes the original label data and pseudo-label data of another part of nodes in the full-scale target nodes determined by a node classifier, and the node classifier is trained based on the feature data and original label data of the part of nodes. An edge set obfuscation module, configured to add the multiple connection edges to the original edge set to obtain an obfuscated edge set for jointly training a graph learning model with the first party.
[0031] According to a ninth aspect, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, causes the computer to execute the method provided in any one of the first aspect to the fourth aspect.
[0032] According to a tenth aspect, there is provided a computing device including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method provided in any one of the first aspect to the fourth aspect is implemented.
[0033] In summary, by using the above methods and devices disclosed in the embodiments of this specification, first construct full-scale label data using the federated label propagation algorithm, and then construct a completely homogeneous subgraph using the full-scale label data, so as to add the connection edges sampled based on the completely homogeneous subgraph to the original edge set to complete obfuscation. The graph learning model obtained by jointly training via the obfuscated edge set can not only effectively resist graph reconstruction attacks but also significantly maintain the usability of the graph learning model. Description of the Drawings
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0035] Figure 1 Schematic diagram of the architecture of the multi-party joint training graph learning model disclosed in the embodiments of this specification;
[0036] Figure 2 Schematic diagram of the interaction of a federated label propagation algorithm in any round of iteration disclosed in the embodiments of this specification;
[0037] Figure 3 Schematic diagram of the interaction of another federated label propagation algorithm in any round of iteration disclosed in the embodiments of this specification;
[0038] Figure 4Schematic diagram of the device structure for jointly training a graph learning model disclosed in the embodiments of this specification. This device is integrated into the business party;
[0039] Figure 5 Schematic diagram of the device structure for jointly training a graph learning model disclosed in the embodiments of this specification. This device is integrated into the service party;
[0040] Figure 6 Schematic diagram of another device structure for jointly training a graph learning model disclosed in the embodiments of this specification. This device is integrated into the business party;
[0041] Figure 7 Schematic diagram of another device structure for jointly training a graph learning model disclosed in the embodiments of this specification. This device is integrated into the service party. Detailed implementation manners
[0042] The following describes the solution provided in this specification in conjunction with the accompanying drawings.
[0043] As mentioned above, FedIGL has the risk of graph topology data leakage due to graph reconstruction attacks. For better understanding, a brief introduction to the basic framework of FedIGL is given first. Taking two-party joint learning as an example, Party A is the entity with node features and labels, and Party B is the platform with the graph topology required by Party A. In this setting, FedIGL realizes the following process:
[0044] 1) Encryption and transmission: Party A encrypts the feature vectors of the nodes and transmits them to Party B.
[0045] 2) Message passing: Party B uses the adjacency matrix and the deployed graph learning model to perform message passing operations on the encrypted feature vectors to obtain the embedding representation of the nodes. It should be understood that the adjacency matrix is determined based on the graph topology, and any element in the i-th row and j-th column indicates whether there is a connection edge between the i-th node and the j-th node.
[0046] 3) Embedding return: The embedding representation output by the graph learning model is sent back to Party A for prediction.
[0047] Based on the exemplary introduction of FedIGL, the graph reconstruction attack can be understood: After the privacy thief obtains the embedding representation of the nodes through the graph learning model, the privacy thief can use the distance between the node embeddings, such as the Euclidean distance, to infer whether there is a connection between two nodes, thereby achieving the purpose of reconstructing the topology, which leads to the privacy leakage of the graph topology data.
[0048] To address graph reconstruction attacks, the industry has proposed random obfuscation strategies and random walk obfuscation strategies. Both of these strategies are privacy protection strategies for the anonymous release of graph data. The former determines whether an edge should be added or removed between a pair of nodes through random probability, which has the problem of significantly reducing the usability of graph learning models. The latter walks between nodes and randomly adds / deletes edges among multi-hop neighbors to achieve the purpose of privacy protection. It has a strong correlation with the dataset. However, for graph datasets with little neighbor relationship, it will also significantly reduce the usability of graph learning models.
[0049] The applicant's research found that the main reason why the above two strategies cannot maintain model usability is that the pure random addition and deletion of the original edge set cannot maintain the homophily level of the obfuscated edge set. It should be noted that homophily was originally a sociological concept that describes the tendency of individuals to establish connections and relationships with others who are similar to themselves in various dimensions.
[0050] In the context of graph data, by introducing homophily, the possibility of nodes with the same label being connected or close to each other is quantified. Specifically, by introducing homophily into the edge set of graph data, it can measure the proportion of a type of edge in the edge set where the connected nodes have the same label. See the following formula (1) for this, which can reflect the degree of closeness between the edge and the classification label.
[0051]
[0052] Among them, Measure the homophily of the edge set ε, (v, w) represents the connecting edge between node v and node w, y v and y w respectively represent the labels of node v and node w, |·| represents the number of elements in the set, and ∧ represents logical AND.
[0053] Intuitively, homogeneous connections in the graph topology promote the flow of useful information between nodes with similar labels, enabling the model to more effectively identify representative features. Therefore, the higher the homophily of the edge set, the better the classification performance of the nodes.
[0054] Based on the above observations and analyses, the applicant proposes to design an improved obfuscation strategy based on homophily measurement. This strategy includes: first calculating the homophily of the edge set, and then adding new connecting edges that can maintain or increase the homophily of the edge set to the original edge set according to the calculation results, so as to effectively maintain the usability of the model while obfuscating the edge set to resist replay attacks.
[0055] However, the implementation of the improved obfuscation strategy faces the following challenges:
[0056] a. Lack of full labels. In the node classification task, the reality is that the business requester (or simply referred to as the business side) only has labels for some of the target nodes, while calculating homogeneity requires labels for all target nodes (or the full set of target nodes). It should be understood that in this article, the target nodes refer to the nodes that the business side is concerned about, such as target users or target products.
[0057] b. Distribution of label data and edge data. In federated learning, it is crucial to consider data distribution and protect data privacy. However, calculating homogeneity requires simultaneous access to label data and edge data, which may conflict with privacy restrictions.
[0058] Furthermore, the applicant proposes a federated label propagation algorithm in the improved solution to address the above challenges. Specifically, Figure 1 The implementation architecture of the improved solution is shown. Under this architecture:
[0059] First, in step S110, the first party ( Figure 1 indicated as the business side in the figure) and the second party ( Figure 1 indicated as the service side in the figure) jointly execute the federated label propagation algorithm based on the original labels of some of the target nodes and the original edge set of all target nodes respectively. Thus, the first party can obtain the full labels, which include the pseudo-labels of the other part of the nodes. In this way, the above challenges a and b can be overcome. Then, in step S120, the first party constructs multiple connected edges where the connected nodes have the same label based on the full labels and provides them to the second party in step S130. Next, in step S140, the second party adds the multiple connected edges to the original edge set to complete the obfuscation of the edge set while maintaining the homogeneity of the edge set. Thereby, the first party and the second party can jointly perform FedIGL learning based on the features of all target nodes and the original labels of some target nodes and the obfuscated edge set to obtain a trained graph learning model, which can resist graph reconstruction attacks and maintain high availability.
[0060] It should be noted that the first party and the second party in this article are generally two independent parties, and both parties can be implemented as any device, platform, server, or device cluster with computing and processing capabilities. In addition, the first party and the business side, as well as the second party and the service side, can be used interchangeably in this article, although the first party is not limited to the business side and the second party is not limited to the service side.
[0061] Next, the data held by the business side and the service side respectively will be introduced first, and then the above implementation steps in the improved solution will be elaborated.
[0062] 1) The business side holds the original label data of some of the nodes in the full set of target nodes.
[0063] It can be understood that the full set of target nodes can be determined through joint negotiation between the business side and the service side. For example, the two parties execute a Private Set Intersection (PSI) protocol to obtain the intersection between the set of business objects in the business side and the set of business objects involved in the massive graph topology data of the service side, and then use the elements in the intersection as the full set of target nodes.
[0064] The business objects corresponding to the target nodes can be users, commodities, or events. Exemplarily, the events can specifically be transaction events, login events, access events, etc. The original label data is adapted to the business objects. For example, assuming the business object is a user or an event, the original label data can be user risk labels or event risk labels. Another example is that assuming the business object is a commodity, the original label data can be the usage scenario category or popularity level of the commodity, etc.
[0065] Typically, the business side also holds the feature data of the full set of target nodes. At this time, the business side can also be referred to as the feature side. In one example, the feature data includes the portrait features and network behavior features of users (such as shopping preferences, transaction preferences, or social preferences, etc.). In another example, the feature data includes the selling price, sales volume, origin, shelf life, etc. of commodities. In yet another example, the feature data includes the occurrence time, occurrence location (such as IP location), transaction amount, trading parties, etc. of transaction events.
[0066] Of course, the improved solution also supports this atypical scenario: the business side only holds label data, and the feature data is held by other parties.
[0067] 2) The service side holds the original edge set of the full set of target nodes.
[0068] It should be understood that connection edges are established between target nodes due to a predetermined association relationship. For example, there are connection edges between user nodes due to social relationships. Another example is that there are connection edges between commodity nodes purchased by the same user. Another example is that there are connection edges between transaction events involving the same trading party. In addition, any edge in the original edge set can be represented by the two nodes it connects; the service side can also be referred to as the topology side.
[0069] The data in the business side and the service side are introduced above. Next, the above-mentioned implementation steps in the improved solution are introduced in detail. It should be understood that all parties including the business side and the service side in the text can be implemented as any device, equipment, platform, or server cluster with computing and processing capabilities, etc. Figure 1 The detailed introduction of the method steps is as follows:
[0070] First, in step S110, the business side and the service side jointly execute the federated label propagation algorithm. The input of the algorithm includes the original label data of some nodes in the business side and the original edge set of all target nodes in the service side. The output of the algorithm includes the full label data of all target nodes. It should be understood that the full label data includes the pseudo-label data of another part of the nodes in all target nodes except some nodes.
[0071] To help understanding, the following separately introduces how the federated label propagation algorithm overcomes the above two challenges a) and b).
[0072] To overcome the above challenge a, that is, to overcome the lack of full labels, a label propagation algorithm for graph structure is proposed (note that the federated framework is not introduced here for the time being). This algorithm is very efficient and can control the label propagation to be completed within linear time.
[0073] Given the original edge set ε and the original label data The label propagation algorithm includes two basic links, propagation and update. The algorithm process is similar to the message passing mechanism commonly used in graph algorithms.
[0074] 1) Propagation:
[0075] For the original edge set ε of all target nodes, self-loop edges of each node in all target nodes can be added to it, and the obtained edge set is denoted as To help achieve the full transmission of labels subsequently. Then, is converted into the corresponding adjacency matrix
[0076] For the original label data of some target nodes Determine the one-hot encoding vectors corresponding to the original labels of each node in this part of the target nodes. Exemplarily, for the original label data The number of label categories involved is denoted as K. Then, assuming that the original label of a certain node is the j-th label category, its corresponding one-hot encoding vector can be determined as a K-dimensional vector in which the j-th element is a predetermined positive number (such as 1) and other elements are all 0. Thus, the label matrix with N (the number of nodes of all target nodes) rows and K columns can be initialized using the determined one-hot encoding vectors For the other part of the target nodes without labels, their corresponding one-hot encoding vectors in the label matrix are initialized as all-zero vectors.
[0077] Furthermore, for the t-th round of propagation, the following formula can be used for calculation:
[0078]
[0079] Where, Indicates the propagation result that has not been updated, and
[0080] 2) Update:
[0081] It should be noted that from the analysis in (2), the essence of propagation is that nodes that are neighbor nodes among all target nodes propagate labels to each other. The result of propagation is that for the i-th row (the i-th one-hot encoded vector) corresponding to the i-th node in after propagation (implemented by multiplying the adjacency matrix ), the i-th row in
[0082] is obtained, and the j-th element in it is the number of neighbor nodes with the j-th label category among its neighbor nodes.
[0083]
[0084] where I max includes the index of the maximum value in each row of
[0085] corresponding to the most likely label for each node. max Then, by converting the index data I
[0086]
[0087] where represents the one-hot encoded form of I max back to a one-hot encoded vector, the following updated label matrix is obtained:
[0088] By repeating the above propagation and update process until a predetermined number of iterations is reached, or until convergence, for example, the label matrix no longer changes, an iterated or updated label matrix can be obtained, which corresponds to the determined all-label data. It should be noted that for a certain target node with an original label, after executing the label propagation algorithm, there is a situation where the label changes. The changed label has a higher confidence level, which to a certain extent realizes the error correction of the original label.
[0089] The above introduction is about the label propagation algorithm proposed to address the challenge a of lacking all labels.
[0090] In response to Challenge b above, that is, the label data and edge data required by the above label propagation algorithm are distributed among different parties, a response method b1 is proposed. The above calculation is entrusted to a server of a trusted third party (Trusted Third Server, TTS), and homomorphic encryption is used to achieve data privacy protection. It should be understood that TTS can be a computing platform controlled by an authoritative institution or an institution with credibility. Alternatively, it can also be a cloud service provider that provides a trusted execution environment.
[0091] The following provides the pseudocode of Algorithm 1 (that is, the federated label propagation algorithm in Response Method b1).
[0092]
[0093] The following combines Figure 2 to introduce the pseudocode in Algorithm 1. The federated label propagation algorithm involves several (referring to one or more) rounds of iterative updates of the label matrix. Any iterative update in the t-th round includes Figure 2 the following interaction steps shown in
[0094] Step S201, the label client uses the public key pk generated by itself to encrypt the current plaintext of the label matrix to obtain the corresponding ciphertext of the label matrix and send it to TTS.
[0095] It should be noted that for the generation of the public key pk and its corresponding private key sk, and the sharing of the public key pk, reference can be made to lines 1-3 in Algorithm 1. It can be understood that all encryption operations in Algorithm 1 are completed using the public key pk.
[0096] For the current plaintext of the label matrix In a possible case, the current round is the first round. At this time is can assign the in the above text to (see the description related to formula (2)), and then encrypt it. For this case, reference can be made to lines 9, 10, and 11.
[0097] In another possible case, the current round is not the first round. At this time, the current plaintext of the label matrix is the plaintext of the label matrix after the previous round of iteration. At this time, it can be directly encrypted and sent to TTS. For this case, reference can be made to the sending in line 14 Temporarily do not need to look at the calculation formula therein.
[0098] As described above, TTS can obtain the ciphertext of the label matrix provided by the label client in this round of iteration.
[0099] Step S202, TTS performs homomorphic ciphertext multiplication on the ciphertext of the label matrix and the ciphertext of the adjacency matrix to obtain the ciphertext of the result matrix and returns it to the label client.
[0100] It should be noted that for the acquisition of the ciphertext of the adjacency matrix in TTS , reference can be made to Figure 2 the steps S200 shown therein, as well as lines 4 - 6 of the code.
[0101] In addition, for the execution of this step, reference can be made to lines 17 and 18 of the code. It should be understood that TTS uses the received from the label client in this round of iteration as the
[0102] in the calculation formula in line 17.
[0103] As described above, the label client can obtain the ciphertext of the result matrix returned by TTS in this round of iteration Step S203, the label client uses the private key sk to decrypt the ciphertext of the result matrix to obtain the plaintext of the result matrix and uses the plaintext of the result matrix
[0104] to obtain the updated current plaintext of the label matrix.
[0105] Furthermore, after the label client obtains the updated through this round of iteration in this step, it can increment the counting variable t by 1, that is, t = t + 1. After that, when it is determined that t < T, a new round of iteration can be continued; otherwise, the iteration is terminated because the finally updated plaintext of the label matrix has been obtained
[0106] As described above, the federated label propagation algorithm (Algorithm 1) introducing a trusted third - party server (TTS) is introduced. By executing Algorithm 1, attacks under a malicious model can be resisted, where the malicious model refers to an attacker (such as a service provider) deviating from the protocol at will, actively tampering with data or deceiving other participants (such as a business party).
[0107] The applicant also proposes another response method b1, which includes executing a federated label propagation algorithm that does not introduce TTS (hereinafter referred to as Algorithm 2). Algorithm 2 can resist attacks under the Semi-honest Model, where the Semi-honest Model means that the attacker follows the protocol process but tries to obtain additional information from the process.
[0108] The following combines Figure 3 to introduce the implementation of Algorithm 2. Algorithm 2 involves several rounds of iterative updates of the label matrix, and the iterative update of any t-th round includes Figure 3 the following interaction steps shown in
[0109] Step S301, the business party uses the public key pk generated by itself to encrypt the current plaintext of the label matrix to obtain the corresponding ciphertext of the label matrix and sends it to the service party.
[0110] Step S302, the service party performs homomorphic ciphertext multiplication on the ciphertext of the label matrix and the ciphertext of the adjacency matrix to obtain the ciphertext of the result matrix and returns it to the business party.
[0111] Step S303, the business party uses the private key sk to decrypt the ciphertext of the result matrix to obtain the plaintext of the result matrix and uses the plaintext of the result matrix to obtain the updated current plaintext of the label matrix
[0112] It should be noted that the main difference between Algorithm 2 and Algorithm 1 is that the service party executes the steps originally performed by TTS in Algorithm 1. In addition, for Figure 3 the introduction of each step in
[0113] Above, in Figure 1 the step S110 shown, the business party, based on the original labels of some nodes, and the service party, based on the original edge set of all target nodes, obtain the updated plaintext of the label matrix through federated calculation which indicates the label data of all target nodes, that is, all label data.
[0114] Then, in step S120, the business party constructs multiple connecting edges based on the all label data, where any two nodes connected by a connecting edge have the same label.
[0115] Specifically, after determining the full set of label data required for calculating homogeneity, K completely homogeneous subgraphs corresponding to K label categories are generated, where all nodes in each subgraph have the same label and there are connecting edges established between every two of them.
[0116] Further, in one implementation, all the connecting edges corresponding to the K completely homogeneous subgraphs can be directly used as the above-mentioned multiple connecting edges.
[0117] In another implementation, considering that using all the connecting edges of the K completely homogeneous subgraphs may occupy too much storage space during the training process. To solve this problem, it is proposed that partial connecting edges can be randomly sampled from each homogeneous subgraph respectively to jointly form the above-mentioned multiple connecting edges, or, first, the K edge sets {ε h} K corresponding to the K completely homogeneous subgraphs are merged, and then the union ∪ε h is randomly sampled to obtain the above-mentioned multiple connecting edges.
[0118] Regarding the above-mentioned random sampling ratio, in one embodiment, it can be set manually. In another embodiment, referring to the definition formula of the homogeneity of the edge set, that is, the above formula (1), a calculation formula for restricting the homogeneity metric value can be designed, such as the following formula (5), so as to obtain the sampling ratio.
[0119]
[0120] where ρ represents the homogeneity metric value set manually, |∪ε h | represents the total number of connecting edges in the union ∪ε h , p represents the probability that each edge in the union ∪ε h is retained, N represents the total number of nodes in the full set of target nodes, N(N - 1) / 2 represents the number of all possible connecting edges between N nodes, represents rounding up.
[0121] It should be understood that after determining the sampling ratio p, some connecting edges can be removed from each edge set ε h or the union ∪ε h according to 1 - p, so that the retained connecting edges are classified into the above-mentioned multiple connecting edges, or, some connecting edges can also be sampled from each edge set ε h or the union ∪ε h according to p and classified into the above-mentioned multiple connecting edges.
[0122] As described above, the business party can obtain multiple connecting edges sampled based on homogeneous subgraphs. It should be understood that at least a part of these multiple connecting edges are not in the original edge set ε.
[0123] Next, in step S130, the business party sends the multiple connection edges to the service party.
[0124] After that, in step S140, the service party adds the multiple connection edges to the original edge set ε to obtain the obfuscated edge set ε ′ , which is used to jointly train a graph learning model with the business party.
[0125] It should be understood that graph learning models refer to models that learn based on graph structures, including graph neural networks (Graph Neural Networks, GNNs), graph embedding methods (such as Node2Vec, DeepWalk), and other graph-based machine learning methods.
[0126] For multiple parties (i.e., multiple participating parties) jointly training a graph learning model, when the business party also holds the feature data of all target nodes, the multiple parties can only include the business party and the service party. Otherwise, the multiple parties can also include other parties that hold feature data. Additionally, for the specific implementation details of joint training, reference can be made to existing related technologies and will not be elaborated here.
[0127] In summary, for the method of multiple-party joint training of a graph learning model disclosed in the embodiments of this specification, first, the full-scale label data is constructed using the newly designed federated label propagation algorithm, and then the fully homogeneous subgraph is constructed using the full-scale label data. Thus, the connection edges sampled based on the fully homogeneous subgraph are added to the original edge set to complete the obfuscation. The graph learning model obtained by joint training via the obfuscated edge set can not only effectively resist graph reconstruction attacks but also significantly maintain the usability of the graph learning model.
[0128] According to an embodiment of another aspect, for the determination of the above full-scale label data, the applicant proposes that when the business party also holds the feature data of all target nodes, it can also be completed by the business party training a node classifier locally. Specifically, first, the business party can train a node classifier in a supervised learning manner based on the original label data and feature data of some target nodes; then, use the trained node classifier to label another part of the target nodes. Specifically, for each node in the other part of the target nodes, input its features into the node classifier, and thus use the predicted classification result as the pseudo-label of the node. Thus, the expansion of label data is realized, and the original label data and pseudo-label data together form the full-scale label data.
[0129] Further, after the business party determines the full-scale label data, for the construction of the fully homogeneous subgraph and the sampling of connection edges and other execution steps, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.
[0130] Corresponding to the above joint training method, embodiments of this specification also disclose a joint training device. Specifically as follows:
[0131] Figure 4 The schematic training device 400 is integrated into the business party, which holds the original label data of some nodes in the full amount of target nodes, and the training device 400 includes the following functional modules:
[0132] The federated label propagation module 410 is configured to, based on the original label data, through federated computing with the service party based on the original edge set, obtain the full amount of label data, including the pseudo-label data of another part of the nodes in the full amount of target nodes. The connection edge construction module 420 is configured to construct multiple connection edges based on the full amount of label data, where any two nodes connected by a connection edge have the same label. The connection edge sending module 430 is configured to send the multiple connection edges to the service party for it to add the multiple connection edges to the original edge set to obtain a confused edge set for jointly training the graph learning model with the business party.
[0133] In some embodiments, the business party also holds the feature data of the full amount of target nodes.
[0134] In some embodiments, the federated computing performed by the federated label propagation module 410 includes several rounds of iterative updates for the plaintext of the label matrix, and the updated plaintext of the label matrix corresponds to the full amount of label data. Any round of iterative update includes:
[0135] Encrypt the current plaintext of the label matrix using the public key of the business party to obtain the ciphertext of the label matrix; any i-th row or column in the plaintext of the label matrix is the one-hot encoding corresponding to the label of the i-th node in the full amount of target nodes. Send the ciphertext of the label matrix to a trusted third party for it to perform homomorphic ciphertext multiplication processing on the ciphertext of the label matrix and the ciphertext of the adjacency matrix, and the ciphertext of the adjacency matrix is determined by the service party encrypting the adjacency matrix corresponding to the original edge set using the public key. Use the private key corresponding to the public key to decrypt the result matrix ciphertext of the homomorphic ciphertext multiplication processing received from the trusted third party to obtain the result matrix plaintext. Update the current plaintext of the label matrix using the result matrix plaintext.
[0136] In some embodiments, the federated computing performed by the federated label propagation module 410 includes several rounds of iterative updates for the plaintext of the label matrix, and the updated plaintext of the label matrix corresponds to the full amount of label data. Any round of iterative update includes:
[0137] Encrypt the current plaintext label matrix using the public key of the business party to obtain the ciphertext label matrix; any i-th row or column in the plaintext label matrix is the one-hot encoding corresponding to the label of the i-th node in the full set of target nodes. Send the ciphertext label matrix to the service party for it to perform homomorphic plaintext-ciphertext multiplication processing on the ciphertext label matrix and the adjacency matrix corresponding to the original edge set. Use the private key corresponding to the public key to decrypt the result matrix ciphertext of the homomorphic plaintext-ciphertext multiplication processing received from the service party to obtain the result matrix plaintext. Update the current plaintext label matrix using the result matrix plaintext.
[0138] Further, in some specific embodiments, the federated label propagation module 410 is further configured to: when any round of iterative update is the first round, before encrypting the current plaintext label matrix using the public key of the business party, initialize the rows or columns in the plaintext label matrix corresponding to the partial nodes based on the original label data; and, initialize the rows or columns in the plaintext label matrix corresponding to the other partial nodes using all-zero vectors.
[0139] In some specific embodiments, the federated label propagation module 410 is configured to update the current plaintext label matrix using the result matrix plaintext, specifically including: for any i-th row in the plaintext label matrix, determine the column number j where the largest element in this row is located, and update the value in the j-th column of this row to a predetermined positive value, and update the values in other columns to 0.
[0140] In some embodiments, the original label data involves K category labels; the connection edge construction module 420 is specifically configured to: based on the full set of label data, divide the full set of target nodes into K groups corresponding to the K category labels; for each group of nodes, establish connection edges between each pair of nodes in it to obtain the corresponding connection edge subset; sample based on the K connection edge subsets to obtain the multiple connection edges.
[0141] Further, in some specific embodiments, the connection edge construction module 420 is configured to sample based on the K connection edge subsets to obtain the multiple connection edges, specifically including: determine the union of the K connection edge subsets; calculate the sampling ratio, which is inversely proportional to the number of connection edges in the union and directly proportional to the number of nodes in the full set of target nodes; sample the ratio of connection edges from the union as the multiple connection edges.
[0142] Figure 5 The schematic training device 500, which is integrated in the service party, holds the original edge set of the full set of target nodes, and Figure 4 has a collaborative relationship with the
[0143] The federated label propagation module 510 is configured to perform federated computing with the business party based on the original edge set and the original label data, so that the business party obtains the full amount of label data, including the pseudo-label data of another part of the nodes in the full amount of target nodes. The connection edge receiving module 520 is configured to receive, from the business party, multiple connection edges constructed based on the full amount of label data, where any two nodes connected by a connection edge have the same label. The edge set obfuscation module 530 is configured to add the multiple connection edges to the original edge set to obtain an obfuscated edge set for jointly training a graph learning model with the business party.
[0144] Figure 6 The schematic training device 600 is integrated into the business party, which holds the feature data of the full amount of target nodes and the original label data of some of the nodes, and the training device 600 includes the following functional modules:
[0145] The classifier training module 610 is configured to train a node classifier based on the feature data and the original label data of the part of the nodes, and use the node classifier to determine the pseudo-label data of another part of the nodes in the full amount of target nodes. The connection edge construction module 620 is configured to construct multiple connection edges based on the original label data and the pseudo-label data, where any two nodes connected by a connection edge have the same label. The connection edge sending module 630 is configured to send the multiple connection edges to the service party for it to add the multiple connection edges to the original edge set to obtain an obfuscated edge set for jointly training a graph learning model with the business party.
[0146] Figure 7 The schematic training device 700 is integrated into the service party, which holds the original edge set of the full amount of target nodes and Figure 6 has a collaborative relationship with the training device 600 shown, and the training device 700 includes the following functional modules:
[0147] The connection edge receiving module 710 is configured to receive, from the business party, multiple connection edges constructed based on the full amount of label data, where any two nodes connected by a connection edge have the same label; wherein, the full amount of label data includes the original label data and the pseudo-label data of another part of the nodes in the full amount of target nodes determined by the node classifier, and the node classifier is trained based on the feature data and the original label data of the part of the nodes. The edge set obfuscation module 720 is configured to add the multiple connection edges to the original edge set to obtain an obfuscated edge set for jointly training a graph learning model with the business party.
[0148] It should be noted that for the introduction of the above functional units, reference can also be made to the relevant introduction of the process method in the foregoing embodiments.
[0149] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having stored thereon a computer program, which, when executed on a computer, causes the computer to execute Figure 1 , Figure 2 or Figure 3 the method described.
[0150] According to an embodiment of yet another aspect, there is also provided a computing device including a memory and a processor, where the memory stores executable code, and when the processor executes the executable code, it implements Figure 1 , Figure 2 or Figure 3 the method described.
[0151] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0152] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the original label data of some nodes in the full set of target nodes, and a second party that holds the original edge set of the full set of target nodes; The method is applied to the first party and includes: Based on the original label data, and based on the original edge set with the second party, obtaining the full - scale label data through federated computing, which includes the pseudo - label data of another part of the nodes in the full - scale target nodes; Based on the full - scale label data, constructing multiple connecting edges, where any two nodes connected by a connecting edge have the same label; Sending the multiple connecting edges to the second party for it to add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
2. The method according to claim 1, wherein the first party also holds the feature data of the full - scale target nodes.
3. The method according to claim 1, wherein the federated computing includes several rounds of iterative updates for the plaintext of the label matrix, and the updated plaintext of the label matrix corresponds to the full - scale label data. Any round of iterative update includes: Encrypting the current plaintext of the label matrix using the public key of the first party to obtain the ciphertext of the label matrix; Any i - th row or column in the plaintext of the label matrix is the one - hot encoding corresponding to the label of the i - th node in the full - scale target nodes; Sending the ciphertext of the label matrix to a trusted third party for it to perform homomorphic ciphertext multiplication processing on the ciphertext of the label matrix and the ciphertext of the adjacency matrix, and the ciphertext of the adjacency matrix is determined by the second party encrypting the adjacency matrix corresponding to the original edge set using the public key; Decrypting the result matrix ciphertext of the homomorphic ciphertext multiplication processing received from the trusted third party using the private key corresponding to the public key to obtain the result matrix plaintext; Updating the current plaintext of the label matrix using the result matrix plaintext.
4. The method according to claim 1, wherein the federated computing includes several rounds of iterative updates for the plaintext of the label matrix, and the updated plaintext of the label matrix corresponds to the full - scale label data. Any round of iterative update includes: Encrypting the current plaintext of the label matrix using the public key of the first party to obtain the ciphertext of the label matrix; Any i - th row or column in the plaintext of the label matrix is the one - hot encoding corresponding to the label of the i - th node in the full - scale target nodes; Sending the ciphertext of the label matrix to the second party for it to perform homomorphic plain - ciphertext multiplication processing on the ciphertext of the label matrix and the adjacency matrix corresponding to the original edge set; Decrypting the result matrix ciphertext of the homomorphic plain - ciphertext multiplication processing received from the second party using the private key corresponding to the public key to obtain the result matrix plaintext; Updating the current plaintext of the label matrix using the result matrix plaintext.
5. The method according to claim 3 or 4, wherein When any round of iterative update is the first round, before encrypting the current plaintext of the label matrix using the public key of the first party, the method further includes: Initializing the rows or columns in the plaintext of the label matrix corresponding to the partial nodes based on the original label data; Initializing the rows or columns in the plaintext of the label matrix corresponding to the other part of the nodes using all - zero vectors.
6. The method according to claim 3 or 4, wherein Updating the current plaintext of the label matrix using the result matrix plaintext includes: For any i-th row in the plaintext of the label matrix, determine the column number j where the largest element in this row is located, update the value in the j-th column of this row to a predetermined positive value, and update the values in other columns to 0.
7. The method according to claim 1, wherein, The original label data involves K category labels; among them, based on the full label data, construct multiple connecting edges, including: Based on the full label data, divide the full target nodes into K groups corresponding to the K category labels; For each group of nodes, establish a connecting edge between each pair of nodes to obtain a corresponding subset of connecting edges; Sample based on the K subsets of connecting edges to obtain the multiple connecting edges.
8. The method according to claim 7, wherein Sampling based on the K subsets of connecting edges to obtain the multiple connecting edges, including: Determine the union of the K subsets of connecting edges; Calculate the sampling ratio, which is inversely proportional to the number of connecting edges in the union and directly proportional to the number of nodes in the full target nodes; Sample the ratio of connecting edges from the union as the multiple connecting edges.
9. A method for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the original label data of some nodes in the full set of target nodes, and a second party that holds the original edge set of the full set of target nodes; The method is applied to the second party, including: Based on the original edge set, perform federated computing with the first party based on the original label data, so that the first party obtains the full label data, including the pseudo-label data of another part of the nodes in the full target nodes; Receive from the first party multiple connecting edges constructed based on the full label data, where any two nodes connected by a connecting edge have the same label; Add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
10. A method for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the feature data of all target nodes and the original label data of some of the nodes, and also include a second party that holds the original edge set of all the target nodes; The method is applied to the first party, including: Based on the feature data and original label data of the partial nodes, train a node classifier and use the node classifier to determine the pseudo-label data of another part of the nodes in the full target nodes; Based on the original label data and pseudo-label data, construct multiple connecting edges, where any two nodes connected by a connecting edge have the same label; Send the multiple connecting edges to the second party for it to add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
11. A method for multi-party joint training of a graph learning model, where the multi-party includes a first party that holds the feature data of the full target nodes and the original label data of some of the nodes, and also includes a second party that holds the original edge set of the full target nodes; the method is applied to the second party, including: Receive from the first party multiple connecting edges constructed based on the full label data, where any two nodes connected by a connecting edge have the same label; where the full label data includes the original label data and the pseudo-label data of another part of the nodes in the full target nodes determined by a node classifier, and the node classifier is trained based on the feature data and original label data of the partial nodes; Add the multiple connecting edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
12. An apparatus for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the original label data of some nodes in the full set of target nodes, and a second party that holds the original edge set of the full set of target nodes; The device is integrated into the first party, including: A federated label propagation module, configured to, based on the original label data, perform federated computation with the second party based on the original edge set to obtain full-scale label data, including pseudo-label data of another part of the nodes in the full-scale target nodes; A connection edge construction module, configured to construct a plurality of connection edges based on the full-scale label data, where any two nodes connected by a connection edge have the same label; A connection edge sending module, configured to send the plurality of connection edges to the second party for it to add the plurality of connection edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
13. An apparatus for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds the original label data of some nodes in the full set of target nodes, and a second party that holds the original edge set of the full set of target nodes; The device is integrated into the second party and includes: A federated label propagation module, configured to perform federated computation with the first party based on the original edge set and the original label data, such that the first party obtains full-scale label data, including pseudo-label data of another part of the nodes in the full-scale target nodes; A connection edge receiving module, configured to receive from the first party a plurality of connection edges constructed based on the full-scale label data, where any two nodes connected by a connection edge have the same label; An edge set confusion module, configured to add the plurality of connection edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
14. A device for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds feature data of full-scale target nodes and original label data of some of the nodes, and also include a second party that holds the original edge set of the full-scale target nodes; the device is integrated into the first party and includes: A classifier training module, configured to train a node classifier based on the feature data and original label data of the some nodes, and use the node classifier to determine pseudo-label data of another part of the nodes in the full-scale target nodes; A connection edge construction module, configured to construct a plurality of connection edges based on the original label data and the pseudo-label data, where any two nodes connected by a connection edge have the same label; A connection edge sending module, configured to send the plurality of connection edges to the second party for it to add the plurality of connection edges to the original edge set to obtain a confused edge set for jointly training a graph learning model with the first party.
15. A device for jointly training a graph learning model by multiple parties, where the multiple parties include a first party that holds feature data of full-scale target nodes and original label data of some of the nodes, and also include a second party that holds the original edge set of the full-scale target nodes; the device is integrated into the second party and includes: A connection edge receiving module, configured to receive from the first party a plurality of connection edges constructed based on full-scale label data, where any two nodes connected by a connection edge have the same label; where the full-scale label data includes the original label data and pseudo-label data of another part of the nodes in the full-scale target nodes determined by a node classifier, and the node classifier is trained based on the feature data and original label data of the some nodes. An edge set obfuscation module, configured to add the plurality of connection edges to the original edge set to obtain an obfuscated edge set for jointly training a graph learning model with the first party.
16. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed in a computer, the computer is made to execute the method according to any one of claims 1-11.
17. A computing device includes a memory and a processor, wherein, Executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-11 is implemented.