Personalized Graph Federated Learning Method, System and Storage Medium Based on Signature Clustering

By introducing signature clustering and personalization factors into graph federated learning, the problem of poor clustering learning effect caused by graph data heterogeneity in traditional methods is solved, and more efficient and accurate graph federated learning training is achieved, improving model performance and generalization capabilities.

CN117556919BActive Publication Date: 2025-06-20HANGZHOU DBAPPSECURITY CO LTD

Patent Information

Application Number
CN202311364523.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-20
Publication Date
2025-06-20
Estimated Expiration
2043-10-20

AI Technical Summary

Technical Problem

Traditional federated learning methods encounter poor clustering learning effects and data heterogeneity when dealing with non-independent and homogeneous graph data. Especially under the unique nature of graph data, existing methods are difficult to effectively solve data heterogeneity between and within clients.

Method used

A personalized graph federated learning method based on signature clustering is proposed. The client receives the model sent by the server and trains and generates a signature. The server clusters according to the signature and performs federated averaging within the cluster cluster. The client combines personalized factors for local training to improve the effectiveness of clustering learning and solves the problem of data heterogeneity.

Benefits of technology

Through signature clustering and personalization factors, the clustering training efficiency and accuracy of graph federated learning are effectively improved, the data heterogeneity problem between and within the client is solved, and the performance and generalization capabilities of the global model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117556919B_ABST
    Figure CN117556919B_ABST
Patent Text Reader

Abstract

The present invention discloses a personalized graph federated learning method, system and storage medium based on signature clustering. The method includes the steps: a client receives a model sent by a server for training and aggregates according to the parameters of the model; generates a signature of the client through function mapping; adds a personalized factor to the local model and uploads the model layer parameters and signature except the personalized factor to the server; the server clusters according to the signatures of each client and performs federated averaging within the clustering clusters; the client receives the federated averaged model parameters sent by the server and performs local training in combination with the personalized factor. The present invention solves the problems of poor clustering learning effect and data heterogeneity between and within clients in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning and privacy protection, and particularly relates to a personalized graph federated learning method, system and storage medium based on signature clustering. Background Art

[0002] Federated learning (FL) enables distributed machine learning while addressing privacy and security concerns by keeping data local. This process involves collaborative training among clients and gradient weight aggregation by a server. However, non-independent and identically distributed (non-IID) data poses challenges to traditional federated learning algorithms, leading to slow convergence and reduced accuracy during training. Additionally, most federated learning algorithms mainly target Euclidean data such as images and sounds, and the research on graph data is still in its early stages.

[0003] Graph federated learning encounters challenges similar to traditional federated learning, and the unique nature of graph data introduces additional complexities. On the one hand, graph data from different domains are usually distributed across different clients and exhibit significant feature and structural heterogeneity. Existing solutions typically rely on indirect statistical information of the data, such as mean, variance, or gradients. However, these indirect statistical information cannot fully represent the clients, thus limiting the effectiveness of clustering learning. On the other hand, existing federated learning methods often overlook the heterogeneity (also known as feature drift) of graph data within the same client and domain. Although some methods have been proposed to handle feature drift in Euclidean data, these methods may not achieve satisfactory results on graph data and may even degrade model performance.

[0004] To improve the effectiveness of clustering learning and address the heterogeneity issues of graph data within and between clients, a personalized graph federated learning method, system and storage medium based on signature clustering are proposed. Summary of the Invention

[0005] Embodiments of the present invention propose a personalized graph federated learning method, system and storage medium based on signature clustering to at least solve the problems of poor clustering learning effectiveness and data heterogeneity between and within clients in related technologies.

[0006] According to an embodiment of the present invention, a personalized graph federated learning method based on signature clustering is proposed, including:

[0007] A client receives a model sent by a server for training and aggregates according to the parameters of the model;

[0008] Generate a signature of the client through function mapping;

[0009] Add a personalized factor to the local model and upload the model layer parameters and signatures except the personalized factor to the server;

[0010] The server clusters according to the signatures of each client and performs federated averaging within the cluster;

[0011] The client receives the federated averaged model parameters sent by the server and performs local training in combination with the personalized factor.

[0012] In an exemplary embodiment, the client receives the model sent by the server for training and aggregates according to the parameters of the model, including the steps of:

[0013] The client receives the initial global model sent by the server and trains according to the model;

[0014] Extract the parameters of the last two layers of the client's model in the last round;

[0015] Use the mean pooling method to aggregate the data into a unified dimension to obtain a consistent dimension matrix.

[0016] In an exemplary embodiment, generating the signature of the client through function mapping includes the steps of:

[0017] Take the consistent dimension matrix aggregated in the last round as the input of the hash function;

[0018] Calculate the hash signature of the client according to the hash function.

[0019] In an exemplary embodiment, adding a personalized factor to the local model includes the steps of:

[0020] Regard each graph as a different client;

[0021] Use GraphNorm as the personalized factor for each graph;

[0022] Retain the personalized factor locally and exclude it from the shared parameters of federated learning.

[0023] In an exemplary embodiment, the server clusters according to the signatures of each client, including the steps of:

[0024] Calculate the similarity between clients according to the signatures of the representative clients;

[0025] Cluster according to the similarity between clients using the k-means clustering algorithm or the hierarchical clustering tree algorithm to obtain multiple clustering results;

[0026] After determining the clustering results, the clients assigned to a specific group will perform synchronous training within the group.

[0027] In an exemplary embodiment, calculating the similarity between clients according to the signatures representing the clients includes the steps of:

[0028] Calculating the cosine similarity according to the cosine value of the angle between the client signature vectors;

[0029] Calculating the Euclidean distance according to the distance between the signature vectors;

[0030] Calculating the Hamming distance according to the difference of the signature vector bits;

[0031] Calculating the Pearson correlation coefficient according to the linear correlation of the signature vectors;

[0032] Calculating the similarity between clients according to the cosine similarity and / or Euclidean distance and / or Hamming distance and / or Pearson correlation coefficient.

[0033] In an exemplary embodiment, performing federated averaging inside the clustering clusters is to take the average of the model parameters inside the clustering using the FedAvg method.

[0034] In an exemplary embodiment, the client receives the federated averaged model parameters sent by the server and performs local training in combination with the personalized factor, including the steps of:

[0035] The server sends the federated averaged model parameters to each client;

[0036] Each client updates the local model parameters except the personalized factor according to the sent model parameters;

[0037] The client performs a new round of GNNs training according to the personalized factor and the updated local model parameters.

[0038] A computer-readable storage medium stores a computer program for electronic data exchange, wherein the computer program causes a computer to execute the above method.

[0039] According to another embodiment of the present invention, a personalized graph federated learning system based on signature clustering is provided, including:

[0040] A processor;

[0041] A memory;

[0042] And

[0043] One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the processor, and the programs cause a computer to execute the above method.

[0044] The advantages of the personalized graph federated learning method, system, and storage medium based on signature clustering of the present invention are as follows:

[0045] (1) The client uses the mean pooling method to aggregate the data into a unified dimension to obtain a consistent dimension matrix and thereby obtain the hash signature representing the client. Compared with the traditional local local model training method, the hash mechanism protects the privacy and security of the data.

[0046] (2) Calculate the similarity between clients according to the cosine similarity and / or Euclidean distance and / or Hamming distance of the client signatures. Compared with the traditional technical solution that only evaluates the model similarity according to the cosine similarity or Hamming distance, it can effectively improve the efficiency and accuracy of clustering training.

[0047] (3) Perform clustering according to the similarity between clients using the k-means clustering algorithm or hierarchical clustering tree algorithm and perform federated averaging within the cluster. Compared with the traditional federated learning method, it can effectively solve the heterogeneity problem of the same or different domain data existing between different clients and improve the performance of the global model.

[0048] (4) Add a personalized factor to the client, and at the same time, the personalized factor does not participate in federated learning, so that the representations generated by the local data after passing through the model are more similar in distribution. Compared with the traditional federated learning method, while meeting the requirements of the local GNN, it further promotes the convergence of the local model and improves the generalization ability of the global model. Description of the Drawings

[0049] Figure 1 is a flowchart of a personalized graph federated learning method based on signature clustering according to an embodiment of the present invention;

[0050] Figure 2 is a flowchart of sub-step S01 according to an embodiment of the present invention;

[0051] Figure 3 is a flowchart of sub-step S02 according to an embodiment of the present invention;

[0052] Figure 4 is a flowchart of sub-step S03 according to an embodiment of the present invention;

[0053] Figure 5 is a flowchart of sub-step S04 according to an embodiment of the present invention;

[0054] Figure 6 is a flowchart of sub-step S041 according to an embodiment of the present invention;

[0055] Figure 7 is a flowchart of sub-step S05 according to an embodiment of the present invention;

[0056] Figure 8 It is a schematic structural diagram of a personalized graph federated learning system based on signature clustering according to an embodiment of the present invention. Specific implementation manners

[0057] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the invention, but do not limit the invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0058] A personalized graph federated learning method based on signature clustering according to an embodiment of the present invention, the flowchart is as Figure 1 shown, including the steps:

[0059] Step S01, the client receives the model sent by the server for training and aggregates according to the parameters of the model;

[0060] Step S02, generate the signature of the client through function mapping;

[0061] Step S03, add a personalized factor to the local model and upload the model layer parameters and signatures except the personalized factor to the server;

[0062] Step S04, the server clusters according to the signatures of each client and performs federated averaging within the clustering clusters;

[0063] Step S05, the client receives the model parameters after federated averaging sent by the server and performs local training in combination with the personalized factor.

[0064] In an exemplary embodiment, the step S01, the flowchart is as Figure 2 shown, including:

[0065] Step S011, the client receives the initial global model sent by the server and trains according to the model;

[0066] Step S012, extract the parameters of the last two layers of the model of the client in the last round;

[0067] Step S013, use the mean pooling method to aggregate the data into a unified dimension to obtain a consistent dimension matrix.

[0068] In this embodiment, the real dataset usually exhibits various characteristics, including different types and domains, and contains a large amount of data. To address this challenge, VGAC can reconstruct and generate new graphical data to achieve generalization and functionality across different domains. Additionally, it utilizes a variational inference method for effectively processing large graph datasets. Therefore, VGAC is selected as the base model. The encoder is constructed as a direct three-layer graph convolutional network (GCN) for encoding the topological structure a and node features X. Clustering can effectively handle data heterogeneity in federated learning. Clustering federated learning treats each client as an entity and represents it with specific features. After clustering on the server side, model averaging is performed within the classes. However, in the client-specific representation step, due to privacy issues, direct program methods such as statistical information are rarely used. Instead, methods such as model gradients and loss changes are used. Unfortunately, these information-based methods cannot accurately reflect the relative heterogeneity degree among clients, thus limiting the performance of clustering.

[0069] To address this challenge and explore the possibility of low-dimensional data representation, aggregation and hashing operations are employed to preserve the basic features of the clients while ensuring data privacy. Considering the data scale drift between different datasets, the calculation is simplified and the communication cost is minimized by achieving a unified length representation for all clients. To this end, the parameters of the last two layers of the model in the last round of the client are extracted, and the data representation of the client at the last round of time is aggregated into a unified dimension, which is based on the graph dimension and denoted as

[0070]

[0071] where T i represents the aggregation matrix for data aggregation using mean pooling, and n i is the number of graphs on the i-th client, represents the unified representation of client i after aggregation.

[0072] In an exemplary embodiment, step S02 has a flowchart as Figure 3 shown, and includes:

[0073] Step S021: Use the consistent dimension matrix obtained by the last-round aggregation as the input of the hash function;

[0074] Step S022: Calculate the hash signature of the client according to the hash function.

[0075] In this embodiment, the low-dimensional vector matrix generated by the client (i.e., the consistent dimension matrix obtained by the last-round aggregation) is regarded as the input, and the hash signature of each client is calculated.

[0076]

[0077] Among them, HASH represents a special mapping function.

[0078] In an exemplary embodiment, for the step S03, the flowchart is as Figure 4 shown and includes:

[0079] Step S031: Treat each graph as a different client;

[0080] Step S032: Use GraphNorm as the personalization factor for each graph;

[0081] Step S033: Locally retain the personalization factor and exclude it from the shared parameters of federated learning;

[0082] Step S034: Upload the model layer parameters and signatures except the personalization factor to the server.

[0083] In this embodiment, treating each graph as a different client aims to personalize the training process of the GNN within each client so that they can better adapt to the specific characteristics of the graph data. To achieve this goal, a personalization factor is introduced to address the data heterogeneity problem within the client. Use GraphNorm as the personalization factor for each graph. The specific representation of this method is as follows:

[0084]

[0085] Where γ j and β j are affine parameters, the same as other normalization methods. The learnable parameter α plays a crucial role in controlling information retention and enhancing the expressive power of the graph neural network (GNN). Upload the model layer parameters and signatures except the personalization factor to the server. By locally retaining the personalization factor and excluding it from the shared parameters of federated learning, the individual graph data can adapt to its specific distribution representation. This method can enable the overall local model to reach a faster convergence speed, thus effectively solving the problem of heterogeneous graph data within the client. The personalization factor ensures that different local graph data can learn approximate distribution representations, thereby improving the model performance and reducing the impact of data heterogeneity within the client.

[0086] In an exemplary embodiment, for the step S04, the flowchart is as Figure 5 shown and includes:

[0087] Step S041: Calculate the similarity between clients according to the signatures representing the clients;

[0088] Step S042: Cluster using the k-means clustering algorithm or the hierarchical clustering tree algorithm based on the similarity between clients to obtain multiple clustering results;

[0089] Step S043: After determining the clustering results, the clients assigned to a specific group will perform synchronous training within that group;

[0090] Step S044: Use the FedAvg method to average the model parameters within the clusters.

[0091] In this embodiment, it is assumed that there is a trusted central server. The global model will send the initially generated model to each client at the beginning. After the clients complete training, the server collects the local model parameters of each trained client, calculates the federated average model parameters based on the clustering of the local model parameters, integrates them to form a new global model, sends the global model to each client, and continues the next round of training until the model converges.

[0092] In an exemplary embodiment, the sub-step S041, the flowchart is as Figure 6 shown, including:

[0093] S0411: Calculate the cosine similarity based on the cosine value of the angle between the client signature vectors;

[0094] S0412: Calculate the Euclidean distance based on the distance between the client signature vectors;

[0095] S0413: Calculate the Hamming distance based on the difference in the bits of the client signature vectors;

[0096] S0414: Calculate the Pearson correlation coefficient based on the linear correlation of the client signature vectors;

[0097] S0415: Calculate the similarity between clients based on the cosine similarity and / or Euclidean distance and / or Hamming distance and / or Pearson correlation coefficient.

[0098] In this embodiment, calculating the cosine similarity based on the cosine value of the angle between the client signature vectors is to calculate the cosine value of the angle as the cosine similarity, which is the ratio of the vector product to the product of the moduli between two models. The cosine similarity is represented by c.

[0099] Calculate the Euclidean distance based on the distance between the client signature vectors. The Euclidean distance is represented by the variable d.

[0100] Calculate the Hamming distance between clients based on the difference in the hash signature bits of the clients. The Hamming distance is represented by the variable n. That is, compare each bit of the vectors. If they are different, the Hamming distance is incremented by 1. The Hamming distance is obtained by summing up the differences in the bits.

[0101] Calculating the Pearson correlation coefficient based on the linear correlation of the client signature vectors refers to obtaining the Pearson correlation coefficient between two client signature vectors according to the Pearson correlation coefficient calculation formula, and the Pearson correlation coefficient is represented by p.

[0102] Calculating the similarity between clients based on the cosine similarity and / or Euclidean distance and / or Hamming distance and / or Pearson correlation coefficient refers to calculating the similarity between clients according to the positive correlation relationship between the cosine similarity and / or Euclidean distance and / or Hamming distance and / or Pearson correlation coefficient between models, and the similarity between clients is represented by e.

[0103] In Table A, A1 to A15 represent different implementation manners for calculating the similarity between clients. For the sake of convenience of expression, the cosine similarity c, Euclidean distance d, Hamming distance n, and Pearson correlation coefficient p are calculated by the method described in any of the above implementation manners.

[0104] Table A Different implementation manners for calculating the similarity between clients

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111] According to the similarity e between clients calculated in Table A, the k-means clustering algorithm or hierarchical clustering tree algorithm is used for clustering to obtain multiple clustering results; after determining the clustering results, the clients assigned to a specific group will perform synchronous training within the group; specifically for federated learning within the assigned cluster, the FedAvg method is used to take the average value of the model parameters within the cluster, which is expressed as:

[0112]

[0113] Among them, represents the number of graphs in the local dataset of client m, represents the total number of graphs of all clients in cluster C.

[0114] In an exemplary embodiment, the sub-step S05, the flowchart is as Figure 7 shown, and includes:

[0115] Step S051: The server distributes the federally averaged model parameters to each client;

[0116] Step S052: Each client updates the local model parameters except for the personalized factors according to the distributed model parameters;

[0117] Step S053: The client performs a new round of GNNs training based on the personalized factors and the updated local model parameters.

[0118] In this embodiment, the server distributes the federally averaged model parameters to each client. Each client updates the local model parameters except for the personalized factors according to the distributed model parameters, and performs a new round of GNNs training based on the personalized factors and the updated local model parameters.

[0119] A computer-readable storage medium according to an embodiment of the present invention stores a computer program for electronic data exchange, wherein the computer program causes a computer to execute the above method.

[0120] A personalized graph federated learning system based on signature clustering according to an embodiment of the present invention has a structural schematic diagram as Figure 8 shown, and includes:

[0121] A processor;

[0122] A memory;

[0123] And

[0124] One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the processor, and the programs cause a computer to execute the above method.

[0125] Of course, those of ordinary skill in the art in the technical field should recognize that the above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. As long as it is within the scope of the present invention, changes and variations of the above embodiments will fall within the protection scope of the present invention.

Claims

1. A personalized graph federated learning method based on signature clustering, characterized in that, Including: The client receives the model sent by the server for training and aggregates according to the parameters of the model; The client receives the model sent by the server for training and aggregates according to the parameters of the model, including the steps of: the client receives the initial global model sent by the server and trains according to the model; extracts the parameters of the last two layers of the client's last round of the model; uses the mean pooling method to aggregate the data into a unified dimension to obtain a consistent dimension matrix; Generates the signature of the client through function mapping; Adds a personalized factor to the local model and uploads the model layer parameters and signature except the personalized factor to the server; The adding of the personalized factor to the local model includes the steps of: regarding each graph as a different object; using GraphNorm as the personalized factor for each graph; retaining the personalized factor locally and excluding it from the shared parameters of federated learning; The server clusters according to the signatures of each client and performs federated averaging within the clustering clusters; The client receives the federated averaged model parameters sent by the server and performs local training in combination with the personalized factor; The client receives the federated averaged model parameters sent by the server and performs local training in combination with the personalized factor, including the steps of: the server sends the federated averaged model parameters to each client; Each client updates the local model parameters except the personalized factor according to the sent model parameters; The client performs a new round of GNNs training according to the personalized factor and the updated local model parameters.

2. The personalized graph federated learning method based on signature clustering according to claim 1, characterized in that, The generating of the signature of the client through function mapping includes the steps of: Taking the consistent dimension matrix obtained by the last round of aggregation as the input of the hash function; Calculating the hash signature of the client according to the hash function.

3. The personalized graph federated learning method based on signature clustering according to claim 1, characterized in that, The server clustering according to the signatures of each client includes the steps of: Calculating the similarity between clients according to the signatures of the representative clients; Performing clustering using the k-means clustering algorithm or the hierarchical clustering tree algorithm according to the similarity between clients to obtain multiple clustering results; After determining the clustering results, the clients assigned to a specific group will perform synchronous training within the group.

4. The personalized graph federated learning method based on signature clustering according to claim 3, characterized in that, The calculating of the similarity between clients according to the signatures of the representative clients includes the steps of: Calculating the cosine similarity according to the cosine value of the angle between the client signature vectors; Calculating the Euclidean distance according to the distance between the client signature vectors; Calculating the Hamming distance according to the difference of the bits of the client signature vectors; Calculating the Pearson correlation coefficient according to the linear correlation of the client signature vectors; Calculating the similarity between clients according to the cosine similarity and / or the Euclidean distance and / or the Hamming distance and / or the Pearson correlation coefficient.

5. The personalized graph federated learning method based on signature clustering according to claim 1, characterized in that, The performing of federated averaging within the clustering clusters is to take the average value of the model parameters within the clustering using the FedAvg method.

6. A computer-readable storage medium storing a computer program for electronic data exchange, wherein, The computer program causes the computer to execute the method according to any one of claims 1-5.

7. A personalized graph federated learning system based on signature clustering, characterized in that Including: A processor; A memory; And One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the processor, and the program causes the computer to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Decentralized federated graph neural network recommendation method for local bipartite graph

    CN116796059A

  • Metaverse-based shared information privacy protection method and related apparatus

    WO2023141809A1

Cited By

  • A personalized federated learning method and system based on risk control aggregation

    CN122635588A