Drug Molecular Property Prediction Method Based on Federal Security Collaboration

Optimizing the prediction model of drug molecular properties through personalized local differential privacy, K-Means clustering and knowledge distillation technology, the problems of data privacy protection and prediction accuracy in federal learning are solved, and efficient drug properties prediction under heterogeneous and unbalanced data are achieved.

CN116580785BActive Publication Date: 2025-07-25GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310575391.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-22
Publication Date
2025-07-25
Estimated Expiration
2043-05-22

AI Technical Summary

Technical Problem

In the federal learning paradigm, privacy protection of drug molecular data faces challenges, and prior art is difficult to improve the accuracy of drug properties prediction while protecting data privacy, especially in the case of heterogeneous data distribution and imbalance in quantity.

Method used

The personalized local differential privacy method is used to perturb the drug molecular data, and combined with K-Means clustering and knowledge distillation technology, the model training process is optimized, and the noise impact is mitigated through comparative learning, and the data privacy protection and prediction accuracy are achieved.

Benefits of technology

On the premise of protecting data privacy, the accuracy and efficiency of drug molecular properties prediction are improved, the problems of heterogeneity and quantity imbalance of data distribution are solved, and the learning ability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580785B_ABST
    Figure CN116580785B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting drug molecular properties based on federated security collaboration, which is characterized by the following steps: 1) dividing the client graph dataset; 2) training the local client model; 3) perturbing using personalized local differential privacy; 4) calculating similarity; 5) training the server model; 6) anti-noise training of the local model. This method can not only protect the privacy data in the graph federated learning paradigm, but also has little impact on the performance of the model and can improve the accuracy of predicting drug properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and privacy protection, and specifically, to a method for predicting drug molecular properties based on federated secure collaboration. Background Art

[0002] A chemical molecular network is a data form that describes the molecular structure and properties, and mainly includes information such as atoms, bonds, and chemical properties in the molecule. By modeling graph data with elements such as atom types and chemical bonds, the chemical molecular network can be applied to research fields such as molecular property prediction and drug design. For example, in the process of drug research and development, it is necessary to determine whether a drug molecule has properties such as the ability to inhibit cancer cell growth, mutagenicity, or toxicity. However, traditional deep learning methods are difficult to obtain satisfactory results when learning these non-Euclidean space molecular graph data.

[0003] Graph Neural Networks (GNNs) is a deep learning method that can process graph-structured data and can extract more effective potential feature information. It is widely used in processing, analyzing, and understanding data with complex topological structures. For example, in graph classification tasks such as molecular property prediction, the true properties of a molecule can be predicted by learning the relationship between the molecular structure and its properties. However, now more and more sensitive data is used in the training and analysis of graph data, which poses a great challenge to the privacy protection of molecular data. In addition, in the real world, molecular data is often not stored in a centralized manager, but is usually distributed among multiple local ends, which further exacerbates the risk of data privacy.

[0004] Federated learning is an emerging distributed machine learning method that can achieve data collaboration and cooperation between different molecular data centers, thereby improving the learning effect of the overall algorithm. To ensure data privacy, each data owner needs to retain the data locally and upload the model parameters or intermediate results of the data to the central server for iterative update, so as to achieve the goal of making the local model more accurate when predicting molecular properties. In the federated learning paradigm, since the server is honest but curious, existing means can deduce the privacy information of the original data through model parameters or intermediate results. Therefore, the security of the training data must be fully considered.

[0005] To address the risk of drug molecule data privacy leakage involved in the federated learning paradigm, in recent years, Local Differential Privacy (LDP) has become an effective privacy protection measure in both industry and academia. Under strict mathematical proofs, the LDP method aims to retain data utility as much as possible while ensuring data privacy, and thus can be widely applied to enhance the protection effect of local private graph data in federated learning. Regarding molecular graph data, differential privacy can also be used to protect the privacy of features and structures. For example, this can be achieved by protecting feature information through the multi-bit mechanism in local differential privacy and perturbing label or structural information using the random response mechanism. However, in drug molecule prediction tasks, not all data is sensitive. For example, when predicting whether a drug molecule has anti-cancer effects, only the data of drug molecules with such effects is valuable, so leaking the information of these data will seriously damage the interests of relevant R & D institutions. However, in the federated learning paradigm, the institutions participating in R & D may have data with different sensitivities. If the same noise is applied to all data, not only will the protection effect be poor, but the data utility will also be greatly damaged. Therefore, a new privacy protection mechanism for graph federated learning is needed to protect sensitive sample data while maintaining the prediction accuracy of drug molecule properties as much as possible. Summary of the Invention

[0006] The object of the present invention is to provide a method for predicting drug molecule properties based on federated secure collaboration in view of the deficiencies of the prior art. This method can not only protect privacy data in the graph federated learning paradigm, but also has little impact on the performance of the model and can improve the accuracy of predicting drug properties.

[0007] The technical solution for achieving the object of the present invention is as follows:

[0008] A method for predicting drug molecule properties based on federated secure collaboration includes the following steps:

[0009] 1) Divide the client graph dataset G: Divide the graph dataset G of drug molecules for each client. The dataset distributions among all clients are heterogeneous and unbalanced in quantity. Each client's dataset contains data in the field of drug molecules. Among them, 80% of the data is used for training and 20% is used for testing. The process of dividing the client graph dataset G is as follows: First, obtain the initial graph dataset G from the crowdsourcing platform, where G = (V, E, X, Y), V is the set of nodes, representing all nodes in the graph dataset G, E is the set of edges, representing the connection relationships between nodes in the graph dataset G, X is the feature matrix of the graph dataset G, and Y is the label set composed of each graph label. The parameter α of the Dirichlet distribution is used to adjust the imbalance degree of the quantity among different clients;

[0010] 2) Training of the local client model: The local client learns the molecular graph representation low-dimensional vector embedding and the probability value logits of the molecular property prediction category through the private dataset. That is, the divided local private drug molecule training dataset is trained through the graph isomorphism network GIN model to learn the graph representation embedding. These graph representations are used as the input of the classifier model based on the feedforward neural network and are used to learn the probability scores logits of the drug molecule property prediction category. Specifically: For the feature matrix X and adjacency matrix A=(V, E) of all graphs, the graph low-dimensional embedding representation embedding is learned based on the graph isomorphism network GIN model:

[0011] That is, the node representations h of all nodes at layer L on the graph are learned through the message passing mechanism v , as shown in formula (1):

[0012]

[0013] Finally, after pooling the representations of the nodes on the graph, the graph representation h G is obtained, that is, the low-dimensional vector embedding, as shown in formula (2):

[0014] h G = readout({h v ; v ∈ V}) (2),

[0015] For the obtained graph representation, a classifier model based on the fully connected layer network is used to predict the drug molecule properties, and the prediction result is the probability z of the drug molecule belonging to the category c , that is, "logits", as shown in formula (3):

[0016] z c = f c (W c ; h G ) (3);

[0017] 3) Perturbation using personalized local differential privacy: After adding noise to desensitize the embedding, logits, and labels of molecular data, they are uploaded to the central server. In the federated learning paradigm, the true labels of private drug molecular properties, as well as the learned graph representation embedding and the probability logits of the property category, need to be protected for privacy before being uploaded to the server. Since in the task of predicting drug molecular properties, not all molecular sample data is sensitive. For example, drugs considered to have no anti-cancer effect are not sensitive. Therefore, the leakage of information of these samples will not cause too many conflicts of interest. Using the method of personalized local differential privacy, more privacy protection adapted to drug molecular training data is carried out on the locally uploaded data. This can not only maintain the privacy protection of sample data but also reduce the noise input to the uploaded data. For discrete vector data such as drug labels, perturbation is performed based on the random response method, and by only reversing non-private labels, the true sensitive labels cannot be inferred from the finally uploaded labels. For continuous vector data such as the learned graph vector embedding and prediction probability value logits, the muti-bit mechanism is used for perturbation. To better protect locally sensitive sample data, the method of personalized local differential privacy is used to adaptively process the sample data for the current drug molecular prediction scenario, so as to achieve privacy protection. Specifically:

[0018] 3-1) Denote each vector in the embedding or logits as x i,j , select m dimensions of the vector for perturbation. When the privacy budget is set to ∈, use Bernoulli sampling to encode the vector x i,j into a new vector as shown in formula (4):

[0019]

[0020] The encoding by Bernoulli sampling perturbs the original one, but the perturbed vector is biased. Therefore, the perturbed vector is transformed so that the transformed vector x′ i,j is statistically ensured to be an unbiased estimate, as shown in formula (5):

[0021]

[0022] 3-2) When dealing with discrete data labels, the random response mechanism is used for perturbation. And in the optimized random response mechanism for the current drug molecular prediction scenario, the sensitive labels are not reversed. Instead, according to the situation of non-sensitive sample labels, with probability they are reversed, and with probability Keep it unchanged. This method can better control the amount of noise added by differential privacy compared to the traditional random response mechanism, reduce the interference of noise on data, thereby protecting data privacy and improving the accuracy of data analysis;

[0023] 4) Similarity calculation: The server calculates the similarity of the uploaded graph representations (embeddings). It clusters the homogeneous embeddings into different clusters. The distribution of the graph representations (embeddings) uploaded by the clients is also heterogeneous. The server uses the K-Means algorithm to calculate the similarity of the graph representations of each client based on the perturbed graph representation vectors (embeddings) uploaded by each client, and clusters the homogeneous client data together into clusters according to this similarity. Specifically, after all the graph representations (embeddings) and predicted probability values (logits) uploaded by the clients are uploaded to the central server, the K-Means clustering algorithm is used to process the graph representations, that is, first calculate the similarity of the graph representations {h G1 , h G2 ,..., h GK}, and then perform the clustering operation to divide the K clients into T clusters {C1, C2,..., C T};

[0024] 5) Training of the server model: Respectively use the embeddings in the clusters as the input of the classifier, and return the re-learned logits to the local users. That is, use the clustered embeddings as the input of the models in these clusters, and combine the training method of knowledge distillation to further optimize the accuracy of the model probability prediction. Finally, return the predicted probability values (logits) learned in each cluster as knowledge to the local clients. Specifically, for each cluster Ci, use the graph representations (embeddings) of the clients in each cluster as the input to train the classifier model within the cluster and learn the predicted probability z s , as shown in formula (6):

[0025] z s = f s (W s ; h G ) (6),

[0026] To better optimize the classifier model f s in this cluster, and then learn more accurate predicted probability values z s , use the KL divergence to construct a knowledge distillation loss function by combining the predicted probability values z s learned by the classifier model in the cluster with the predicted probability values z c uploaded by the clients for optimization, as shown in formula (7): For better optimization of the classifier model f s in this cluster, and then to learn more accurate predicted probability values z s , use the KL divergence to construct a knowledge distillation loss function by combining the predicted probability values z s learned by the classifier model in the cluster with the predicted probability values z c uploaded by the clients for optimization, as shown in formula (7):

[0027]

[0028] Finally, the more accurate predicted probability value z in all clusters learned by the server is s returned to the local client to implement knowledge distillation to guide the training of the local client model. This method can effectively improve the accuracy of the local client model in predicting molecular properties and achieve more efficient and secure data processing and analysis while protecting data privacy;

[0029] 6) Anti-noise training of the local model: The local user uses the returned logits as positive and negative samples respectively and trains based on contrastive learning to mitigate the negative impact of noise on the local client model. There is still a certain degree of noise in the knowledge logits learned by the server. When these logits are returned to the local client for retraining in the form of knowledge distillation, it may damage the performance of the local model. To mitigate the impact of noise on the local model training, a contrastive learning method is used to participate in the training of the local model. Specifically: The local client uses the returned predicted probability value to construct a contrastive learning loss term, that is, the global predicted probability value z in this cluster s and the predicted probability value z learned by the local model c are regarded as positive samples, while the global predicted probabilities in other clusters and the predicted probability value z of the local model c are regarded as negative samples, thereby constructing the loss function l of the client contrastive learning CL , as shown in formula (8):

[0030]

[0031] Optimizing this loss function can improve the graph representation learning ability of the local client model, thereby more accurately predicting the properties of molecules. Through this method, more accurate and efficient distributed data processing and analysis can be achieved on the basis of protecting data privacy.

[0032] In a real scenario, the molecular graph data distributions among different clients are usually non-independent and identically distributed (Non-IID). In addition, due to device heterogeneity and different drug R & D progress at each local end, the amount of data uploaded to the server is also unbalanced. These problems make it difficult to improve the accuracy of predicting drug molecular properties in the federated learning paradigm. In addition, existing means can reverse the sensitive information of the local original data. This technical solution uses local differential privacy to enhance the privacy protection effect of federated learning.

[0033] This technical solution has the following advantages:

[0034] 1. This technical solution uses knowledge distillation as the basis for data communication in the graph federated learning paradigm, successfully solving the problems of heterogeneous client data distribution and quantity imbalance in the process of drug molecule research and development. At the same time, in the privacy-preserving federated learning paradigm, this technical solution can still effectively predict the properties of drug molecules;

[0035] 2. This technical solution analyzes the privacy leakage risks in graph federated learning and uses the method of personalized local differential privacy to protect users' private data while reducing the negative impact of noise on the data;

[0036] 3. This technical solution uses the K-Means clustering algorithm to partition homogeneous graph data into the same cluster. The advantage of doing this is that when these homogeneous data are used as the input of the classifier model, the model can converge better;

[0037] 4. This technical solution uses the contrastive learning method at the local end to mitigate the impact of noisy global logits on the local client model during the data exchange process.

[0038] This method can not only protect the private data in the graph federated learning paradigm, but also has little impact on the performance of the model. Even if the local data is denoised and desensitized, the local model can still obtain a high prediction accuracy of drug molecule properties. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flowchart of the method for the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0040] The following further elaborates on the content of the present invention in conjunction with the drawings and embodiments, but does not limit the present invention.

[0041] Embodiment:

[0042] See Figure 1 , a method for predicting the properties of drug molecules based on federated security collaboration, including the following steps:

[0043] 1) Divide the client graph dataset G: Divide the graph dataset G of drug molecules for each client. The dataset distributions among all clients are heterogeneous and numerically imbalanced. The dataset of each client contains data in the field of drug molecules. Among them, 80% of the dataset is used for training and 20% for testing. The process of dividing the client graph dataset G is as follows: First, obtain the initial graph dataset G from the crowdsourcing platform, where G = (V, E, X, Y), V is the set of nodes, representing all nodes in the graph dataset G, E is the set of edges, representing the connection relationships between nodes in the graph dataset G, X is the feature matrix of the graph dataset G, and Y is the label set composed of various graph labels. Use the parameter α of the Dirichlet distribution to adjust the imbalance degree of the quantity between different clients;

[0044] 2) Training of the local client model: The local client learns the low-dimensional vector embedding of the molecular graph representation and the probability value logits of the molecular property prediction category through the private dataset, that is, the divided local private drug molecule training dataset is trained through the Graph Isomorphism Network (GIN) model, so as to learn the graph representation embedding. These graph representations are used as the input of the classifier model based on the feed-forward neural network and are used to learn the probability scores logits of the drug molecule property prediction category. Specifically: For the feature matrix X and adjacency matrix A = (V, E) of all graphs, learn the low-dimensional embedding representation embedding of the graph based on the GIN model:

[0045] That is, learn the node representations h of all nodes at L layers on the graph through the message passing mechanism v , as shown in formula (1):

[0046]

[0047] Finally, after pooling the representations of the nodes on the graph, obtain the graph representation h G , that is, the low-dimensional vector embedding, as shown in formula (2):

[0048] h G = readout({h v ; v ∈ V}) (2),

[0049] For the obtained graph representation, use a classifier model based on a fully connected layer network to predict the drug molecule properties. The prediction result is the probability z of the drug molecule belonging to a category c , that is, "logits", as shown in formula (3):

[0050] z c = f c (W c ; h G ) (3);

[0051] 3) Perturbation using personalized local differential privacy: After adding noise to desensitize the embedding, logits, and labels of molecular data, they are uploaded to the central server. In the federated learning paradigm, the true labels of private drug molecular properties, as well as the learned graph representation embedding and the probabilities logits of the property categories, need to be protected by privacy before being uploaded to the server. Since in the task of predicting drug molecular properties, not all molecular sample data is sensitive. For example, drugs considered to have no anti-cancer effect are not sensitive. Therefore, the leakage of information of these samples will not cause excessive conflicts of interest. Using the method of personalized local differential privacy, more privacy protection adapted to drug molecular training data is performed on the locally uploaded data. This can not only maintain the privacy protection of sample data but also reduce the noise input to the uploaded data. For discrete vector data such as drug labels, perturbation is performed based on the method of random response, and by only reversing non-private labels, the true sensitive labels cannot be inferred from the finally uploaded labels. For continuous vector types such as the learned graph vector embedding and the predicted probability values logits, the muti-bit mechanism is used for perturbation. To better protect locally sensitive sample data, the method of personalized local differential privacy is used to adaptively process the sample data for the current drug molecular prediction scenario, thereby achieving privacy protection. Specifically:

[0052] 3-1) Denote each vector in the embedding or logits as x i,j , select m dimensions of the vector for perturbation. When the privacy budget is set to ∈, use Bernoulli sampling to encode the vector x i,j into a new vector as shown in formula (4):

[0053]

[0054] The encoding by Bernoulli sampling perturbs the original, but the perturbed vector is biased. Therefore, the perturbed vector is transformed so that the transformed vector x′ i,j is statistically ensured to be an unbiased estimate, as shown in formula (5):

[0055]

[0056] 3-2) When dealing with discrete data labels, the random response mechanism is used for perturbation. And in the optimized random response mechanism for the current drug molecular prediction scenario, sensitive labels are not reversed. Instead, according to the situation of non-sensitive sample labels, with probability they are reversed, and with probability Keep it unchanged. This method can better control the amount of noise added by differential privacy compared to the traditional random response mechanism, reduce the interference of noise on data, thereby protecting data privacy and improving the accuracy of data analysis;

[0057] 4) Similarity calculation: The server calculates the similarity of the uploaded graph representations (embeddings). It clusters the homogeneous embeddings into different clusters. The embeddings distributed by the client are also heterogeneous. The server uses the K-Means algorithm to calculate the similarity of the graph representations of each client based on the perturbed graph representation vectors (embeddings) uploaded by each client, and clusters the homogeneous client data together into clusters. Specifically, after all the graph representations (embeddings) and prediction probability values (logits) uploaded by the clients are uploaded to the central server, the K-Means clustering algorithm is used to process the graph representations. That is, first, the similarity of the graph representations {h G1 , h G2 ,..., h GK} is calculated, and then clustering operations are performed to divide the K clients into T clusters {C1, C2,..., C T};

[0058] 5) Training of the server model: The embeddings in the clusters are used as the input of the classifier respectively, and the re-learned logits are returned to the local users. That is, the clustered embeddings are used as the input of the models in these clusters, and combined with the training method of knowledge distillation to further optimize the accuracy of the model probability prediction. Finally, the predicted probability values (logits) learned in each cluster are returned to the local client as knowledge. Specifically, for each cluster Ci, the graph representations (embeddings) of the clients in each cluster are used as the input to train the classifier model within the cluster and learn the predicted probability z s of the model within the cluster, as shown in formula (6):

[0059] z s = f s (W s ; h G ) (6),

[0060] To better optimize the classifier model f s in this cluster, and then learn more accurate predicted probability values z s , the KL divergence is used to construct a knowledge distillation loss function by combining the predicted probability value z s learned by the classifier model in the cluster with the predicted probability value z c uploaded by the client for optimization, as shown in formula (7): For optimization, as shown in formula (7):

[0061]

[0062] Finally, return the more accurate predicted probability value z in all clusters learned by the server side s to the local client to implement knowledge distillation to guide the training of the local client model. This method can effectively improve the accuracy of the local client model in predicting molecular properties and achieve more efficient and secure data processing and analysis while protecting data privacy;

[0063] 6) Anti-noise training of the local model: The local user takes the returned logits as positive and negative samples respectively and trains based on contrastive learning to mitigate the negative impact of noise on the local client model. There is still a certain degree of noise in the knowledge logits learned by the server side. When these logits are returned to the local client for retraining in the way of knowledge distillation, it may damage the performance of the local model. To mitigate the impact of noise on the local model training, a contrastive learning method is used to participate in the training of the local model. Specifically: The local client constructs a contrastive learning loss term by using the returned predicted probability value, that is, taking the global predicted probability value z s in this cluster and the predicted probability value z c learned by the local model as positive samples, while the global predicted probabilities in other clusters and the predicted probability value z c of the local model are regarded as negative samples, thus constructing the loss function l CL of client-side contrastive learning, as shown in formula (8):

[0064]

[0065] In this column, optimizing this loss function can enhance the graph representation learning ability of the local client model, thus predicting the properties of molecules more accurately. Through this method, more accurate and efficient distributed data processing and analysis can be achieved on the basis of protecting data privacy, providing more powerful support for scientific research and industrial exploration.

Claims

1. A method for predicting the properties of drug molecules based on federated security collaboration, characterized in that, The steps are as follows: 1) Divide the client graph dataset G: Divide the graph dataset G of drug molecules for each client. The dataset distributions among all clients are heterogeneous and imbalanced in quantity. The dataset of each client contains data in the field of drug molecules. Among them, 80% of the dataset is used for training and 20% is used for testing. The process of dividing the client graph dataset G is as follows: First, obtain the initial graph dataset G from the crowdsourcing platform, where G = (V, E, X, Y), V is the set of nodes, representing all nodes in the graph dataset G, E is the set of edges, representing the connection relationships between nodes in the graph dataset G, X is the feature matrix of the graph dataset G, and Y is the label set composed of each graph label. The parameter α of the Dirichlet distribution is used to adjust the imbalance degree of the quantity between different clients; 2) Training of the local client model: The local client learns the low-dimensional vector embedding of the molecular graph representation and the probability value logits of the molecular property prediction category through the private dataset, that is, the divided local private drug molecule training dataset is trained through the graph isomorphism network GIN model to learn the graph representation embedding. These graph representations are used as the input of the classifier model based on the feed-forward neural network and are used to learn the probability scores logits of the drug molecule property prediction category. Specifically: For the feature matrix X and the adjacency matrix A = (V, E) of all graphs, the low-dimensional embedding representation embedding of the graph is learned based on the graph isomorphism network GIN model; That is, the node representations \(h\) of all nodes at layer \(L\) on the graph are learned through the message passing mechanism v , as shown in Equation (1): After obtaining the representation of the nodes on the final pooling graph, the graph representation h is obtained G , that is, the low-dimensional vector embedding, as shown in Equation (2): h G = readout({h v ; v ∈ V}) (2), For the obtained graph representation, a classifier model based on a fully connected layer network is used to predict the properties of drug molecules, and the prediction result is the probability z of the drug molecule belonging to a certain category c , that is, "logits", as shown in formula (3): z c = f c (W c ; h G ) (3); 3) Perturbation using personalized local differential privacy: The embedding, logits, and the labels of the molecular data are uploaded to the central server after adding noise for desensitization. For continuous vectors such as the learned graph vector embedding and the prediction probability value logits, the mut-bit mechanism is used for perturbation. The method of personalized local differential privacy is used to adaptively process the sample data for the current drug molecule prediction scenario. Specifically: 3-1) Denote each vector in the embedding or logits as x i,j , select m dimensions of the vector for perturbation. When the privacy budget is set to ∈, use Bernoulli sampling to encode the vector x i,j into a new vector as shown in formula (4): The encoding after Bernoulli sampling perturbs the original, but the perturbed vector is biased. Therefore, the perturbed vector is transformed so that the transformed vector x′ i,j is statistically guaranteed to be an unbiased estimate, as shown in Equation (5): 3-2) When dealing with discrete data labels, a random response mechanism is adopted for perturbation. And in the optimized random response mechanism for the current drug molecule prediction scenario, sensitive labels are not inverted. Instead, according to the situation of non-sensitive sample labels, inversion is carried out with probability and kept unchanged with probability ; 4) Similarity calculation: The server calculates the similarity of the uploaded graph representations (embeddings), clusters the homogeneous embeddings into different clusters. The distribution of the graph representations (embeddings) uploaded by the client is also heterogeneous. The server uses the K-Means algorithm to calculate the similarity of the graph representations of each client based on the perturbed graph representation vectors (embeddings) uploaded by each client, and clusters the homogeneous client data together to form clusters. Specifically, after all the graph representations (embeddings) and prediction probability values (logits) uploaded by the clients are uploaded to the central server, the K-Means clustering algorithm is used to process the graph representations, that is, first calculate the similarity of the graph representations {h G1 ,h G2 ,...,h GK}, and then perform the clustering operation to divide the K clients into T clusters {C1, C2,..., C T}; 5) Training of the server model: The embeddings in the clusters are respectively used as the inputs of the classifier, and the re-learned logits are returned to the local users. That is, the clustered embeddings are used as the inputs of the models in these clusters, and combined with the training method of knowledge distillation to further optimize the accuracy of model probability prediction. Finally, the predicted probability values logits learned in each cluster are returned as knowledge to the local client. Specifically, for each cluster Ci, the graph representation embeddings of the clients in each cluster are used as the inputs to train the classifier model within the cluster and learn the predicted probability z of the model within the cluster s , as shown in formula (6): z s = f s (W s ; h G ) (6), Use the KL divergence to construct the prediction probability value z learned by the classifier model in the cluster s and the prediction probability value z uploaded by the client c together to construct the loss function of knowledge distillation for optimization, as shown in Equation (7): Finally, return the more accurate predicted probability value z in all clusters learned by the server side s to the local client; 6) Anti-noise training of the local model: The method of contrastive learning is used to participate in the training of the local model. Specifically, the local client uses the returned predicted probability values to construct the contrastive learning loss term, that is, the global predicted probability value z in this cluster s and the predicted probability value z learned by the local model c are regarded as positive samples, while the global predicted probabilities in other clusters and the predicted probability value z of the local model c are regarded as negative samples, thereby constructing the loss function l of client-side contrastive learning CL , as shown in formula (8):