Federated graph learning method and system based on feature fusion and temperature scaling distillation

Through the federated graph learning method of feature fusion and temperature scaling distillation, the GNN model and encoder parameters are optimized, which solves the problem of inaccurate node classification caused by data heterogeneity and improves the accuracy of node classification.

CN120316580BActive Publication Date: 2025-09-16CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510446081.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-09-16
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In federated graph learning, data heterogeneity leads to poor node classification performance, especially the deviation of neighborhood information of a few nodes, which affects the classification accuracy.

Method used

Using a method based on feature fusion and temperature scaling distillation, the client determines the initial GNN model and structural proxy set through local social graph data, optimizes model parameters and encoders, and the server performs parameter aggregation and iterative optimization until the model converges, alleviating the problem of data heterogeneity.

Benefits of technology

The accuracy of node classification is improved, the deviation of neighborhood information of minority nodes during GNN model training is reduced, and the overall performance of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316580B_ABST
    Figure CN120316580B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of federated graph neural networks, and provides a federated graph learning method and system based on feature fusion and temperature scaling distillation, wherein the client obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value according to the first soft target, and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value. The client obtains a predicted label distribution based on the optimized GNN model, calculates a second target loss value according to the predicted label distribution, and optimizes the initial feature structure encoder parameters and the initial structure proxy set in the direction of reducing the second target loss value. The server aggregates the feature structure encoder parameters and the class-level structure proxy average values ​​respectively to obtain the target feature structure encoder parameters and the target structure proxy, thereby alleviating the data heterogeneity problem and improving the accuracy of node classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of federated graph neural network technology, and in particular to a federated graph learning method and system based on feature fusion and temperature scaling distillation. Background Art

[0002] Graph Neural Networks (GNNs) demonstrate exceptional performance in processing graph-structured data. Through message passing, they efficiently aggregate attribute information about node neighbors in a graph, enabling the successful learning of highly expressive node representations. However, in practice, large amounts of graph data are generated by multiple different data owners. Due to privacy constraints and commercial competition, this data cannot be centrally trained on GNNs.

[0003] Federated Learning (FL) is a distributed learning solution that enables multiple data owners to collaborate and jointly train models without sharing their private data. However, because data is not independently and identically distributed (IID) across different clients, federated learning suffers from data heterogeneity. In Federated Graph Learning (FGL), the data heterogeneity problem is even more severe, and it also faces the unique challenge of a small number of nodes having high heterogeneity. Specifically, the neighbors of a small number of nodes mostly come from other categories, which leads to deviations in the neighborhood information obtained by the small number of nodes during GNN model training. This, in turn, makes the embedding of the small number of nodes insufficient, which has a negative impact on the classification performance of the nodes.

[0004] Therefore, how to solve the problem of data heterogeneity in FGL that has an adverse effect on the classification performance of nodes is an urgent problem that those skilled in the art need to solve. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a federated graph learning method and system based on feature fusion and temperature scaling distillation to solve the problem in the prior art that data heterogeneity has an adverse effect on the classification performance of nodes; that is, an embodiment of the present invention can improve the accuracy of node classification.

[0006] According to one aspect of the present invention, a federated graph learning method based on feature fusion and temperature scaling distillation is provided, and the federated graph learning method based on feature fusion and temperature scaling distillation includes: the client determines an initial GNN model and an initial structure proxy set based on local social graph data, and determines initial feature structure encoder parameters based on the initial structure proxy set; the client obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value based on the first soft target, and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model; the client obtains a predicted label distribution based on the optimized GNN model, calculates a second target loss value based on the predicted label distribution, and calculates a second target loss value based on reducing the first target loss value. The initial feature structure encoder parameters and the initial structure proxy set are optimized in the direction of the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set; each client uploads the class-level structure proxy average and feature structure encoder parameters to the server, wherein the feature structure encoder parameters include the initial feature structure encoder parameters and the optimized feature structure encoder parameters, and the class-level structure proxy average includes the initial class-level structure proxy average and the optimized class-level structure proxy average; the server aggregates based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and aggregates based on the class-level structure proxy average to obtain target structure proxy; iterate the above process until the GNN model converges or reaches a preset model accuracy.

[0007] In one embodiment, the client determines an initial GNN model and an initial structure proxy set based on local social graph data, and determines initial feature structure encoder parameters based on the initial structure proxy set, including: the client determines the initial GNN model based on the local social graph data; the client determines the initial structure proxy set based on the local social graph data, and determines the initial feature structure encoder parameters based on the initial structure proxy set.

[0008] In one embodiment, the client obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value based on the first soft target, and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model, including: the client converts and combines the initial features in the initial feature structure encoder parameters to generate a first soft target; the client generates a first loss value based on the true label in the initial feature structure encoder parameters, generates a second loss value based on the first soft target, calculates the first loss value and the second loss value to obtain a first target loss value; and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model.

[0009] In one embodiment, the client obtains a predicted label distribution based on the optimized GNN model, calculates a second target loss value according to the predicted label distribution, and optimizes the initial feature structure encoder parameters and the initial structure proxy set in the direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set, including: the client obtains a predicted label distribution based on the optimized GNN model; the client generates a third loss value based on the true label and the second soft target in the initial feature structure encoder parameters, generates a fourth loss value based on the predicted label distribution, calculates the third loss value and the fourth loss value to obtain the second target loss value; optimizes the initial feature structure encoder parameters and the initial structure proxy set in the direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set.

[0010] In one embodiment, the client converts and combines the initial features in the initial feature structure encoder parameters to generate a first soft target, where the first soft target is:

[0011]

[0012] Among them, softmax is the flexible maximum function, W cls is the weight matrix of the combination operation, b cls is the bias vector of the classifier, is the concatenation of vectors, is the embedding vector, The local social graph data is an initial structure agent, k is the kth client, and i is the i-th marked node.

[0013] In one embodiment, the first loss value and the second loss value are calculated to obtain a first target loss value, and the first target loss value is:

[0014]

[0015] Among them, λ1 is a predefined hyperparameter, is the first loss value, is the second loss value.

[0016] In one embodiment, the third loss value and the fourth loss value are calculated to obtain a second target loss value, and the second target loss value is:

[0017]

[0018] Among them, λ2 is a predefined hyperparameter, is the third loss value, is the fourth loss value.

[0019] In one embodiment, the server performs aggregation based on the feature structure encoder parameters to obtain target feature structure encoder parameters, where the target feature structure encoder parameters are:

[0020]

[0021] Among them, |V ( k ) | is the number of nodes in the local social graph data of client k, |V| is the total number of nodes in the local social graph data of all clients, is the characteristic structure encoder parameter, and L is the number of clients.

[0022] In one embodiment, the target structure agent is obtained by aggregating the average value of the class-level structure agent, and the target structure agent is:

[0023]

[0024] in, is the number of nodes of category c in the local social graph data in the client k, A c is the total number of nodes of category c in the local social graph data of all the clients, A proxy average is provided for the class-level structure in the local social graph data.

[0025] According to another aspect of the present invention, a federated graph learning system based on feature fusion and temperature scaling distillation is provided, and the federated graph learning system based on feature fusion and temperature scaling distillation includes: a server and at least one client, wherein the client is used to determine an initial GNN model and an initial structure proxy set based on local social graph data, and determine initial feature structure encoder parameters based on the initial structure proxy set; the client is also used to obtain a first soft target based on the initial feature structure encoder parameters, calculate a first target loss value based on the first soft target, and optimize the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model; the client is also used to obtain a predicted label distribution based on the optimized GNN model, calculate a second target loss value based on the predicted label distribution, and calculate the second target loss value based on the reduced The initial feature structure encoder parameters and the initial structure proxy set are optimized in the direction of minimizing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set; each of the clients is also used to upload the class-level structure proxy average and feature structure encoder parameters to the server, wherein the feature structure encoder parameters include the initial feature structure encoder parameters and the optimized feature structure encoder parameters, and the class-level structure proxy average includes the initial class-level structure proxy average and the optimized class-level structure proxy average; the server is used to perform aggregation based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and to perform aggregation based on the class-level structure proxy average to obtain target structure proxy; the server and the client are used to iterate the above process until the GNN model converges or reaches a preset model accuracy.

[0026] In summary, in an embodiment of the present invention, a client determines an initial GNN model and an initial structure proxy set based on local social graph data, and determines initial feature structure encoder parameters based on the initial structure proxy set. The client obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value based on the first soft target, and optimizes the model parameters in the initial GNN model in a direction that reduces the first target loss value to obtain an optimized GNN model. The client obtains a predicted label distribution based on the optimized GNN model, calculates a second target loss value based on the predicted label distribution, and optimizes the initial feature structure encoder parameters and the initial structure proxy set in a direction that reduces the second target loss value to obtain optimized feature structure encoder parameters and an optimized structure proxy set. Each client uploads a class-level structure proxy average value and feature structure encoder parameters to a server. The server aggregates based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and aggregates based on the class-level structure proxy average value to obtain a target structure proxy, thereby alleviating the problem of data heterogeneity, reducing the problem of deviation in neighborhood information obtained by minority class nodes during GNN model training, and improving the accuracy of node classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Further details, features and advantages of the present invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0028] Figure 1 A schematic diagram of the process of the federated graph learning method based on feature fusion and temperature scaling distillation disclosed in an embodiment of the present application is shown;

[0029] Figure 2 Shown Figure 1 The schematic diagram of the step flow of step S110 is shown;

[0030] Figure 3 Shown Figure 2 The structural diagram of the characteristic structure encoder in step S112 is shown;

[0031] Figure 4 Shown Figure 1 The schematic diagram of the step flow of step S120 is shown;

[0032] Figure 5 Shown Figure 1 The schematic diagram of the step flow of step S130 is shown;

[0033] Figure 6 A structural diagram of a federated graph learning system based on feature fusion and temperature scaling distillation disclosed in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0034] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0035] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0036] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0037] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0038] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0039] It should be noted that the execution entity of the federated graph learning method based on feature fusion and temperature scaling distillation provided in the embodiments of the present invention can be one or more electronic devices, which is not limited by the present invention. Among them, the electronic device can be a terminal (i.e., a client) or a server. If the execution entity includes multiple electronic devices, and the multiple electronic devices include at least one terminal and at least one server, the federated graph learning method based on feature fusion and temperature scaling distillation provided in the embodiments of the present invention can be jointly executed by the terminal and the server. Accordingly, the terminals mentioned here can include but are not limited to: smartphones, tablets, laptops, desktop computers, smart watches, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, etc.

[0040] Based on the above description, an embodiment of the present invention proposes a federated graph learning method based on feature fusion and temperature scaling distillation. The federated graph learning method based on feature fusion and temperature scaling distillation can be executed by the electronic device (terminal or server) mentioned above; or, the federated graph learning method based on feature fusion and temperature scaling distillation can be jointly executed by the terminal and the server. For the sake of convenience, the following description will be given by taking the example of an electronic device executing the federated graph learning method based on feature fusion and temperature scaling distillation.

[0041] See also Figure 1 , which is a flow chart of the federated graph learning method based on feature fusion and temperature scaling distillation disclosed in the embodiment of the present application. The federated graph learning method based on feature fusion and temperature scaling distillation solves the problem of adverse effects on node classification performance caused by data heterogeneity in FGL, thereby improving the accuracy of node classification. It should be noted that the federated graph learning method based on feature fusion and temperature scaling distillation in the embodiment of the present application is not limited to Figure 1 The steps and order in the flowchart shown. According to different needs, the steps in the flowchart shown can be added, removed, or changed in order. In the embodiment of the present application, Figure 1 As shown, the process of the federated graph learning method based on feature fusion and temperature scaling distillation includes at least the following steps.

[0042] S110: The client determines an initial GNN model and an initial structure proxy set based on local social graph data, and determines initial feature structure encoder parameters based on the initial structure proxy set.

[0043] like Figure 2 As shown, in an embodiment of the present invention, Figure 2 The step S110 includes at least the following steps:

[0044] S111. The client determines the initial GNN model based on local social graph data.

[0045] In an embodiment of the present invention, a social network graph is a graphical representation used to show the social relationships between people. It consists of nodes (vertices) and edges connecting nodes. Each user is a node, and an edge indicates that there is a direct relationship between the two nodes (two people) it connects (such as family relationships, work relationships, social relationships, or communication relationships, or friend relationships on chat platforms such as WeChat). If there is an edge between two nodes, it means that the corresponding two users have a direct relationship. Tags can be user interest tags, professional tags, etc. Features can include time series features, geographic location features and graph structure association features. Among them, time series features are the time distribution of users uploading photos (such as peak hours, interval regularity), geographic location features can be the geocoding of the photo shooting location or user registration location (such as longitude and latitude) or the geographic aggregation characteristics of user groups, and graph structure association features can include adjacency and community divisions. Among them, adjacency is the relationship between users such as attention, comments, and common groups, which is represented by the adjacency matrix (such as adj_full.npz). Community division is the user community affiliation identified based on the graph clustering algorithm (such as "professional photographers" and "amateurs").

[0046] Each client k has local social graph data G ( k ) =(V ( k ) ,E ( k ) ,X ( k ) ), where V ( k ) is the node set of the client, E (k) is an edge set, is the node feature matrix, d x is the number of node features. Local social graph data is from legitimate sources.

[0047] Each client k randomly initializes a personalized GNN model to obtain the initial GNN model f(θ (k) ), where θ(k) As parameters, the initial GNN model f(θ (k) ) will mark the node set Each marked node in Output predicted label distribution. GNN models can include Graph Convolutional Network (GCN), Simple Graph Convolutional Network (SGCN), and GraphSAGE (Graph Sample and Aggregator) models. Taking the GraphSAGE model as an example, the expression of the GraphSAGE model is as follows:

[0048]

[0049] in, is the feature representation of node v in the first layer, is a nonlinear activation function, W (1) is the learnable weight parameter of the first layer, Mean is the mean of the aggregated neighbor node features, is the input feature of node u, and N(v) is the neighbor node feature of node v.

[0050] S112: The client determines the initial structure proxy set based on local social graph data, and determines the initial feature structure encoder parameters based on the initial structure proxy set.

[0051] In the embodiment of the present invention, the initial structure proxy set S( 0 ) is a matrix, in which the initial structure agent in the local social graph data is represented in vector form, using Represents the initial class structure agent in the local social graph data, where each row s j ∈S is the d of the jth node class s Dimensional structure agent. For each of the local social graph data belongs to the jth type of labeled node Each marked node are associated with real labels Initial Structure Agent For s j , where each client has many marked nodes, and the i-th marked node in the k-th client belongs to category j. However, for the server, each row in the structural proxy matrix represents a type of marked node, the j-th row is the j-th type of marked node, and the i-th marked node of the k-th client corresponds to the j-th category of the server.

[0052] The feature structure encoder g is based on the initial features of the labeled nodes in the local social graph data. and initial structure agent As input, generate the first soft target of the marked node Among them, such as Figure 3 As shown, the feature structure encoder 131 may include an embedding layer emb, a classifier cls, and a projection layer proj.

[0053] Since the feature structure encoder can only generate soft targets for labeled nodes, but in order to better train the initial GNN model, it is necessary to obtain the soft targets of the unlabeled nodes in the local social graph data, so a projection layer proj is added. The projection layer proj has the same structure as the classifier cls, but the projection layer proj is only based on the embedding vector Soft targets can be generated. For unlabeled nodes or Based on embedding vector A second soft target can be generated Second soft target As shown in the following formula:

[0054]

[0055] Among them, softmax is the flexible maximum function, W proj is the weight matrix, is the embedding vector, b proj is the bias vector of the projection layer proj, k is the kth client, and i is the i-th marked node.

[0056] According to the second soft target and the initial class structure agent S to generate the structure agent of the unlabeled nodes Structural Agents for Unlabeled Nodes As shown in the following formula:

[0057]

[0058] in, is the second soft target, and S is the initial class structure agent in the local social graph data.

[0059] S120. The client obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value according to the first soft target, and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model.

[0060] like Figure 4 As shown, in an embodiment of the present invention, Figure 4The step S120 at least includes the following steps:

[0061] S121: The client converts and combines initial features in the initial feature structure encoder parameters to generate a first soft target.

[0062] In an embodiment of the present invention, for a marked node in the local social graph data The labeled nodes are embedded in the embedding layer emb The initial characteristics Convert to low-dimensional embedding vector Embedding vector As shown in the following formula:

[0063]

[0064] Among them, W emb is the weight matrix, is the initial feature, b emb is the bias vector of the embedding layer emb.

[0065] The embedding vector is passed through the classifier cls and initial structure agent Combine to generate the first soft target First soft target As shown in the following formula:

[0066]

[0067] Among them, softmax is the flexible maximum function, W cls is the weight matrix of the combination operation (such as weighted summation or fully connected layer), b cls is the bias vector of the classifier cls, is the concatenation of vectors, is the embedding vector, is an initial structure proxy in the local social graph data.

[0068] S122. The client generates a first loss value based on the true label in the initial feature structure encoder parameter, generates a second loss value based on the first soft target, calculates the first loss value and the second loss value, and obtains a first target loss value.

[0069] In the embodiment of the present invention, during the training process of the initial GNN model, the initial GNN model f(θ (k) ) in θ (k) The update of is initially achieved by minimizing the cross entropy loss. For each labeled node in the local social graph data The calculation formula for the first loss value (also known as the cross entropy loss value) is as follows:

[0070]

[0071] Among them, CE(·,·) is the cross entropy loss, is the true label, To predict the label distribution, is a marked node in the local social graph data, is a set of labeled nodes in the local social graph data.

[0072] Since most of the neighbor nodes of a few nodes in federated graph learning come from different categories, the domain information obtained through the message passing mechanism is biased, resulting in insufficient embedding vector representation. Therefore, the knowledge distillation term is introduced. Specifically, the feature structure encoder is based on the marked node Generate the first soft target And by minimizing the first soft objective and predicted label distribution The difference between them is used to guide the initial GNN model f(θ ( k ) ) in θ ( k ) Local training. By minimizing the first soft objective and predicted label distribution This goal is achieved by calculating the KL (Kullback-Leibler) divergence between the two models. T is the temperature parameter, which is used to control the smoothness of the softmax function. When T is large, the softmax output is smoother and provides richer inter-category information; when T is small, the softmax output is sharper and closer to the one-hot distribution. By introducing the temperature parameter, the knowledge transfer between models can be better balanced and the distillation effect can be improved. The second loss value (also known as knowledge distillation loss) as follows:

[0073]

[0074] Among them, KL(·||·) is the KL divergence, The first soft target, is the predicted label distribution, T is the temperature parameter, is a marked node in the local social graph data, is a set of labeled nodes in the local social graph data.

[0075] The first target loss value is calculated based on the first loss value and the second loss value (also known as the first overall loss value), the first target loss value As shown in the following formula:

[0076]

[0077] Among them, λ1 is a predefined hyperparameter, which is used to control the second loss value The contribution to the first target loss value.

[0078] When λ1 = 0, it is equivalent to training the GNN model separately on each client, and the knowledge distillation term is not used to regularize the training.

[0079] S123. Optimize the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model.

[0080] In an embodiment of the present invention, a stochastic gradient descent algorithm is used to update the model parameters in the initial GNN model. The gradient descent method is as follows:

[0081]

[0082] Among them, η1 is the first learning rate, is the loss function for the parameter θ ( k ) gradient.

[0083] S130. The client obtains a predicted label distribution based on the optimized GNN model, calculates a second target loss value according to the predicted label distribution, and optimizes the initial feature structure encoder parameters and the initial structure proxy set in the direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set.

[0084] like Figure 5 As shown, in an embodiment of the present invention, Figure 5 The step S130 at least includes the following steps:

[0085] S131. The client obtains a predicted label distribution based on the optimized GNN model.

[0086] In this embodiment of the present invention, the predicted label distribution of each marked node is obtained by the optimized GNN model. Predicted label distribution As shown in the following formula:

[0087]

[0088] Among them, f ( k ) () is the optimized GNN model, is the initial feature, G( k ) is the local map data, θ ( k ) As a parameter.

[0089] S132. The client generates a third loss value based on the true label and the second soft target in the initial feature structure encoder parameter, generates a fourth loss value based on the predicted label distribution, calculates the third loss value and the fourth loss value, and obtains a second target loss value.

[0090] In an embodiment of the present invention, the client generates a third loss value based on the true label and the second soft target in the initial feature structure encoder parameter The third loss value As shown in the following formula:

[0091]

[0092] Among them, CE(·,·) is the cross entropy loss, is the true label, is the second soft target, is a marked node in the local social graph data, is a set of labeled nodes in the local social graph data.

[0093] The third loss value Weighing the true label and the second soft target By minimizing this difference, the second soft target generated by the feature structure encoder is Closer to the true label This improves the GNN model's ability to predict node labels.

[0094] Generate a fourth loss value based on the predicted label distribution The fourth loss value As shown in the following formula:

[0095]

[0096] Among them, KL(·||·) is the KL divergence, To predict the label distribution, is the first soft target, T is the temperature parameter, is a marked node in the local social graph data, is a set of labeled nodes in the local social graph data.

[0097] The second target loss value is calculated based on the third loss value and the fourth loss value (also known as the second overall loss value), the second target loss value As shown in the following formula:

[0098]

[0099] Among them, λ2 is a predefined hyperparameter.

[0100] S133. Optimize the initial feature structure encoder parameters and the initial structure proxy set in a direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set.

[0101] In an embodiment of the present invention, a stochastic gradient descent algorithm is used to update the initial feature structure encoder parameters and the initial structure proxy set. The gradient descent method is as follows:

[0102]

[0103] Among them, η2 is the second learning rate, η3 is the third learning rate, is the gradient of the loss function with respect to the parameters of the initial feature structure encoder, is the gradient of the loss function with respect to the initial structure proxy set.

[0104] S140. Each client uploads the class-level structure proxy average value and feature structure encoder parameters to the server, wherein the feature structure encoder parameters include initial feature structure encoder parameters and optimized feature structure encoder parameters, and the class-level structure proxy average value includes initial class-level structure proxy average value and optimized class-level structure proxy average value.

[0105] In an embodiment of the present invention, during each stage of optimizing the GNN model, feature structure encoder parameters, and structure proxy set, the structure proxy S(k) is a continuous vector maintained for each class. After local training is completed, each client uploads the class-level structure proxy average and feature structure encoder parameters to the server. The feature structure encoder parameters may include initial feature structure encoder parameters and optimized feature structure encoder parameters, and the class-level structure proxy average may include initial class-level structure proxy average and optimized class-level structure proxy average.

[0106] S150. The server aggregates based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and aggregates based on the class-level structure agent average to obtain a target structure agent.

[0107] In the embodiment of the present invention, the server is based on the characteristic structure encoder parameters Aggregation is performed to obtain the target feature structure encoder parameter π(t), which is as follows:

[0108]

[0109] Among them, |V ( k ) | is the number of nodes in the local social graph data of client k, |V| is the total number of nodes in the local social graph data of all clients, and K is the number of clients.

[0110] Aggregate based on the average value of the class-level structural proxy to obtain the target structural proxy For each category c, the target structure agent As shown in the following formula:

[0111]

[0112] in, is the number of nodes of category c in the local social graph data in client k, is the total number of nodes of category c in the local social graph data of all clients.

[0113] S160, iterate the above process until the GNN model converges or reaches the preset model accuracy.

[0114] In summary, in the federated graph learning method based on feature fusion and temperature scaling distillation of the present application, the client determines the initial GNN model and the initial structure proxy set based on the local social graph data, and determines the initial feature structure encoder parameters based on the initial structure proxy set. The client obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value based on the first soft target, and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model. The client obtains a predicted label distribution based on the optimized GNN model, and calculates a second target based on the predicted label distribution. The target loss value is set, and the initial feature structure encoder parameters and the initial structure proxy set are optimized in the direction of reducing the second target loss value to obtain the optimized feature structure encoder parameters and the optimized structure proxy set. Each client uploads the class-level structure proxy average value and feature structure encoder parameters to the server. The server aggregates based on the feature structure encoder parameters to obtain the target feature structure encoder parameters, and aggregates based on the class-level structure proxy average value to obtain the target structure proxy, thereby alleviating the data heterogeneity problem, reducing the problem of deviation in neighborhood information obtained by minority class nodes during the GNN model training process, and improving the accuracy of node classification.

[0115] See also Figure 6 , which is a schematic diagram of the structure of the federated graph learning system based on feature fusion and temperature scaling distillation disclosed in the embodiment of this application. In one embodiment, Figure 6 As shown, the present application provides a federated graph learning system 100 based on feature fusion and temperature scaling distillation. The federated graph learning system 100 based on feature fusion and temperature scaling distillation may include at least: a server 110 and at least one client 130. Information exchange occurs between the server 110 and the client 130.

[0116] The client 130 is configured to determine an initial GNN model and an initial structure proxy set based on local social graph data, and to determine initial feature structure encoder parameters based on the initial structure proxy set. The client 130 may include at least a lightweight feature structure encoder 131 and a GNN model 133.

[0117] The client 130 is also used to obtain a first soft target based on the initial feature structure encoder parameters, calculate a first target loss value based on the first soft target, and optimize the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model.

[0118] The client 130 is also used to obtain a predicted label distribution based on the optimized GNN model, calculate a second target loss value according to the predicted label distribution, and optimize the initial feature structure encoder parameters and the initial structure proxy set in the direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set.

[0119] Each client 130 is also used to upload the class-level structure agent average and feature structure encoder parameters to the server 110, wherein the feature structure encoder parameters include the initial feature structure encoder parameters and the optimized feature structure encoder parameters, and the class-level structure agent average includes the initial class-level structure agent average and the optimized class-level structure agent average.

[0120] The server 110 is configured to perform aggregation based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and perform aggregation based on the class-level structure agent average to obtain a target structure agent.

[0121] The server 110 and the client 130 are used to iterate the above process until the GNN model converges or reaches a preset model accuracy.

[0122] In summary, in the federated graph learning system based on feature fusion and temperature scaling distillation of the present application, the client 130 determines the initial GNN model and the initial structure proxy set based on the local social graph data, and determines the initial feature structure encoder parameters based on the initial structure proxy set. The client 130 obtains a first soft target based on the initial feature structure encoder parameters, calculates a first target loss value based on the first soft target, and optimizes the model parameters in the initial GNN model in the direction of reducing the first target loss value to obtain an optimized GNN model. The client 130 obtains a predicted label distribution based on the optimized GNN model, and calculates a second target based on the predicted label distribution. The target loss value is obtained by optimizing the initial feature structure encoder parameters and the initial structure proxy set in the direction of reducing the second target loss value, and each client 130 uploads the class-level structure proxy average value and the feature structure encoder parameters to the server 110. The server 110 aggregates based on the feature structure encoder parameters to obtain the target feature structure encoder parameters, and aggregates based on the class-level structure proxy average value to obtain the target structure proxy, thereby alleviating the problem of data heterogeneity, reducing the problem of deviation in neighborhood information obtained by minority class nodes during the GNN model training process, and improving the accuracy of node classification.

[0123] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "example," "specific example," "one implementation," "a preferred implementation," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0124] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A federated graph learning method based on feature fusion and temperature scaling distillation, characterized by: The method comprises: The client is based on local social graph data G (k) =(V (k) ,E (k) ,X (k) ) Determine the initial GNN model and the initial structure proxy set, and determine the initial feature structure encoder parameters based on the initial structure proxy set, where V (k) is the set of client nodes, each user is a node, E (k) is an edge set, where an edge indicates a direct relationship between the two nodes it connects. (k) It is a node feature matrix, which includes time series features, geographic location features, and graph structure association features. The time series features are the time distribution of user uploaded photos, the geographic location features are the geocode of the photo shooting location or the user registration location, or the geographic clustering features of the user group. The graph structure association features include adjacency and community division. The adjacency is the attention, comments, and common groups between users. The community division is the user community affiliation identified based on the graph clustering algorithm. The client obtains a first soft target based on the initial feature structure encoder parameters, and calculates a first target loss value based on the first soft target, wherein the first target loss value is calculated based on the first soft target, including: the client generates a first loss value based on the true label in the initial feature structure encoder parameters, generates a second loss value based on the first soft target, calculates the first loss value and the second loss value to obtain the first target loss value, and optimizes the model parameters in the initial GNN model in a direction of reducing the first target loss value to obtain an optimized GNN model, wherein the label is the user's interest label and occupation label; The client obtains a predicted label distribution based on the optimized GNN model, and calculates a second target loss value according to the predicted label distribution, wherein the second target loss value is calculated according to the predicted label distribution, including: the client generates a third loss value based on the true label and a second soft target in the initial feature structure encoder parameter, generates a fourth loss value based on the predicted label distribution, calculates the third loss value and the fourth loss value to obtain the second target loss value, and optimizes the initial feature structure encoder parameters and the initial structure proxy set in a direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set; Each of the clients uploads a class-level structure proxy average value and a feature structure encoder parameter to the server, wherein the feature structure encoder parameter includes the initial feature structure encoder parameter and the optimized feature structure encoder parameter, and the class-level structure proxy average value includes the initial class-level structure proxy average value and the optimized class-level structure proxy average value; The server aggregates based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and aggregates based on the class-level structure agent average to obtain a target structure agent; Iterate the above process until the GNN model converges or reaches the preset model accuracy.

2. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 1 is characterized in that: The client is based on local social graph data G (k) =(V (k) ,E (k) ,X (k) ) Determine the initial GNN model and initial structure agent set, including: The client determines the initial GNN model based on the local social graph data; The client determines the initial structure proxy set based on the local social graph data, and determines the initial feature structure encoder parameters based on the initial structure proxy set.

3. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 2 is characterized in that: The client obtains a first soft target based on the initial feature structure encoder parameters, including: The client converts and combines the initial features in the initial feature structure encoder parameters to generate the first soft target.

4. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 3 is characterized in that: The client converts and combines the initial features in the initial feature structure encoder parameters to generate the first soft target, where the first soft target is: Among them, softmax(·) represents the flexible maximum function, W cls is the weight matrix of the combination operation, b cls is the bias vector of the classifier, is the concatenation of vectors, is the embedding vector, is the initial structure proxy in the local social graph data, k is the kth client, and i is the i-th marked node.

5. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 1 is characterized in that: The first loss value and the second loss value are calculated to obtain the first target loss value, where the first target loss value is: Among them, λ1 is a predefined hyperparameter, is the first loss value, is the second loss value.

6. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 1, characterized in that: The third loss value and the fourth loss value are calculated to obtain the second target loss value, where the second target loss value is: Among them, λ2 is a predefined hyperparameter, is the third loss value, is the fourth loss value.

7. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 1, characterized in that: The server aggregates the characteristic structure encoder parameters to obtain target characteristic structure encoder parameters, where the target characteristic structure encoder parameters are: Among them, |V (k) | is the number of nodes in the local social graph data of client k, |V| is the total number of nodes in the local social graph data of all clients, is the characteristic structure encoder parameter, and K is the number of clients.

8. The federated graph learning method based on feature fusion and temperature scaling distillation according to claim 1, characterized in that: The target structure agent is obtained by aggregating the average value of the class-level structure agent, and the target structure agent is: in, is the number of nodes of category c in the local social graph data in client k, A c is the total number of nodes of category c in the local social graph data of all clients, A proxy average is provided for the class-level structure in the local social graph data.

9. A federated graph learning system based on feature fusion and temperature scaling distillation, characterized by: The system includes: a server and at least one client, wherein: The client is used to (k) =(V (k) ,E (k) ,X (k) ) Determine the initial GNN model and the initial structure proxy set, and determine the initial feature structure encoder parameters based on the initial structure proxy set, where V (k) is the set of client nodes, each user is a node, E (k) is an edge set, where an edge indicates a direct relationship between the two nodes it connects. (k) It is a node feature matrix, which includes time series features, geographic location features, and graph structure association features. The time series features are the time distribution of user uploaded photos, the geographic location features are the geocode of the photo shooting location or the user registration location, or the geographic clustering features of the user group. The graph structure association features include adjacency and community division. The adjacency is the attention, comments, and common groups between users. The community division is the user community affiliation identified based on the graph clustering algorithm. The client is further configured to obtain a first soft target based on the initial feature structure encoder parameters, and calculate a first target loss value based on the first soft target, wherein the first target loss value is calculated based on the first soft target, including: the client generating a first loss value based on the true label in the initial feature structure encoder parameters, generating a second loss value based on the first soft target, calculating the first loss value and the second loss value to obtain the first target loss value, and optimizing the model parameters in the initial GNN model in a direction of reducing the first target loss value to obtain an optimized GNN model, wherein the label is the user's interest label and occupation label; The client is further configured to obtain a predicted label distribution based on the optimized GNN model, and calculate a second target loss value based on the predicted label distribution, wherein the second target loss value is calculated based on the predicted label distribution, including: the client generating a third loss value based on the true label and a second soft target in the initial feature structure encoder parameter, generating a fourth loss value based on the predicted label distribution, calculating the third loss value and the fourth loss value to obtain the second target loss value, and optimizing the initial feature structure encoder parameters and the initial structure proxy set in a direction of reducing the second target loss value to obtain optimized feature structure encoder parameters and optimized structure proxy set; Each of the clients is further configured to upload a class-level structure proxy average value and feature structure encoder parameters to the server, wherein the feature structure encoder parameters include the initial feature structure encoder parameters and the optimized feature structure encoder parameters, and the class-level structure proxy average value includes the initial class-level structure proxy average value and the optimized class-level structure proxy average value; The server is configured to perform aggregation based on the feature structure encoder parameters to obtain target feature structure encoder parameters, and perform aggregation based on the class-level structure agent average to obtain a target structure agent; The server and the client are used to iterate the above process until the GNN model converges or reaches a preset model accuracy.

Citation Information

Patent Citations

  • Federal map learning method based on knowledge distillation and automatic driving method

    CN115907001A

  • Comparative longitudinal federal map learning method based on adaptive multi-kernel clustering

    CN119358029A