Graph structure data processing method based on double graph neural network knowledge distillation

Through the knowledge distillation method based on dual graph neural network, the knowledge alignment and transmission is used to align and transmit knowledge with KL divergence and attention mechanism, the problem of high computational cost, high memory usage and low accuracy when processing graph structure data is solved, and the generalization performance and expression ability of the student model are improved.

CN120124680APending Publication Date: 2025-06-10BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510197006.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing graph neural networks have high computational cost, high memory usage and low accuracy when processing graph structure data. The student model is susceptible to the teacher model and its own excessive smoothing, resulting in a decline in overall performance.

Method used

The knowledge distillation method based on dual graph neural network is adopted, and the teacher model and student model are constructed, the knowledge alignment and transmission is performed using the calculation of KL divergence and attention mechanism, node weights are dynamically allocated, and the student model is trained by constructing a mixed loss function.

Benefits of technology

It realizes more effective knowledge transfer from teacher model to student model, avoids excessive smoothing problems in student model during training, and improves the generalization performance and expression ability of student model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124680A_ABST
    Figure CN120124680A_ABST
Patent Text Reader

Abstract

The invention provides a graph structure data processing method based on double graph neural network knowledge distillation, and relates to the technical field of graph structure data processing, and the method comprises the steps: inputting a graph structure data training set into a teacher model and a student model, and obtaining the knowledge distillation loss of an output layer; obtaining layer-by-layer knowledge distillation loss by utilizing an attention-based self-adaptive layer-by-layer knowledge alignment mechanism; utilizing an attention-based adaptive interlayer knowledge transfer mechanism to obtain interlayer knowledge distillation loss; and constructing a mixed loss function based on the knowledge distillation loss of the output layer, the layer-by-layer knowledge distillation loss and the interlayer knowledge distillation loss, training the student model to obtain a trained student model, processing the graph structure data set to obtain a graph structure data processing result, and storing the graph structure data processing result. And graph structure data processing based on double graph neural network knowledge distillation is completed. According to the method, the problems of high calculation cost, high memory occupation and low accuracy of graph neural network processing graph structure data are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of graph structure data processing, and particularly relates to a graph structure data processing method based on dual graph neural network knowledge distillation. Background Art

[0002] Graph Neural Network (GNN) extends the advantages of Convolutional Neural Network (CNN) to non-Euclidean data and has been widely applied in multiple fields. With the increasing demand, GNN models have evolved from simple Graph Convolutional Network (GCN) to more complex large models such as Metapath Aggregated Graph Neural Network (MAGNN) and Graph Foundation Model (GFM). However, the irregularity of graph structure data and the complexity of the models will lead to an increase in inference latency in real-time applications, such as in recommendation systems, image search, and spam detection. Therefore, the Knowledge Distillation (KD) technology of graph neural network can reduce the computational cost and memory occupancy while ensuring high accuracy by transferring the knowledge of complex models to lightweight models, which is especially suitable for real-time scenarios.

[0003] Existing methods mostly focus on the node embedding alignment of the final layer and do not fully utilize the information of the intermediate layer, and even neglect it to a certain extent. In addition, the student model is vulnerable to the influence of the teacher model and its own over-smoothing, resulting in a decline in overall performance. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, a graph structure data processing method based on dual graph neural network knowledge distillation provided by the present invention solves the problems of high computational cost, high memory occupancy, and low accuracy of graph neural network in processing graph structure data.

[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is: a graph structure data processing method based on dual graph neural network knowledge distillation, comprising:

[0006] S1: Construct a teacher model and a student model;

[0007] S2: Input the graph structure data training set into the teacher model and the student model, and obtain the knowledge distillation loss of the output layer by calculating the KL divergence;

[0008] S3: Use the attention-based adaptive layer-by-layer knowledge alignment mechanism to match each layer of the student model with the most critical corresponding layer in the teacher model to obtain the layer-by-layer knowledge distillation loss;

[0009] S4: Use the attention-based adaptive inter-layer knowledge transfer mechanism to dynamically allocate the weights of each node in the student model to obtain the inter-layer knowledge distillation loss;

[0010] S5: Construct a hybrid loss function based on the knowledge distillation loss of the output layer, the layer-by-layer knowledge distillation loss, and the inter-layer knowledge distillation loss, and train the student model to obtain a trained student model; wherein, the trained student model is used to process the graph structure data set to obtain a graph structure data processing result, and complete the graph structure data processing based on the dual graph neural network knowledge distillation.

[0011] Further, the expression of the knowledge distillation loss of the output layer is:

[0012]

[0013] wherein, L output represents the knowledge distillation loss of the output layer, V represents any node, L CE represents the cross-entropy loss, represents the predicted value of the student model, y v represents the true label, λ represents the weighting parameter, L KL represents the KL divergence loss, z v represents the soft target generated by the teacher model, y v,i represents the probability of the true label on class i, represents the predicted probability of the student model on class i, KL represents the KL divergence function, softmax represents the softmax function, ρ represents the temperature parameter, and || represents the concatenation operation.

[0014] Further, the S3 includes:

[0015] Generate an importance matrix based on the teacher model and the student model;

[0016] For each layer in the student model, use the importance matrix to select the teacher layer for alignment to obtain the aligned teacher model and student model;

[0017] Extract features from the aligned teacher model and student model to obtain the teacher model edge feature embedding and the student model edge feature embedding respectively;

[0018] Using the teacher model edge feature embedding and the student model edge feature embedding, alignment loss is calculated to obtain the layer-by-layer knowledge distillation loss.

[0019] Further, the expression of the importance matrix is:

[0020]

[0021] where A represents the importance matrix, softmax represents the softmax function, T p represents the feature matrix obtained after linear projection of the average node embedding matrix of each layer of the teacher model, represents the feature matrix obtained after linear projection of the average node embedding matrix of each layer of the teacher model, d a represents the dimension of the projection space.

[0022] Further, the expression of the layer-by-layer knowledge distillation loss is:

[0023]

[0024] where L layer-wise represents the layer-by-layer knowledge distillation loss, j represents the j-th layer of the student model, s l represents the number of layers of the student model, (u, v) represents any edge, E represents the edge set, Sim represents the similarity measurement function, represents the edge feature embedding of the i * -th layer of the teacher model that is the most important for the j-th layer of the student model, represents the edge feature embedding of the j-th layer of the student model.

[0025] Further, the S4 includes:

[0026] By dynamically allocating the neighbor node weights of the student model, attention scores are obtained;

[0027] The attention scores are normalized to obtain attention coefficients;

[0028] Using the attention coefficients, by calculating the nodes in each layer, dissimilarity vectors are obtained;

[0029] Using the dissimilarity vectors to calculate the transfer loss to obtain the inter-layer knowledge distillation loss.

[0030] Further, the expression of the attention score is:

[0031]

[0032] where, denotes the attention score, and LeakyReLU denotes the LeakyReLU function. denotes the attention vector of the l-th layer, and || denotes the concatenation operation. and denotes the output of the hidden layer of the l-th layer;

[0033] The expression of the attention coefficient is:

[0034]

[0035] where denotes the attention coefficient, denotes the attention score of node i to neighbor node j, exp denotes the exponential form of the natural constant e, k denotes any neighbor, N ( i ) denotes the neighbor set of node i, denotes the attention score of node i to neighbor node k;

[0036] The expression of the dissimilarity vector is:

[0037]

[0038] where denotes the dissimilarity vector, V denotes the node set, denotes the dissimilarity vector of the l-th layer;

[0039] The expression of the inter-layer knowledge distillation loss is:

[0040]

[0041] where L layer-inter denotes the inter-layer knowledge distillation loss, i denotes any node, I denotes the indicator function, denotes the dissimilarity vector of the l-th layer, denotes the dissimilarity vector of the (l + 1)-th layer, τ (l) denotes the threshold of the l-th layer, denotes the Euclidean norm, denotes the average dissimilarity vector of the l-th layer, denotes the average dissimilarity vector of the (l + 1)-th layer.

[0042] Furthermore, the expression of the hybrid loss function is:

[0043] L ALL = L output + α · L layer-wise + β · L layer-inter ;

[0044] Among them, L ALL represents the hybrid loss function, L output represents the knowledge distillation loss of the output layer, α and β represent the balance parameters, and L layer-wise represents the layer-by-layer knowledge distillation loss, and L layer-inter represents the inter-layer knowledge distillation loss.

[0045] The beneficial effects of the present invention are as follows: The processor obtains a trained student model by using the knowledge distillation method of the dual graph neural network, realizes more effective knowledge transfer from the teacher model to the student model, and at the same time avoids the over-smoothing problem that occurs in the training process of the student model, thereby improving the generalization performance and expression ability of the student model. Description of the Drawings

[0046] This specification will further illustrate in the manner of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0047] Figure 1 is an exemplary flowchart of a graph structure data processing method based on knowledge distillation of a dual graph neural network shown in some embodiments of this specification. Detailed Embodiments

[0048] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0049] Embodiment

[0050] Figure 1 is an exemplary flowchart of a graph structure data processing method based on knowledge distillation of a dual graph neural network shown in some embodiments of this specification. As Figure 1 shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.

[0051] S1: Construct a teacher model and a student model.

[0052] The teacher model is a graph neural network model used to guide the training of the student model. For example, the teacher model can include GCN model, GCNII model, GAT model, APPNP model, etc.

[0053] In some embodiments, the processor can pre-train the GCN model, GCNII model, GAT model, and APPNP model to obtain a teacher model.

[0054] The student model is an initial graph neural network model for completing graph-structured data processing tasks. For example, the student model can include a shallow GCN model and a lightweight GAT model, etc.

[0055] In some embodiments, the processor can select a graph neural network model with the same architecture as the teacher model as the student model according to the actual requirements of graph-structured data processing tasks, such as tasks based on recommendation systems, image search, and spam detection.

[0056] S2: Input the graph-structured data training set into the teacher model and the student model, and calculate the knowledge distillation loss of the output layer by computing the KL divergence.

[0057] The graph-structured data training set is a graph-structured data set for training the student model. For example, the image training set can include the Cora dataset with 2708 nodes, 5429 edges, 1433 node features, and 7 categories.

[0058] In some embodiments, the processor can use the student model to perform a node classification task on the graph-structured data training set, that is, classify the nodes into predefined categories according to the node features and graph structure to obtain a classification result, and train the student model.

[0059] The knowledge distillation loss of the output layer is a loss function that comprehensively considers the similarity between the soft labels output by the teacher model and the logical output of the student model and the difference between the output of the student model and the true labels.

[0060] In some embodiments, the processor can weight and combine the KL divergence loss and the cross-entropy loss to form the knowledge distillation loss of the output layer.

[0061] The KL divergence loss is based on the soft labels of the output layer of the teacher model and the logical output values of the student model.

[0062] The cross-entropy loss is the loss value between the true labels on the training set and the logical output values of the student model.

[0063] In some embodiments, the processor can use the teacher model to generate soft targets for each node in the graph-structured data training set, use the student model to output the predicted values of each node in the image training set, and calculate the loss using the true labels, predicted values, and soft targets of the nodes to obtain the knowledge distillation loss of the output layer.

[0064] In some embodiments, the expression of the knowledge distillation loss of the output layer can be:

[0065]

[0066] Among them, L output represents the knowledge distillation loss of the output layer, V represents any node, and L CE represents the cross-entropy loss, represents the predicted value of the student model, y v represents the true label, λ represents the weighting parameter, and L KL represents the KL divergence loss, z v represents the soft target generated by the teacher model, y v,i represents the probability of the true label on class i, represents the predicted probability of the student model on class i, KL represents the KL divergence function, softmax represents the softmax function, ρ represents the temperature parameter, and || represents the concatenation operation.

[0067] The KL divergence loss and the cross-entropy loss are weighted and combined to form the final knowledge distillation loss of the output layer, ensuring that the student model can balance accuracy and compliance with the output of the teacher model.

[0068] S3: Use the attention-based adaptive layer-by-layer knowledge alignment mechanism to match each layer of the student model with the most critical corresponding layer in the teacher model to obtain the layer-by-layer knowledge distillation loss.

[0069] The layer-by-layer knowledge distillation loss is the knowledge distillation loss generated after each layer of the student model is aligned with the teacher model.

[0070] In some embodiments, the processor can implement S3 based on the following steps: generate an importance matrix based on the teacher model and the student model; for each layer in the student model, use the importance matrix to select the teacher layer for alignment to obtain the aligned teacher model and student model; perform feature extraction on the aligned teacher model and student model to obtain the teacher model edge feature embedding and the student model edge feature embedding respectively; use the teacher model edge feature embedding and the student model edge feature embedding to calculate the alignment loss to obtain the layer-by-layer knowledge distillation loss.

[0071] The importance matrix is a matrix used to analyze the importance scores of each layer of the teacher model for each layer in the student model.

[0072] In some embodiments, the processor can construct the importance matrix based on the number of layers of the teacher model and the student model and the importance scores of each layer of the teacher model corresponding to each layer of the student model. For example, the processor can take the average value of the node embedding matrix along the node dimension to calculate the average node embedding of each layer of the teacher and student networks, and represent these tensors as and Convert T a Each layer in is converted to Generate Obtain the importance matrix; where T a Represents the embeddings of each layer of the teacher model, R represents the real vector space, and t l Represents the number of layers of the teacher model, and d t Represents the node feature dimension of the teacher model, and S a Represents the embeddings of each layer of the student model, and s l Represents the number of layers of the student model, and d s Represents the node feature dimension of the student model.

[0073] In some embodiments, the expression of the importance matrix can be:

[0074]

[0075] Where A represents the importance matrix, softmax represents the softmax function, and T p Represents the feature matrix obtained after linear projection of the average node embedding matrix of each layer of the teacher model, Represents the feature matrix obtained after linear projection of the average node embedding matrix of each layer of the teacher model, and d a Represents the dimension of the projection space.

[0076] The aligned teacher model and student model are the teacher model and student model after each layer of the student model is matched one by one with the teacher model at the maximum index layer.

[0077] In some embodiments, the processor can index the layers in the teacher model based on each layer of the student model using the importance matrix to obtain the maximum index value as the matching teacher model layer:

[0078]

[0079] Where i * Represents the maximum index value, Represents the maximum value index function, and A ij Represents the value of the i-th layer of the teacher model in the importance matrix for the j-th layer of the student model.

[0080] The teacher model edge feature embedding is the information in the teacher model that reflects the structural features of each edge.

[0081] The student model edge feature embedding is the information in the student model that reflects the structural features of each edge.

[0082] In some embodiments, the processor may calculate the loss while following the LSP (Local Structure Preserving) principle based on the aligned teacher model and student model, ensuring that the structural features of each edge are preserved, and obtaining the edge feature embedding of the teacher model and the edge feature embedding of the student model.

[0083] In some embodiments, the expression of the edge feature embedding may be:

[0084] h (u,v) = mean(h u + h v );

[0085] where h (u,v) represents the edge feature embedding, mean represents the averaging operation, h u represents the feature embedding of node u, and h v represents the feature embedding of node v.

[0086] In some embodiments, the edge feature embedding of the teacher model and the edge feature embedding of the student model may be denoted as: and

[0087] In some embodiments, the processor realizes the optimal alignment between each layer of the student model and the key layer of the teacher model by selecting the corresponding relationship with the highest score in the importance matrix, and constructs the layer-by-layer knowledge distillation loss.

[0088] In some embodiments, the expression of the layer-by-layer knowledge distillation loss may be:

[0089]

[0090] where L layer-wise represents the layer-by-layer knowledge distillation loss, j represents the j-th layer of the student model, s l represents the number of layers of the student model, (u, v) represents any edge, E represents the edge set, Sim represents the similarity measurement function, represents the edge feature embedding of the i * -th layer of the teacher model that is the most important for the j-th layer of the student model, represents the edge feature embedding of the j-th layer of the student model.

[0091] The similarity measurement function may include cosine similarity and Euclidean distance, etc.

[0092] S4: Utilize the attention-based adaptive inter-layer knowledge transfer mechanism to dynamically allocate the weights of each node in the student model, and obtain the inter-layer knowledge distillation loss.

[0093] The inter-layer knowledge distillation loss reflects the knowledge distillation loss between the layers of the student model.

[0094] The inter-layer knowledge distillation loss can be used to ensure that the student model can maintain the shallow inductive bias while avoiding being affected by the over-smoothing trends of the teacher model and itself.

[0095] In some embodiments, the processor can implement S4 based on the following steps: obtaining an attention score by dynamically allocating the weights of the neighbor nodes of the student model; normalizing the attention score to obtain an attention coefficient; using the attention coefficient to calculate a dissimilarity vector by calculating the nodes in each layer; and calculating a transfer loss using the dissimilarity vector to obtain the inter-layer knowledge distillation loss.

[0096] The attention score reflects the importance score of the node for its neighbor nodes.

[0097] In some embodiments, the expression of the attention score can be:

[0098]

[0099] where, represents the attention score, LeakyReLU represents the LeakyReLU function, represents the attention vector of the l-th layer, || represents the concatenation operation, and represent the hidden layer output of the l-th layer.

[0100] The attention coefficient is the normalized attention score.

[0101] In some embodiments, the expression of the attention coefficient can be:

[0102]

[0103] where, represents the attention coefficient, represents the attention score of node i for neighbor node j, exp represents the exponential form of the natural constant e, k represents any neighbor, N ( i ) represents the neighbor set of node i, represents the attention score of node i for neighbor node k.

[0104] The dissimilarity vector is a vector that measures the overall dissimilarity between the nodes in each layer.

[0105] In some embodiments, the expression of the dissimilarity vector can be:

[0106]

[0107] Among them, represents the dissimilarity vector, V represents the set of nodes, represents the dissimilarity vector of the l-th layer.

[0108] In some embodiments, the expression of the inter-layer knowledge distillation loss can be:

[0109]

[0110] Among them, L layer-inter represents the inter-layer knowledge distillation loss, i represents any node, I represents the indicator function, represents the dissimilarity vector of the l-th layer, represents the dissimilarity vector of the (l + 1)-th layer, τ (l) represents the threshold of the l-th layer, represents the Euclidean norm, represents the average dissimilarity vector of the l-th layer, represents the average dissimilarity vector of the (l + 1)-th layer.

[0111] By introducing self-distillation regularization to control the dissimilarity between adjacent layers of the student model and prevent the student model from being overly smoothed; adopting the GAT (Graph Attention Network) mechanism to calculate the attention weights of each layer of nodes in the student model to neighbor nodes to evaluate the importance of nodes to neighbor nodes; calculating the feature differences of each layer of nodes in the student model through dissimilarity measurement and imposing regularization constraints on nodes exceeding the preset threshold to ensure that the model maintains feature differences during the knowledge distillation process.

[0112] S5: Construct a hybrid loss function based on the knowledge distillation loss of the output layer, the layer-by-layer knowledge distillation loss, and the inter-layer knowledge distillation loss, and train the student model to obtain a trained student model; wherein, the trained student model is used to process the graph structure data set to obtain a graph structure data processing result, completing the graph structure data processing based on dual graph neural network knowledge distillation.

[0113] The hybrid loss function is the final loss function used to train the student model.

[0114] In some embodiments, the processor can synthesize the knowledge distillation loss of the output layer, the layer-by-layer knowledge distillation loss, and the inter-layer knowledge distillation loss to obtain the hybrid loss function.

[0115] In some embodiments, the expression of the result of the hybrid loss function can be:

[0116] L ALL = L output + α·Llayer-wise +β·L layer-inter ;

[0117] Wherein, L ALL represents the result of the hybrid loss function, L output represents the knowledge distillation loss of the output layer, α and β represent balance parameters, L layer-wise represents the layer-by-layer knowledge distillation loss, L layer-inter represents the inter-layer knowledge distillation loss.

[0118] In some embodiments, the student model can be trained using a labeled image training set. For example, the labeled image training set can be input into the initial student model, and a hybrid loss function can be constructed based on the labels and the results of the initial student model. The parameters of the initial student model can be iteratively updated based on the hybrid loss function by gradient descent or other methods. When a preset condition is satisfied, the model training is completed, and the trained student model is obtained. Among them, the preset condition can be that the loss function converges, the number of iterations reaches a threshold, etc.

[0119] In some embodiments, the processor can use the trained student model to process the image data set to obtain the graph structure data processing result, and complete the graph structure data processing based on the knowledge distillation of the dual graph neural network.

[0120] In some embodiments of the present specification, the processor uses the knowledge distillation method of the dual graph neural network to obtain a trained student model, realizes more effective knowledge transfer from the teacher model to the student model, and at the same time avoids the over-smoothing problem that occurs during the training of the student model, thereby improving the generalization performance and expression ability of the student model.

Claims

1. A graph structure data processing method based on dual graph neural network knowledge distillation, characterized in that: include: S1: Construct teacher model and student model; S2: Input the graph structure data training set into the teacher model and the student model, and obtain the knowledge distillation loss of the output layer by calculating the KL divergence; S3: Using the attention-based adaptive layer-by-layer knowledge alignment mechanism, each layer of the student model is matched with the most critical corresponding layer in the teacher model to obtain the layer-by-layer knowledge distillation loss; S4: Using the attention-based adaptive inter-layer knowledge transfer mechanism, the weights of each node in the student model are dynamically allocated to obtain the inter-layer knowledge distillation loss; S5: Construct a mixed loss function based on the knowledge distillation loss of the output layer, the layer-by-layer knowledge distillation loss and the inter-layer knowledge distillation loss, train the student model, and obtain a trained student model; wherein the trained student model is used to process the graph structure data set to obtain a graph structure data processing result, and complete the graph structure data processing based on dual graph neural network knowledge distillation.

2. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 1 is characterized in that: The expression of the knowledge distillation loss of the output layer is: Among them, L output represents the knowledge distillation loss of the output layer, V represents any node, and L CE represents the cross entropy loss, represents the predicted value of the student model, y v represents the true label, λ represents the weighting parameter, L KL represents the KL divergence loss, z v represents the soft target generated by the teacher model, y v,i represents the probability of the true label on category i, represents the predicted probability of the student model on category i, KL represents the KL divergence function, softmax represents the softmax function, ρ represents the temperature parameter, and || represents the connection operation.

3. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 1 is characterized in that: The S3 includes: Based on the teacher model and the student model, generating an importance matrix; For each layer in the student model, using the importance matrix, select a teacher layer for alignment to obtain an aligned teacher model and a student model; Performing feature extraction on the aligned teacher model and student model to obtain edge feature embedding of the teacher model and edge feature embedding of the student model respectively; The teacher model edge feature embedding and the student model edge feature embedding are used to calculate the alignment loss to obtain the layer-by-layer knowledge distillation loss.

4. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 3 is characterized in that: The expression of the importance matrix is: Among them, A represents the importance matrix, softmax represents the softmax function, T p represents the feature matrix obtained by linearly projecting the average node embedding matrix of each layer of the teacher model, represents the feature matrix obtained by linearly projecting the average node embedding matrix of each layer of the teacher model, d a Represents the dimension of the projection space.

5. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 3 is characterized in that: The expression of the layer-by-layer knowledge distillation loss is: Among them, L layer-wise represents the layer-by-layer knowledge distillation loss, j represents the jth layer of the student model, and s l represents the number of student model layers, (u,v) represents any edge, E represents the edge set, Sim represents the similarity measurement function, Represents the most important teacher model i for the jth layer of the student model * Layer edge feature embedding, represents the edge feature embedding of the j-th layer of the student model.

6. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 1 is characterized in that: The S4 includes: By dynamically allocating weights of neighbor nodes of the student model, an attention score is obtained; Normalizing the attention score to obtain an attention coefficient; Using the attention coefficient, a dissimilarity vector is obtained by calculating the nodes in each layer; The dissimilarity vector is used to calculate the transfer loss to obtain the inter-layer knowledge distillation loss.

7. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 6 is characterized in that: The expression of the attention score is: in, represents the attention score, LeakyReLU represents the LeakyReLU function, represents the attention vector of the lth layer, || represents the connection operation, and represents the hidden layer output of the lth layer; The expression of the attention coefficient is: in, represents the attention coefficient, represents the attention score of node i to neighbor node j, exp represents the exponential form of the natural constant e, k represents any neighbor, N(i) represents the neighbor set of node i, represents the attention score of node i to neighbor node k; The expression of the dissimilarity vector is: in, represents the dissimilarity vector, V represents the node set, represents the dissimilarity vector of the lth layer; The expression of the inter-layer knowledge distillation loss is: Among them, L layer-inter represents the inter-layer knowledge distillation loss, i represents any node, I represents the indicator function, represents the dissimilarity vector of the lth layer, represents the dissimilarity vector of the l+1th layer, τ (l) represents the threshold of the lth layer, represents the Euclidean norm, represents the average dissimilarity vector of the lth layer, represents the average dissimilarity vector of the l+1th layer.

8. The graph structure data processing method based on dual graph neural network knowledge distillation according to claim 1 is characterized in that: The expression of the mixed loss function is: L ALL =L output +α·L layer-wise +β·L layer-inter ; Among them, L ALL represents the mixed loss function, L output represents the knowledge distillation loss of the output layer, α and β represent the balance parameters, and L layer-wise represents the layer-by-layer knowledge distillation loss, L layer-inter represents the inter-layer knowledge distillation loss.

Citation Information

Cited By

  • Generation method of power grid topology analysis model based on graph convolutional network

    CN120832580A