Higher-Order Relationship Knowledge Distillation Method and System Based on Heterogeneous Graph Neural Network

By introducing node-level and relation-level knowledge distillation methods into heterogeneous graph neural networks, the problems of inaccurate data annotation and difficulty in semantic relationship modeling are solved, and the representation ability of heterogeneous graph neural networks and the performance of downstream tasks are improved.

CN115115862BActive Publication Date: 2025-07-04INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210553500.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-07-04
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The existing heterogeneous graph neural network (HGNN) has difficulty modeling the semantic relationship between inaccurate data annotation and different types of nodes, resulting in limited representation ability.

Method used

The advanced relational knowledge distillation method based on heterogeneous graph neural network is adopted. Through node-level knowledge distillation and relationship-level knowledge distillation, the soft label knowledge and advanced semantic relationship knowledge of the teacher model are extracted respectively, and integrated into the advanced relational knowledge training student model.

Benefits of technology

It improves the generalization ability and performance of student models and significantly improves the performance of heterogeneous graph neural networks in different downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115862B_ABST
    Figure CN115115862B_ABST
Patent Text Reader

Abstract

The present invention proposes a high-order relational knowledge distillation method and system based on heterogeneous graph neural networks. The method mainly includes two parts: first-order node-level knowledge distillation and second-order relational-level knowledge distillation, effectively solving the two problems of inaccurate data labels and difficult semantic modeling of heterogeneous high-order relations. Specifically, the method encodes the semantics of individual nodes of a pre-trained heterogeneous teacher model through node-level knowledge distillation; and models the semantic relationships between different types of nodes of the pre-trained heterogeneous teacher model through relational-level knowledge distillation. By integrating node-level knowledge distillation and relational-level knowledge distillation, this high-order relational knowledge distillation method becomes a practical and general training method applicable to any heterogeneous graph neural network, not only improving the performance and generalization ability of the heterogeneous student model, but also ensuring the extraction of node-level and relational-level knowledge of the heterogeneous graph neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph data mining, specifically to the field of heterogeneous graph data mining, and more specifically to a high-order relationship knowledge distillation method and system based on a heterogeneous graph neural network. Background Art

[0002] Heterogeneous graphs are prevalent in academic and industrial fields. Recently, researchers have proposed a large number of heterogeneous graph neural networks (HGNNs), and learning node representations in heterogeneous graphs has become a hot topic in current research. Compared with homogeneous graphs, heterogeneous graph modeling has the advantage of integrating more information. However, how to embed the rich structural and semantic information in heterogeneous graphs into low-dimensional node representations is a severe challenge.

[0003] In recent years, to solve the heterogeneity problems of nodes and edges in heterogeneous graphs, researchers have proposed many HGNN-based methods, mainly divided into metapath-based methods and edge-relation-based methods. To capture the heterogeneity of edges, edge-relation-based methods directly use specific relation matrices to process edge relations of various node types in different metric spaces, such as heterogeneous graph neural network models like RGCN, HGT, and HGConv. However, edge-relation-based methods can only capture the local structural information of heterogeneous graphs. To be able to encode the rich semantic information in heterogeneous graphs, metapath-based methods have been proposed. A metapath is an effective semantic mining tool that can capture more complex and richer high-order semantic information between nodes in heterogeneous graphs. Among them, HAN is a pioneering work of metapath-based methods.

[0004] Although existing HGNNs have achieved good performance, their representation capabilities are limited by: (1) inaccurate data annotation. Generally, the training method of HGNNs belongs to semi-supervised learning, so its performance highly depends on a large amount of high-quality labeled data. However, fuzzy data annotation will become a bottleneck for HGNN modeling; (2) it is difficult to model the semantic relationships between different types of nodes. Although metapaths are used for high-order semantic modeling in heterogeneous graphs, the selection of metapaths in different fields is still challenging because it requires sufficient domain knowledge.

[0005] In recent years, the knowledge distillation (KD) technology in deep learning has shown certain advantages in improving the performance of models. Currently, there are some works attempting to combine knowledge distillation methods and graph neural networks for application. But they are all designed for homogeneous graph neural networks, where each node or edge in the processed data is of the same type. Summary of the Invention

[0006] The object of the present invention is to overcome the two major defects of inaccurate data annotation and difficult semantic relationship modeling faced by HGNN in the above-mentioned prior art, and a high-order relationship knowledge distillation method based on heterogeneous graph neural network is proposed, which includes:

[0007] Step S1: Respectively obtain the heterogeneous graph neural network model of the knowledge to be distilled as the teacher model, obtain the heterogeneous graph neural network model to receive the knowledge as the student model, and obtain the model prediction values of the output layers of the teacher model and the student model and the heterogeneous node embedding representations of the intermediate graph convolutional layers;

[0008] Step S2: Based on the model prediction values of the teacher model and the student model, extract the first-order node-level soft label knowledge of the teacher model through node-level knowledge distillation;

[0009] Step S3: Based on the intermediate graph convolutional layer embedding representations of the teacher model and the student model, extract the second-order relationship-level heterogeneous semantic knowledge of the teacher model through relationship-level knowledge distillation;

[0010] Step S4: Integrate the first-order node-level soft label knowledge and the second-order relationship-level heterogeneous semantic knowledge to obtain high-order relationship knowledge, train the student model based on the high-order relationship knowledge, and use the trained student model for the specified task.

[0011] The high-order relationship knowledge distillation method based on heterogeneous graph neural network, wherein the step S1 includes:

[0012] Obtain a heterogeneous dataset D, which includes n training set samples, and the feature dimension of each sample is d-dimensional; construct teacher model T and student model S with the same configuration, each including 5 layers: input layer, first convolutional layer, second convolutional layer, MLP linear transformation layer and Softmax output layer; the teacher and student neural network parameters are W t and W s , and the activation function RELU used in the convolutional layer is f(x) = max(x, 0);

[0013] The heterogeneous node embedding representations of the intermediate graph convolutional layers of the teacher model and the student model include:

[0014] The input sample feature is h 0 , the expression of the convolutional layer is h, then h t = RELU(W t *h 0 ), h s = RELU(W s *h 0 ); the output expression of the MLP linear transformation layer is z, then the output expressions of the linear transformation layers of the teacher and student models are z t and z s;

[0015] The model prediction values of the teacher model and the student model include: If the expression of the Softmax output layer is p, then p t = Softmax(z t ), p s = Softmax(z s ).

[0016] The described high-order relation knowledge distillation method based on heterogeneous graph neural network, wherein step S2 includes:

[0017] Using the teacher and student model prediction values p t , p s , and using the node-level knowledge distillation method to transfer the soft label knowledge in the teacher model to the student model, obtaining the first-order node-level distillation loss L NKD as the first-order node-level soft label knowledge:

[0018] L NKD = (1 - α)L CE + αL KD

[0019] where are the basic cross-entropy loss and the distillation loss respectively, α is the hyperparameter that balances the cross-entropy loss and the distillation loss, D(·) is the KL metric function; in addition is the sfotmax probability output scaled with the temperature coefficient τ.

[0020] The described high-order relation knowledge distillation method based on heterogeneous graph neural network, wherein step S3 includes:

[0021] Using the teacher and student intermediate convolutional layer embedding representations h t , h s , and using the relation-level knowledge distillation method to transfer the high-order semantic relation knowledge in the teacher model to the student model;

[0022] The correlation matrix MetaCorr of the teacher and student network models is:

[0023]

[0024]

[0025] where k is the total number of heterogeneous node types corresponding to the corresponding heterogeneous dataset D, and i, j represent different types of nodes; is the Gaussian kernel function;

[0026] Perform a non-linear transformation on the intermediate layer embedding, and then apply a shared attention vector q to obtain the attention value of the student model

[0027]

[0028] Among them, W s is the weight matrix of the teacher model, and b s is the bias vector;

[0029] Normalize the attention values and obtain the final attention coefficients through the softmax function

[0030]

[0031] Obtain the second-order relational level knowledge distillation loss L RKD , as the second-order relational level heterogeneous semantic knowledge;

[0032]

[0033] where D is the mean squared error.

[0034] The above-mentioned high-order relational knowledge distillation method based on heterogeneous graph neural network, wherein the step S4 includes:

[0035] Integrate L NKD and L RKD , and obtain the overall loss L of the final high-order relational knowledge distillation scheme as high-order relational knowledge to perform end-to-end training on the student model;

[0036] L = L NKD + βL RKD

[0037] where β is the hyperparameter of L NKD and L RKD .

[0038] The above-mentioned high-order relational knowledge distillation method based on heterogeneous graph neural network, wherein the training set samples include movie names, directors, actors, and movie categories, and the specified task includes inputting the movie name and / or director and / or actor to be classified into the student model to obtain the movie category to which it belongs.

[0039] The present invention also proposes a high-order relational knowledge distillation system based on heterogeneous graph neural network, which includes:

[0040] A model acquisition module, which is used to respectively acquire the heterogeneous graph neural network model of the knowledge to be distilled as the teacher model, acquire the heterogeneous graph neural network model to receive the knowledge as the student model, and acquire the model prediction values and intermediate graph convolutional layer heterogeneous node embedding representations of the output layers of the teacher model and the student model;

[0041] The first knowledge extraction module is used to extract the first-order node-level soft label knowledge of the teacher model through node-level knowledge distillation according to the model prediction values of the teacher model and the student model;

[0042] The second knowledge extraction module is used to extract the second-order relational heterogeneous semantic knowledge of the teacher model through relational knowledge distillation based on the intermediate graph convolutional layer embedding representations of the teacher model and the student model;

[0043] The training module is used to integrate the first-order node-level soft label knowledge and the second-order relational heterogeneous semantic knowledge to obtain high-order relational knowledge, train the student model based on the high-order relational knowledge, and use the trained student model for the specified task;

[0044] The model acquisition module is used for:

[0045] Obtain a heterogeneous dataset D, which includes n training set samples, and the feature dimension of each sample is d-dimensional; construct teacher model T and student model S with the same configuration, each containing 5 layers: an input layer, a first convolutional layer, a second convolutional layer, an MLP linear transformation layer, and a Softmax output layer; the teacher and student neural network parameters are W t and W s , and the activation function RELU used in the convolutional layer is f(x) = max(x, 0);

[0046] The intermediate graph convolutional layer heterogeneous node embedding representations of the teacher model and the student model include:

[0047] The input sample feature is h 0 , the expression of the convolutional layer is h, then h t = RELU(W t *h 0 ), h s = RELU(W s *h 0 ); the output expression of the MLP linear transformation layer is z, then the output expressions of the linear transformation layers of the teacher and student models are z t and z s ;

[0048] The model prediction values of the teacher model and the student model include: the expression of the Softmax output layer is p, then p t = Softmax(z t ), p s = Softmax(z s );

[0049] The first knowledge extraction module is used for:

[0050] Adopt the teacher and student model prediction values p t, p s , the soft label knowledge in the teacher model is transferred to the student model using the node-level knowledge distillation method to obtain the first-order node-level distillation loss L NKD As this first-order node-level soft label knowledge:

[0051] L NKD =(1-α)L CE +αL KD

[0052] where are the basic cross-entropy loss and the distillation loss respectively, α is the hyperparameter that balances the cross-entropy loss and the distillation loss, D(·) is the KL metric function; additionally is the sfotmax probability output convex scaled with the temperature coefficient τ

[0053] This second knowledge extraction module is used for:

[0054] Adopt the intermediate convolutional layer embeddings h t , h s , and use the relation-level knowledge distillation method to transfer the high-order semantic relation knowledge in the teacher model to the student model;

[0055] The correlation matrix MetaCorr of the teacher and student network models is:

[0056]

[0057]

[0058] where k is the total number of heterogeneous node types corresponding to the corresponding heterogeneous dataset D, and i, j represent different types of nodes; is the Gaussian kernel function;

[0059] Perform a non-linear transformation on the intermediate layer embedding, and then apply a shared attention vector q to obtain the attention value of the student model

[0060]

[0061] where W s is the weight matrix of the teacher model, and b s is the bias vector;

[0062] Normalize the attention value, and obtain the final attention coefficient through the softmax function

[0063]

[0064] Obtain the second-order relation-level knowledge distillation loss LRKD as second-order relational heterogeneous semantic knowledge;

[0065]

[0066] where D is the mean squared error;

[0067] The training module is used for:

[0068] Integrate L NKD and L RKD to obtain the overall loss L of the final high-order relational knowledge distillation scheme as high-order relational knowledge for end-to-end training of the student model;

[0069] L = L NKD + βL RKD

[0070] where β is the hyperparameter of L NKD and L RKD

[0071] The high-order relational knowledge distillation system based on the heterogeneous graph neural network, where the training set samples include movie names, directors, actors, and movie categories, and the specified task includes inputting the movie name and / or director and / or actor to be classified into the student model to obtain the movie category to which it belongs.

[0072] The present invention also provides a storage medium for storing a program for executing any one of the high-order relational knowledge distillation methods based on the heterogeneous graph neural network.

[0073] The present invention also provides a client for the high-order relational knowledge distillation system based on the heterogeneous graph neural network.

[0074] The embodiment of the present invention provides a high-order relational knowledge distillation method, which applies knowledge distillation to the heterogeneous graph neural network for the first time, filling the gap in extracting knowledge from the heterogeneous graph model. This scheme combines first-order node-level knowledge distillation and second-order relational knowledge distillation and can be flexibly applied to any HGNN model. Through this scheme, the student model can fully utilize and extract the soft label knowledge and high-order heterogeneous relational knowledge hidden in the HGNN. Therefore, the generalization ability of the student model is improved, and its performance is significantly better than its corresponding teacher model. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 is a schematic flowchart of the high-order relational knowledge distillation method based on the heterogeneous graph neural network provided by the embodiment of the present invention;

[0076] Figure 2Schematic diagram of the principle of the high-order relationship knowledge distillation method based on the heterogeneous graph neural network provided by the embodiments of the present invention;

[0077] Figure 3 Schematic diagram of the architecture of the high-order relationship knowledge distillation method based on the heterogeneous graph neural network provided by the embodiments of the present invention. Detailed implementation manners

[0078] The present invention proposes a high-order relationship knowledge distillation method and system based on the heterogeneous graph neural network. The specific technical solutions are as follows. In this paper, classic heterogeneous datasets such as IMDB (including three heterogeneous node types: movies, directors, and actors), ACM (including three heterogeneous node types: papers, authors, and fields), and DBLP (including four heterogeneous node types: papers, conferences, authors, and keywords) are used as examples for illustration:

[0079] According to the first aspect of the present invention, aiming at the problem of inaccurate data label annotation, a first-order node-level knowledge distillation (NKD) method is introduced to transfer the soft labels of target nodes (such as movies in movie data) to students, providing general supervision information for downstream tasks (such as node classification). The method includes:

[0080] Step S1: Construct teacher and student heterogeneous graph neural network models respectively, and obtain the model prediction values of the output layers of the teacher and student and the heterogeneous node embedding representations of the intermediate graph convolutional layers;

[0081] Step S2: Use the model prediction values of the teacher and student networks obtained in step 1, and transfer the first-order node-level soft label knowledge of the pre-trained teacher model to the student model using node-level knowledge distillation;

[0082] Step S3: Use the intermediate graph convolutional layer embedding representations of the teacher and student networks obtained in step 1, and transfer the second-order relationship-level high-order heterogeneous semantic knowledge of the pre-trained teacher model to the student model using relationship-level knowledge distillation (RKD);

[0083] Step S4: Integrate the node-level knowledge and relationship-level knowledge in steps 2 and 3 to obtain the final high-order relationship knowledge distillation scheme, and then train the student model. By minimizing the loss until the student network converges, a well-trained student model is finally obtained, which can be used for different downstream tasks. The downstream tasks for the fields to which the papers belong in the ACM dataset include tasks such as classification, clustering, and visualization; the downstream tasks for movies in IMDB include tasks such as classification, clustering, and visualization; the downstream tasks for the research neighborhoods of authors in DBLP include tasks such as classification, clustering, and visualization.

[0084] In one embodiment of the present invention, step S1 further includes: inputting a heterogeneous dataset and constructing a teacher and student heterogeneous graph neural network model, and the specific dataset and model settings are as follows

[0085] Prepare a heterogeneous dataset D (such as classic heterogeneous data like IMDB, ACM, DBLP, etc.), there are n training set samples, and the feature dimension of each sample is d-dimensional; construct benchmark teacher and student models T and S with the same configuration, including 5 layers: an input layer, a first convolutional layer, a second convolutional layer, an MLP linear transformation layer, and a Softmax output layer; the neural network parameters of the teacher and student are denoted as W t and W s , and the activation function used in the convolutional layer is RELU, in the form of f(x) = max(x, 0).

[0086] In one embodiment of the present invention, step S1 further includes: calculating the output layer model prediction values of the teacher and student and the heterogeneous node embedding representations in the intermediate graph convolutional layer, and the specific calculation is as follows

[0087] Denote the input sample feature as h 0 , and the expression of the convolutional layer is h, then h t = RELU(W t *h 0 ), h s = RELU(W s *h 0 ); denote the output expression of the MLP linear transformation layer as z, then the probability outputs of the teacher and student models are z t and z s , denote the expression of the Softmax output layer as p, then p t = Softmax(z t ), p s = Softmax(z s ).

[0088] In one embodiment of the present invention, step S2 includes: using the teacher and student model prediction values p t , p s , and using the node-level knowledge distillation method to transfer the soft label knowledge in the teacher model to the student model, obtaining a first-order node-level distillation loss L NKD , and the loss function is

[0089] L NKD = (1 - α)L CE + αL KD

[0090] where They are the basic cross - entropy loss and the distillation loss respectively. Here, \(i\) represents a node, \(\alpha\) is a hyperparameter for balancing the cross - entropy loss and the distillation loss, and \(D(\cdot)\) is the KL metric function. Additionally is the softmax probability output scaled with the temperature coefficient \(\tau\). The larger the hyperparameter \(\tau\), the smoother the probability distribution over classes, which promotes the student model to learn more about the smooth information.

[0091] The step S3 includes: using the intermediate convolutional layers of the teacher and the student to embed the representations \(h\) t , \(h\) s , and using the relational - level knowledge distillation method to transfer the high - order semantic relationship knowledge in the teacher model to the student model.

[0092] In an embodiment of the present invention, the step S3 further includes: in order for the student to fully extract the high - order semantic information hidden in the HGNN from the teacher, a MetaCorr correlation matrix is designed to encode the relational - level knowledge between different types of nodes from the pre - trained teacher model. The MetaCorr of the teacher and student network models is calculated as

[0093]

[0094]

[0095] where \(k\) is the total number of heterogeneous node types corresponding to the corresponding heterogeneous dataset, and \(i\), \(j\) represent different types of nodes; is the Gaussian kernel function, which is used to measure the similarity between two node embedding representations. A larger indicates a greater distance between the two node representations. The reason for using the Gaussian RBF kernel is that the Gaussian RBF is more flexible and powerful in capturing the complex non - linear relationships between nodes. To avoid the curse of dimensionality, a second - order Taylor expansion is selected for .

[0096] Meanwhile, a type - related attention layer is introduced behind the convolutional layer to automatically learn the importance of different node types. First, a non - linear transformation is performed on the intermediate layer embedding, and then a shared attention vector \(q\) is applied to obtain the attention values of the student model

[0097]

[0098] where \(W\) s is the weight matrix of the teacher model, and \(b\) s is the bias vector. Then, the attention values are normalized, and the final attention coefficients are obtained through the softmax function

[0099]

[0100] Obviously, the higher α is, the more critical the node is, and α can be dynamically adjusted during model training. Finally, the second-order relationship-level knowledge distillation loss L is obtained. RKD , and the loss function is

[0101]

[0102] where D is the mean squared error (MSE) loss.

[0103] In one embodiment of the present invention, step S4 includes: integrating the node-level knowledge distillation loss L NKD and the relationship-level knowledge distillation loss L RKD , to obtain the overall loss L of the final high-order relationship knowledge distillation scheme, and the loss function is

[0104] L = L NKD + βL RKD

[0105] where β is a hyperparameter for balancing the first-order node-level knowledge distillation and the second-order relationship-level knowledge distillation.

[0106] According to the overall loss L, the student model can be trained end-to-end. By minimizing the loss L until the student network converges, a well-trained student model can be obtained, and thus the student model can be used for different downstream tasks.

[0107] According to a second aspect of the present invention, there is provided a computer-readable storage medium storing one or more computer programs, which are used to implement the high-order relationship knowledge distillation method based on heterogeneous graph neural networks of the present invention when executed.

[0108] According to a third aspect of the present invention, there is provided a computing system, including: a storage device, and one or more processors; wherein, the storage device is used to store one or more computer programs, and the computer programs are used to implement the high-order relationship knowledge distillation method based on heterogeneous graph neural networks of the present invention when executed by the processor.

[0109] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given, and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows.

[0110] It can be seen from the background technology that the performance of the existing HGNN model is limited to: (1) inaccurate data annotation; (2) insufficient high-order relationship semantic modeling. Inspired by the successful application of knowledge distillation technology in deep learning, which has shown certain advantages in improving the performance of the model, some work has attempted to combine the knowledge distillation method with the graph neural network for application. However, these methods are designed for homogeneous graph neural networks, and each node or edge in the processed data is of the same type.

[0111] In response to the above two problems faced by HGNN, the inventors conducted research and designed a high-order relationship knowledge distillation method for heterogeneous graph neural networks to improve the performance of the student heterogeneous graph neural network model. Generally speaking, the method of the present invention is as Figure 1 shown. Step S1: Based on the constructed teacher and student heterogeneous graph neural network models, obtain the model prediction values of the output layers of the teacher and student and the heterogeneous node embedding representations of the intermediate graph convolutional layers respectively; Step S2: Then, based on the obtained model prediction values of the teacher and student networks, use node-level knowledge distillation to transfer the first-order node-level soft label knowledge of the pre-trained teacher model to the student model; Step S3: Then, based on the embedding representations of the intermediate graph convolutional layers of the teacher and student, use relationship-level knowledge distillation to transfer the second-order relationship-level high-order heterogeneous semantic knowledge of the pre-trained teacher model to the student model; Step S4: Finally, integrate the previous node-level knowledge and relationship-level knowledge to obtain the final high-order relationship knowledge, and train the student model. By minimizing the student loss, a well-trained student model can be obtained, which can then be used for different downstream tasks.

[0112] The present invention will be described in detail below with reference to the accompanying drawings. Figure 2 It shows a high-order relationship knowledge distillation method based on a heterogeneous graph neural network provided by the present invention. Figure 3 It shows a high-order relationship knowledge distillation system based on a heterogeneous graph neural network composed of a teacher model and a student model according to an embodiment of the present invention. The method includes the following 4 steps:

[0113] Step S1 : Construct teacher and student heterogeneous graph neural network models respectively, and obtain the model prediction values of the output layers of the teacher and student and the heterogeneous node embedding representations of the intermediate graph convolutional layers.

[0114] According to an embodiment of the present invention, input heterogeneous data D and construct T and S: Prepare the heterogeneous data set D, with n training set samples, and the feature dimension of each sample is d-dimensional; construct benchmark teacher and student models T and S with the same configuration (see Figure 3 ), including 5 layers: an input layer, a first convolutional layer, a second convolutional layer, an MLP linear transformation layer, and a Softmax output layer; the neural network parameters of the teacher and student are denoted as Wt and W s For the convolutional layer, the activation function used is RELU, in the form of f(x) = max(x, 0).

[0115] Calculate the predicted values p of the output layer models of T and S and the heterogeneous node embedding representations h of the intermediate graph convolutional layer. The specific calculation is as follows: Denote the input sample features as h 0 , and the expression of the convolutional layer is h, then h t = RELU(W t *h 0 ), h s = RELU(W s *h 0 ); Denote the output expression of the MLP linear transformation layer as z. Then the probability outputs of the teacher and student models are z t and z s respectively. Denote the expression of the Softmax output layer as p, then p t = Softmax(z t ), p s = Softmax(z s ).

[0116] Step S2: Use p t , p s obtained in Step 1, and transfer the first-order node-level soft label knowledge of the pre-trained T to S using node-level knowledge distillation to obtain the first-order node-level distillation loss L NKD , and the loss function is

[0117] L NKD = (1 - α)L CE + αL KD

[0118] where are the basic cross-entropy loss and the distillation loss respectively, α is the hyperparameter that balances the cross-entropy loss and the distillation loss, and D(·) is the KL metric function; in addition is the sfotmax probability output scaled with the temperature coefficient τ. The larger τ is, the smoother the probability distribution over classes is, which promotes the student model to learn more about the smooth information.

[0119] Step S3: Use the heterogeneous node embedding representations h of the intermediate graph convolutional layers of the T and S networks obtained in Step 1, and transfer the second-order relationship-level high-order heterogeneous semantic knowledge of the pre-trained T to the S model using relationship-level knowledge distillation.

[0120] Among them, in order for students to fully extract the high-order semantic information hidden in the heterogeneous graph neural network from teachers, a MetaCorr correlation matrix is designed to encode the relationship-level knowledge between different types of nodes from the pre-trained T. The MetaCorr calculation of the T and S network models is as follows

[0121]

[0122]

[0123] where k is the total number of heterogeneous node types corresponding to the corresponding heterogeneous dataset, and i, j represent different types of nodes; is the Gaussian kernel function, which is used to measure the similarity between the embedded representations of two nodes. Among them, the larger indicates that the distance between the two node representations is farther. The reason for using the Gaussian RBF kernel is that the Gaussian RBF is more flexible and powerful in capturing the complex nonlinear relationships between various nodes. To avoid the curse of dimensionality, for a second-order Taylor expansion is selected.

[0124] At the same time, a type-related attention layer is introduced behind the convolutional layer to automatically learn the importance of different node types. First, a nonlinear transformation is performed on the intermediate layer embedding, and then a shared attention vector q is applied to obtain the attention value of the student model

[0125]

[0126] where W s is the weight matrix of the T model, and b s is the bias vector. Then, the attention value is normalized, and the final attention coefficient is obtained through the softmax function

[0127]

[0128] Obviously, the higher the α, the more critical the node, and α can be dynamically adjusted during the model training process. Finally, the second-order relationship-level knowledge distillation loss L RKD is obtained, and the loss function is

[0129]

[0130] where D is the mean squared error.

[0131] Step S4: Integrate the node-level knowledge L NKD and the relationship-level knowledge L RKD to obtain the overall loss L of the final high-order relationship knowledge distillation scheme. The loss function is

[0132] L = L NKD + βL RKD

[0133] where β is a hyperparameter that balances first-order node-level knowledge distillation and second-order relation-level knowledge distillation.

[0134] Based on the overall loss L, the S model can be trained end-to-end. By minimizing the loss L until S converges, a well-trained student model can be obtained, which can then be used for different downstream tasks.

[0135] To illustrate the effectiveness of the above scheme of the embodiments of the present invention, experiments are carried out for illustration. The experiments are carried out on several classical heterogeneous graph datasets, and the details are as follows:

[0136] I. Datasets

[0137] The experiments involve 3 benchmark datasets, including 2 citation networks (ACM and DBLP) and 1 movie network (IMDB) dataset. Their relevant descriptions are shown in Table 1 below:

[0138] Table 1 Three heterogeneous graph datasets adopted in this scheme

[0139]

[0140] Among them, the meaning of the column of meta-path is the type of meta-path of the corresponding dataset, which is represented by the node types passed in the meta-path.

[0141] II. Benchmark Models

[0142] To verify the effectiveness of the distillation scheme of the present invention, this experiment will be tested on classical heterogeneous models, namely RGCN, HAN, HGT and HGConv heterogeneous graph neural network models.

[0143] III. Experimental Results

[0144] The high-order relation knowledge distillation algorithm designed in the present invention is applied to RGCN, HAN, HGT and HGConv heterogeneous graph neural network models, and node classification is carried out on three datasets of ACM, IMDB and DBLP. The classification metric is Micro-F1. The specific experimental results are shown in Table 2:

[0145] Table 2 Classification effects of this scheme based on various heterogeneous graph neural networks on heterogeneous datasets

[0146]

[0147] It can be found from Table 2 that by using the high-order relationship knowledge distillation scheme involved in the present invention, the performance of the heterogeneous graph neural network has been significantly and consistently improved, with the improvement range being 0.5% - 9.6%.

[0148] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0149] The present invention also proposes a high-order relationship knowledge distillation system based on a heterogeneous graph neural network, which includes:

[0150] A model acquisition module, configured to respectively acquire a heterogeneous graph neural network model of knowledge to be distilled as a teacher model, acquire a heterogeneous graph neural network model to receive knowledge as a student model, and acquire the model prediction values of the output layers of the teacher model and the student model and the heterogeneous node embedding representations of the intermediate graph convolutional layers;

[0151] A first knowledge extraction module, configured to extract the first-order node-level soft label knowledge of the teacher model through node-level knowledge distillation according to the model prediction values of the teacher model and the student model;

[0152] A second knowledge extraction module, configured to extract the second-order relationship-level heterogeneous semantic knowledge of the teacher model through relationship-level knowledge distillation based on the intermediate graph convolutional layer embedding representations of the teacher model and the student model;

[0153] A training module, configured to integrate the first-order node-level soft label knowledge and the second-order relationship-level heterogeneous semantic knowledge to obtain high-order relationship knowledge, train the student model based on the high-order relationship knowledge, and use the trained student model for a specified task;

[0154] The model acquisition module is configured to:

[0155] Acquire a heterogeneous dataset D, which includes n training set samples, and the feature dimension of each sample is d-dimensional; construct teacher model T and student model S with the same configuration, each including 5 layers: an input layer, a first convolutional layer, a second convolutional layer, an MLP linear transformation layer, and a Softmax output layer; the neural network parameters of the teacher and the student are W t and W s , and the activation function RELU used in the convolutional layer is f(x) = max(x, 0);

[0156] The heterogeneous node embedding representations of the intermediate graph convolutional layers of the teacher model and the student model include:

[0157] The input sample feature is h0 For the convolutional layer, the expression is h, then h t = RELU(W t *h 0 ), h s = RELU(W s *h 0 ); For the output expression of the MLP linear transformation layer is z, the output expressions of the linear transformation layers of the teacher and student models are z t and z s ;

[0158] The model prediction values of the teacher model and the student model include: For the expression of the Softmax output layer is p, then p t = Softmax(z t ), p s = Softmax(z s );

[0159] The first knowledge extraction module is used for:

[0160] Adopt the model prediction values p t , p s , and use the node-level knowledge distillation method to transfer the soft label knowledge in the teacher model to the student model to obtain the first-order node-level distillation loss L NKD as the first-order node-level soft label knowledge:

[0161] L NKD = (1 - α)L CE + αL KD

[0162] where are the basic cross-entropy loss and the distillation loss respectively, α is the hyperparameter that balances the cross-entropy loss and the distillation loss, D(·) is the KL metric function; in addition is the sfotmax probability output scaled with the temperature coefficient τ;

[0163] The second knowledge extraction module is used for:

[0164] Adopt the intermediate convolutional layer embedding representations h t , h s , and use the relation-level knowledge distillation method to transfer the high-order semantic relation knowledge in the teacher model to the student model;

[0165] The correlation matrix MetaCorr of the teacher and student network models is:

[0166]

[0167]

[0168] where k is the total number of heterogeneous node types corresponding to the corresponding heterogeneous dataset D, and i, j represent nodes of different types; is the Gaussian kernel function;

[0169] Perform a non-linear transformation on the intermediate layer embedding, and then apply a shared attention vector q to obtain the attention value of the student model

[0170]

[0171] where W s is the weight matrix of the teacher model, and b s is the bias vector;

[0172] Normalize the attention value to obtain the final attention coefficient through the softmax function

[0173]

[0174] Obtain the second-order relationship level knowledge distillation loss L RKD , as the second-order relationship level heterogeneous semantic knowledge;

[0175]

[0176] where D is the mean squared error;

[0177] This training module is used for:

[0178] Integrate L NKD and L RKD , and obtain the overall loss L of the final high-order relationship knowledge distillation scheme as high-order relationship knowledge to perform end-to-end training on the student model;

[0179] L = L NKD + βL RKD

[0180] where β is the hyperparameter of L NKD and L RKD .

[0181] The described high-order relationship knowledge distillation system based on heterogeneous graph neural network, where the training set samples include movie names, directors, actors, and movie categories, and the specified task includes inputting the movie name and / or director and / or actor to be classified into the student model to obtain the movie category to which it belongs.

[0182] The present invention also proposes a storage medium for storing a program for executing any one of the high-order relationship knowledge distillation methods based on heterogeneous graph neural networks.

[0183] The present invention also provides a client for the high-order relationship knowledge distillation system based on the heterogeneous graph neural network.

Claims

1. A high-order relationship knowledge distillation method based on heterogeneous graph neural network, characterized in that Including: Step S1: Obtain the heterogeneous graph neural network model of the knowledge to be distilled as the teacher model, obtain the heterogeneous graph neural network model of the knowledge to be received as the student model, and obtain the model prediction values of the output layers of the teacher model and the student model and the heterogeneous node embedding representations of the intermediate graph convolutional layers; Step S2: Based on the model prediction values of the teacher model and the student model, extract the first-order node-level soft label knowledge of the teacher model through node-level knowledge distillation; Step S3: Based on the embedding representations of the intermediate graph convolutional layers of the teacher model and the student model, extract the second-order relationship-level heterogeneous semantic knowledge of the teacher model through relationship-level knowledge distillation; Step S4: Integrate the first-order node-level soft label knowledge and the second-order relationship-level heterogeneous semantic knowledge to obtain high-order relationship knowledge, train the student model based on the high-order relationship knowledge, and use the trained student model for the specified task; The training set samples include movie names, directors, actors, and movie categories, and the specified task includes inputting the movie name and / or director and / or actor to be classified into the student model to obtain the movie category to which it belongs.

2. The high-order relationship knowledge distillation method based on the heterogeneous graph neural network according to claim 1, wherein This step S1 includes: Obtain a heterogeneous dataset D, which includes n training set samples, and the feature dimension of each sample is d-dimensional; construct teacher model T and student model S with the same configuration, each containing 5 layers: an input layer, a first convolutional layer, a second convolutional layer, an MLP linear transformation layer, and a Softmax output layer; the teacher and student neural network parameters are W t and W s , and the activation function RELU used in the convolutional layer is f(x) = max(x, 0); The heterogeneous node embedding representations of the intermediate graph convolutional layers of the teacher model and the student model include: The input sample feature is h 0 , the expression of the convolutional layer is h, then h t = RELU(W t *h 0 ), h s = RELU(W s *h 0 ); The output expression of the MLP linear transformation layer is z, then the output expressions of the linear transformation layers of the teacher and student models are z t and z s ; The model prediction values of the teacher model and the student model include: If the expression of the Softmax output layer is p, then p t = Softmax(z t ), p s = Softmax(z s ).

3. The high-order relationship knowledge distillation method based on the heterogeneous graph neural network according to claim 2, characterized in that This step S2 includes: Using the predicted values p of the teacher and student models t , p s , the soft label knowledge in the teacher model is transferred to the student model using the node-level knowledge distillation method to obtain the first-order node-level distillation loss L NKD As this first-order node-level soft label knowledge: L NKD = (1 - α)L CE + αL KD where are the basic cross-entropy loss and the distillation loss respectively, α is a hyperparameter that balances the cross-entropy loss and the distillation loss, and D(·) is the KL metric function; in addition is the softmax probability output scaled with the temperature coefficient τ.

4. The high-order relationship knowledge distillation method based on heterogeneous graph neural network according to claim 3, wherein This step S3 includes: Use the convolutional layer embedding representation h between the teacher and the student t ,h s , and use the relational knowledge distillation method to transfer the high-order semantic relationship knowledge in the teacher model to the student model; The correlation matrix MetaCorr of the teacher and student network models is: where k is the total number of heterogeneous node types corresponding to the corresponding heterogeneous data set D, and i, j represent nodes of different types; is the Gaussian kernel function; Perform a non-linear transformation on the intermediate layer embedding and then apply a shared attention vector q to obtain the attention value of the student model Among which W s is the weight matrix of the teacher model, and b s is the bias vector; Normalize the attention value and obtain the final attention coefficient through the softmax function Obtain the second-order relational knowledge distillation loss \(L\) RKD , as the second-order relational heterogeneous semantic knowledge; where D is the mean squared error.

5. The high-order relationship knowledge distillation method based on heterogeneous graph neural network according to claim 4, wherein This step S4 includes: Integrated L NKD and L RKD to obtain the overall loss L of the final high-order relationship knowledge distillation scheme as high-order relationship knowledge for end-to-end training of the student model; L = L NKD + βL RKD where β is a hyperparameter of L NKD and L RKD .

6. A high-order relationship knowledge distillation system based on heterogeneous graph neural network, characterized in that, Including: A model acquisition module, which is used to respectively obtain the heterogeneous graph neural network model of the knowledge to be distilled as the teacher model, obtain the heterogeneous graph neural network model of the knowledge to be received as the student model, and obtain the model prediction values of the output layers of the teacher model and the student model and the heterogeneous node embedding representations of the intermediate graph convolutional layers; A first knowledge extraction module, which is used to extract the first-order node-level soft label knowledge of the teacher model through node-level knowledge distillation according to the model prediction values of the teacher model and the student model; A second knowledge extraction module, which is used to extract the second-order relationship-level heterogeneous semantic knowledge of the teacher model through relationship-level knowledge distillation based on the embedding representations of the intermediate graph convolutional layers of the teacher model and the student model; A training module, which is used to integrate the first-order node-level soft label knowledge and the second-order relationship-level heterogeneous semantic knowledge to obtain high-order relationship knowledge, train the student model based on the high-order relationship knowledge, and use the trained student model for the specified task; This model acquisition module is used for: Obtain a heterogeneous dataset D, which includes n training set samples, and the feature dimension of each sample is d-dimensional; construct teacher model T and student model S with the same configuration, each containing 5 layers: an input layer, a first convolutional layer, a second convolutional layer, an MLP linear transformation layer, and a Softmax output layer; the teacher and student neural network parameters are W t and W s , and the activation function RELU used in the convolutional layer is f(x) = max(x, 0); The heterogeneous node embedding representations of the intermediate graph convolutional layers of the teacher model and the student model include: The input sample feature is h 0 , and the expression of the convolutional layer is h, then h t = RELU(W t * h 0 ), h s = RELU(W s * h 0 ); The output expression of the MLP linear transformation layer is z, then the output expressions of the linear transformation layers of the teacher and student models are z t and z s ; The model prediction values of the teacher model and the student model include: If the expression of the Softmax output layer is p, then p t = Softmax(z t ), p s = Softmax(z s ); This first knowledge extraction module is used for: Using the predicted values p of the teacher and student models t , p s , the soft label knowledge in the teacher model is transferred to the student model using the node-level knowledge distillation method to obtain the first-order node-level distillation loss L NKD As this first-order node-level soft label knowledge: L NKD = (1 - α)L CE + αL KD where are the basic cross-entropy loss and the distillation loss respectively, α is the hyperparameter that balances the cross-entropy loss and the distillation loss, and D(·) is the KL metric function; additionally is the softmax probability output scaled with the temperature coefficient τ; This second knowledge extraction module is used for: Using the intermediate convolutional layer of the teacher and student to embed and represent h t , h s , and using the relational knowledge distillation method to transfer the high-order semantic relationship knowledge in the teacher model to the student model; The correlation matrix MetaCorr of the teacher and student network models is: where k is the total number of heterogeneous node types corresponding to the corresponding heterogeneous dataset D, and i, j represent nodes of different types; is the Gaussian kernel function; Perform a non-linear transformation on the intermediate layer embedding and then apply a shared attention vector q to obtain the attention value of the student model Among which W s is the weight matrix of the teacher model, and b s is the bias vector; Normalize the attention value and obtain the final attention coefficient through the softmax function Obtain the second-order relational level knowledge distillation loss L RKD , as the second-order relational level heterogeneous semantic knowledge; where D is the mean squared error; This training module is used for: Integrated L NKD and L RKD , to obtain the overall loss L of the final high-order relation knowledge distillation scheme as high-order relation knowledge for end-to-end training of the student model; L = L NKD + βL RKD where β is the hyperparameter of L NKD and L RKD ; The training set samples include movie names, directors, actors, and movie categories, and the specified task includes inputting the movie name and / or director and / or actor to be classified into the student model to obtain the movie category to which it belongs.

7. A storage medium for storing a program for executing any one of the high-order relationship knowledge distillation methods based on heterogeneous graph neural networks as claimed in claims 1 to 5.

8. A client for the high-order relationship knowledge distillation system based on heterogeneous graph neural network according to claim 6.