Supervised contrast learning method for graph zero sample learning

By adopting the supervised comparison learning method in graph zero sample learning, combining graph diffusion technology and fully connected neural networks, the training of graph convolutional neural networks is optimized, and the problem of insufficient generalization ability in the existing technology is solved, and better unseen type recognition and new category exploration ability are achieved.

CN120147684APending Publication Date: 2025-06-13SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111562.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing graph zero-sample learning method has the problem that insufficient generalization ability leads to the inability to accurately classify unseen data.

Method used

A supervised contrast learning method is adopted to extract class semantic descriptions and node feature matrix, class diagrams and node diagrams are constructed, graph diffusion technology and fully connected neural networks are used for potential representation learning, and the training of graph convolutional neural networks is optimized by combining self-alignment and affinity alignment loss functions.

Benefits of technology

It improves the generalization of the model for unknown types, enhances the ability to identify newly emerging categories, alleviates the distribution offset problem, and improves the overall performance of zero-sample learning of graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147684A_ABST
    Figure CN120147684A_ABST
Patent Text Reader

Abstract

The invention discloses a supervised comparative learning method for graph zero sample learning, which comprises the following steps of: acquiring undirected graph structure data, extracting a category semantic description matrix of graph data, constructing an adjacent matrix of a category by utilizing a k-nearest neighbor method, and acquiring an affinity node set of each node in the graph data through a graph diffusion technology; and projecting the node feature matrix and the category semantic description matrix to the same dimension, and inputting the node and category adjacency matrix into a graph convolutional neural network to obtain a potential representation matrix of the nodes and categories. Constructing a supervised comparative learning objective function based on the potential representation matrixes and the affinity node sets, and using the objective function to guide iterative optimization of the graph convolutional neural network until the model is converged and training is completed; and inputting the feature matrix of the test node and the corresponding adjacent matrix into a network to obtain a potential representation vector of each test node, and calculating the similarity between the potential representation vector and all potential representation vectors without categories to obtain a category label of each test node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to classification technology, and particularly to a supervised contrastive learning method for graph zero-shot learning. Background Art

[0002] Graph-structured data is prevalent in various real-world scenarios, such as social networks and molecular graphs. Understanding the interactions between entities in these networks is crucial. A key task in this field is node classification, whose goal is to classify unlabeled nodes based on a small number of labeled nodes in the graph. Recently, there has been a great deal of interest in using graph neural networks (GNNs) to accomplish this task, and significant success has been achieved. However, graphs typically evolve dynamically with the emergence of nodes and edges, inevitably giving rise to new categories. For example, in citation networks, the publication of new research papers often brings new interdisciplinary topics, which usually involve the re-integration of domain knowledge and cross-innovation, and there may be a lack of prior annotations or relevant data in the citation network. Unfortunately, traditional GNNs usually assume that all categories are fixed and covered by the categories of the labeled nodes. When encountering newly emerging categories, GNNs need to collect a large amount of labeled data of the new categories to achieve satisfactory performance. However, labeling these tags is a time-consuming and costly process.

[0003] For fast annotation, graph zero-shot learning methods have emerged in the prior art. To achieve effective graph zero-shot learning, some previous works have conducted preliminary explorations. DGPN designed a graph-based generalized model based on the obtained category semantic descriptions, utilizing the principles of locality and compositionality; DBiGCN introduced two graph convolutional networks (GCNs) with opposite running directions to promote mutual enhancement; GraphCEN utilized multi-granularity information through two-level contrastive learning to collaboratively optimize representation learning and class assignment.

[0004] Although existing zero-shot learning methods for graphs have achieved good results in practical applications, there are still some inherent problems. On the one hand, these methods mainly focus on establishing connections between categories based only on the category semantic description matrix, ignoring useful attributes of category and node representations, or not fully utilizing the structural information of the graph, resulting in insufficient generalization ability of the model. For example, DGPN focuses on mining relationships between categories, while DBiGCN and GraphCEN focus on integrating node information with category semantics. However, these methods do not fully consider how to jointly learn category and node representations, integrate consistent useful attributes, and also do not fully utilize the graph structure information. On the other hand, existing methods usually rely on seen categories to obtain supervision signals, but lack sufficient exploration of the semantics of unseen categories, resulting in being vulnerable to the problem of distribution shift when applied to unknown categories. For example, the above methods only construct cross-entropy loss through seen categories, ignoring the in-depth mining of the semantics of unseen categories. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the supervised contrastive learning method for graph zero-shot learning provided by the present invention solves the problem that existing zero-shot learning methods cannot accurately classify when facing unseen data due to insufficient generalization ability.

[0006] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0007] Provide a supervised contrastive learning method for graph zero-shot learning, which includes the steps:

[0008] S1. Collect undirected graph structure data And extract the category semantic description matrix S of all categories in the category set of the undirected graph structure data, is the node set, ε is the edge set; X is the node feature matrix;

[0009] S2. Regard each category in the category semantic description matrix as a node to construct a class graph, and use the k-nearest neighbor method to construct the adjacency matrix A of the category; c ; Based on the node adjacency matrix of the undirected graph structure data Use the graph diffusion technology to construct the affinity node set of each node in the undirected graph structure data ;

[0010] S3. Use a fully connected neural network to project the node feature matrix and the category semantic description matrix into the same dimension, and then jointly input the node adjacency matrix and the category adjacency matrix into a graph convolutional neural network to obtain the node and category latent representation matrices Z and O, where each row in Z and O is the latent representation vector of each node and each category respectively;

[0011] S4. Construct a supervised contrastive learning objective function to guide the learning of the graph convolutional neural network based on the node and class latent representation matrices Z and O and the set of affinity nodes;

[0012] S5. Iteratively optimize the graph convolutional neural network using the supervised contrastive learning objective function until the graph convolutional neural network converges, completing the training of the graph convolutional neural network;

[0013] S6. Input the node feature matrix and the corresponding adjacency matrix of the test nodes into the graph convolutional neural network to obtain the latent representation vectors of each test node, and obtain the class labels of each test node according to the similarity between its latent representation vectors and those of all unseen classes.

[0014] Further, step S4 further includes:

[0015] S41. Construct a node-class pair loss function for self-alignment and affinity alignment between nodes and classes based on the node and class latent representation matrices Z and O and the set of affinity nodes;

[0016] S42. Input the node and class latent representation matrices Z and O into the class generator, and obtain the prototype embedding representation W learned from the node latent representation matrix Z by minimizing the node-prototype pair loss function;

[0017] S43. In the class generator, obtain the first node-level features and the first class-level features through the embedding propagation of the prototype embedding representation W and the class latent representation matrix O;

[0018] S44. Perform embedding interpolation on the first node-level features and the first class-level features to obtain the second node-level features and the second class-level features;

[0019] S45. Construct a synthetic node-class pair loss function incorporating uniformity and alignment based on the first node-level features and the first class-level features and the second node-level features and the second class-level features;

[0020] S46. Use the node-class pair loss function, the node-prototype pair loss function, and the synthetic node-class pair loss function to form the supervised contrastive learning objective function.

[0021] Further, the node-class pair loss function has the following expression:

[0022]

[0023] where y i is the true label of node i in the node set A(i) is the set of affinity nodes of node v i and contains v i , za is the latent representation vector of the \(a\)-th node in the node latent representation matrix \(Z\), where the value of \(a\) in the above formula is determined by the set \(A(i)\), that is, \(a\) takes all index values in the set \(A(i)\); is the class \(y\) i 's latent representation vector; \(\text{sim}(\cdot,\cdot)\) is the cosine similarity function; is the set of seen classes; \(o\) j is the latent representation vector of the \(j\)-th class in the set of seen classes in the class latent representation matrix \(O\); the class latent representation matrix \(O\) includes the set of seen classes and the set of unseen classes of all class latent representation matrices; \(\tau\) is the temperature hyperparameter; is the class uniformity; is the feature alignment; \(|\cdot|\) is the cardinality of the set; \(N\) is the total number of nodes in the undirected graph structure data; \(e\) is the natural logarithm.

[0024] Furthermore, the node-prototype pair loss function has the following expression:

[0025]

[0026] where is the trainable feature prototype corresponding to the class \(y\) i ; \(w\) j is the \(j\)-th prototype embedding in the prototype embedding representation \(W\), \(w\) 1 and are the first and the th prototype embeddings respectively.

[0027] Furthermore, the synthetic node-class pair loss function has the following expression:

[0028]

[0029] where is the total number of seen classes in the set of seen classes ; \(e\) is the natural logarithm; \(\text{sim}(\cdot,\cdot)\) is the cosine similarity function; \(w'\) i and \(o'\) i are the first node-level feature and the first class-level feature corresponding to the prototype embedding \(w\) i and the class latent representation vector \(o\) i respectively; \(o\) i and \(w\) i in \(w''\) i and \(o''\) i are \(w'\) i and \(o'\) respectivelyi The corresponding second node-level feature and the second class-level feature; τ is the temperature hyperparameter.

[0030] Furthermore, the expressions for calculating the first class-level feature and the first node-level feature are respectively:

[0031]

[0032] where w j is the j-th prototype embedding in the prototype embedding representation W; o j is the latent representation vector of the j-th class in the set of seen classes; ω i,j is the weight for controlling the update of o j and w j ; are both cosine similarities, w j’ and o j’ are respectively the j'-th prototype embedding in the prototype embedding representation W and the latent representation vector of the j'-th class in the set of seen classes;

[0033] The expressions for calculating the second class-level feature and the second node-level feature are respectively:

[0034] o″ i = αo i +(1 - α)o′ i w″ i = αw i +(1 - α)w′ i

[0035] where o″ i is the second class-level feature corresponding to o′ i ; w″ i is the second node-level feature corresponding to w′ i ; α is the balance parameter.

[0036] Furthermore, the expression of the contrastive learning objective function is:

[0037]

[0038] where, are respectively the node-class pair loss function, the node-prototype pair loss function, and the synthetic node-class pair loss function; ∈ and η are both balance hyperparameters.

[0039] Furthermore, the method for extracting the semantic description matrix S of all classes in the class set of the graph structure data includes:

[0040] S11. Collect the class set Text materials related to each category, and preprocess them;

[0041] S12. Input the preprocessed text materials into the Word2Vec model for training to obtain the vector representation of each word;

[0042] S13. For each category, aggregate the corresponding word vectors to obtain the category semantic description vector of each category;

[0043] S14. Use all the category semantic description vectors to form the semantic description matrix S of the category set. Each row in the category semantic description matrix represents the category semantic description vector of a category.

[0044] Furthermore, use the k-nearest neighbor method to construct the adjacency matrix of the categories, and then based on the adjacency matrix of the nodes in the undirected graph structure data Use the graph diffusion technology to construct the affinity node set of each node in the undirected graph structure data The method for including:

[0045] S21. For each category semantic description vector in the category semantic description matrix, calculate its distance from other category semantic description vectors, and select the k nearest categories as its neighbors to obtain the adjacency matrix A of the categories c ;

[0046] S22. Standardize the adjacency matrix of the node v in the undirected graph structure data i , and then based on the standardized node adjacency matrix, calculate the graph diffusion matrix F:

[0047]

[0048] where η ∈ (0, 1) is the transmission probability; is the standardized node adjacency matrix; I is the identity matrix;

[0049] S23. Select the nodes corresponding to the indices of the K maximum values in the i-th row of the graph diffusion matrix F as the affinity node set of the node v i .

[0050] Furthermore, when the undirected graph structure data is a citation network, the nodes in the graph structure data represent academic papers, the edges represent the citation relationships between the nodes, and having an edge means having a citation relationship; the feature matrix of each node contains content features related to the paper.

[0051] The beneficial effects of the present invention are:

[0052] 1) This solution optimizes the effect of representation learning to improve the generalization of the model to unseen types by extracting the semantic descriptions of each category in the category set, projecting the category semantic descriptions and the node feature matrix to the same dimension through dimensionality reduction, and then fusing the graph structure information and the semantic descriptions of the categories into the learned representations through a graph neural network.

[0053] 2) This solution incorporates discriminative attributes and graph structure information into the learned representations by combining graph diffusion techniques, including two supervision signals to regularize feature representation learning, namely category uniformity and feature alignment. On the one hand, category uniformity encourages the semantic representations of different categories to be evenly distributed and preserves the maximum information on the unit hypersphere, which is crucial for generalization to unseen categories. On the other hand, feature alignment requires the node features and the category semantics of the same category to be as close as possible. When the node features are fully aligned with the category semantics, the model can more accurately associate new nodes with the correct categories according to the similarity of node semantics under the guarantee of category uniformity, thus improving the generalization ability of the graph zero-shot learning task to unseen categories.

[0054] 3) To effectively address the distribution shift problem, this solution develops a class generator to synthesize the features of unseen classes through the embedding propagation and interpolation of seen classes, thereby providing additional supervision signals and well guaranteeing the generalization to newly emerging classes. Description of the Drawings

[0055] Figure 1 It is a flowchart of a supervised contrastive learning method for graph zero-shot learning.

[0056] Figure 2 It is a schematic block diagram of a supervised contrastive learning method for graph zero-shot learning.

[0057] Figure 3 It is a comparison chart of the predicted classification accuracies under different hyperparameter settings on different datasets in the embodiments of the present invention. (a) and (b) are respectively the comparison charts of the predicted accuracies of zero-shot node classification using the method of the present invention under different k n and topk settings, and (c) and (d) are respectively the comparison charts of the predicted accuracies of zero-shot node classification using the method of the present invention under different ∈ and η settings.

[0058] Figure 4Schematic diagram of a case study on alignment and uniformity in an embodiment of the present invention; (a) and (b) are respectively the distribution diagrams of the cosine distances between the node representations and the corresponding category semantic representations in the training set and the test set; (c) to (e) are respectively the distribution diagrams of the node representations obtained by training on the training set of the embodiment of the present invention and the competing models DGPN and DBiGCN on the unit circle; (f) to (h) are respectively the distribution diagrams of the node representations obtained by training on the test set of the embodiment of the present invention and the competing models DGPN and DBiGCN on the unit circle.

[0059] The specific implementation manner is as follows

[0060] The following describes the specific implementation manner of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manner. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0061] Refer to Figure 1 , Figure 1 which shows a flowchart of a supervised contrastive learning method for graph zero-shot learning; as shown in Figure 1 and Figure 2 , this method S includes steps S1 to S6.

[0062] In step S1, collect undirected graph structure data and extract the category semantic description matrix S of all categories in the category set of the undirected graph structure data, where is the node set, v i represents the i-th node, i = 1,..., N, is the edge set, is the node feature matrix, d f is the feature dimension.

[0063] Taking the citation network as an example, the nodes in the graph represent academic papers; whether there is an edge connection between nodes represents whether there is a citation relationship between papers, and an edge means there is a citation relationship; the feature matrix of each node contains content features related to the paper, such as abstract, keywords or other information. Through such a definition, the graph can not only represent the citation structure between nodes, but also capture the content features of nodes and the propagation manner of these features in the network.

[0064] The adjacency matrix of the graph is represented by A ∈ {0, 1} N×N , where A ij = 1 means there is an edge between nodes v i and v j , and if there is no edge, then A ij= 0. Assume that the set of categories to which the nodes in the graph belong is divided into the set of seen categories and the set of unseen categories where it satisfies In zero-shot learning on graphs, the categories of labeled nodes only come from the set of seen categories while the categories of unlabeled nodes only come from the set of unseen categories All node features and seen node labels can participate in training.

[0065] In addition, in zero-shot learning on graphs, it is crucial to obtain high-quality category semantic information, which helps capture the semantic relationships between categories and thus improve the model's recognition ability for unseen categories.

[0066] When implementing, the method for preferably extracting the semantic description matrix S of all categories in the category set of graph structure data in this solution includes:

[0067] S11. Collect the text materials related to each category in the category set and preprocess them;

[0068] S12. Input the preprocessed text materials into the Word2Vec model for training to obtain the vector representation of each word;

[0069] S13. For each category, aggregate its corresponding word vectors to obtain the category semantic description vector of each category;

[0070] S14. Use all the category semantic description vectors to form the semantic description matrix S of the category set, and each row in the category semantic description matrix represents the category semantic description vector of a category.

[0071] The category semantic description matrix (CSD) S can be used in subsequent zero-shot learning tasks as the representation of category semantics. The CSD should have sufficient expressive power to reflect the complex relationships between categories. where |·| represents the cardinality of the set, is the category set, and d c is the dimension of the CSD.

[0072] In step S2, each category in the category semantic description matrix is regarded as a node to construct a category graph, and the k-nearest neighbor method is used to construct the adjacency matrix A of the category c ; based on the node adjacency matrix of the undirected graph structure data use the graph diffusion technology to construct the affinity node set of each node in the undirected graph structure data ;

[0073] In one embodiment of the present invention, the k-nearest neighbor method is used to construct the adjacency matrix of categories, and then based on the adjacency matrix of nodes in the undirected graph structure data The method for constructing the affinity node set of each node in the undirected graph structure data by using the graph diffusion technology includes:

[0074] S21. For each category semantic description vector in the category semantic description matrix, calculate its distance from other category semantic description vectors, and select the k nearest categories as its neighbors to obtain the adjacency matrix A of categories c ; In the adjacency matrix A c , if category p and category q are k-nearest neighbors of each other, then set the values at positions (p,q) and (q,p) in matrix A c to 1, indicating that there is a connection relationship between them; otherwise set to 0, indicating no connection. In this way, the adjacency matrix A c reflects the semantic adjacency relationship between categories.

[0075] S22. Normalize the adjacency matrix of node v in the undirected graph structure data i , and then based on the normalized node adjacency matrix, calculate the graph diffusion matrix F:

[0076]

[0077] where η ∈ (0,1) is the teleportation probability; is the normalized node adjacency matrix; I is the identity matrix;

[0078] The values in the i-th row of matrix F can reflect the influence between node v i and all other nodes. The larger the value, the greater the influence between nodes and the higher the affinity;

[0079] S23. Select the nodes corresponding to the indices of the K largest values in the i-th row of the graph diffusion matrix F as the affinity node set of node v i .

[0080] In step S3, a fully connected neural network is used to project the node feature matrix and the category semantic description matrix into the same dimension (specifically: by inputting two tuples {X,A} and {S,A c} into a one-layer fully connected neural network to project X and S into the same dimension), and then jointly input the node adjacency matrix and the category adjacency matrix into a graph convolutional neural network to obtain the node and category latent representation matrices Z and O; z i(i = 1, …, N) represents the latent representation vector of each node (each row in the latent representation matrix of the nodes), represents the latent representation vector of each category (each row in the latent representation matrix of the categories, also called the category semantic representation vector), and d′ is the dimension of the representation.

[0081] In step S4, according to the node and category latent representation matrices Z and O and the affinity node set, a supervised contrastive learning objective function for guiding the learning of the graph convolutional neural network is constructed;

[0082] In an embodiment of the present invention, step S4 further includes:

[0083] S41. According to the node and category latent representation matrices Z and O and the affinity node set, construct a node-category pair loss function for self-alignment and affinity alignment between nodes and categories

[0084]

[0085] where y i is the true label of node i in the node set A(i) is the affinity node set of node v i and contains v i z a is the latent representation vector of the a-th node in the node latent representation matrix Z. In the above formula, the value of a is determined by the set A(i), that is, a takes all index values in the set A(i); is the latent representation vector of category y i ; sim(·, ·) is the cosine similarity function; is the set of seen categories; o j is the latent representation vector of the j-th category in the set of seen categories in the category latent representation matrix O; the category latent representation matrix O includes the set of seen categories and the set of unseen categories of all category latent representation matrices; τ is the temperature hyperparameter; is the category uniformity; is the feature alignment; |·| is the cardinality of the set; N is the total number of nodes in the undirected graph structure data; e is the natural logarithm.

[0086] Minimizing the loss encourages the representations of the visible classes to be evenly distributed in the embedding space and promotes the positive enhancement pairs (i.e., the class semantics and node features of the same category, and the features of their affinity nodes) to have similar representations in the embedding space, so as to achieve good separation of the class semantic representations and maximum information retention of the node representations.

[0087] S42. Input the node and class latent representation matrices \(Z\) and \(O\) into the class generator. By minimizing the node-prototype pair loss function, obtain the prototype embedding representation \(W\) learned from the node latent representation matrix \(Z\).

[0088] During implementation, this solution preferably uses the node-prototype pair loss function The expression of which is:

[0089]

[0090] where, is the trainable feature prototype corresponding to class \(y\); \(w_j\) i is the \(j\)-th prototype embedding in the prototype embedding representation \(W\), j and \(w_i\) and 1 and are the first and the -th prototype embeddings respectively.

[0091] S43. In the class generator, through the embedding propagation of the prototype embedding representation \(W\) and the class latent representation matrix \(O\), obtain the first node-level feature and the first class-level feature:

[0092]

[0093] where, \(w_j\) j is the \(j\)-th prototype embedding in the prototype embedding representation \(W\); \(o_j\) j is the latent representation vector of the \(j\)-th class in the set of seen classes; \(\omega\) i,j is the weight controlling the update of j \(o_j\) j and \(w_j\); are both cosine similarities, \(w_{j'}\) j’ and \(o_{j'}\) j’ are the \(j'\)-th prototype embedding in the prototype embedding representation \(W\) and the latent representation vector of the \(j'\)-th class in the set of seen classes respectively.

[0094] By minimizing the loss, the representations of the original seen classes and the feature prototypes are approximately uniformly distributed, and the negative correlation between the defined weights and the similarities ensures the expansion of new classes and features. Therefore, the correlation relationships between different classes are established, and the cross-class transfer ability is enhanced, which can alleviate the domain transfer problem.

[0095] S44. Perform embedding interpolation on the first node-level feature and the first class-level feature to obtain the second node-level feature and the second class-level feature:

[0096] \(o''\) i \(=\alpha o\)i +(1 - α)o′ i w″ i = αw i +(1 - α)w′ i

[0097] wherein, o″ i is the second - class - level feature corresponding to o′ i ; w″ i is the second - node - level feature corresponding to w′ i ; α is a balance parameter.

[0098] S45. Construct a synthetic node - class pair loss function incorporating uniformity and alignment based on the first - node - level feature, the first - class - level feature, the second - node - level feature, and the second - class - level feature

[0099]

[0100] wherein, is the total number of seen classes in the set of seen classes ; e is the natural logarithm; sim(·, ·) is the cosine similarity function; w′ i and o′ i are the first - node - level feature and the first - class - level feature corresponding to the prototype embedding w i and the latent representation vector o of the class i respectively; o i and w i in w″ i and o″ i are the second - node - level feature and the second - class - level feature corresponding to w′ i and o′ i respectively; τ is the temperature hyper - parameter.

[0101] S46. Use the node - class pair loss function, the node - prototype pair loss function, and the synthetic node - class pair loss function to form a supervised contrastive learning objective function

[0102]

[0103] wherein, are the node - class pair loss function, the node - prototype pair loss function, and the synthetic node - class pair loss function respectively; ∈ and η are both balance hyper - parameters.

[0104] In step S5, use the supervised contrastive learning objective function to iteratively optimize the graph convolutional neural network until the graph convolutional neural network converges, completing the training of the graph convolutional neural network;

[0105] In step S6, the node feature matrix and the corresponding adjacency matrix of the test nodes are input into the graph convolutional neural network to obtain the latent representation vectors of each test node. According to the similarity between its latent representation vector and the latent representation vectors of all unseen classes, the class label of each test node is obtained. This class comes from the set of unseen classes That is:

[0106]

[0107] where sim(·,·) is the cosine similarity, z test and o y are the representations learned for the test nodes and unseen classes respectively by the graph convolutional neural network.

[0108] To illustrate the effect of the supervised contrastive learning method of this solution, the following is a comparison experiment of 13 (graph) zero-shot learning algorithms in the prior art:

[0109] Introduction to the methods of the prior art: (1) RandomGuess (the most common baseline method, randomly assigning unseen classes to unlabeled nodes on the graph), (2) DAP[1], (3) DAP(CNN)[1], (4) ESZS[2], (5) ZS-GCN[3], (6) ZS-GCN(CNN)[3], (7) WDVSc[4], (8) Hyperbolic-ZSL[5], (9) AREN[6], (10) RGEN[7], (11) DGPN[8], (12) DBiGCN[9], (13) GraphCEN

[10] ; where (2)-(10) are zero-shot learning algorithms, (11)-(13) are graph zero-shot learning algorithms, and the 13 methods in the prior art correspond to the following references (1)-(13) respectively.

[0110] This embodiment conducts a comparative experiment on three widely recognized citation datasets; including: (1) Cora, (2) CiteSeer, (3) C-M10M; where each node corresponds to a publication, and the edge represents the citation relationship between two linked publications. Regarding the setting of the division of seen / unseen classes, that is, the labels of the same class only belong to one of the training set, validation set, or test set; the test set must contain class labels that do not appear in the training set and validation set; the number of classes in each set is finally ensured to add up to the total number of classes in the dataset. The specific division details can be referred to Table 1:

[0111] Table 1 Division details of the dataset

[0112]

[0113] In addition, the present invention uses two types of CSDs: TEXT-CSDs (default) and LABEL-CSDs, which are generated by Bert-Tiny, to provide different category semantics.

[0114] The method of the present invention uses the model proposed by PyTorch and conducts all experiments on NVIDIA GeForce RTX 3090. In all experiments, the method of the present invention uses a three-layer GCN as the backbone encoder and trains the model from scratch by randomly initializing the network parameters without using any pre-training. The method of the present invention adopts grid search to adjust the hyperparameters. The hyperparameters {∈, η} of the loss are set as follows: {0.1, 1} for the Cora dataset, {0.01, 0.5} for the Citeseer dataset, and {1, 1} for the C-M10M dataset. In addition, the learning rate is selected from {1e-3, 1e-4, 1e-5}; the hidden layer dimension is selected from {32, 64, 128, 256}; the number of neighbors k in the k-nearest neighbors n is selected from ; the top k (the K maximum values) of the number of affinity nodes is selected from {1, 10, 50, 100, 200}. To evaluate the performance, the method of the present invention uses the accuracy on the test set as the main metric in the experiment.

[0115] Analysis of the main experimental results:

[0116] The method of the present invention is compared with 13 advanced graph zero-shot learning methods on three datasets: Cora, CiteSeer, and C-M10M. The experimental task is zero-shot node classification. Under each method, the node features and labeled node labels in each dataset are used for training, and the prediction accuracy is evaluated on the test set (nodes belonging to categories not seen during training). Under the dataset partitioning method in Table 1, the prediction accuracies corresponding to these two partitioning cases can be obtained. For specific reference, see Table 2.

[0117] Table 2 Prediction accuracy

[0118]

[0119] From the prediction accuracy in Table 2, it can be seen that:

[0120] The method of the present invention (ours in Table 2) is always superior to other baseline methods under different category partitioning settings of all three datasets. In particular, under category partitioning I of the C-M10M dataset, the method of the present invention achieves a significant improvement of 9.19% compared to the closest competing baseline, and under category partitioning II of the Citeseer dataset, the improvement is 4.37%. This shows that the method of the present invention has excellent model generalization ability in graph zero-shot learning.

[0121] Zero-shot learning methods in the visual field usually perform worse than graph zero-shot learning methods (DPGN, DBiGCN, and GraphCEN). This may be attributed to their limited ability to explore the relationship information between nodes and capture the complex features of the graph.

[0122] On all datasets, the method of the present invention far exceeds the existing graph zero-shot learning methods (DPGN, DBiGCN, and GraphCEN), which demonstrates the superiority of modeling consistency and alignment. In addition, synthesizing the node features and class semantics of unseen classes helps enhance the generalization ability and further promotes the transfer of knowledge from seen classes to unseen classes.

[0123] Analysis of ablation experiment results

[0124] In this part, an ablation study is conducted to verify the importance of the three key components in the method of the present invention. The following are the comparison variants: Ours w / o The method of the present invention removes the supervised contrastive learning loss of the seen class categories. Ours w / o The method of the present invention removes the supervised contrastive learning loss of the feature prototypes. Ours w / o The method of the present invention removes the supervised contrastive learning loss of synthesizing unseen classes.

[0125] The zero-shot node classification task is performed under the above three variant models, and experiments are conducted on three datasets: Cora, CiteSeer, and C-M10M. The experimental results are evaluated by the prediction classification accuracy on the test set. The comparison results of different variants and the method of the present invention are shown in Table 3.

[0126] Table 3 Comparison results of different variants and the method of the present invention

[0127]

[0128] It can be clearly seen from Table 3 that the method of the present invention achieves the highest performance on all three datasets under different class division settings. Therefore, it is very necessary to construct a joint framework to capture the uniformity and alignment of both seen and unseen classes simultaneously.

[0129] In addition, it can also be seen that different components play different roles on different datasets. Taking class division I as an example, on the Cora dataset, Ours w / o shows the worst results, indicating the importance of the class generator. While on the Citeseer dataset, is the most important. The change in performance can be attributed to the inherent property differences of the datasets, which leads to different focuses of the method of the present invention in capturing information. This further confirms the importance of each component in the method of the present invention.

[0130] In addition, in Table 4, the present invention records the training running times (unit: seconds) of the method of the present invention and three variant models under class partition I. Under each variant, exclusive operations related to the removed loss during training are omitted. Combining with Table 3, it can be observed that the method of the present invention achieves better prediction performance with only a slight increase in computational cost.

[0131] Table 4 Running times of different ablation experiments under class partition I on three datasets

[0132]

[0133] Sensitivity analysis:

[0134] This part discusses the influence of the hyperparameters of the method of the present invention, that is, the number of neighbors k in k-nearest neighbors n , the number of affinity nodes topk (the K maximum values), and the balance weights (loss hyperparameters) ∈ and η, where k n is considered to take values in {1, 2, 3, 4, 5}, topk (the K maximum values) is considered to take values in {1, 10, 50, 100, 200}, ∈ is considered to take values in {0.01, 0.1, 1, 10}, and η is considered to take values in {0.5, 1, 10, 50}.

[0135] k n and the influence of topk (the K maximum values): Figure 3 (a) and 3(b) show the comparison of the prediction accuracies of zero-shot node classification using the method of the present invention under different settings of k n and topk (the K maximum values). It can be observed that on the Cora dataset, when topk (the K maximum values) is fixed, the performance initially improves and then decreases as k n increases, indicating that there are significant differences in class semantics in Cora, and forced connections will blur the distinctions. However, when k n is fixed and topk (the K maximum values) increases to 100, the results continue to improve, while further increase leads to a decrease. This may be due to the introduction of noise by selecting too many affinity nodes. On the other hand, increasing k n is beneficial to the Citeseer dataset because it better captures the relationships between classes, while the influence of topk (the K maximum values) on the Citeseer dataset is relatively small.

[0136] The influence of ∈ and η: Figure 3(c) and 3(d) show the comparison of the prediction accuracy of zero-shot node classification using the method of the present invention under different settings of ∈ and η. It can be seen that on the Cora dataset, when η is fixed and relatively small, increasing ∈ has a positive impact on the prediction of the model to a certain extent, highlighting the importance of feature prototypes. However, further increasing ∈ may have a negative impact in some cases. Similarly, when ∈ is fixed and η increases, the model performance also improves, showing greater stability. This is consistent with the ablation study, that is, deleting will result in the most significant performance drop. However, on the Citeseer dataset, the effects of these parameters show the opposite phenomenon. When both η and ∈ are very small, the model achieves the best results. This is consistent with the ablation study, indicating that in Citeseer, is the main factor affecting performance. Excessive η and ∈ will mask 's effect, resulting in performance degradation.

[0137] 4) Discussion on different CSDs:

[0138] The method of the present invention constructs a category adjacency matrix A based on the CSD matrix (category semantic description matrix) S c . Since different CSDs can provide different semantic information for the category graph, in graph zero-shot learning, the model performance relying on label semantics may be significantly affected by the selected CSD. Therefore, the present invention analyzes two types of CSDs, namely TEXT-CSDs obtained based on text descriptions about categories and LABEL-CSDs obtained based on category labels.

[0139] Considering these two CSDs and performing the zero-shot node classification task based on the method of the present invention, the prediction accuracy on the test set is obtained as shown in Table 5.

[0140] Table 5 shows the comparison of the prediction accuracy of zero-shot node classification under different types of CSD matrices

[0141]

[0142] It can be seen from Table 5 that TEXT-CSDs are generally better than LABEL-CSDs, which is consistent with the previous research results. This highlights the stronger ability of natural language in capturing the relationship information between label semantics. In addition, regardless of the type of CSD used, the method of the present invention achieves the best performance, further demonstrating the superiority of modeling uniformity and alignment, as well as the generalization ability of the category generator.

[0143] 5) Case study:

[0144] To verify the advantages of the alignment and uniformity captured by the method of the present invention, the present invention compares it with the competing models DGPN and DBiGCN on the C-M10M dataset, as Figure 4 shown. Figure 4 (a) and 4(b) respectively show the distributions of the cosine distances between the trained node representations and the corresponding class semantic representations, which are used to show the alignment on the training set and the test set.

[0145] Obviously, the present invention has the smallest mean value in the distance distribution, indicating the best alignment. In addition, DGPN shows better alignment than DBiGCN on the training set, but the opposite is true on the test set. This is because DGPN models the consistency between node features and class semantics during the training process, but lacks the connection between seen and unseen classes. On the other hand, DBiGCN effectively captures the relationships between classes, thereby improving the generalization ability.

[0146] For uniformity, the present invention uses principal component analysis (PCA) and Gaussian kernel density estimation to visualize (using PCA to reduce the trained node representations to two dimensions) the distribution of the trained node representations on the unit circle, as Figure 4 (c)-(h) shown. Taking Figure 4 (c) as an example, it can be clearly seen that the method of the present invention shows higher uniformity at different positions on the unit circle, with features concentrated within classes and scattered between classes, which proves the distinguishability of the learned node representations.

[0147] References (1)-(13) are respectively:

[0148] [1] C.H. Lampert, H. Nickisch, and S. Harmeling, “Attribute-based classification for zero-shot visual object categorization,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 3, pp. 453–465, 2013.

[0149] [2] B. Romera-Paredes and P. Torr, “An embarrassingly simple approach to zero-shot learning,” in International Conference on Machine Learning. PMLR, 2015, pp. 2152–2161.

[0150] [3] X. Wang, Y. Ye, and A. Gupta, “Zero-shot recognition via semantic embeddings and knowledge graphs,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6857–6866.

[0151] [4] Z. Wan, D. Chen, Y. Li, X. Yan, J. Zhang, Y. Yu, and J. Liao, “Transductive zero-shot learning with visual structure constraint,” in Advances in Neural Information Processing Systems, vol. 32, 2019.

[0152] [5] S. Liu, J. Chen, L. Pan, C.-W. Ngo, T.-S. Chua, and Y.-G. Jiang, “Hyperbolic visual embedding learning for zero-shot recognition,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 9273–9281.

[0153] [6] G.-S. Xie, L. Liu, X. Jin, F. Zhu, Z. Zhang, J. Qin, Y. Yao, and L. Shao, “Attentive region embedding network for zero-shot learning,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 9384–9393.

[0154] [7]G.-S. Xie, L. Liu, F. Zhu, F. Zhao, Z. Zhang, Y. Yao, J. Qin, and L. Shao, “Region graph embedding network for zero-shot learning,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16. Springer, 2020, pp. 562–580.

[0155] [8]Z. Wang, J. Wang, Y. Guo, and Z. Gong, “Zero-shot node classification with decomposed graph prototype network,” in Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, 2021, pp. 1769–1779.

[0156] [9]Q. Yue, J. Liang, J. Cui, and L. Bai, “Dual bidirectional graph convolutional networks for zero-shot node classification,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022, pp. 2408–2417.

[0157]

[10] W. Ju, Y. Qin, S. Yi, Z. Mao, K. Zheng, L. Liu, X. Luo, and M. Zhang, “Zero-shot node classification with graph contrastive embedding network,” Transactions on Machine Learning Research, 2023。

Claims

1. A supervised contrastive learning method for zero-shot learning of graphs, characterized in that: Includes steps: S1. Collect undirected graph structure data And extract the category semantic description matrix S of all categories in the category set of undirected graph structure data, is the node set, ε is the edge set; X is the node feature matrix; S2. Treat each category in the category semantic description matrix as a node to construct a class graph, and use the k-nearest neighbor method to construct the category adjacency matrix A c ; Based on undirected graph structure data The node adjacency matrix of The set of affinity nodes for each node in; S3, using a fully connected neural network to project the node feature matrix and the category semantic description matrix to the same dimension, and then jointly input the node adjacency matrix and the category adjacency matrix into the graph convolutional neural network to obtain the node and category potential representation matrices Z and O, where each row in Z and O is the potential representation vector of each node and each category respectively; S4. Construct a supervised contrastive learning objective function to guide the learning of graph convolutional neural network based on the node and category potential representation matrices Z and O and the affinity node set; S5. Use the supervised contrast learning objective function to iteratively optimize the graph convolutional neural network until the graph convolutional neural network reaches convergence, completing the training of the graph convolutional neural network; S6. Input the node feature matrix and the corresponding adjacency matrix of the test node into the graph convolutional neural network to obtain the potential representation vector of each test node, and obtain the category label of each test node based on the similarity between the potential representation vector of each test node and all unseen categories.

2. The supervised contrastive learning method for graph zero-shot learning according to claim 1, characterized in that: Step S4 further comprises: S41, constructing a node-category pair loss function for self-alignment and affinity alignment between nodes and categories according to the node and category potential representation matrices Z and O and the affinity node set; S42, input the node and category potential representation matrices Z and O into the category generator, and obtain the prototype embedding representation W learned from the node potential representation matrix Z by minimizing the node-prototype pair loss function; S43, in the category generator, by embedding propagation of the prototype embedding representation W and the category latent representation matrix O, obtaining a first node-level feature and a first category-level feature; S44, embedding and interpolating the first node-level feature and the first class-level feature to obtain a second node-level feature and a second class-level feature; S45, constructing a synthetic node-class pair loss function adding uniformity and alignment according to the first node-level feature and the first class-level feature and the second node-level feature and the second class-level feature; S46. A node-category pair loss function, a node-prototype pair loss function and a synthetic node-category pair loss function are used to form a supervised contrastive learning objective function.

3. The supervised contrastive learning method for graph zero-shot learning according to claim 2, characterized in that: Node-class pair loss function The expression is: Among them, y i For node set The true label of node i in the graph, A(i) is the true label of node v i The affinity node set of v i , z a is the potential representation vector of the ath node in the node potential representation matrix Z. In the above formula, the value of a is determined by the set A(i), that is, a takes all index values ​​in the set A(i); For category y i The potential representation vector of ; sim(·,·) is the cosine similarity function; is the set of categories that have been seen; o j is the potential representation vector of the jth category in the seen category set in the category potential representation matrix O; the category potential representation matrix O includes the seen category set and the set of unseen categories All categories of potential representation matrix; τ is the temperature hyperparameter; is the category uniformity; is feature alignment; |·| is the cardinality of the set; N is the total number of nodes in the undirected graph structure data; e is the natural logarithm.

4. The supervised contrastive learning method for graph zero-shot learning according to claim 3, characterized in that: Node-Prototype Pair Loss Function The expression is: in, For category y i The corresponding trainable feature prototype; w j is the prototype embedding representing the j-th prototype embedding in W, w1 and The first and Prototype embedding.

5. The supervised contrastive learning method for graph zero-shot learning according to claim 2, characterized in that: Synthetic node-class pair loss function The expression is: in, Collection of seen categories The total number of classes seen in ; e is the natural logarithm; sim(·,·) is the cosine similarity function; w′ i and o′ i The prototype embedding w i and the potential representation vector o of the category i The corresponding first node level features and first class level features; o i and w i In w″ i and o″ i are w′ i and o′ i The corresponding second node-level features and second class-level features; τ is the temperature hyperparameter.

6. The supervised contrastive learning method for zero-shot learning of graphs according to claim 5, characterized in that: The expressions for calculating the first class-level features and the first node-level features are: Among them, w j is the prototype embedding representation of the j-th prototype embedding in W; o j is the potential representation vector of the jth category in the seen category set; ω i,j To control j and w j Updated weights; and are cosine similarities, w j’ and j’ are the j'th prototype embedding in the prototype embedding representation W and the potential representation vector of the j'th category in the seen category set; The expressions for calculating the second class-level features and the second node-level features are: oh i =ao i +(1―α)o′ i w″ i =αw i +(1―α)w′ i Among them, o″ i o′ i The corresponding second-class feature; w″ i is w′ i The corresponding second node level feature; α is the balance parameter.

7. The supervised contrastive learning method for graph zero-shot learning according to any one of claims 2 to 6, characterized in that: Supervised Contrastive Learning Objective Function The expression is: in, and They are the node-category pair loss function, the node-prototype pair loss function, and the synthetic node-category pair loss function; ∈ and η are both balancing hyperparameters.

8. The supervised contrastive learning method for graph zero-shot learning according to claim 1, characterized in that: The method of extracting the semantic description matrix S of all categories in the category set of graph structure data includes: S11, collect text data related to each category in category set c and preprocess them; S12, input the preprocessed text data into the Word2Vec model for training to obtain the vector representation of each word; S13, for each category, aggregate the corresponding word vectors to obtain the category semantic description vector of each category; S14. All the category semantic description vectors are used to form a semantic description matrix S of the category set. Each row in the category semantic description matrix represents a category semantic description vector of a category.

9. The supervised contrastive learning method for graph zero-shot learning according to claim 1, characterized in that: The k-nearest neighbor method is used to construct the category adjacency matrix, and then based on the undirected graph structure data The adjacency matrix of the nodes in the graph is constructed using graph diffusion technology to construct undirected graph structure data The methods for finding the affinity node set for each node in the include: S21. For each category semantic description vector in the category semantic description matrix, calculate its distance with other category semantic description vectors, and select the k categories with the closest distance as its neighbors to obtain the category adjacency matrix A. c ; S22. For undirected graph structure data Middle node v i The adjacency matrix is ​​normalized, and then the graph diffusion matrix F is calculated based on the standardized node adjacency matrix: Among them, η∈(0,1) is the transmission probability; is the standardized node adjacency matrix; I is the identity matrix; S23. Select the node corresponding to the index of the K largest values ​​in the i-th row of the graph diffusion matrix F as the node v i The set of affinity nodes.

10. The supervised contrastive learning method for graph zero-shot learning according to claim 1, characterized in that: When the undirected graph structure data is a citation network, the nodes in the graph structure data represent academic papers, the edges represent the citation relationships between nodes, and the presence of edges represents the presence of citation relationships; the feature matrix of each node contains content features related to the paper.