Multi-graph representation learning method based on double-layer optimization

By introducing a two-layer optimization method in multiple graph representation learning, learning self-expression matrix and graph structure, the problem of difficulty in capturing global relationships and graph structure updates in the prior art is solved, and the node classification performance is significantly improved.

CN119992151APending Publication Date: 2025-05-13UESTC (SHENZHEN) ADVANCED RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411430278.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing multi-graph representation learning method is difficult to effectively capture the global positive and negative relationship between nodes, and the update of the graph structure is uneven, which leads to edge starvation problems and affects model performance.

Method used

A multi-graph representation learning method based on double-layer optimization is adopted. The global positive and negative relationship between nodes is captured by learning the self-expression matrix, and the local relationship of the graph structure is supplemented, and the self-expression matrix parameters and representation learning parameters are updated respectively to avoid edge hunger problems.

Benefits of technology

The performance of the multi-graph representation classification task in downstream nodes is significantly improved, the quality of node representation and the uniformity of graph structure are improved, and the edge hunger problem is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992151A_ABST
    Figure CN119992151A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of graph representation learning, and provides a multi-graph representation learning method based on double-layer optimization. The method mainly aims at solving the problems that in existing multi-graph representation learning, the global positive and negative relation of nodes cannot be effectively captured, and training is unbalanced in the graph structure learning process. Through a double-layer optimization framework, a global positive and negative relationship between nodes is learned through a self-expression matrix at an inner layer, a plurality of graph structures are weighted and aggregated by adopting an attention mechanism at an outer layer, and the graph structures and the self-expression matrix are aggregated through a graph convolutional neural network, so that the performance of a node classification task is improved. Besides, the parameters of the self-expression matrix and the expression learning module are respectively updated through a double-layer optimization training mode, so that the side starvation problem is solved, and the performance of multiple graphs in downstream tasks is remarkably improved. The method can be widely applied to various fields of social media analysis, community anomaly detection, recommendation systems and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of graph representation learning, and more specifically, relates to a multiple graph representation learning method based on double-layer optimization. Background Art

[0002] Graphs are composed of nodes and edges, which are used to represent a certain structure between entities. Nodes usually represent abstract entities, and edges represent the relationship between entities. Multigraphs are composed of multiple graphs that share node features, and each graph reflects a specific relationship between nodes. Multigraphs are usually used to model entities with multiple connections in reality, and have extremely wide practical applications, such as social media analysis, community anomaly detection, and recommendation systems. The premise for studying and analyzing multigraphs is to effectively represent multigraphs, that is, graph representation learning. The purpose of graph representation learning is to mine the original topological information from the given graph data and embed the topological structure information into a low-dimensional vector space. However, traditional graph representation learning methods generally focus on a single graph, lack the ability to effectively handle the complex relationships of multigraphs, and it is difficult to obtain effective multigraph node representations. In order to solve this problem, various multigraph representation learning methods have been proposed in recent years. Their goal is to mine the hidden information in multigraphs to learn low-dimensional node representations.

[0003] Existing methods for learning multi-graph representations can be roughly divided into two categories, namely node feature-free methods and node feature-based methods. The former generally uses graph structure information to obtain node representations, while the latter generally uses node features to obtain node representations. However, these methods ignore the discriminative information in node features, thereby reducing the quality of node representations. Therefore, recent work has proposed to simultaneously consider node feature information and structural information in multi-graphs to learn node representations.

[0004] The main difficulty of existing multiple graph representation learning methods is that the original graph structure inevitably has noise or missing problems. In order to alleviate this problem, existing work adopts the graph structure learning method in traditional graph representation learning to improve model performance, and directly uses the similarity between node features to optimize the graph structure, thereby improving the performance of the model. However, the existing multiple graph representation learning methods have the following main limitations: 1) Using feature similarity to measure whether nodes belong to the same category to optimize the graph structure, it is unable to effectively capture the negative correlation between nodes. In addition, feature similarity focuses on a single pair of nodes and cannot effectively capture the global relationship between each node and all nodes at the same time, thereby ignoring the global positive and negative relationship between node features, resulting in suboptimal model performance. 2) Directly optimizing the parameters of graph structure learning based on the node classification task, the update of the graph structure is supervised by the node classification task, resulting in uneven training of the graph structure. Taking the training scenario of a two-layer graph convolutional neural network as an example, node information can only be propagated within a two-hop range. If there is no labeled node within two hops of an unlabeled node, the edges around the unlabeled node will lack supervised information. Ultimately, the updated graph structure is more inclined to fit the training set data, resulting in edge starvation. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and propose a new multiple graph representation learning framework, namely, multiple graph representation learning based on two-layer optimization. The global positive and negative relationships between nodes are captured by learning a self-expressive matrix based on node features, and supplemented by local relationships in the graph structure. The learning of the self-expressive matrix and the representation learning are regarded as two-layer optimization problems. The parameters of the self-expressive matrix and the parameters of the representation learning are updated respectively through the two-layer optimization, which effectively solves the edge starvation problem and significantly improves the performance of multiple graph representations in downstream node classification tasks.

[0006] This aspect is based on a multi-graph representation learning method with double-layer optimization, and constructs a multi-graph representation learning model consisting of an inner optimization module and an outer optimization module. In the inner optimization module, the learning model first generates a self-expression matrix based on the node features of the graph, and optimizes the self-expression matrix through learning loss to capture the global positive and negative relationships between nodes. In the outer optimization module, the learning model first aggregates the graph structure and the self-expression matrix to obtain the aggregated graph structure, and then extracts features from the graph node features and the aggregated graph structure through the graph convolutional neural network module to generate the final node representation, which is optimized through cross entropy loss, and finally the node classification is completed using the optimized node representation.

[0007] The present invention has the following beneficial effects:

[0008] 1) The present invention utilizes the self-expressive property of the self-expressive matrix to adaptively capture the global positive and negative relationships between nodes, thereby improving the effectiveness of multi-graph representation learning.

[0009] 2) The present invention reduces the negative effects of directly optimizing graph structure learning according to classification tasks by performing graph structure learning and representation learning separately, and improves the learning effect of multiple graph representations.

[0010] 3) This paper introduces two-layer optimization into the field of multi-graph representation learning for the first time, and uses self-expression matrix and representation learning method to optimize the inner and outer layers of the model respectively, which solves the edge starvation problem in multi-graph learning and significantly improves the node classification effect of graph node representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of a specific implementation method of the multi-graph representation learning method based on double-layer optimization of the present invention;

[0012] Figure 2 It is a structural diagram of the double-layer optimization learning module in the present invention. DETAILED DESCRIPTION

[0013] The specific implementation of the present invention is described below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0014] Example

[0015] Figure 1 This is a flowchart of a specific implementation of the multi-graph representation learning method based on double-layer optimization of the present invention. Figure 1 As shown, the specific steps of the multi-graph representation learning method based on double-layer optimization of the present invention include:

[0016] S101: Building a self-expressive learning module:

[0017] Although the relationship between nodes can be captured by node feature similarity, the similarity graph is usually derived by calculating the similarity between each node and its neighbors, which cannot capture the global relationship between nodes, resulting in the loss of discriminant information of the intrinsic structure of the data. Therefore, the present invention uses the self-expression characteristics between node features to adaptively capture the global positive and negative relationships between nodes.

[0018] To this end, the present invention first initializes the similarity graph by the K nearest neighbor (KNN) algorithm, and obtains the feature similarity matrix M between each node pair by calculating the distance between all node pairs, that is,

[0019]

[0020] where x i and xj Yes i and v j Then we select the k nodes with the highest similarity to each node in the distance matrix M as their neighbors, thus obtaining the initial feature similarity graph C.

[0021] In order to further capture the negative relationship between nodes with different labels and the positive relationship between nodes with the same label, the present invention uses the self-expressive characteristics of node features to optimize the similarity graph. Specifically, given feature X = {x1,...,x N} of the node set V, for any v j ∈V, there is a coefficient So that:

[0022] x j =∑x i c ij

[0023] where c ij is corresponding to v i The self-expression coefficient is used to reconstruct v from all other nodes j Each node can be represented as a linear combination of all other nodes with positive or negative coefficients. The present invention optimizes the similarity graph using the following constraints to capture the global positive and negative relationships between nodes with self-expressive properties, namely:

[0024]

[0025] However, the self-expression matrix can easily obtain a trivial solution, namely the identity matrix. Therefore, the present invention further considers two regularization terms in the above objective function to avoid obtaining a trivial solution, namely:

[0026]

[0027] Among them, α and β (r) are two non-negative parameters used to balance the second and third terms. The second term aims to regularize the sparsity of the self-expression matrix, while the third term aims to optimize the self-expression matrix under the guidance of the graph structure to avoid trivial solutions. Afterwards, we let C = (C + C) / 2 to ensure that the learned self-expression matrix is ​​symmetric.

[0028] In this way, each element of the self-expressive matrix is ​​not restricted to be non-negative, thereby adaptively capturing the positive and negative relationship between two nodes. Moreover, the self-expressive matrix relies on all data points to describe each node, unlike previous multi-graph representation learning methods that only calculate the similarity between a node and its nearby nodes. Therefore, the self-expressive matrix also captures the global relationship of nodes.

[0029] S102: Constructing fusion graph learning module:

[0030] In a multigraph, an edge can connect two nodes from different classes, and two nodes in the same class may not be connected. To solve these problems, the present invention fuses the graph structure with the self-expression matrix, uses the positive relationship between two nodes to connect two nodes of the same class, and uses the negative relationship between two nodes to disconnect two nodes of different classes.

[0031] To this end, the present invention first integrates all graph structures in multiple graphs, uses the attention mechanism to learn the weights of different graph structures and aggregates them. Specifically, given multiple graph structures A = {A (1) ,...,A (R)}, we get the fusion graph structure matrix A with graph-level attention mechanism topology ,Right now:

[0032]

[0033] in is the stacking matrix of all graph structures. Ψ1 represents the channel attention layer, whose weight matrix Indicates the importance of different graph structures.

[0034] After the attention mechanism, we further use the attention layer to fuse the graph structure matrix A topology And the self-expression matrix C:

[0035]

[0036] in is the stacked matrix of the graph structure and the self-expression matrix. Ψ2 is the channel attention layer, whose weight matrix It indicates the importance of the fusion graph structure and the self-expression matrix. After that, we further set S = (S + S) / 2 to ensure that the aggregate graph is symmetrical.

[0037] S103: Construct representation learning module:

[0038] Given a node feature matrix X and an aggregate graph S, the present invention uses a graph convolutional layer g: The node representation H is obtained by:

[0039]

[0040] where σ is the activation function and Θ is the weight matrix of the encoder g. is the symmetric normalized graph structure of the aggregate graph, S+wI N The degree matrix, w is the identity matrix I NThe weight of H aggregates the neighbor information from the original topology and node feature space with S. Given the node representation H, the present invention further trains the model by minimizing the cross entropy loss between the true label and the predicted label. To this end, a fully connected layer is used to obtain the predicted class according to the node representation H, that is:

[0041]

[0042] Where W represents the parameters of the fully connected layer, is the category of the node predicted by the classifier. Therefore, the cross entropy loss can be expressed as:

[0043]

[0044] where Y l is a set of node indices with labels, Y l and are the labels and predicted classes of the labeled nodes. Therefore, the label information is used to guide the training process of the cross entropy loss, so that meaningful node representations H can be learned.

[0045] S104: Set the number of iterations t = 1

[0046] S105: Using two-layer optimization to train the learning model

[0047] The present invention optimizes the parameters of graph structure learning and the parameters of representation learning in a double-layer manner. Figure 2 It is a structural diagram of the double-layer optimization learning module in the present invention. Figure 2 As shown, the double-layer optimization learning module in the present invention includes an inner optimization part and an outer optimization part. Specifically, the parameters are represented as θ gcn and θ c , the inner and outer layers can be optimized simultaneously through the following objective function, namely:

[0048] Γ out (θ gcn ,θ c )=L CE ,Γ in (θ gcn ,θ c )=L GL ,

[0049] where Γ out Aims to optimize the parameters of the representation learning module, Γ in It aims to optimize the parameters of the self-expressive learning module. In particular, the above objective function can be expressed as the following two-level optimization problem, namely:

[0050]

[0051] For the above two-layer optimization method, usually the internal optimization dynamic T steps are first expanded, and then the super gradient (external gradient) is calculated based on the expanded dynamic. The present invention performs the internal optimization process by stochastic gradient descent method, that is:

[0052]

[0053] where

[0054]

[0055] where ψ t represents the update scheme of the tth step of the internal optimization process, T represents the total number of iterations of the internal optimization, η t is the learning rate of internal optimization. After T steps, the internal parameters can be expressed as:

[0056]

[0057] in represents the composite dynamics operation of the entire iteration. Therefore, the external optimization is:

[0058]

[0059] In order to update the external parameters (i.e., θ gcn ), we need to calculate the super gradient in Operation t Obviously it depends directly on θ gcn In addition, the operation ψ t pass Indirectly depends on θ gcn ,and According to θ gcn Updated. Therefore, the super gradient It can be calculated by the chain rule, that is:

[0060]

[0061] In formula (1.16), the first term represents the direct gradient, which can be directly calculated. The second term represents the indirect gradient, which is difficult to obtain directly, especially the Jacobian determinant. To solve this problem, we further expand it to:

[0062]

[0063] Finally, by superimposing T rounds of gradients, we get θ gcn .

[0064] S106: Determine whether t<T, where T represents the maximum number of predicted iterations. If so, proceed to step S107; otherwise, proceed to step S108.

[0065] S107: Let t=t+1, use the current round of θ c ,θ gcn The parameters of the self-expression learning module and the representation learning module are updated, and the process returns to step S105.

[0066] S108: Node classification using node representation:

[0067] Node representations are extracted from the final representation learning module and used to classify the nodes.

[0068] In order to better illustrate the technical effect of the present invention, the present invention is experimentally verified using specific examples. In this experimental verification, two citation multi-graph datasets are used, namely the Association for Computing Machinery Citation Dataset (ACM) and the Digital Library and Information System Citation Dataset (DBLP). The ACM dataset contains papers in the field of computer science and their citation relationships, such as information such as authors, conferences, journals, etc., totaling more than one million citations. The DBLP dataset contains papers in the field of computer science and their author information, such as conferences, journals, and publication years, totaling more than one million papers. And two movie multi-graph datasets, namely the Internet Movie Database dataset (IMDB) and the Freebase dataset (Freebase). The IMDB dataset contains detailed information about movies, such as actors, directors, ratings, etc., totaling millions of movies and related data. The Freebase dataset contains a wide range of knowledge graph data, such as the relationships between entities such as movies, people, and places, totaling hundreds of millions of facts and entities.

[0069] In order to fully reflect the advantages of the present invention, the present invention is compared with 2 single-view methods and 7 multi-graph methods. The single-view graph methods include 2 baseline methods, namely graph convolutional networks (GCN) and graph attention networks (GAT). The multi-graph methods include heterogeneous attention networks (HAN), graph tensor networks (GTN), multi-graph neural networks (MAGNN), multi-graph neural networks-adaptive aggregation (MAGNN-AC), heterogeneous graph self-supervised learning (HGSL), multi-graph deep relations (MGDCR) and multi-level graph convolutional networks (MHGCN).

[0070] The present invention is implemented by PyTorch and trained on an NVIDIA RTX3090 GPU. Table 1 is a statistical table of classification accuracy of the present invention and the comparative method under different tasks in this embodiment.

[0071]

[0072] Table 1

[0073] As shown in Table 1, it can be seen from the results in Table 1 that the present invention has achieved the best results in the node classification tasks of the four data sets, proving that the correlation learned from the node features of the present invention can provide supplementary information for the graph structure and help learn distinguishable node representations, thereby verifying the effectiveness of the present invention.

[0074] Although the above describes the illustrative specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.

Claims

1. A multi-graph representation learning method based on two-layer optimization, characterized in that: The following steps are involved: S1: Building a self-expression learning module: Although the relationship between nodes can be captured by node feature similarity, the similarity graph is usually derived by calculating the similarity between each node and its neighbors, which cannot capture the global relationship between nodes, resulting in the loss of discriminant information of the intrinsic structure of the data. Therefore, the present invention uses the self-expression characteristics between node features to adaptively capture the global positive and negative relationships between nodes. To this end, the present invention first initializes the similarity graph by the K nearest neighbor (KNN) algorithm, and obtains the feature similarity matrix M between each node pair by calculating the distance between all node pairs, that is, where x i and x j Yes i and v j Then we select the k nodes with the highest similarity to each node in the distance matrix M as their neighbors, thus obtaining the initial feature similarity graph C. In order to further capture the negative relationship between nodes with different labels and the positive relationship between nodes with the same label, the present invention uses the self-expressive characteristics of node features to optimize the similarity graph. Specifically, given feature X = {x1,...,x N } of the node set V, for any v j ∈V, there is a coefficient So that: x j =∑x i c ij where c ij is corresponding to v i The self-expression coefficient is used to reconstruct v from all other nodes j Each node can be represented as a linear combination of all other nodes with positive or negative coefficients. The present invention optimizes the similarity graph using the following constraints to capture the global positive and negative relationships between nodes with self-expressive properties, namely: However, the self-expression matrix can easily obtain a trivial solution, namely the identity matrix. Therefore, the present invention further considers two regularization terms in the above objective function to avoid obtaining a trivial solution, namely: Among them, α and β (r) are two non-negative parameters used to balance the second and third terms. The second term aims to regularize the sparsity of the self-expression matrix, while the third term aims to optimize the self-expression matrix under the guidance of the graph structure to avoid trivial solutions. Afterwards, we let C = (C + C) / 2 to ensure that the learned self-expression matrix is ​​symmetric. In this way, each element of the self-expressive matrix is ​​not restricted to be non-negative, thereby adaptively capturing the positive and negative relationship between two nodes. Moreover, the self-expressive matrix relies on all data points to describe each node, unlike previous multi-graph representation learning methods that only calculate the similarity between a node and its nearby nodes. Therefore, the self-expressive matrix also captures the global relationship of nodes. S2: Constructing fusion graph learning module: In a multigraph, an edge can connect two nodes from different classes, and two nodes in the same class may not be connected. To solve these problems, the present invention fuses the graph structure with the self-expression matrix, uses the positive relationship between two nodes to connect two nodes of the same class, and uses the negative relationship between two nodes to disconnect two nodes of different classes. To this end, the present invention first integrates all graph structures in multiple graphs, uses the attention mechanism to learn the weights of different graph structures and aggregates them. Specifically, given multiple graph structures A = {A (1) ,...,A (R) }, we get the fusion graph structure matrix A with graph-level attention mechanism topology ,Right now: in is the stacking matrix of all graph structures. Ψ1 represents the channel attention layer, whose weight matrix Indicates the importance of different graph structures. After the attention mechanism, we further use the attention layer to fuse the graph structure matrix A topology And the self-expression matrix C: in is the stacked matrix of the graph structure and the self-expression matrix. Ψ2 is the channel attention layer, whose weight matrix It indicates the importance of the fusion graph structure and the self-expression matrix. After that, we further set S = (S + S) / 2 to ensure that the aggregate graph is symmetrical. S3: Building a representation learning module: Given a node feature matrix X and an aggregate graph S, the present invention uses a graph convolutional layer g: The node representation H is obtained by: Where σ is the activation function and Θ is the weight matrix of the encoder g. is the symmetric normalized graph structure of the aggregate graph, S+wI N The degree matrix, w is the identity matrix I N The weight of H aggregates the neighbor information from the original topology and node feature space with S. Given the node representation H, the present invention further trains the model by minimizing the cross entropy loss between the true label and the predicted label. To this end, a fully connected layer is used to obtain the predicted class based on the node representation H, that is: Where W represents the parameters of the fully connected layer, is the category of the node predicted by the classifier. Therefore, the cross entropy loss can be expressed as: where Y l is a set of node indices with labels, Y l and are the labels and predicted classes of the labeled nodes. Therefore, the label information is used to guide the training process of the cross entropy loss, so that meaningful node representations H can be learned. S4: Set the number of iterations t = 1 S5: Using two-layer optimization to train the learning model The present invention optimizes the parameters of graph structure learning and representation learning in a two-layer manner. Specifically, the parameters are represented as θ gcn and θ c , the inner and outer layers can be optimized simultaneously through the following objective function, namely: C out (i gcn ,i c )=L CE ,C in (i gcn ,i c )=L GL , where Γ out Aims to optimize the parameters of the representation learning module, Γ in It aims to optimize the parameters of the self-expressive learning module. In particular, the above objective function can be expressed as the following two-level optimization problem, namely: For the above two-level optimization methods, one usually first unfolds the inner optimization dynamics T steps, and then computes the super gradients (outer gradients) based on the unfolded dynamics. The present invention performs an internal optimization process by stochastic gradient descent [Amari, 1993], namely: where ψ t represents the update scheme of the tth step of the internal optimization process, T represents the total number of iterations of the internal optimization, η t is the learning rate of internal optimization. After T steps, the internal parameters can be expressed as: in represents the composite dynamics operation of the entire iteration. Therefore, the external optimization is: In order to update the external parameters (i.e., θ gcn ), we need to calculate the super gradient in Operation t Obviously it depends directly on θ gcn In addition, the operation ψ t pass Indirectly depends on θ gcn ,and According to θ gcn Updated. Therefore, the super gradient It can be calculated by the chain rule, that is: In formula (1.16), the first term represents the direct gradient, which can be directly calculated. The second term represents the indirect gradient, which is difficult to obtain directly, especially the Jacobian determinant. To solve this problem, we further expand it to: Finally, by superposition of T round gradients, we get θ gcn . S6: Determine whether t<T, where T represents the maximum number of predicted iterations. If so, proceed to step S107; otherwise, proceed to step S108. S7: Let t = t + 1, using the current round of θ c ,θ gcn The parameters of the self-expression learning module and the representation learning module are updated, and the process returns to step S105. S8: Node classification using node representation: Node representations are extracted from the final representation learning module and used to classify the nodes.