Graph classification method and system based on sub-graph integration and position awareness
By combining a node- and graph-based subgraph extraction strategy with graph convolutional networks and anchor position awareness, we optimize node representation, solve the problem of insufficient global information integration in graph classification, and achieve higher graph classification accuracy and precise recognition in specific fields.
Patent Information
- Application Number
- CN202510493924.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2025-09-05
AI Technical Summary
Existing graph neural networks find it difficult to effectively integrate global graph information in graph classification tasks, resulting in a separation between local features and global information, which limits the accuracy of graph classification.
A node-based and graph-based subgraph extraction strategy is adopted, combined with a graph convolutional network to encode substructure features and fuse them through an attention mechanism. An anchor-based method is used to calculate the relative position information of nodes, and the graph information bottleneck mechanism is used to optimize node representation and remove redundant information.
It significantly improves the accuracy of graph classification by an average of 3.3%, and shows high potential value in specific fields such as drug activity prediction and identification of fake accounts on social networks.
Smart Images

Figure CN120597019A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graph classification, and in particular relates to a graph classification method and system based on subgraph integration and location awareness. Background Art
[0002] In recent years, graphs have been widely used for representation in various fields, including social networks, recommendation systems, and bioinformatics analysis. With the development of deep learning, graph neural networks (GNNs) have become the most popular method for graph data analysis due to their powerful inductive ability and high efficiency. Learning node embeddings through GNNs has achieved remarkable results in tasks such as node classification, graph classification, link prediction, and community detection. Classic GNNs usually adopt a message passing method to update the features of each central node by aggregating the features of neighboring nodes. This aggregation implicitly integrates part of the structural information and features of the graph into the embedding of each node. However, classic GNNs methods mainly focus on the relationship between nodes and their neighbors, without explicitly utilizing the subgraph structure.
[0003] Subgraph structure contains important topological information. It is crucial for graph classification tasks. Subgraph extraction strategies can generally be divided into two types: node-based strategies and graph-based strategies. Node selection strategies extract subgraphs by determining the central node or key node of the subgraph. For example, MOSGSL selects the top k nodes with the largest degree as the subgraph center and uses a breadth-first search algorithm to extract a subgraph for each central node. The study introduced a pattern-driven structure guidance module to capture key subgraph-level structural patterns and enhance personalized structural learning. SAGIN uses the return probability based on random walks to encode the self-subgraph structure rooted in the node, and then uses structural information to enhance the expressive power of GNNs. AdaSNN[9] designed a subgraph detection module based on reinforcement learning to adaptively search for subgraphs without heuristic assumptions or predefined rules. SupCosine samples subgraphs by simulating a diffusion process. It is called a cascade of ordered node sequences. Structural reasoning is then used to recover the enhanced graph from the cascade. GCP generates enhanced subgraph pairs by maintaining the structural similarity of the original graph; MSSGCL uses multi-scale subgraph sampling to construct global-local view comparisons, providing rich self-supervisory signals. However, existing methods have two significant flaws: first, the generalized subgraph selection strategy leads to the loss of specific structural information. Taking molecular graphs as an example, their specific structures such as functional groups and molecular fragments contain rich semantics, but existing methods fail to effectively pre-extract these key structures; second, the global position relationship modeling of nodes is insufficient. Although existing methods capture local features through subgraph extraction, they ignore the relative position of nodes in the entire graph and cross-subgraph associations, making it difficult for the model to identify the semantic consistency of the same node in different substructures. More seriously, the separation of local features and global information directly restricts the performance improvement of graph classification tasks.
[0004] Through the above analysis, the problems and defects of the existing technology are as follows:
[0005] Graph neural networks (GNNs) have made significant progress in graph classification tasks. However, existing GNNs have limitations in recognizing fine-grained substructures. In addition, most subgraph extraction methods in graph classification tasks cannot effectively integrate global graph information, thus limiting the accuracy of graph classification. Summary of the Invention
[0006] In response to the problems existing in the prior art, the present invention provides a graph classification method based on subgraph integration and location awareness.
[0007] The present invention is implemented as follows: a graph classification method based on subgraph integration and location awareness includes:
[0008] Step 1: Node-based and graph-based subgraph extraction strategies are adopted to capture different substructures in the graph, and their structural features are encoded using a graph convolutional network, which is then fused using an attention mechanism.
[0009] Step 2: Use an anchor-based method to calculate the relative position information of the nodes and embed it into the node representation to capture the global position feature;
[0010] Step 3: Use the graph information bottleneck mechanism to optimize the node representation and remove redundant information irrelevant to the classification task.
[0011] Furthermore, the graph information bottleneck:
[0012] The representation h learned from the original data contains the maximum amount of information related to the predicted target y (i.e., maximizing I(h; y)) while filtering out redundant information irrelevant to the prediction task (i.e., minimizing I(x; h)). Its objective function can be formally expressed as:
[0013] L IB =I(y;h)-βI(x;h) (1)
[0014] where I(;) represents mutual information; KL divergence can be used to calculate mutual information
[26] :
[0015]
[0016] Where P(y,h) is the joint probability distribution of y and h, P(y) and P(h) are the marginal probability distributions of y and h respectively; similarly, I(x;h) can be calculated.
[0017] Furthermore, the subgraph:
[0018] (1) Subgraph type selection and definition
[0019] To cover a wide range of graph topologies, this paper selects four typical subgraph types: ring subgraph, tree subgraph, clique subgraph, and cut subgraph. These subgraph types can effectively capture the cyclic structure, hierarchical structure, dense connections, and structural information of key regions in the graph. In an undirected graph g = (V, E), the following are the definitions of the four subgraphs of graph g:
[0020] 1) Ring subgraph: A ring subgraph is composed of a subset of vertices. The closed loop formed satisfies |V ′ |≥3, and there is a simple path between any two vertices u and v in V', and the starting and ending vertices of the path are the same, forming a closed loop;
[0021] 2) Cluster subgraph: A cluster subgraph is composed of a subset of vertices. The complete subgraph is such that there exists (u,v)∈U between any two vertices u and v in U.
[0022] 3) Tree subgraph: t represents a tree subgraph; for any u, v∈V, there is a unique path from u to v, and there is no cycle in the tree t;
[0023] 4) Cut subgraph: a cut subgraph Cut(v) b A connected subgraph is generated by removing some edges from a graph. The process of generating a cut subgraph is as follows: first, the edge betweenness centrality (EBC) of each edge in the graph is calculated, which measures the importance of the edge's connectivity in the graph. Then, starting with the edge with the largest EBC, the edges are removed one by one until the graph is split into b connected blocks.
[0024] (2) Substructure coding
[0025] Obtain different types of subgraphs Cs from a graph = {c1, c2, ...c m}, Ts={t1,t2,...t n}, Q s ={q1,q2,...q o}, Cut={Cut1, Cut2, ...Cut a Where Cs represents the set of rings, Ts represents the set of trees, Qs represents the set of cliques, and Cut represents the set of cut subgraphs. For all subgraph types, shared GCN is used to capture features. Taking tree subgraphs as an example, for each tree subgraph t∈Ts, node features are first learned through GCN:
[0026] h t =GCN(t)(3)
[0027] Among them, h t is the embedding representation of the tree substructure t learned by GCN; then, the features of all the tree substructures in the graph are added together to obtain the tree embedding representation H of graph g Ts ∈R n×d In this process, since the number of nodes in different tree subgraphs may be different, the node filling is used to expand the dimension of its embedding vector to the same as Figure 1 The same length; similarly, the ring embedding H Cs ∈R n×d , cluster embedded in H Qs ∈R n×d and cut-embed H Cut ∈R n×d ; In order to integrate different types of substructure embeddings, the attention mechanism att(H Ts ,H Cs ,H Qs,H Cut ) to learn their corresponding importance (α Ts ,α Cs ,α Qs ,α Cut );
[0028] (α Ts , α Cs , α Qs , α Cut )=att(H Ts ,H Cs ,H Qs ,H Cut ) (4)
[0029] Here α Ts , α Cs , α Qs , α Cut Represent the embedding H Ts ,H Cs ,H Qs ,H Cut The attention value of ; the nonlinear transformation of the embedding is performed and the attention value is obtained as follows:
[0030] ω Ts =W·(H Ts ) T +b (5)
[0031] Where W is the weight matrix and b is the bias vector; similarly, the embedding matrix ω of each subgraph can be obtained Ts 、ω Qs and ω Cut ; Then, the attention values are normalized using the softmax function to obtain the final weights:
[0032]
[0033] where α Ts The larger it is, the more important the corresponding embedding is; similarly, α Cs =softmax(ω Cs ), α Qs =softmax(ω Qs ), α Cut =softmax(ω Cut ) Then, these four embeddings are combined to get the final embedding H SE :
[0034] H SE =α Ts ⊙H Ts +α Cs ⊙H Cs +α Qs ⊙HQs +α Cut ⊙H Cut (7)
[0035] Among them, α Ts ⊙H Ts Indicates that α Ts After the broadcast, Ts Bit-by-bit multiplication, H SE ∈R n×d is the fine-grained node embedding representation output by the SE module.
[0036] Furthermore, the anchor point:
[0037] 1) Anchor point selection
[0038] By selecting some important nodes with high centrality or representativeness, the model can better understand and distinguish the global topological structure of the graph; in this section, the node selection strategy is used to select the first k nodes as anchor nodes P = {p1, p2, p3…p k}, these nodes provide a reference for the relative positions of other nodes; the node selection strategy combines GCN and MLP to evaluate the importance of nodes from multiple perspectives and retain the most representative nodes;
[0039]
[0040] H mlp =σ(HW mlp ) (9)
[0041] S=H gcn +H mlp (10)
[0042] Where σ(·) is the activation function; W gcn and W mlp ∈R d×1 is the parameter matrix used for linear transformation of node features;
[0043] Because larger scores in vector S correspond to more important nodes, the k nodes with the highest scores are selected as anchor nodes;
[0044] 2) Location-aware computing
[0045] Here, location awareness can be viewed as a form of embedding method; inspired by P-GNNs
[19] , this paper measures location information by calculating location similarity; specifically, the shortest path distance of a node relative to an anchor node is calculated and encoded as a location embedding s(v,p i ):
[0046]
[0047] Among them, P i is one of the anchor nodes, is the importance score of the anchor node p, d(v,P i ) is the shortest path distance between nodes v and pi; s(v,P i ) represents the position similarity between them;
[0048] Next, a more comprehensive location-aware node embedding is generated by combining node features and location similarity;
[0049]
[0050] in, It is the node representation obtained by encoding the graph through GCN. represents the node embedding of node v after GCN, Represents node P i Node embedding after GCN, yes and connection;
[0051] Finally, by aggregating the embeddings of different anchor nodes i , we can get the location-aware node embedding H AGG The specific formula is as follows:
[0052] H AGG =Mean(F(v,P1),F(v,P2),...,F(v,P k )) (14)
[0053] Where Mean(·) represents the bitwise average; thus, position embedding is added to each node, and the node embedding representation H is finally obtained. PE ∈R n×d .
[0054] Furthermore, the graph information bottleneck mechanism is used to optimize node representation:
[0055] Consistent representation learning uses the information bottleneck GIB mechanism as an auxiliary optimization objective. The GIB objective function can be written as:
[0056] L IB =γI(H SE ;H G )+(1-γ)I(H PE ;H G ) (15)
[0057] Among them, I(H SE ;H G ) is to minimize H SE and HG The mutual information between PE ;H G ) is to minimize H PE and H G The mutual information between them; γ is a learnable parameter used to adjust the influence of each part; calculate I(H SE ;H G ) and I(H LE ;H G ) can be estimated using KL divergence;
[0058]
[0059] Similarly, for I(H PE ;H G ):
[0060]
[0061] Where D KL (||) represents the KL divergence between two probability distributions, P(H PE ;H G ) represents the joint probability distribution, P(H PE )、P(H G ) represent H PE and H G The marginal probability distribution of ;
[0062] Joint loss function.
[0063] Furthermore, the joint loss function is:
[0064] Embed the substructure into H SE With position-aware embedding H PE Fusion is performed to generate the final node embedding representation H End To flexibly adapt to different task requirements, an adaptive learning mechanism is introduced to dynamically adjust the degree of fusion of the two embeddings by learning weights. In this way, the final graph-level embedding representation Z not only comprehensively considers the structural information of the graph, but also effectively preserves the position information of the nodes.
[0065] H END =β1H SE +(1-β1)H PE (18)
[0066] Z=Readout(H End ) (19)
[0067] Among them, Z represents the fused graph embedding representation, and Readout(.) represents the averaging operation, which is a hyperparameter;
[0068] The final fused graph-level embedding representation Z is used for downstream graph classification tasks; the embedding is classified through a fully connected layer:
[0069]
[0070] Among them, F is the number of categories, y i,j is the one-hot encoding of the real category, The class probability predicted by the model;
[0071] The information bottleneck loss and cross entropy loss are weighted and combined to obtain the final loss function:
[0072] L=β2L ce +(1-β2)L IB (twenty two)
[0073] Among them, β2 is a hyperparameter used to adjust the weight of the information bottleneck loss in the overall loss function; the model parameters are optimized using the backpropagation algorithm to minimize the loss function.
[0074] Another object of the present invention is to provide a graph classification system based on subgraph integration and location awareness, including:
[0075] A fusion module that uses node-based and graph-based subgraph extraction strategies to capture different substructures in the graph, encodes their structural features using a graph convolutional network, and then fuses them using an attention mechanism;
[0076] A calculation module is used to calculate the relative position information of the nodes using an anchor-based method and embed it into the node representation to capture the global position feature;
[0077] The removal module is used to optimize node representation by utilizing the graph information bottleneck mechanism and remove redundant information that is irrelevant to the classification task.
[0078] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the graph classification method based on subgraph integration and location awareness.
[0079] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the graph classification method based on subgraph integration and location awareness.
[0080] Another object of the present invention is to provide an information data processing terminal, which is used to implement the graph classification system based on subgraph integration and location awareness.
[0081] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0082] First, this paper proposes a graph classification method based on subgraph integration and position awareness (SIPA). This method uses two strategies to explicitly extract and integrate subgraph information, while utilizing position-aware embedding to capture global graph information. First, node-based and graph-based subgraph extraction strategies are adopted to capture different substructures in the graph, and a graph convolutional network is used to encode their structural features, which are then fused using an attention mechanism. Next, an anchor-based method is used to calculate the relative position information of the nodes and embed it into the node representation to capture global position features. Finally, the graph information bottleneck mechanism is used to optimize the node representation and remove redundant information that is irrelevant to the classification task. Extensive experiments demonstrate the effectiveness of the proposed method, with an average improvement of 3.3% in graph classification accuracy compared to mainstream baselines.
[0083] This paper proposes a graph classification method based on subgraph integration and position awareness (SIPA). This method explicitly extracts and integrates subgraph information through two strategies and captures global graph information using position-aware embedding. The framework consists of three main modules: a subgraph embedding (SE) module, a position-aware embedding (PE) module, and a consistent representation learning module. In the SE module, two subgraph selection strategies are adopted: a node-based strategy and a graph-based strategy. The node-based strategy extracts specific substructures, including cycles, trees, and cliques, to capture different types of fine-grained substructures in the graph. The graph-based subgraph strategy uses cut subgraphs for extraction, extracting substructure from a holistic perspective by selectively removing edges from the original graph. To enhance global representation capabilities, the PE module samples anchor nodes and aggregates the position information associated with these anchor nodes into each node. Finally, consistent representation learning uses the graph information bottleneck (GIB) mechanism as an auxiliary optimization objective, which regularizes the mutual information to remove irrelevant information from the task. In summary, the main contributions of this paper are as follows.
[0084] (1) A subgraph embedding module is designed. This module can capture both local and global substructure information, which is conducive to better learning of substructure information.
[0085] (2) A location-aware module is proposed, which selects anchor nodes through a node sampling strategy and integrates the location information of anchor nodes into the node representation, thereby enhancing the global expressiveness of the graph.
[0086] (3) Comprehensive experiments show that the proposed SIPA achieves better graph classification than mainstream graph classification methods on five graph classification datasets.
[0087] Second, graph classification methods based on subgraph integration and location awareness can significantly reduce the time required to predict drug activity by accurately identifying molecular structures. In social networks, location-aware modules can improve the accuracy of identifying fake accounts, and this technology has high potential value in the graph analysis market.
[0088] The graph classification method based on subgraph integration and position awareness breaks through the traditional separation problem of local features and global information. The existing technology extracts subgraphs through a single strategy, resulting in the loss of key molecular structure information. The present invention creatively proposes a dynamic adaptation mechanism for subgraph types, defines exclusive extraction rules for specific structures, and achieves accurate capture of specific fields.
[0089] The field of graph neural networks has long harbored a technical bias that global position information is difficult to effectively encode. Mainstream methods such as GraphSAGE and GAT rely on neighborhood aggregation strategies. This paper pioneers a groundbreaking anchor importance assessment system, integrating GCN feature responses with MLP structural analysis to establish a multi-dimensional evaluation metric for node centrality. This nonlinear mapping of shortest path distances to spatial weights exponentially represents the positional differences of edge nodes relative to core anchors.
[0090] Third, the graph classification method based on subgraph integration and location awareness proposed in this paper, which integrates the detailed modeling of graph structure information with the precise embedding of node positions, breaks through the limitations of existing graph neural networks in capturing structural information and modeling inter-node position information. It specifically solves the following key technical problems and achieves significant technological progress:
[0091] The present invention introduces ring, tree, cluster and shear Figure 4 The team identified typical structural subgraph types, established a shared encoding channel for these subgraph types, and extracted embedding features for each subgraph type using a graph convolutional network (GCN). Incorporating an attention mechanism to assess and weight cross-subgraph importance, they constructed a unified substructure ensemble embedding (H^SE). This effectively improved the ability to express complex graph structures and addressed the uneven performance of existing GNN models when handling different local structures.
[0092] The present invention introduces an anchor mechanism. By selecting representative nodes with high centrality or strong representation ability, the shortest path distance of any node relative to the anchor point is calculated. The feature representation of the node and the anchor point is combined to construct a position similarity embedding (H^PE). This captures the relative position information of the node in the entire graph, thus breaking through the limitation of traditional GCN based only on local propagation of the adjacency matrix and achieving a more comprehensive modeling of the global topological structure of the graph.
[0093] This paper introduces a learnable adaptive weighting mechanism during the fusion phase. By dynamically adjusting the weight ratio between the structural embedding H^SE and the positional embedding H^PE, it generates a task-aware node representation (H^End). This ensures that the model has strong generalization capabilities across different tasks and graph types. Furthermore, the fused graph-level representation Z preserves both structural and positional characteristics, providing a more robust feature foundation for graph classification.
[0094] This paper introduces a graph information bottleneck mechanism during the training phase. By minimizing the mutual information I(h;G) between the structural representation, the position representation, and the full graph representation, it enhances representation sparsity and reduces redundant interference. It also jointly optimizes the information bottleneck loss and the cross-entropy loss to construct a joint loss function L, improving the model's discriminative ability and generalization performance. This mechanism significantly improves classification accuracy on multiple public datasets, validating the model's significant technical advancements in noise suppression and feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 This is a flow chart of a graph classification method based on subgraph integration and location awareness provided by an embodiment of the present invention.
[0096] Figure 2 This is a structural block diagram of a graph classification system based on subgraph integration and location awareness provided by an embodiment of the present invention.
[0097] Figure 3 The overall framework of the SIPA method provided by the embodiment of the present invention includes three main modules: subgraph embedding (SE), position-aware embedding (PE) and consistent representation learning graph.
[0098] Figure 4 This is a diagram showing the impact of different sampling strategies on graph classification accuracy provided by an embodiment of the present invention.
[0099] Figure 5 This is a diagram showing the impact of different subgraph fusion strategies on graph classification accuracy provided by an embodiment of the present invention.
[0100] Figure 6 This is a graph showing how the graph classification accuracy changes with the number of GCN layers, provided by an embodiment of the present invention.
[0101] Figure 7 This is a graph showing how the graph classification accuracy changes with the embedding dimension, as provided by an embodiment of the present invention.
[0102] Figure 8 This is a graph showing how the accuracy of graph classification changes with the learning rate, as provided by an embodiment of the present invention.
[0103] Figure 9 This is a graph showing how the graph classification accuracy changes with the number of anchor nodes, as provided by an embodiment of the present invention.
[0104] Figure 10 This is a graph showing how the accuracy of graph classification changes with the number of cut subgraphs, as provided by an embodiment of the present invention.
[0105] Figure 11 This is a visualization diagram of the GCN provided by the embodiment of the present invention for classifying graphs of different data sets.
[0106] Figure 12 This is a visualization diagram of SIPA used for classification of graphs of different data sets provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0107] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0108] like Figure 1 As shown, an embodiment of the present invention provides a graph classification method based on subgraph integration and location awareness, comprising the following steps:
[0109] S101 adopts node-based and graph-based subgraph extraction strategies to capture different substructures in the graph, encodes their structural features using a graph convolutional network, and then fuses them using an attention mechanism;
[0110] S102, using an anchor-based method to calculate the relative position information of the nodes and embed it into the node representation to capture the global position feature;
[0111] S103, using the graph information bottleneck mechanism to optimize the node representation and remove redundant information irrelevant to the classification task.
[0112] The graph information bottleneck provided by the embodiment of the present invention:
[0113] The representation h learned from the original data contains the maximum amount of information related to the predicted target y (i.e., maximizing I(h; y)) while filtering out redundant information irrelevant to the prediction task (i.e., minimizing I(x; h)). Its objective function can be formally expressed as:
[0114] L IB =I(y;h)-βI(x;h) (1)
[0115] where I(;) represents mutual information; KL divergence can be used to calculate mutual information
[26] :
[0116]
[0117] Where P(y,h) is the joint probability distribution of y and h, P(y) and P(h) are the marginal probability distributions of y and h respectively; similarly, I(x;h) can be calculated.
[0118] The subgraph provided by the embodiment of the present invention is:
[0119] (1) Subgraph type selection and definition
[0120] To cover a wide range of graph topologies, this paper selects four typical subgraph types: ring subgraph, tree subgraph, clique subgraph, and cut subgraph. These subgraph types can effectively capture the cyclic structure, hierarchical structure, dense connections, and structural information of key regions in the graph. In an undirected graph g = (V, E), the following are the definitions of the four subgraphs of graph g:
[0121] 1) Ring subgraph: A ring subgraph is composed of a subset of vertices. The closed loop formed satisfies |V ′ |≥3, and there is a simple path between any two vertices u and v in V', and the starting and ending vertices of the path are the same, forming a closed loop;
[0122] 2) Cluster subgraph: A cluster subgraph is composed of a subset of vertices. The complete subgraph is such that there exists (u,v)∈U between any two vertices u and v in U.
[0123] 3) Tree subgraph: t represents a tree subgraph; for any u, v∈V, there is a unique path from u to v, and there is no cycle in the tree t;
[0124] 4) Cut subgraph: a cut subgraph Cut(v) b A connected subgraph is generated by removing some edges from a graph. The process of generating a cut subgraph is as follows: first, the edge betweenness centrality (EBC) of each edge in the graph is calculated, which measures the importance of the edge's connectivity in the graph. Then, starting with the edge with the largest EBC, the edges are removed one by one until the graph is split into b connected blocks.
[0125] (2) Substructure coding
[0126] Obtain different types of subgraphs Cs from a graph = {c1, c2, ...c m}, Ts={t1,t2,...t n}, Q s ={q1,q2,...q o}, Cut={Cut1, Cut2, ...Cut aWhere Cs represents the set of rings, Ts represents the set of trees, Qs represents the set of cliques, and Cut represents the set of cut subgraphs. For all subgraph types, shared GCN is used to capture features. Taking tree subgraphs as an example, for each tree subgraph t∈Ts, node features are first learned through GCN:
[0127] h t =GCN(t)(3)
[0128] Among them, h t is the embedding representation of the tree substructure t learned by GCN; then, the features of all the tree substructures in the graph are added together to obtain the tree embedding representation H of graph g Ts ∈R n×d In this process, since the number of nodes in different tree subgraphs may be different, the node filling is used to expand the dimension of its embedding vector to the same as Figure 1 The same length; similarly, the ring embedding H Cs ∈R n×d , cluster embedded in H Qs ∈R n×d and cut-in H Cut ∈R n×d ; In order to integrate different types of substructure embeddings, the attention mechanism att(H Ts ,H Cs ,H Qs ,H Cut ) to learn their corresponding importance (α Ts ,α Cs ,α Qs ,α Cut );
[0129] (α Ts , α Cs , α Qs , α Cut )=att(H Ts ,H Cs ,H Qs ,H Cut ) (4)
[0130] Here α Ts , α Cs , α Qs , α Cut Represent the embedding H Ts ,H Cs ,H Qs ,H Cut The attention value of ; the nonlinear transformation of the embedding is performed and the attention value is obtained as follows:
[0131] ω Ts =W·(H Ts) T +b (5)
[0132] Where W is the weight matrix and b is the bias vector; similarly, the embedding matrix ω of each subgraph can be obtained Ts 、ω Qs and ω Cut ; Then, the attention values are normalized using the softmax function to obtain the final weights:
[0133]
[0134] where α Ts The larger it is, the more important the corresponding embedding is; similarly, α Cs =softmax(ω Cs ), α Qs =softmax(ω Qs ), α Cut =softmax(ω Cut ) Then, these four embeddings are combined to get the final embedding H SE :
[0135] H SE =α Ts ⊙H Ts +α Cs ⊙H Cs +α Qs ⊙H Qs +α Cut ⊙H Cut (7)
[0136] Among them, α Ts ⊙H Ts Indicates that α Ts After the broadcast, Ts Bit-by-bit multiplication, H SE ∈R n×d is the fine-grained node embedding representation output by the SE module.
[0137] Anchor points provided by embodiments of the present invention:
[0138] 1) Anchor point selection
[0139] By selecting some important nodes with high centrality or representativeness, the model can better understand and distinguish the global topological structure of the graph; in this section, the node selection strategy is used to select the first k nodes as anchor nodes P = {p1, p2, p3…p k}, these nodes provide a reference for the relative positions of other nodes; the node selection strategy combines GCN and MLP to evaluate the importance of nodes from multiple perspectives and retain the most representative nodes;
[0140]
[0141] H mlp =σ(HW mlp ) (9)
[0142] S=H gcn +H mlp (10)
[0143] Where σ(·) is the activation function; W gcn and W mlp ∈R d×1 is the parameter matrix used for linear transformation of node features;
[0144] Because larger scores in vector S correspond to more important nodes, the k nodes with the highest scores are selected as anchor nodes;
[0145] 2) Location-aware computing
[0146] Here, location awareness can be viewed as a form of embedding method; inspired by P-GNNs
[19] , this paper measures location information by calculating location similarity; specifically, the shortest path distance of a node relative to an anchor node is calculated and encoded as a location embedding s(v,p i ):
[0147]
[0148] Among them, P i is one of the anchor nodes, S Pi is the importance score of the anchor node p, d(v,P i ) is the shortest path distance between nodes v and pi; s(v,P i ) represents the position similarity between them;
[0149] Next, a more comprehensive location-aware node embedding is generated by combining node features and location similarity;
[0150]
[0151] in, It is the node representation obtained by encoding the graph through GCN. represents the node embedding of node v after GCN, Represents node P i Node embedding after GCN, yes and connection;
[0152] Finally, by aggregating the embeddings of different anchor nodes i, we can get the location-aware node embedding H AGG The specific formula is as follows:
[0153] H AGG =Mean(F(v,P1),F(v,P2),...,F(v,P k )) (14)
[0154] Where Mean(·) represents the bitwise average; thus, position embedding is added to each node, and the node embedding representation H is finally obtained. PE ∈R n×d .
[0155] The embodiment of the present invention provides an optimization method for node representation using a graph information bottleneck mechanism:
[0156] Consistent representation learning uses the information bottleneck GIB mechanism as an auxiliary optimization objective. The GIB objective function can be written as:
[0157] L IB =γI(H SE ;H G )+(1-γ)I(H PE ;H G ) (15)
[0158] Among them, I(H SE ;H G ) is to minimize H SE and H G The mutual information between PE ;H G ) is to minimize H PE and H G The mutual information between them; γ is a learnable parameter used to adjust the influence of each part; calculate I(H SE ;H G ) and I(H LE ;H G ) can be estimated using KL divergence;
[0159]
[0160] Similarly, for I(H PE ;H G ):
[0161]
[0162] Where D KL (||) represents the KL divergence between two probability distributions, P(H PE ;H G ) represents the joint probability distribution, P(H PE )、P(HG ) represent H PE and H G The marginal probability distribution of ;
[0163] Joint loss function.
[0164] The joint loss function provided by the embodiment of the present invention is:
[0165] Embed the substructure into H SE With position-aware embedding H PE Fusion is performed to generate the final node embedding representation H End To flexibly adapt to different task requirements, an adaptive learning mechanism is introduced to dynamically adjust the degree of fusion of the two embeddings by learning weights. In this way, the final graph-level embedding representation Z not only comprehensively considers the structural information of the graph, but also effectively preserves the position information of the nodes.
[0166] H END =β1H SE +(1-β1)H PE (18)
[0167] Z=Readout(H End ) (19)
[0168] Among them, Z represents the fused graph embedding representation, and Readout(.) represents the averaging operation, which is a hyperparameter;
[0169] The final fused graph-level embedding representation Z is used for downstream graph classification tasks; the embedding is classified through a fully connected layer:
[0170]
[0171] Among them, F is the number of categories, y i,j is the one-hot encoding of the real category, The class probability predicted by the model;
[0172] The information bottleneck loss and cross entropy loss are weighted and combined to obtain the final loss function:
[0173] L=β2L ce +(1-β2)L IB (twenty two)
[0174] Among them, β2 is a hyperparameter used to adjust the weight of the information bottleneck loss in the overall loss function; the model parameters are optimized using the backpropagation algorithm to minimize the loss function.
[0175] like Figure 2As shown, an embodiment of the present invention provides a graph classification system based on subgraph integration and location awareness, including:
[0176] A fusion module that uses node-based and graph-based subgraph extraction strategies to capture different substructures in the graph, encodes their structural features using a graph convolutional network, and then fuses them using an attention mechanism;
[0177] A calculation module is used to calculate the relative position information of the nodes using an anchor-based method and embed it into the node representation to capture the global position feature;
[0178] The removal module is used to optimize node representation by utilizing the graph information bottleneck mechanism and remove redundant information that is irrelevant to the classification task.
[0179] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the graph classification method based on subgraph integration and location awareness.
[0180] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the graph classification method based on subgraph integration and location awareness.
[0181] Another object of the present invention is to provide an information data processing terminal, which is used to implement the graph classification system based on subgraph integration and location awareness.
[0182] The present invention is specifically implemented:
[0183] 1. A graph classification method based on Subgraph Integration and Position Awareness (SIPA). This method explicitly extracts and integrates subgraph information through two strategies and captures global graph information using position-aware embedding. The framework consists of three main modules: a subgraph embedding (SE) module, a position-aware embedding (PE) module, and a consistent representation learning module. In the SE module, two subgraph selection strategies are adopted: a node-based strategy and a graph-based strategy. The node-based strategy extracts specific substructures, including cycles, trees, and cliques, to capture different types of fine-grained substructures in the graph. The graph-based subgraph strategy uses cut subgraphs for extraction, extracting substructure from a holistic perspective by selectively removing edges from the original graph. To enhance global representation capabilities, the PE module samples anchor nodes and aggregates the position information associated with these anchor nodes into each node. Finally, consistent representation learning uses the Graph Information Bottleneck (GIB) mechanism as an auxiliary optimization objective, which regularizes the mutual information to remove irrelevant information from the task.
[0184] 2Related Work
[0185] 2.1 Graph Neural Networks
[0186] Graph Neural Networks (GNNs) have demonstrated powerful capabilities in graph representation learning. Gilmer et al.
[13] proposed a message passing neural network (MPNN) framework that iteratively aggregates the feature representations of neighbors to update the feature representation of the central node. The updated representation is then passed to a readout function for graph-level classification tasks. Both Graph Convolutional Networks (GCNs) [4] and Graph Attention Networks (GATs) [5] follow the message passing paradigm. The difference is that GCNs use normalized convolutions to aggregate neighborhood features, while GATs use attention coefficients. GraphSage [6] randomly samples a fixed number of neighbors for each node to reduce computational complexity.
[0187] In recent years, a number of innovative studies have combined graph neural networks (GNNs) with advanced mathematical tools to address the challenges of graph learning: PGOT
[14] innovatively introduced the fused Gromov-Wasserstein distance from Optimal Transport Theory (OT) as a graph structure similarity metric, explicitly modeled graph structure features through the OT framework, and achieved an end-to-end interpretable learning process while improving classification transparency; TP-GNN
[15] specifically designed a time-aware graph representation learning architecture for continuous dynamic graph scenarios, effectively capturing the time-varying characteristics of graph structure through a dynamic neighbor aggregation mechanism; G-Prompt
[16] improved graph classification accuracy by optimizing the graph-level prompt generation module. CAMA
[17] aims to attack graph classification models by generating adversarial samples containing global graph-level information at the node level. ICL
[18] integrates information theory and causal reasoning. It introduces a mutual information objective to reduce irrelevant features, enhancing the robustness and interpretability of GNNs.
[0188] 2.2 Position information in graphics
[0189] To enhance the expressive power of graph neural networks (GNNs), researchers have deeply explored the representation mechanism of relative distances between nodes. Position information plays a key role in distinguishing topologically isomorphic substructures in graphs. Recent studies have focused on various strategies for incorporating position encoding into GNNs: P-GNN
[19] generates position embeddings by calculating the shortest path distance between each node and the selected anchor node, and then fuses it with the GNN representation and inputs it into the classification layer; GraphReach
[20] optimizes the anchor node selection strategy through a greedy hill climbing algorithm to maximize reachability, but its high computational complexity limits its application on large-scale graphs; AdaSNN
[21] uses Laplacian matrices or random walks to generate position encodings, and implements task-oriented dynamic adaptive updates through learnable parameters; PO-GNN
[22] proposes a time-aware position encoding framework for dynamic graphs, incorporating temporal relationship dependencies into the embedding calculation process; SPGNN
[23] achieves position awareness through the joint optimization of local neighborhood information aggregation and global position encoding; CP-GNN
[24] innovatively divides neighborhood nodes into three directions: left / self / right, and implements multi-directional position relationship modeling through the attention mechanism.
[0190] In this paper, we sample nodes through a node selection strategy and aggregate the location information associated with these anchor nodes into the representation of each node. This process enriches the node representation by introducing global location information, allowing nodes to contain not only local neighborhood information but also the global structure of the graph.
[0191] 2.3 Graph Information Bottleneck
[0192] The Information Bottleneck (IB)
[25] is a core principle in information theory. It uses mutual information as a regularization term to establish a balance between raw data and generalization ability. The core goal of this theory is to force the representation h learned from the raw data to contain the maximum amount of information related to the predicted target y (i.e., maximizing I(h; y)) while filtering out redundant information that is irrelevant to the prediction task (i.e., minimizing I(x; h)). Its objective function can be formally expressed as:
[0193] L IB =I(y;h)-βI(x;h) (1)
[0194] where I(;) represents mutual information. KL divergence can be used to calculate mutual information
[26] :
[0195]
[0196] where P(y,h) is the joint probability distribution of y and h, and P(y) and P(h) are the marginal probability distributions of y and h, respectively. Similarly, I(x;h) can be calculated.
[0197] In recent years, some studies have extended the information bottleneck principle to graph structure learning and utilized feature denoising. LGCN-SGIB
[27] maximizes the mutual information between channels of the same and different modalities, enhancing node classification performance in downstream tasks. LGG-DMRL
[28] applies the information bottleneck principle to identify shared representations across views and combines them with view-specific feature representations. This promotes complementarity and completeness. Yu et al.
[29] addressed the subgraph identification problem by estimating mutual information in irregular graph data.
[0198] This paper minimizes the embeddings between the substructure and the original graph, as well as the embeddings between the position information and the original graph. This strategy ensures that the embeddings of the substructure and position information become similar to those of the original graph. By bringing these embeddings closer, the model can better capture the intrinsic relationships between the subgraph structure, position information, and the overall graph.
[0199] In this section, the proposed SIPA graph classification method is introduced in detail. The overall framework is as follows Figure 3 As shown in the figure. The framework consists of three modules: (1) Substructure Embedding Module (SE). Substructures are extracted according to node-based and graph-based strategies, these substructures are encoded, and attention fusion is used to generate fine-grained node representations. (2) Position-Aware Embedding Module (PE). PE adopts a node selection strategy to obtain anchor nodes, and then summarizes the position information of these nodes relative to the anchor nodes into the node embedding to generate a global representation. (3) Consistent Representation Learning Module. The most relevant information about the original graph is extracted from the SE and PE representations through the GIB mechanism, thereby generating optimized graph embeddings for downstream graph classification tasks. Next, these three modules are introduced in detail.
[0200] 3.1 Problem Definition and Notation
[0201] A graph can be represented as g = (V, E), where V is the set of nodes and E is the set of edges in the graph. The adjacency matrix and attribute matrix of the graph are denoted as A∈Rn×n and X∈Rn×d, where n is the number of nodes and d is the dimension of the node attributes. The set of neighbors of node v is denoted as N(v). Given a graph containing M nodes with class labels The dataset of the graph, where y i It is g i The goal of the graph classification problem is to learn a projection function from the graph space to the category label space, and then use the projection function to predict the category label of the new graph. For convenience, Table 1 lists the definitions of some important symbols used.
[0202] Table 1 Problem definition and notation
[0203]
[0204]
[0205] 3.2 Substructure Embedding Module
[0206] Substructures contain key information and usually provide more discriminative information than a single node or edge. The key to utilizing substructures is to transmit information directly at the substructure level, rather than being limited to the interaction of a single node. This can generate richer contextual information for each node, and therefore performs well in graph classification tasks. In order to deeply explore the substructure information in the graph, this chapter designs a substructure embedding module (SE) to extract diverse structural features from the input graph. The module first generates different types of substructures through two substructure extraction strategies, node-based and graph-based, and then encodes these substructures using an encoder. Finally, the substructure embedding is integrated through the attention mechanism to generate fine-grained node representations. The goal of the SE module is to extract rich local structural features from the input graph to capture the local topological information of the graph.
[0207] 3.2.1 Subgraph Type Selection and Definition
[0208] To cover a wide range of graph topologies, this paper selects four typical subgraph types: ring subgraph, tree subgraph, clique subgraph, and cut subgraph. These subgraph types can effectively capture the structure of loops, hierarchies, dense connections, and key regions in the graph. In an undirected graph g = (V, E), the following are the four subgraph definitions of graph g:
[0209] 1) Ring subgraph: A ring subgraph is composed of a subset of vertices. The closed loop formed satisfies |V ′ |≥3, and there is a simple path between any two vertices u and v in V', and the starting and ending vertices of the path are the same, forming a closed loop.
[0210] 2) Cluster subgraph: A cluster subgraph is composed of a subset of vertices. The complete subgraph is formed, satisfying that there exists (u,v)∈U between any two vertices u and v in U.
[0211] 3) Tree subgraph: t represents a tree subgraph. For any u, v ∈ V, there is a unique path from u to v, and there is no cycle in the tree t.
[0212] 4) Cut subgraph: a cut subgraph Cut(v) bA connected subgraph is generated by removing some edges from a graph. The process of generating a cut subgraph is as follows: first, the edge betweenness centrality (EBC) of each edge in the graph is calculated. This metric measures the importance of the edge's connectivity in the graph. Then, starting with the edge with the largest EBC, the edges are removed incrementally until the graph is partitioned into b connected blocks.
[0213] 3.2.2 Substructure Encoding
[0214] Obtain different types of subgraphs Cs from a graph = {c1, c2, ...c m}, Ts={t1,t2,...t n}, Q s ={q1,q2,...q o}, Cut={Cut1, Cut2, ...Cut a Where Cs represents the set of rings, Ts represents the set of trees, Qs represents the set of cliques, and Cut represents the set of cut subgraphs. For all subgraph types, a shared GCN is used to capture features. Taking the tree subgraph as an example, for each tree subgraph t∈Ts, the node features are first learned by GCN:
[0215] h t =GCN(t)(3)
[0216] Among them, h t is the embedding representation of the tree substructure t learned by GCN. Then, the features of all the tree substructures in the graph are added together to obtain the tree embedding representation H of the graph g. Ts ∈R n×d In this process, since the number of nodes in different tree subgraphs may be different, the node filling is used to expand the dimension of its embedding vector to the same as Figure 1 Similarly, we get the ring embedding H Cs ∈R n×d , cluster embedded in H Qs ∈R n×d and cut-embed H Cut ∈R n×d In order to integrate different types of substructure embeddings, the attention mechanism att(H Ts ,H Cs ,H Qs ,H Cut ) to learn their corresponding importance (α Ts ,α Cs ,α Qs ,α Cut ).
[0217] (α Ts , α Cs , α Qs , αCut )=att(H Ts ,H Cs ,H Qs ,H Cut ) (4)
[0218] Here α Ts , α Cs , α Qs , α Cut Represent the embedding H Ts ,H Cs ,H Qs ,H Cut The attention value of . Perform nonlinear transformation on the embedding and get the attention value as follows:
[0219] ω Ts =W·(H Ts ) T +b (5)
[0220] Where W is the weight matrix and b is the bias vector. Similarly, the embedding matrix ω of each subgraph can be obtained Ts 、ω Qs and ω Cut The attention values are then normalized using the softmax function to obtain the final weights:
[0221]
[0222] where α Ts The larger α is, the more important the corresponding embedding is. Cs =softmax(ω Cs ), α Qs =softmax(ω Qs ), α Cut =softmax(ω Cut ) Then, these four embeddings are combined to get the final embedding H SE :
[0223] H SE =α Ts ⊙H Ts +α Cs ⊙H Cs +α Qs ⊙H Qs +α Cut ⊙H Cut (7)
[0224] Among them, α Ts ⊙H Ts Indicates that α Ts After the broadcast, Ts Bit-by-bit multiplication, H SE ∈Rn×d is the fine-grained node embedding representation output by the SE module.
[0225] 3.3 Location Awareness Module
[0226] The topological structure of a graph is not only determined by the connection relationship between nodes and edges, but is also closely related to the relative positions between nodes. Position information helps the model understand the overall layout of nodes in the graph and their relative positional relationships in the graph by explicitly encoding the relative positions of nodes. Although the SE module extracts local structural information, it cannot capture the relationship between distant but relatively important nodes in the graph. To solve this problem, a position-aware (PE) module is introduced. This module captures the global topological features of the graph by combining relative position information, and can distinguish the consistency of global positions even if nodes appear in different substructures. Specifically, the anchor nodes in the graph are identified through a node selection strategy, and then the shortest path distance of each node relative to the anchor node is calculated and converted into a position embedding.
[0227] 3.3.1 Anchor point selection
[0228] By selecting some important nodes with high centrality or representativeness, the model can better understand and distinguish the global topological structure of the graph. In this section, the node selection strategy is used to select the first k nodes as anchor nodes P = {p1, p2, p3…p k}, these nodes provide a reference for the relative positions of other nodes. The node selection strategy combines GCN and MLP to evaluate the importance of nodes from multiple perspectives and retain the most representative nodes.
[0229]
[0230] H mlp =σ(HW mlp ) (9)
[0231] S=Hgcn+H mlp (10)
[0232] Where σ(·) is the activation function. gcn and W mlp ∈R d×1 is the parameter matrix used for linear transformation of node features.
[0233] Because larger scores in vector S correspond to more important nodes, the k nodes with the highest scores are selected as anchor nodes.
[0234] 3.2.2 Location-Aware Computing
[0235] Here, location awareness can be considered as a form of embedding method. Inspired by P-GNNs
[19] , this paper measures location information by calculating location similarity. Specifically, the shortest path distance of a node relative to an anchor node is calculated and encoded as a location embedding s(v,p i ):
[0236]
[0237] Among them, P i is one of the anchor nodes, S Pi is the importance score of the anchor node p, d(v,P i ) is the shortest path distance between nodes v and pi. i ) represents the position similarity between them.
[0238] Next, a more comprehensive location-aware node embedding is generated by combining the node’s features and location similarity.
[0239]
[0240]
[0241] in, It is the node representation obtained by encoding the graph through GCN. represents the node embedding of node v after GCN, Represents node P i Node embedding after GCN, yes and connection.
[0242] Finally, by aggregating the embeddings of different anchor nodes i , we can get the location-aware node embedding H AGG The specific formula is as follows:
[0243] H AGG =Mean(F(v,P1),F(v,P2),...,F(v,P k )) (14)
[0244] Where Mean(·) represents the bitwise average. In this way, position embedding is added to each node, and the node embedding representation H is finally obtained. PE ∈R n×d .
[0245] 3.4 Consistent Representation Learning Module
[0246] In this section, consistent representation learning uses the information bottleneck (GIB) mechanism as an auxiliary optimization objective. The core idea of the GIB mechanism is to reduce the noise information in the input data and retain the key information related to the task, thereby improving the performance of the model on a specific task. In this way, GIB can guide the model to focus on the most meaningful features during the learning process, while eliminating redundant or irrelevant information. Therefore, following the information bottleneck strategy, sufficient semantic information is retained in SE and PE, which is more relevant to downstream tasks. Specifically, minimizing H SE 、H PE and input H G Irrelevant redundant information. Therefore, the GIB objective function can be written as:
[0247] L IB =γI(H SE ;H G )+(1-γ)I(H PE ;H G ) (15)
[0248] Among them, I(H SE ;H G ) is to minimize H SE and H G The mutual information between PE ;H G ) is to minimize H PE and H G The mutual information between them. γ is a learnable parameter used to adjust the influence of each part. Calculate I(H SE ;H G ) and I(H LE ;H G ) can be estimated using the KL divergence.
[0249]
[0250] Similarly, for I(H PE ;H G ):
[0251]
[0252] Where D KL (||) represents the KL divergence between two probability distributions, P(H PE ;H G ) represents the joint probability distribution, P(H PE )、P(H G ) represent H PE and H G The marginal probability distribution of .
[0253] 3.5 Joint Loss Function
[0254] Embed the substructure into H SE With position-aware embedding H PE Fusion is performed to generate the final node embedding representation H End To flexibly adapt to different task requirements, an adaptive learning mechanism is introduced to dynamically adjust the degree of fusion of the two embeddings by learning weights. In this way, the final graph-level embedding representation Z not only comprehensively considers the structural information of the graph, but also effectively preserves the position information of the nodes.
[0255] H END =β1H SE +(1-β1)H PE (18)
[0256] Z=Readout(H End )(19)
[0257] Among them, Z represents the fused graph embedding representation, and Readout(.) represents the averaging operation, which is a hyperparameter.
[0258] The final fused graph-level embedding representation Z is used for downstream graph classification tasks. This embedding is classified through a fully connected layer:
[0259]
[0260] Among them, F is the number of categories, y i,j is the one-hot encoding of the real category, The class probabilities predicted by the model.
[0261] The information bottleneck loss and cross entropy loss are weighted and combined to obtain the final loss function:
[0262] L=β2L ce +(1-β2)L IB (twenty two)
[0263] Where β2 is a hyperparameter used to adjust the weight of the information bottleneck loss in the overall loss function. The model parameters are optimized using the backpropagation algorithm to minimize the loss function.
[0264] 3.6 Algorithm Steps
[0265] Algorithm 3.1 describes a graph classification method that combines substructure embeddings with position awareness. First, a preliminary node representation is obtained by extracting and encoding the tree, cycle, clique, and cut subgraphs of the graph. Next, the embeddings of these subgraphs are fused using an attention mechanism, and anchor nodes are selected to obtain a position-aware embedding. Finally, the substructure embeddings and position-aware embeddings are combined, optimized using a joint loss function, and a classification model is used to make the final prediction, thus completing the graph classification task.
[0266]
[0267]
[0268] Experimental results and analysis
[0269] This section selects six commonly used graph classification datasets: NCI1, NCI109, PTC, ENZYMES, IMDB-BIN, and PROTEINS. The goal is to compare SIPA with state-of-the-art graph classification methods. These experiments cover graphs of chemical molecules, biological networks, and social networks. The performance of the proposed model is comprehensively evaluated by comparing it with classic GNN methods, models based on contrastive learning, and substructure-based methods. Furthermore, ablation experiments are conducted to validate the contribution of each module, and parameter sensitivity experiments are conducted to explore the impact of hyperparameters on task performance.
[0270] 4.1 Experimental Details
[0271] A 10-fold cross-validation was performed on each dataset, and the average and standard deviation of the obtained results were used to evaluate the performance of the model. For each dataset, a set of hyperparameters was set, limiting the maximum number of epochs to 500, and early-stopping was used to search for appropriate hyperparameters. Training was stopped if the training loss of the current epoch did not decrease for 20 consecutive times, and the average classification accuracy and standard error were finally reported. The batch size of the present invention was set to 8, and Adam was used as the optimizer. All experiments were trained and evaluated on an Intel Core i7-12700 CPU with 16GB of memory. To ensure fair experimental comparison, the results of the baseline methods were directly quoted from the best performance indicators reported in the original papers, while the experimental results of the classic GCN, GAT, and GraphSAGE were obtained by local experimental reproduction. The specific implementation process is as follows: First, feature learning is performed on each node in the graph using GCN, GAT, or GraphSAGE to generate a node-level embedding representation; then, the node-level embedding is aggregated into a global graph-level representation through a readout function; finally, based on the graph-level representation, MLP and Softmax are used for classification prediction.
[0272] In this experiment, the cycle_basis function of the NetworkX library is used to extract the cycle, and the length of the cycle is set to [3, 5]. The find_cliques function of the NetworkX library is used to find all the clusters in the graph and filter out the largest cluster. The depth-first algorithm is used to extract the tree structure, and the depth of the tree is set to 3, the number of cut subgraphs b is in the range of [2, 3, 4, 5, 6], the feature dimension is 32, the number of anchor nodes k is in the range of [2, 4, 8, 16], and the number of GCN layers l is set in the range of [1, 2, 3, 4].
[0273] 4.2 Comparison of Classification Accuracy
[0274] To validate SIPA's strong performance in graph classification, this section compares it with a GNN baseline, a contrastive learning-based model, and a subgraph architecture. Table 2 reports the average accuracy and standard deviation. The baseline papers are listed in Sections 1 and 2. A "-" indicates that the original paper did not provide classification results. The best results are highlighted in bold.
[0275] Table 2 Comparison of graph classification results with different models
[0276]
[0277]
[0278] Experimental results show that compared with other methods, this method achieves the best graph classification accuracy in all five datasets. The main reasons for the improvement of SIPA's graph classification accuracy are as follows: (1) It uses two substructure extraction strategies, node-based and graph-based, to extract substructure information, which improves the local expression ability of the graph. (2) It introduces position information to capture the long-distance dependency between anchor nodes and other nodes, which improves the global expression ability. (3) It uses the graph information bottleneck mechanism to retain task-related features and suppress the interference of noise information, thereby improving the task relevance of the model. In contrast, other methods such as MSSGCL, GCP and MOSGSL fail to fully consider the influence of multiple substructures, which limits their adaptability to complex graph structures. MSSGCL and GCP mainly rely on a single subgraph extraction strategy, resulting in inferior performance to SIPA when processing different types of graph data. Although MOSGSL attempts to use multi-scale substructure learning, due to its relatively fixed design, it fails to fully adapt to diverse graph structures. Although LCC4GC performs well in processing ring and cluster structures and can effectively identify these specific substructures in the graph, it mainly relies on the compression and edge centralization of specific structures and lacks global structure perception. The design of this method focuses more on the processing of certain specific types of graphs and cannot achieve the same effect on more complex or diverse graph structures, resulting in certain deficiencies in graph diversity and global structural adaptability.
[0279] 4.3 Ablation Experiment
[0280] In order to comprehensively evaluate the contribution of each module of the proposed SIPA model to the performance, this section sets up and implements three sets of ablation experiments to verify the effectiveness of the three main modules, the impact of the anchor node sampling strategy, and the advantages and disadvantages of the four types of subgraph fusion strategies.
[0281] (1) Effectiveness Verification of the Three Main Modules: This experiment verifies the independent contribution of each module by gradually removing the three core modules of the model: SE, PE, and consistent representation learning module. To this end, SIPA has three variants: a) SIPA w / oSE removes the substructure module, b) SIPA w / oPE removes the position-aware embedding module, and c) SIPA w / oGIB removes the graph information bottleneck module. The experimental results are shown in Table 3.
[0282] Table 3 Comparison of graph classification accuracy of SIPA and its three variants
[0283]
[0284] Removing any component from the SIPA model results in a significant drop in classification performance. The magnitude of the performance drop varies across different datasets. On the ENZYMS dataset, the performance drop is approximately 3%. The largest drop, approximately 4.6%, is observed on the PTC dataset. For a single dataset, such as IMDB-B, the drop is even greater after removing the SE module, indicating that incorporating substructure information is more important for downstream classification tasks in this dataset.
[0285] (2) Anchor node sampling strategy analysis: In order to deeply analyze the impact of anchor node sampling strategy on graph classification tasks, this section conducts a comparative experiment between the anchor node sampling strategy and the random sampling strategy in the SIPA model. In the comparative experiment, the SIPA model adopts the MLP+GCN selection strategy to replace the traditional random sampling strategy. Other modules and experimental settings remain the same, such as Figure 4 shown.
[0286] Experimental results show that compared with the random sampling strategy, the MLP+GCN selection strategy significantly improves the classification performance on multiple datasets. This shows that in graph classification tasks, the selection strategy combining MLP and GCN can better capture the complex relationships between nodes and the structural information of the graph, thereby improving classification accuracy.
[0287] (3) Analysis of four types of substructure fusion strategies: In order to further study the effects of different substructure fusion strategies, the experiment compared the attention fusion strategy with the traditional summation strategy. Figure 5As shown in the figure, on most datasets, the attention fusion strategy generally achieves higher classification accuracy than the summation strategy. By introducing the attention mechanism, the model can dynamically adjust the information weight of each subgraph based on its contribution to the graph classification task, thereby accurately identifying and enhancing useful features and improving classification performance.
[0288] 4.4 Parameter Sensitivity Analysis
[0289] This section explores the impact of the model's hyperparameters on graph classification performance, and sets different hyperparameters on four datasets: IMDB-BIN, PTC, NCI1, and ENZYMES.
[0290] (1) Number of GCN layers l: Figure 6 As shown in Figure 2, when setting the number of GCN layers in SIPA, the optimal number of convolutional layers for GCN generally tends to be small. The best accuracy is achieved when l = 3 for the IMDB-BIN and ENZYMES datasets, l = 2 for the PTC dataset, and l = 4 for NCI1.
[0291] (2) Embedding dimension d: Figure 7 The impact of different embedding dimensions d on graph classification. The results show that classification performance is relatively optimal when embedding dimension 128 is used on the NCI1 and IMDB-B datasets, when embedding dimension 32 is used on the PTC dataset, and when embedding dimension 64 is used on the ENZYMES dataset. This indicates that the choice of embedding dimension is closely related to the characteristics of the dataset. The graphs in the PTC dataset are small, with a limited number of nodes and edges, and their topological characteristics are relatively simple. In this case, a lower embedding dimension d = 32 can effectively represent the structural characteristics of the graph, while reducing noise and redundant information and avoiding overfitting caused by high dimensionality.
[0292] (3) Learning rate: Figure 8 Different datasets have different sensitivities to the learning rate. The datasets PTC, NCI1, and ENZYMES have the same sensitivity to the learning rate of 10. -4 The classification performance is the best when the learning rate is equal to 10. -5 The classification performance is optimal.
[0293] (4) Number of anchor nodes k: As the hyperparameter k increases, the performance of the model gradually improves, but the performance tends to be stable after a certain critical value. In order to explore the impact of the number of anchor nodes, k is set in different intervals [2, 4, 8, 16]. Figure 9The experimental results show that as k increases, the performance of the model continues to improve until it reaches a balance point. Therefore, the present invention ultimately selected k = 4 as the optimal value in the experiment on the PTC and ENZYMES datasets, and k = 8 as the optimal value on the NCI1 and IMDB-B datasets.
[0294] (5) Number of cut subgraphs b: To explore the impact of the number of cut subgraphs b on the classification performance of the model, experiments were conducted on five benchmark datasets (IMDB, NCI1, ENZYMES, PTC, and PROTEINS). The results are as follows: Figure 10 As shown in Figure 2 . Experimental results show that the choice of b has a significant impact on model accuracy, and the optimal b value varies across different datasets. When b is small, the size of the cropped subgraph is large, preserving more global structural information, but may not fully capture local features, resulting in lower classification accuracy. As b increases, the size of the cropped subgraph gradually decreases, allowing the model to better capture local structural features and significantly improving classification accuracy. When b is too large, the size of the cropped subgraph is too small, which may lead to excessive fragmentation of local structural information and loss of important global contextual information, thereby reducing classification accuracy.
[0295] 4.5 Graph Classification Visualization
[0296] To evaluate the performance of GCN and SIPA in graph classification tasks, this section conducted experiments on three classic datasets: PTC, IMDB-BIN, and PROTEINS. The feature distributions generated by the two models were compared using t-SNE visualization. Experiment 11 shows that in the PTC dataset, the graph feature distribution is relatively dispersed, exhibiting a distinct clustered branching structure, demonstrating that GCN is able to capture local patterns in the data. In the IMDB-BIN dataset, the graph feature distribution forms a complex, entangled pattern with relatively blurred boundaries between regions, reflecting the potential limitations of GCN in handling complex graph structures. In the PROTEINS dataset, the graph feature distribution is more dense, primarily concentrated in the central region and radiating outward, indicating that GCN has a certain degree of centralization in its feature representation of protein structure graphs. Experiment 12 shows that in the PTC dataset, the graph feature distribution also exhibits a branching structure, but is more compact than that of GCN, with clearer boundaries between branches, indicating that SIPA is able to better balance local and global features. In the IMDB-BIN dataset, the graph feature distribution exhibits a clustered pattern, which is more concentrated than the GCN results, indicating that SIPA has a stronger feature integration capability when processing complex graph structures. In the PROTEINS dataset, the graph feature distribution exhibits an irregular, sheet-like distribution, which is significantly different from the radial distribution of GCN, reflecting that SIPA has a more diverse feature representation of protein structure graphs. A comprehensive comparison shows that the feature distribution generated by SIPA is more compact and has clear boundaries, especially in the PTC and IMDB-BIN datasets. In contrast, the feature distribution of GCN is more dispersed, especially showing certain limitations when processing complex graph structures. This difference in feature representation ability directly affects the classification performance of the model.
[0297] This paper proposes a graph classification method SIPA that combines substructure embedding and node position awareness. SIPA first extracts subgraphs through two different strategies to capture the diverse substructures in the graph, and uses a graph convolutional network to encode its structural features, and effectively fuses the information of different subgraphs through the attention mechanism. Then, anchor nodes are introduced to calculate the relative position information of the nodes, and this information is embedded in the node representation to better capture the global position features. Finally, the graph information bottleneck mechanism is used to optimize the node representation and remove redundant information that is irrelevant to the classification task. This method not only effectively learns local structural information, but also enhances the perception of the relative position information of nodes. Experimental results show that SIPA outperforms the existing baseline models on five datasets, verifying its superiority in graph classification tasks.
[0298] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0299] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A graph classification method based on subgraph integration and location awareness, characterized in that: The following steps are involved: Step 1: Node-based and graph-based subgraph extraction strategies are adopted to capture different substructures in the graph, and their structural features are encoded using a graph convolutional network, which is then fused using an attention mechanism. Step 2: Use an anchor-based method to calculate the relative position information of the nodes and embed it into the node representation to capture the global position feature; Step 3: Use the graph information bottleneck mechanism to optimize the node representation and remove redundant information irrelevant to the classification task.
2. The graph classification method based on subgraph integration and location awareness as claimed in claim 1, characterized in that: The graph information bottleneck: The representation h learned from the original data contains the maximum amount of information related to the predicted target y (i.e., maximizing I(h; y)) while filtering out redundant information irrelevant to the prediction task (i.e., minimizing I(x; h)). Its objective function can be formally expressed as: L IB =I(y;h)-βI(x;h) (1) where I(;) represents mutual information; KL divergence can be used to calculate mutual information [26]: Where P(y,h) is the joint probability distribution of y and h, P(y) and P(h) are the marginal probability distributions of y and h respectively; similarly, I(x;h) can be calculated.
3. The graph classification method based on subgraph integration and location awareness as claimed in claim 1, characterized in that: The subgraph: (1) Subgraph type selection and definition To cover a wide range of graph topologies, this paper selects four typical subgraph types: ring subgraph, tree subgraph, clique subgraph, and cut subgraph. These subgraph types can effectively capture the cyclic structure, hierarchical structure, dense connections, and structural information of key regions in the graph. In an undirected graph g = (V, E), the following are the definitions of the four subgraphs of graph g: 1) Ring subgraph: A ring subgraph is composed of a subset of vertices. The closed loop is formed if |V′|≥3 and there is a simple path between any two vertices u and v in V′, and the starting and ending vertices of the path are the same, forming a closed loop; 2) Cluster subgraph: A cluster subgraph is composed of a subset of vertices. The complete subgraph is such that there exists (u,v)∈U between any two vertices u and v in U. 3) Tree subgraph: t represents a tree subgraph; for any u, v∈V, there is a unique path from u to v, and there is no cycle in the tree t; 4) Cut subgraph: a cut subgraph Cut(v) b A connected subgraph is generated by removing some edges from a graph. The process of generating a cut subgraph is as follows: first, the edge betweenness centrality (EBC) of each edge in the graph is calculated, which measures the importance of the edge's connectivity in the graph. Then, starting with the edge with the largest EBC, the edges are removed one by one until the graph is split into b connected blocks. (2) Substructure coding Obtain different types of subgraphs Cs from a graph = {c1, c2, ...c m }, Ts={t1,t2,...t n }, Q s ={q1,q2,...q o }, Cut={Cut1, Cut2, ...Cut a Where Cs represents the set of rings, Ts represents the set of trees, Qs represents the set of cliques, and Cut represents the set of cut subgraphs. For all subgraph types, shared GCN is used to capture features. Taking tree subgraphs as an example, for each tree subgraph t∈Ts, node features are first learned through GCN: h t =GCN(t) (3) Among them, h t is the embedding representation of the tree substructure t learned by GCN; then, the features of all the tree substructures in the graph are added together to obtain the tree embedding representation H of graph g Ts ∈R n×d In this process, since the number of nodes in different tree subgraphs may be different, node padding is used to expand the dimension of the embedding vector to the same length as the graph when adding. Similarly, the ring embedding H is obtained. Cs ∈R n×d , cluster embedded in H Qs ∈R n×d and cut-embed H Cut ∈R n×d ; In order to integrate different types of substructure embeddings, the attention mechanism att(H Ts ,H Cs ,H Qs ,H Cut ) to learn their corresponding importance (α Ts ,α Cs ,α Qs ,α Cut ); (a Ts ,a Cs ,a Qs ,a Cut )=att(H Ts ,H Cs ,H Qs ,H Cut ) (4) Here α Ts , α Cs , α Qs , α Cut Represent the embedding H Ts ,H Cs ,H Qs ,H Cut The attention value of ; the nonlinear transformation of the embedding is performed and the attention value is obtained as follows: ω Ts =W·(H Ts ) T +b (5) Where W is the weight matrix and b is the bias vector; similarly, the embedding matrix ω of each subgraph can be obtained Ts 、ω Qs and ω Cut ; Then, the attention values are normalized using the softmax function to obtain the final weights: where α Ts The larger it is, the more important the corresponding embedding is; similarly, α Cs =softmax(ω Cs ), α Qs =softmax(ω Qs ), α Cut =softmax(ω Cut ) Then, these four embeddings are combined to get the final embedding H SE : H SE =a Ts ⊙H Ts +a Cs ⊙H Cs +a Qs ⊙H Qs +a Cut ⊙H Cut (7) Among them, α Ts ⊙H Ts Indicates that α Ts After the broadcast, Ts Bit-by-bit multiplication, H SE ∈R n×d is the fine-grained node embedding representation output by the SE module.
4. The graph classification method based on subgraph integration and location awareness as claimed in claim 1, characterized in that: The anchor point: 1) Anchor point selection By selecting some important nodes with high centrality or representativeness, the model can better understand and distinguish the global topological structure of the graph; in this section, the node selection strategy is used to select the first k nodes as anchor nodes P = {p1, p2, p3…p k }, these nodes provide a reference for the relative positions of other nodes; the node selection strategy combines GCN and MLP to evaluate the importance of nodes from multiple perspectives and retain the most representative nodes; H mlp =σ(HW mlp )(9) S=H gcn +H mlp (10) Where σ(·) is the activation function; W gcn and W mlp ∈R d×1 is the parameter matrix used for linear transformation of node features; Because larger scores in vector S correspond to more important nodes, the k nodes with the highest scores are selected as anchor nodes; 2) Location-aware computing Here, location awareness can be viewed as a form of embedding method; inspired by P-GNNs[19], this paper measures location information by calculating location similarity; specifically, the shortest path distance of a node relative to an anchor node is calculated and encoded as a location embedding s(v,p i ): Among them, P i is one of the anchor nodes, is the importance score of the anchor node p, d(v,P i ) is the shortest path distance between nodes v and pi; s(v,P i ) represents the position similarity between them; Next, a more comprehensive location-aware node embedding is generated by combining node features and location similarity; in, It is the node representation obtained by encoding the graph through GCN. represents the node embedding of node v after GCN, Represents node P i Node embedding after GCN, yes and connection; Finally, by aggregating the embeddings of different anchor nodes i , we can get the location-aware node embedding H AGG The specific formula is as follows: H AGG =Mean(F(v,P1),F(v,P2),...,F(v,P k )) (14) Where Mean(·) represents the bitwise average; thus, position embedding is added to each node, and the node embedding representation H is finally obtained. PE ∈R n×d .
5. The graph classification method based on subgraph integration and location awareness as claimed in claim 1, characterized in that: The use of graph information bottleneck mechanism to optimize node representation: Consistent representation learning uses the information bottleneck GIB mechanism as an auxiliary optimization objective. The GIB objective function can be written as: L IB =γI(H SE ;H G )+(1-γ)I(H PE ;H G )(15) Among them, I(H SE ;H G ) is to minimize H SE and H G The mutual information between PE ;H G ) is to minimize H PE and H G The mutual information between them; γ is a learnable parameter used to adjust the influence of each part; calculate I(H SE ;H G ) and I(H LE ;H G ) can be estimated using KL divergence; Similarly, for I(H PE ;H G ): Where D KL (||) represents the KL divergence between two probability distributions, P(H PE ;H G ) represents the joint probability distribution, P(H PE )、P(H G ) represent H PE and H G The marginal probability distribution of ; Joint loss function.
6. The graph classification method based on subgraph integration and location awareness as claimed in claim 5, characterized in that: The joint loss function: Embed the substructure into H SE With position-aware embedding H PE Fusion is performed to generate the final node embedding representation H End To flexibly adapt to different task requirements, an adaptive learning mechanism is introduced to dynamically adjust the degree of fusion of the two embeddings by learning weights. In this way, the final graph-level embedding representation Z not only comprehensively considers the structural information of the graph, but also effectively preserves the position information of the nodes. H END =β1H SE +(1-β1)H PE (18) Z=Readout(H End ) (19) Among them, Z represents the fused graph embedding representation, and Readout(.) represents the averaging operation, which is a hyperparameter; The final fused graph-level embedding representation Z is used for downstream graph classification tasks; the embedding is classified through a fully connected layer: Among them, F is the number of categories, y i,j is the one-hot encoding of the real category, The class probability predicted by the model; The information bottleneck loss and cross entropy loss are weighted and combined to obtain the final loss function: L=β2L ce +(1-β2)L IB (22) Among them, β2 is a hyperparameter used to adjust the weight of the information bottleneck loss in the overall loss function; the model parameters are optimized using the backpropagation algorithm to minimize the loss function.
7. A graph classification system based on subgraph integration and location awareness that implements the graph classification method based on subgraph integration and location awareness as described in any one of claims 1 to 6, characterized in that: The graph classification system based on subgraph integration and location awareness includes: A fusion module that uses node-based and graph-based subgraph extraction strategies to capture different substructures in the graph, encodes their structural features using a graph convolutional network, and then fuses them using an attention mechanism; A calculation module is used to calculate the relative position information of the nodes using an anchor-based method and embed it into the node representation to capture the global position feature; The removal module is used to optimize node representation by utilizing the graph information bottleneck mechanism and remove redundant information that is irrelevant to the classification task.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the graph classification method based on subgraph integration and location awareness as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the graph classification method based on subgraph integration and location awareness as described in any one of claims 1 to 6.
10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the graph classification system based on subgraph integration and location awareness as described in claim 7.
Citation Information
Cited By
Graph node representation method based on hyper-spherical hierarchical contrast learning
CN121145923A
Position-aware hyperspectral image classification method based on subgraph convolutional network
CN121415126A
A position-aware subgraph convolution network-based hyperspectral image classification method
CN121415126B