Method and system for enhancing graph neural network capability and generalization by utilizing graph retrieval

By constructing resource graphs and implementing inverse importance sampling strategies to generate resource subgraph collections, searching the query graphs of user problems and disseminating knowledge, the problem of insufficient generalization ability of graph neural networks between different modes, fields and tasks is solved, and the ability and generalization of graph neural networks is greatly improved.

CN120197645APending Publication Date: 2025-06-24PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145665.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-13
Filing Date
2025-02-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Graph neural networks lack generalization capabilities between different modes, domains and tasks, and it is difficult to effectively utilize the retrieved context information in a dynamically changing environment, and lack a prompt mechanism without adjustment to support the retrieved knowledge.

Method used

A method is proposed to enhance the ability and generalization of graph neural networks by using graph retrieval. By constructing resource graphs and implementing inverse importance sampling strategies and self-center graph addition strategies, a resource sub-graph collection is generated, query graph retrieval of user problems, disseminate knowledge, and fine-tuning of pre-trained graph neural networks.

Benefits of technology

Significantly improves the capabilities and generalization of graph neural networks, can show superiority across tasks and across datasets, and achieve considerable performance without additional fine-tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197645A_ABST
    Figure CN120197645A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for enhancing graph neural network capability and generalization by utilizing graph retrieval, and belongs to the technical field of graph neural networks. The method comprises the following steps: constructing a resource graph, and executing an inverse importance sampling strategy and a self-center graph adding strategy based on the resource graph to generate a resource sub-graph set; generating a query graph of the user question, and performing retrieval in the resource sub-graph set based on the query graph to obtain topK resource sub-graphs which are most matched with the query graph; and spreading knowledge of the topK resource sub-graphs which are most matched with the query graph to a central node of the query graph, and performing fine tuning on a pre-trained graph neural network based on the query graph after knowledge injection and the tag of the user question. According to the method, the capability and generalization of the graph neural network can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of graph neural networks, and relates to a method and system for enhancing the capabilities and generalization of graph neural networks by graph retrieval. Background Art

[0002] Graph neural networks (GNNs) have recently attracted extensive attention in both academia and industry, thanks to their powerful capabilities in modeling complex, real-world data, including social, biochemical, and traffic-related domains. GNNs utilize message-passing mechanisms, going beyond traditional node embedding methods, enabling them to capture the intricate relationships within data through complex architectures and advanced graph representation learning techniques. However, the generalization ability of GNNs across different modalities, domains, and tasks remains a less explored challenge. This stands in stark contrast to the remarkable success of GPTs in natural language processing and Sora in computer vision, presenting a crucial new frontier for further research and domain development in graph data generalization.

[0003] In graph learning tasks, providing necessary context is crucial for graph generalization, such as retrieving shopping contexts similar to those shown in the graph. Therefore, the generalization ability and prediction accuracy of the model are enhanced by retrieving necessary context during the graph learning process. Retrieval-augmented generation (RAG) represents a prominent methodology that significantly enhances the functionality of language models by integrating a dynamic retrieval mechanism during the generation process. RAG not only enriches accurate and reliable content but also reduces factual errors, addressing challenges such as incorrect answers, hallucinations, and limited interpretability in knowledge-intensive tasks, thus eliminating the need to update model parameters and enabling generalization even in unseen scenarios.

[0004] However, how to enable retrieval-augmented generation for graph learning, i.e., retrieving a user's historical purchase behavior to enhance recommendation capabilities and identifying fraud crimes by searching for similar fraud-related behavior, remains unexplored and faces the following challenges C1 and C2.

[0005] C1. The first challenge is how to utilize the retrieved context, namely features (X) and labels (Y), in the GNNs model in a dynamically changing scenario. Previous studies such as PRODIGY adopted the concept of in-context learning (ICL) by constructing consistent and static task graphs for each specific task or dataset. These task graphs determine labels by calculating similarities using hidden vectors and adopt few-shot learning methods. However, PRODIGY's reliance on a fixed set of examples as rules may be insufficient to address and generalize various scenarios encountered in the real-world environment, especially in a dynamically changing environment. This is the problem because the system mainly focuses on teaching the direct mapping paradigm from input to output (X→Y), rather than truly integrating the input (X) and output (Y) data into the analysis. Compared with RAG, PRODIGY has difficulty incorporating external information (X and Y) related to data nodes, which is crucial for enriching the learning process of graph-based systems.

[0006] C2. Additionally, developing a non-tuning prompt mechanism to support the retrieved knowledge and applicable to seamlessly switching between unseen scenarios and multi-tasks is challenging. Many initiatives have been carried out in the field of graph pre-training. However, the challenge of designing a plug-and-play RAG module that seamlessly interfaces with the pre-trained model still exists. According to previous insights from investigations on graph prompts, the knowledge obtained by RAG can be conveniently injected into the prompts in a plug-and-play manner. Summary of the Invention

[0007] To strive to address the above two challenges, the present invention proposes a method for enhancing the capabilities and generalization of graph neural networks using graph retrieval. Through the general retrieval-enhanced graph learning framework (RAGraph) proposed by the present invention, the capabilities and generalization of graph neural networks can be effectively improved.

[0008] To achieve the above object, the technical solution of the present invention includes the following content.

[0009] A method for enhancing the capabilities and generalization of graph neural networks using graph retrieval, the method comprising:

[0010] Construct a resource graph, and perform an inverse importance sampling strategy and a self-centered graph augmentation strategy based on the resource graph to generate a set of resource subgraphs;

[0011] Generate a query graph for the user's question, and retrieve in the set of resource subgraphs based on the query graph to obtain the topK resource subgraphs that best match the query graph;

[0012] Propagate the knowledge of the topK resource subgraphs that best match the query graph to the central node of the query graph, and fine-tune the pre-trained graph neural network based on the query graph after knowledge injection and the label of the user's question.

[0013] Further, an inverse importance sampling strategy and a self-centered graph augmentation strategy are executed based on the resource graph to generate a set of resource subgraphs, including:

[0014] Executing an inverse importance sampling strategy based on the resource graph to generate a number of self-centered graphs;

[0015] Executing a self-centered graph augmentation strategy for each self-centered graph to obtain a corresponding resource subgraph.

[0016] Further, executing an inverse importance sampling strategy based on the resource graph to generate a number of self-centered graphs, including:

[0017] Combining a web centrality algorithm and degree centrality to calculate the importance of each node in the resource graph;

[0018] Inverting and normalizing the importance of nodes to obtain the sampling probability of each node;

[0019] Performing node sampling in the resource graph based on the sampling probability to obtain the main node of the self-centered graph;

[0020] Obtaining the k-hop neighbors of the main node of the self-centered graph in the resource graph;

[0021] Obtaining a self-centered graph according to the main node of the self-centered graph and its k-hop neighbors.

[0022] Further, executing a self-centered graph augmentation strategy for each self-centered graph to obtain a corresponding resource subgraph, including:

[0023] Calculating the average inverted importance of nodes in the self-centered graph;

[0024] Obtaining the number of augmentations according to the average inverted importance;

[0025] Augmenting the self-centered graph at least once according to the number of augmentations to obtain the resource subgraph corresponding to the self-centered graph; wherein, the augmentation techniques used in each augmentation include: node deletion, adding Gaussian noise, node interpolation, and edge rewriting. Node deletion randomly discards some nodes according to the sampling probability of nodes. Adding Gaussian noise adds Gaussian noise to node features. Node interpolation creates new nodes by linearly combining the features of two nodes and updates the weights of edges accordingly. Edge rewriting reconnects edges according to the average sampling probability of the involved nodes.

[0026] Further, retrieving in the set of resource subgraphs based on the query graph to obtain the top K resource subgraphs that best match the query graph, including:

[0027] Convert the resource sub-graph into key-value pairs; where the key information includes: the historical information, environmental information, structural encoding, and hidden embedding of the main node in the resource sub-graph, and the value information includes: the task-specific output vector and hidden embedding;

[0028] Obtain the information of the central node in the query graph; where the information of the central node in the query graph includes: the historical information, environmental information, structural encoding, and hidden embedding of the central node in the query graph;

[0029] By calculating the similarity between the information of the central node in the query graph and each key information, obtain the top K resource sub-graphs that best match the query graph.

[0030] Further, calculating the similarity between the information of the central node in the query graph and each key information includes:

[0031] Respectively obtain the historical temporal feature t(v m , m , time , environment ) of the historical information of the main node v in the resource sub-graph and the historical temporal feature t(v c ) of the historical information of the central node v in the query graph, and input the historical temporal feature t(v m ) and the historical temporal feature t(v m ) into the time similarity metric model to obtain the time similarity S c ) between the main node v m ) and the central node v c m ; where the time similarity metric model is established based on the exponential decay function; time

[0032] Based on the environmental information of the main node v c in the resource sub-graph and the environmental information of the central node in the query graph, calculate the intersection and union ratio of the neighbor sets in the query graph and the resource sub-graph to obtain the environmental similarity S c between the main node v m and the central node v environment c ;

[0033] Respectively, for the structural encoding of the main node v c and the structural encoding of the central node v m , calculate the position-aware encoding s c of the main node v c and the position-aware encoding s m of the central node v m , and based on the cosine similarity of the position-aware encoding s c and the position-aware encoding s m , obtain the structural similarity S structure ;

[0034] Calculate the hidden embeddings of the computing resource subgraph and the query graph, and obtain the semantic feature similarity S based on the cosine distance between the hidden embeddings. semantic ;

[0035] Based on the time similarity S time , the environment similarity S environment , the structure similarity S structure and the semantic feature similarity S semantic , obtain the similarity between the information of the central node in the query graph and the key information.

[0036] Furthermore, based on the structure encoding of the main node v c in the resource subgraph, calculate the position-aware encoding s c of the main node v c , including:

[0037] Randomly select several nodes in the resource subgraph to obtain the set of anchor nodes VS;

[0038] Based on the hop distance between nodes, calculate the distance similarity dis(v c and each anchor node v w ), v c , v w ) where v w ∈ VS;

[0039] Define the maximum effective hop number hyperparameter disq;

[0040] When the hop distance between the main node v c and the anchor node v w is less than the maximum effective hop number hyperparameter disq, calculate the normalized distance centroid feature d2c between the main node v c and the anchor node v w based on the distance similarity dis(v c , v w ); when the hop distance between the main node v c and the anchor node v w is greater than the maximum effective hop number hyperparameter disq, set the normalized distance centroid feature d2c between the main node v c and the anchor node v w to zero;

[0041] Based on the normalized distance centroid feature d2c between the main node v c and all anchor nodes v w , obtain the position-aware encoding s c of the main node v c .

[0042] Further, propagating the knowledge of the resource subgraphs of the top K most matching query graphs to the central node of the query graph includes:

[0043] Aggregating the task-specific output vectors and hidden embeddings of neighbor nodes in the resource subgraph to the main node of the resource subgraph through a pre-trained GNN network to obtain the aggregated task-specific output vectors and hidden embeddings;

[0044] Aggregating the aggregated task-specific output vectors and hidden embeddings in the resource subgraphs of the top K most matching query graphs to the central node of the query graph through a pre-trained GNN network.

[0045] Further, inferring the user question in the test set based on the fine-tuned graph neural network to obtain the answer to the user question.

[0046] A system for enhancing the ability and generalization of a graph neural network using graph retrieval, the system includes:

[0047] A module for constructing a retrieval resource subgraph, configured to construct a resource graph and generate a set of resource subgraphs based on the resource graph by performing an inverse importance sampling strategy and a self-centered graph augmentation strategy;

[0048] A resource subgraph query and retrieval module, configured to generate a query graph for a user question and retrieve in the set of resource subgraphs based on the query graph to obtain the resource subgraphs of the top K most matching query graphs;

[0049] A model training module, configured to propagate the knowledge of the resource subgraphs of the top K most matching query graphs to the central node of the query graph, and fine-tune a pre-trained graph neural network based on the query graph after knowledge injection and the label of the user question.

[0050] Compared with the prior art, the present invention has at least the following beneficial effects.

[0051] · The framework RAGraph proposed by the present invention is the first to integrate RAG with pre-trained GNNs. By constructing a key-value vector library of the resource graph, RAGraph realizes explicit plug-and-play access to pre-trained GNNs, and achieves considerable performance even without fine-tuning, demonstrating its superiority in cross-task and cross-dataset capabilities.

[0052] ● The RAGraph of the present invention adopts a classical message passing mechanism and introduces a well-designed prompt mechanism to integrate knowledge. This method effectively integrates the knowledge X and Y retrieved from the resource graph into the pre-trained GNNs model, enhancing the accuracy and relevance of the model output.

[0053] ● The present invention has been extensively tested on static and dynamic graphs for a variety of graph tasks (nodes, edges, and graphs). The results verify the effectiveness of the model of the present invention, showing significant improvements over the existing state-of-the-art baseline models in both fine-tuning and non-fine-tuning scenarios, especially in cross-dataset validation. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Frame diagram of a method for enhancing the capabilities and generalization of graph neural networks using graph retrieval.

[0055] Figure 2 On the ENZYME and PROTEIN datasets, the system performance varies with the hyperparameters k and topK. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] Retrieval-augmented generation has been well applied to large language models. Specifically, RAG integrates an external knowledge retrieval component and prompt engineering into a pre-trained language model to enhance factual consistency, thereby improving the reliability and interpretability of LLM responses. Traditional RAG methods use a retriever model to retrieve relevant documents from a wide knowledge base, which are then further processed by a reader model - mainly an LLM. In addition, some research focuses on fine-tuning the reader LLM by applying prompt tuning of retrieved knowledge or using RAG API calls. Although RAG has achieved considerable success in the field of natural language processing, it has also been applied to tasks involving joint visual and text retrieval, code retrieval, audio retrieval, and video retrieval. Although there are examples of applying RAG to structured data such as knowledge graphs, these mainly utilize the text information of knowledge graph nodes to enhance language or graph models. In contrast, there has been no significant research on using RAG on structured graphs without text information to enhance pre-trained GNNs. The present invention aims to similarly extend this successful method to graph data, enhance the capabilities of pre-trained GNNs, and can be adapted to various tasks and different graphs by integrating a plug-and-play RAG module without additional fine-tuning.

[0058] In addition, inspired by the application of pre-trained models and prompt learning in natural language processing, recently, learning on graphs has been divided into pre-training models on large-scale graph data, with or without labels, and then fine-tuning the model parameters through prompts to adapt to diverse downstream tasks. Adopting the prompt mechanism in graph learning represents a promising way to overcome the limitations of traditional graph representation methods, achieving a balance between flexibility and expressiveness. For example, VNT uses virtual nodes as prompts to optimize the application of pre-trained graph models. GraphPrompt introduces task-specific readout mechanisms to customize models for various tasks, while GraphPro implements spatial and temporal-based gating mechanisms applicable to dynamic recommendation systems. In addition, PRODIGY constructs task graphs (prompts) and data graphs to enhance the ICL ability of the model. Drawing on the success of graph prompt learning, the present invention aims to inject retrieved knowledge into pre-trained GNNs through prompts to support downstream tasks.

[0059] The present invention draws inspiration from the success of RAG on LLMs and the application of ICL on GNNs. By slicing from a resource graph, a resource graph vector library is constructed, where the keys of the library store key information, including environmental, historical, structural, and semantic details, while the node features and label information (task-specific output vectors) are stored as values. For downstream tasks, the key values of the query nodes will be used to retrieve the resource graph through key similarity, and the stored features (X) and labels (Y) will be structurally aggregated to provide the necessary knowledge to the query nodes, rather than the mapping paradigm, to address challenge C1 in the background art. In the design of the prompt mechanism, the present invention first passes the features and task-specific outputs from the resource graph to its main node (the central node of the resource graph) through message passing. Subsequently, the features aggregated from the neighbors of the main node and the query node, as well as the task-specific output of the main node, are aggregated to the query node. This process can be parameter-free, indicating that the model of the present invention can be applied across different tasks and datasets without fine-tuning for downstream tasks, effectively addressing challenge C2 in the background art.

[0060] The proposed General Retrieval-Augmented Graph Learning framework (RAGraph) of the present invention can operate on any graph, whether additional fine-tuning is required or not. Initially, the present invention elucidates the method for constructing a resource graph. Subsequently, the present invention details the resource graph retrieval process. Finally, the training and inference processes are introduced in detail. These processes utilize the resource graphs retrieved from two propagation perspectives - internal propagation and cross-propagation - and handle two types of information: hidden embeddings and task-specific output vectors. These two techniques are the noise-trainable method and the parameter-free method respectively. It can be seen that in RAGraph, the present invention focuses on multi-level graph tasks. For consistency, the present invention defines the graph as a dynamic graph and regards the static graph as a special case within this framework. The following definition provides a detailed description of the resource graph, including the definitions of keys and values used in RAGraph. In addition, inspired by GraphPrompt, the present invention unifies node-level, edge-level, and graph-level tasks into a coherent framework and precisely processes downstream tasks using query graphs. Among them, the node-level task is to classify each node on the graph, such as judging the possibility of each node in a financial transaction network being financial fraud; the graph-level task is to classify the entire graph, such as predicting the properties of a protein given the graph structure of the protein in a protein prediction task; the edge-level task is to predict whether an edge will be formed between two nodes on the graph, such as predicting whether there will be a purchase behavior between a user and a commodity in a commodity trading network.

[0061] Figure 1 It is a framework diagram of a method for enhancing the capabilities and generalization of graph neural networks using graph retrieval, which includes three parts: constructing a retrieval resource subgraph, querying and retrieving the resource subgraph, and model training and inference.

[0062] I. Constructing a retrieval resource subgraph.

[0063] In graph-based learning, nodes with higher connectivity - usually nodes with higher degrees - tend to be more important, meaning that their information has been more widely learned during graph pre-training. On the contrary, those less important nodes - nodes in the long tail - are often overlooked. This problem is particularly evident in large language models implementing RAG because the prevalence of common knowledge obscures the long-tail knowledge that RAG aims to utilize. To address this problem, the present invention uses an inverse importance sampling strategy to construct a resource graph, offsetting this bias by preferentially sampling and enhancing resource graphs that emphasize long-tail knowledge.

[0064] In one embodiment, the process of constructing a retrieval resource subgraph includes the following steps 1.1 to 1.4.

[0065] Step 1.1: Construct a resource graph.

[0066] For more accurate reasoning, the present invention proposes that the resource graph is a resource graph associated with subsequent reasoning tasks. However, due to the good transferability of the general retrieval enhanced graph learning framework (RAGraph) of the present invention, resources in related fields can also be used to construct this resource graph.

[0067] Step 1.2: Execute the inverse importance sampling strategy based on the resource graph to generate a number of self-centered graphs.

[0068] The inverse importance sampling strategy calculates the importance of each node in the resource graph by combining the PageRank algorithm (web page centrality algorithm) and degree centrality. In order to preferentially sample nodes with lower importance, the present invention reverses and normalizes the importance of the nodes, thereby obtaining the sampling probability of each node. During the sampling process, for each main node, the present invention generates its k-hop neighbors to form a self-centered graph. In order to enhance the representativeness and diversity of the generated graph, the present invention adopts data augmentation techniques commonly used in contrastive learning.

[0069] Step 1.3: Execute the resource subgraph enhancement strategy for each self-centered graph to obtain the corresponding resource subgraph.

[0070] The self-centered graph enhancement strategy aims to enhance the diversity and representativeness of the self-centered graph through various techniques. First, the present invention calculates the average reversed importance of the nodes in the self-centered graph, which determines the number of enhancements. The number of enhancements is adjusted by a proportional constant. For the nodes in the graph, the present invention adopts the following enhancement techniques in each enhancement:

[0071] 1. Node deletion: Randomly delete some nodes according to the sampling probability of the nodes.

[0072] 2. Adding Gaussian noise: Add Gaussian noise to the node features to increase the diversity of the data.

[0073] 3. Node interpolation: Create new nodes by linearly combining the features of two existing nodes and update the edge weights accordingly.

[0074] 4. Edge rewriting: Reconnect the edges according to the average sampling probability of the involved nodes.

[0075] These techniques work together to enhance the complexity and diversity of the generated graph.

[0076] Step 1.4: Key-value pair construction.

[0077] After the sampling and enhancement processes are completed, the generated resource subgraphs are converted into key-value pairs to construct a key-value resource graph vector database for storage. Specifically, the present invention collects the historical information (including timestamps), environmental information (such as neighbor nodes), structural encoding of each master node, and the hidden embedding obtained by processing the ego-centric graph through a frozen prediction neural network, and stores this information as keys at the master node positions of the ego-centric graph. In addition, the present invention stores the task-specific output vector and the hidden embedding as values at each node. To efficiently store and retrieve these key-value pairs, the present invention uses the FASS vector library.

[0078] II. Resource Subgraph Query Retrieval.

[0079] After constructing the key-value resource graph vector database, the present invention performs a retrieval process for subtasks based on four types of sub-similarities between the master node v m in the resource graph and the central node v c in the query graph. The final similarity score is a weighted combination of these factors, and the topK resource graphs are selected as the retrieval results:

[0080] S(v c , v m ) = w × [S time (v c , v m ), S structure (v c , v m ), S environment (v c , v m )S semantic (v c , v m )] T ,

[0081] where w = [w1, w2, w3, w4] are hyperparameterized weights corresponding to temporal, structural, environmental, and semantic similarities respectively. Using this combined similarity, the present invention sorts and retrieves the topK ego-centric graphs to obtain a subset of ego-centric graphs that best match the query based on the combined criteria

[0082] For the temporal similarity S time , the present invention establishes a temporal similarity metric model using an exponential decay function: setting the decay coefficient η > 0, the temporal similarity c between node v m and node v This function can smoothly handle the impact of temporal differences on similarity. Among them, t(v c ) represents the historical temporal characteristics of the historical information of the master node in the resource subgraph, t(v m) represents the historical temporal features of querying the historical information of the central node in the graph.

[0083] For the environmental similarity S environment , the present invention calculates the intersection - union ratio of the neighbor sets in the query graph and the resource sub - graph based on the Jaccard similarity, and takes the obtained intersection - union ratio as the environmental similarity S environment =|N(v c )∩N(v m )| / |N(v c )∪N(v m )|. Wherein, N(v c ) represents the nodes in the resource sub - graph, and N(v m ) represents the nodes in the query graph.

[0084] For the semantic feature similarity S semantic , the hidden embeddings of the resource sub - graph and the query graph are extracted by a pre - trained graph neural network, and the semantic feature similarity S is obtained by using the cosine similarity semantic .

[0085] For the structural similarity S structure , the present invention first obtains the position - aware encodings of the central node in the query graph and the main node in the resource sub - graph, and then directly applies the cosine similarity of the position - aware encodings to obtain the structural similarity S structure .

[0086] In one embodiment, the position - aware encoding first gives a randomly selected set of anchor nodes where V is the set of nodes in the query graph or the resource sub - graph. The present invention calculates the hop - count distance between nodes as the minimum path length. For any node v u ∈V and anchor node v w ∈VS, define its distance similarity as dis(v u ,v w ). By introducing the maximum effective hop - count hyper - parameter disq, a normalized distance centroid feature d2c is constructed: when the actual hop - count is less than disq, d2c(v u ,v w ) = 1 / (dis(v u ,v w ) + 1); when it exceeds the threshold, it takes a zero value. This design converts the structural features in the non - Euclidean space into an Euclidean space expression, while reducing the matrix dimension and retaining rich topological information.

[0087] Node v uThe structural features are constructed by aggregating the distance centroid features with all anchor nodes, forming a structural matrix S of dimension n×∣VS∣. According to the theoretical framework of P-GNNs, the scale of the anchor node set is set to the order of log2(n). This method replaces the traditional identity coding with distance coding, significantly improving the expression ability of structural features while ensuring computational efficiency.

[0088] III. Model Training and Inference.

[0089] The model training of the present invention includes two parts: knowledge propagation and model fine-tuning.

[0090] 3.1: Knowledge Propagation.

[0091] After retrieving the topK ego-centric graphs, the knowledge of the task-specific output vector O and the hidden embedding H will be propagated from these ego-centric graphs to the main node (intra-ego-centric graph propagation), and then to the central node v in the query graph c (inter-query-ego-centric graph propagation). This propagation is carried out using the message passing mechanism of graph neural networks (GNNs). The main node v in each ego-centric graph m is connected according to the similarity score S(v c ,v c ,v m ) with the central node v in the query graph. The connection weight determines the influence of each ego-centric graph, ensuring that graphs with higher similarity have greater influence. This process can be implemented using parameter-free or learnable methods. For the learnable method, the parameters of the GNN are different.

[0092] Intra-resource graph propagation: Within each resource subgraph, the information z is propagated from the neighbors to the main node through a pre-trained GNN. The task-specific output vector O and the hidden embedding H of the neighbors are aggregated and transmitted to the main node.

[0093] Resource graph - query graph propagation: Next, the information will be aggregated from the resource subgraphs to the query graph. Specifically, during the propagation process, the information z of the neighbors and the main node in the resource subgraphs is propagated to the central node through the same pre-trained GNN. For the central node v Q in the query graph G c , the GNN aggregates the hidden embedding H of the main node v m retrieved from the ego-centric graphs and propagates it to the central node of the query graph; when propagating the task-specific output vector O, only the information of the main node is passed to the central node.

[0094] In the case where the propagation mechanism is learnable, the attention mechanism can be adjusted on the edges. In the parameter-free case - without learnable weights - the attention on the edges is determined based on the edge weights of the previous resource graphs.

[0095] 3.2: Model Fine-tuning.

[0096] At the knowledge fusion layer, the aggregated hidden embedding H of the central node v c is processed by the decoder of the pre-trained graph neural network (GNN) to obtain the output vector O. Then, this output vector is combined with the aggregated task-specific output vector in a weighted manner to generate the final output for the downstream task.

[0097] For the same task, the decoder can be directly used to generate the output. For different tasks, the decoder can be masked, allowing the model to use the pre-computed embeddings without additional training. Additionally, the decoder can be fine-tuned to better meet the specific requirements of each task, providing flexibility and optimized performance. This approach ensures that the model effectively integrates and utilizes information from the resource graph and the query graph, enhancing its effectiveness in various downstream tasks through the aggregated task-specific output vector.

[0098] Among them, when fine-tuning the decoder, the present invention adopts the same prompt loss function L as the base model (such as GraphPro, GraphPrompt) prompt . However, to alleviate the common noise retrieval problem in traditional RAG - that is, retrieving highly relevant but irrelevant data - the present invention enhances the training process by adding noise data, thereby improving the robustness of the model. This method is inspired by relevant literature. Specifically, the present invention implements two types of noise integration strategies:

[0099] · Intra-self-centered graph noise: When constructing the self-centered graph, the present invention deliberately introduces some nodes that are irrelevant to the current graph to complement other enhancement techniques.

[0100] · Self-centered graph noise: Throughout the training process, the present invention not only retrieves the top K most relevant self-centered graphs but also deliberately includes some of the bottom K least relevant self-centered graphs to introduce noisy knowledge.

[0101] The integration of these noise elements aims to improve the model's ability to distinguish relevant information from irrelevant information, significantly enhancing its robustness and overall performance in downstream tasks through noise training. However, during the inference phase, the present invention does not introduce noise.

[0102] 3.3: Model Inference.

[0103] The present invention enhances the functionality and generalization of graph neural networks (GNNs) by improving graph retrieval technology. The application scenarios of this technology are extremely extensive, especially having significant advantages in fields that require highly personalized and data-driven decision-making. For example, in e-commerce recommendation systems, RAGraph can quickly recommend products to users dynamically by retrieving the user's historical purchase behavior and the behavior patterns of similar users. This method not only saves the time and cost of retraining or fine-tuning the network but also improves the relevance of recommendations and user satisfaction.

[0104] In the field of healthcare, especially in case diagnosis and drug discovery, RAGraph can utilize its powerful graph similarity retrieval ability to identify historical cases or drug molecules similar to the current case, thereby assisting doctors and researchers in making more accurate decisions in the case of scarce labeled data.

[0105] In the financial field, this technology can be applied to prevent fraud activities. By analyzing past fraud cases and behavior patterns, it can identify potential risky activities in real time, thus protecting consumers and enterprises from fraud.

[0106] Social media platforms can also utilize RAGraph technology to provide more customized content recommendations by analyzing users' interaction data and content preferences, enhancing the user experience and increasing user stickiness and activity on the platform.

[0107] In summary, the implementation of the RAGraph patent not only improves the efficiency and effectiveness of GNNs but also realizes the innovation of data-driven decision-making in multiple industries through its advanced graph retrieval function, demonstrating broad application prospects.

[0108] A series of experiments were conducted below to evaluate the performance of the present invention against state-of-the-art baselines for three-level graph tasks on three dynamic datasets and five static datasets.

[0109] Experimental settings.

[0110] Datasets: Four static datasets (PROTEINS, COX2, ENZYMES, and BZR) were used for graph classification and node classification, and three dynamic datasets (TAOBAO, KOUBEI, and AMAZON) were used for link prediction.

[0111] Methods and baselines: Three versions of the proposed framework RAGraph of the present invention were considered: 1) RAGraph / NF, which means using the plug-and-play RAGraph without fine-tuning on the training set; 2) RAGraph / FT, which adopts prompt tuning of RAG on the training set; 3) RAGraph / NFT, which applies noisy prompt tuning of RAG on the training set. For the baselines of dynamic graphs, LightGCN, SGL, MixGCF, SimGCL, GraphPro, and GraphPro+PRODIGY were selected. For static graphs, GCN, GraphSAGE, GAT, GIN, GraphPrompt, and GraphPrompt+PRODIGY were selected as baselines.

[0112] Settings and evaluation: A training resource split was established, and the remaining data remained unseen during fine-tuning. For static graphs, the split was based on node splitting with a ratio of 50%:30%; while for dynamic graphs, the split was based on snapshot splitting, with historical snapshots as resource graphs. For fair comparison, for the methods using PRODIGY and RAGraph, the training set was used to fine-tune the model while retrieving the resource graph to prevent information leakage and overfitting; during testing, the combined training and resource graphs were retrieved. For other methods, for fairness, fine-tuning was directly performed on the combined training and resource sets. For the evaluation of static graphs, referring to GraphPrompt, pre-trained GNNs were used to perform node-level and graph-level tasks within the k-shot classification framework. For dynamic graphs, following GraphPro, pre-trained GNNs were used on a part of the large dataset, and fine-tuning and testing were performed on subsequent snapshots. In addition, unsupervised pre-training of GraphPro and GraphPrompt was performed on other datasets in similar domains to avoid information leakage. For classification tasks, accuracy was used as the evaluation metric; for link prediction tasks, the standard metrics Recall@k and nDCG@k (k = 20) were used, consistent with existing methods.

[0113] Experimental results.

[0114] The results of three graph tasks for static and dynamic graphs are shown in Tables 1 and 2. From the reported accuracies, the following observations can be made:

[0115] Beyond state-of-the-art methods: First, the present invention outperforms all baselines in almost all three graph tasks, demonstrating the effectiveness of RAGraph in knowledge transfer from pre-training to downstream tasks compared to traditional GNNs (such as GCN and GraphSAGE). It achieves the highest average accuracy in almost all tasks of ENZYMES, with at least a 5.19% improvement for static graphs and up to a 1.81‰ improvement compared to the best baseline PRODIGY / FT for dynamic graphs. It can be seen that by integrating hidden embeddings and task-specific output vectors, RAGraph can understand more knowledge than simply learning from the X→Y paradigm. Second, compared with the models of PRODIGY / NF and RAGraph / NF, the noise training in noise prompt tuning also improves the robustness of the model, avoiding the impact of a large amount of noise on information aggregation within the query graph.

[0116] Powerful retrieval-augmented performance on unseen datasets: It is observed that PRODIGY / NF and RAGraph / NF are better than Vanilla / NF, indicating that retrieving knowledge is indeed effective when testing unseen datasets. Additionally, the difference between PRODIGY / NF and PRODIGY / FT is much greater than that of RAGraph, which also shows that the simple ICL learning paradigm is insufficient, and RAGraph can achieve acceptable results without complex fine-tuning even on unseen downstream datasets.

[0117]

[0118] Table 1

[0119]

[0120] Table 2

[0121] Hyperparameter experiments.

[0122] To explore the impact of different hyperparameters on RAGraph, the varying effects of the self-centered graph hop number k (selected from the list [1, 2, 3, 4, 5]) and the number of connected self-centered graphs topK (selected from the list [1, 5, 10, 15, 30, 50]) were also analyzed separately to verify the sensitivity of the model.

[0123] Figure 2The left figure in [Figure 0] shows the relationship between the resource graph hop count k and the accuracy. It can be observed that as k increases, the amount of knowledge retrieved grows exponentially. However, excessive knowledge accumulation not only fails to improve the accuracy but also introduces more irrelevant noise, imposing a burden on GNNs. Notably, the accuracy shows a trend of first increasing and then decreasing as k increases. This pattern indicates that at lower k values, the retrieved information often includes isolated and less useful knowledge. In contrast, at higher k values, GNNs struggle to handle extensive inference chains, resulting in the utilization of complex and rich information, and its performance is even worse than that of the baseline model.

[0124] Figure 2 The right figure in [Figure 0] shows the impact of different numbers of resource graph topK on the accuracy. Similar to the previous chart, increasing topK indicates that excessive knowledge can hinder the understanding ability of GNNs. Conversely, a smaller topK leads to insufficient knowledge to improve performance on downstream tasks.

[0125] In summary, graph neural networks (GNNs) have become an important tool for interpreting relational data in various fields, but they often struggle to generalize to unseen graph data that is significantly different from the training instances. The present invention introduces external graph data into a general graph-based model to improve the model's generalization ability in unseen scenarios. At the top layer of the present invention's framework is a resource graph vector library that captures key attributes such as features and task-specific label information. During the inference process, RAGraph can cleverly retrieve similar resource graphs based on key similarities in downstream tasks and integrate the retrieved data through a message-passing hint mechanism to enrich the learning context. Extensive experimental evaluations show that RAGraph significantly outperforms state-of-the-art graph learning methods on multiple tasks (such as node classification, link prediction, and graph classification), whether on dynamic datasets or static datasets. In addition, extensive tests confirm that RAGraph continuously maintains high performance without the need for task-specific fine-tuning, highlighting its adaptability, robustness, and wide applicability.

[0126] The above implementation is only used to illustrate the technical solution of the present invention and not to limit it. Those of ordinary skill in the art can modify or equivalently replace the technical solution of the present invention without departing from the scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A method for enhancing the capability and generalization of graph neural networks using graph retrieval, characterized in that: The method comprises: Constructing a resource graph, and executing an inverse importance sampling strategy and a self-center graph addition strategy based on the resource graph to generate a resource subgraph set; Generate a query graph of the user's question, and search the resource subgraph set based on the query graph to obtain the topK resource subgraphs that best match the query graph; The knowledge of the topK resource subgraphs that best match the query graph is propagated to the central node of the query graph, and the pre-trained graph neural network is fine-tuned based on the query graph after knowledge injection and the label of the user question.

2. The method according to claim 1, characterized in that Based on the resource graph, an inverse importance sampling strategy and a self-center graph addition strategy are executed to generate a resource subgraph set, including: Executing an inverse importance sampling strategy based on the resource graph to generate a plurality of self-centered graphs; The self-centered graph addition strategy is executed on each self-centered graph to obtain the corresponding resource subgraph.

3. The method according to claim 2, characterized in that An inverse importance sampling strategy is executed based on the resource graph to generate several self-centered graphs, including: Combining the webpage centrality algorithm and degree centrality to calculate the importance of each node in the resource graph; The importance of the nodes is inverted and normalized to obtain the sampling probability of each node; Execute node sampling in the resource graph based on the sampling probability to obtain a main node of the self-centered graph; Obtaining k-hop neighbors of the main node of the self-centered graph in the resource graph; According to the main node and k-hop neighbors of the egocentric graph, the egocentric graph is obtained.

4. The method according to claim 2, characterized in that: Execute the self-centered graph addition strategy for each self-centered graph to obtain the corresponding resource subgraph, including: Calculate the average inversion importance of nodes in the autocentric graph; According to the average reversal importance, the number of enhancements is obtained; The autocentric graph is enhanced at least once according to the number of enhancements to obtain a resource subgraph corresponding to the autocentric graph; wherein the enhancement techniques used in each enhancement include: node loss, adding Gaussian noise, node interpolation and edge rewriting, wherein the node loss is to randomly discard some nodes according to the sampling probability of the nodes, the adding Gaussian noise is to add Gaussian noise to the node features, the node interpolation is to create a new node by linearly combining the features of two nodes and updating the edge weights accordingly, and the edge rewriting is to reconnect the edges according to the average sampling probability of the nodes involved.

5. The method according to claim 1, characterized in that Based on the query graph, a search is performed in the resource subgraph set to obtain topK resource subgraphs that best match the query graph, including: Convert the resource subgraph into a key-value pair; the key information includes: the history information, environment information, structure encoding and hidden embedding of the main node in the resource subgraph, and the value information includes: the task-specific output vector and hidden embedding; Acquire information of a central node in the query graph; wherein the information of the central node in the query graph includes: historical information, environmental information, structural coding and hidden embedding of the central node in the query graph; By calculating the similarity between the information of the central node in the query graph and each key information, the top K resource subgraphs that best match the query graph are obtained.

6. The method according to claim 5, characterized in that Calculating the similarity between the information of the central node and each key information in the query graph includes: Get the main node v in the resource subgraph respectively c The historical time series characteristics of historical information t(v c ) and query the central node v in the graph m The historical time series characteristics of historical information t(v m ), and the historical time series feature t(v c ) and the historical time series characteristics t(v m ) is input into the time similarity measurement model to obtain the master node v c With the central node v m The temporal similarity S time ; Wherein, the temporal similarity measurement model is established based on an exponential decay function; Based on the main node v in the resource subgraph c The environment information of the central node in the query graph and the environment information of the central node in the query graph are calculated to obtain the intersection and union ratio of the neighbor sets in the query graph and the resource subgraph to obtain the master node v c With the central node v m The environmental similarity S environment ; Master node v c The structural encoding and central node v m The structure encoding of the main node v c Position-aware coding c and the central node v m Position-aware coding m , and according to the position-aware encoding s c and position-aware codes m The cosine similarity of structure ; Calculate the hidden embedding of the resource subgraph and the hidden embedding of the query graph, and obtain the semantic feature similarity S based on the cosine distance between the hidden embeddings semantic ; Based on the time similarity S time , the environmental similarity S environment , the structural similarity S structure and the semantic feature similarity S semantic , obtain the similarity between the information of the central node in the query graph and the key information.

7. The method according to claim 6, characterized in that Based on the main node v in the resource subgraph c The structure encoding of the main node v c Position-aware coding c ,include: Randomly select several nodes in the resource subgraph to obtain the anchor node set VS; Based on the hop distance between nodes, calculate the master node v c and each anchor node v w The distance similarity dis(v c ,v w ), v w ∈VS; Define the maximum effective hop count hyperparameter disq; On the primary node v c and anchor node v w When the hop count distance between them is less than the maximum effective hop count hyperparameter disq, based on the distance similarity dis(v c ,v w ) Calculate the master node v c and anchor node v w The normalized distance between the centroid features d2c; at the primary node v c and anchor node v w When the hop distance between them is greater than the maximum effective hop hyperparameter disq, the master node v c and anchor node v w The normalized distance between the centroid features d2c is set to zero; Based on the master node v c and all anchor nodes v w The normalized distance between the centroid features d2c, get the main node v c Position-aware coding c .

8. The method according to claim 5, characterized in that Propagate the knowledge of the topK resource subgraphs that best match the query graph to the central node of the query graph, including: Aggregate the task-specific output vectors and hidden embeddings of neighbor nodes in the resource subgraph to the main node of the resource subgraph through the pre-trained GNN network to obtain the aggregated task-specific output vectors and hidden embeddings; Through the pre-trained GNN network, the aggregated task-specific output vectors and hidden embeddings in the topK resource subgraphs that best match the query graph are aggregated to the central node of the query graph.

9. The method according to any one of claims 1 to 8, characterized in that: Based on the fine-tuned graph neural network, the user questions in the test set are inferred to obtain the answers to the user questions.

10. A system for enhancing the capability and generalization of graph neural networks using graph retrieval, characterized in that: The system comprises: Construct a resource subgraph retrieval module for constructing a resource graph, and execute an inverse importance sampling strategy and a self-center graph addition strategy based on the resource graph to generate a resource subgraph set; A resource subgraph query retrieval module is used to generate a query graph for a user's question, and search the resource subgraph set based on the query graph to obtain the top K resource subgraphs that best match the query graph; The model training module is used to propagate the knowledge of the topK resource subgraphs that best match the query graph to the central node of the query graph, and fine-tune the pre-trained graph neural network based on the query graph after knowledge injection and the label of the user question.