Graph question answering method based on text graph subgraph enhancement
Through the method based on text graph subgraph enhancement, using language model coding and subgraph optimization, the contradiction between local structure perception and global reasoning in graph question-and-answer tasks is solved, and stronger generalization capabilities and information extraction effects are achieved. It is suitable for graph question-and-answer tasks in the field of deep learning and graph neural networks.
Patent Information
- Application Number
- CN202510648762.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-26
AI Technical Summary
The existing graph neural networks have problems in the graph question-and-answer task of node-level local structure perception and global graph reasoning capabilities, resulting in insufficient generalization and expression capabilities, which makes it difficult to effectively apply in complex tasks.
Through the method based on text graph subgraph enhancement, the language model is used to encode problem text and node text, calculate the similarity and select candidate nodes, optimize the subgraph structure, generate node-level embedding representations and project them into the large language model space, and build a dual-path instruction fine-tuning mechanism for semantic understanding and generation paths to realize graph instruction adjustment and graph-to-text reconstruction.
It improves the generalization ability of graph question-and-answer tasks, takes into account node-level local structure perception and global graph inference, enhances the cross-task adaptability and information extraction capabilities of the model, and reduces the difficulty of modal alignment.
Smart Images

Figure CN120541204A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and graph neural networks, and more specifically, to a graph question answering method based on text graph subgraph enhancement. Background Art
[0002] Visual question answering (VQA) is a challenging and practical task in the field of artificial intelligence. It is a multimodal task that primarily involves an interdisciplinary research direction involving computer vision and natural language processing. Chart question answering (CQA) is a task within VQA that primarily leverages computer vision and natural language processing techniques to generate or select natural language answers that match the image content based on an input image and a natural language question.
[0003] Graph Neural Networks (GNNs) are a class of deep learning models specialized for processing graph-structured data. As a means of addressing graph question answering tasks, GNNs leverage relationships between nodes within subgraphs to effectively capture local and global patterns in graphs, thereby extracting valuable information. Consequently, GNNs have been successfully applied in various fields, including social network analysis. GNNs utilize powerful message passing and aggregation mechanisms to process graph data and have achieved exceptional performance in tasks such as graph classification. However, GNNs also have limitations. They typically require fine-tuning for specific downstream applications, which complicates multi-task scenarios and results in insufficient generalization. Furthermore, despite their impressive performance in classification tasks, current research is shifting towards addressing more complex user needs. These advanced tasks require models with not only enhanced expressive power but also the ability to effectively share and integrate information across multiple tasks. However, existing GNN designs often struggle to handle these complexities, limiting their potential for application in advanced tasks.
[0004] The prior art discloses a graph question answering method based on graph neural networks, which uses a visual graph neural network and a bidirectional long short-term memory network to extract two modal feature representations, image and text, respectively. The two modal feature representations are aligned and spliced together, and the first stage feature fusion is performed on the cross-modal feature representation to obtain a low-order cross-modal feature representation. The second stage feature fusion is performed on the low-order cross-modal feature representation to obtain a high-order cross-modal feature representation. The high-order cross-modal feature representation is input into a classifier to obtain a prediction result. Through the two-stage feature fusion, the cross-modal features are fully interacted, and the high-order semantic relationship between the graph and the question keywords is better mined. However, this solution is still limited by the inherent problems of GNNs, including insufficient generalization and expression capabilities.
[0005] Recent research is exploring how to apply Large Language Modules (LLMs) to graph-related tasks. LLMs' powerful semantic knowledge is leveraged to enhance the quality of graph-text attributes in GNNs. There are two main approaches to applying LLMs to text-graph tasks. The first is to convert text graphs into sequential text inputs. However, converting graph structures into sequential text results in a significant loss of topological information, and this descriptive approach is inconsistent with the pre-training objectives of LLMs. This mismatch leads to poor performance of LLMs on graph-related tasks. The second approach addresses this limitation by projecting structural information into LLMs using graph encoders. For example, GraphGPT and GNP ensure that LLMs can recognize topological information by training additional graph transformers and embedding them into the latent space of LLMs. While this approach better preserves structural information, it is primarily used for classic graph tasks and still lacks the ability to handle large-scale knowledge graph question answering, posing a contradiction between node-level local structure perception and global graph reasoning. Summary of the Invention
[0006] To address the problem that current methods for graph question answering cannot take into account both node-level local structure perception and global graph reasoning, the present invention proposes a graph question answering method based on text graph subgraph enhancement, which ensures that key information can be extracted when facing small-scale local subgraphs and large-scale global graphs. At the same time, it takes into account both node-level local structure perception and global graph reasoning, improves generalization ability, effectively aligns the subgraph of the text graph with the semantic space of the large language model, reduces the difficulty of modal alignment, and realizes the fusion of the large language model and graph question answering tasks.
[0007] In order to achieve the above technical effects, the technical solutions of the present invention are as follows:
[0008] A graph question answering method based on text graph subgraph enhancement, the method comprising the following steps:
[0009] S1. Obtain the question text and text graph raised by the user, wherein the text graph includes nodes and edges;
[0010] S2. Encode the question text or node text using a language model to obtain a query semantic embedding representation. Input the node text of the text graph into the language model for encoding to obtain a semantic embedding representation. Calculate the similarity between the query semantic embedding representation and the node embedding representation.
[0011] S3. Arrange the similarities in descending order, select the nodes corresponding to the first K arranged similarities, form an initial candidate node set, optimize the initial candidate node set, and output the initial optimized candidate node set;
[0012] S4. Starting from the nodes of the initial optimized candidate node set, perform topological retrieval, dynamically evaluate the marginal benefits of the nodes, optimize the original text graph, and obtain the target subgraph;
[0013] S5. Model the topological relationship of the target subgraph, generate node-level embedding representations, and project the node-level embedding representations into the text semantic space of the large language model to obtain the projection of the target subgraph in the large language model space.
[0014] S6. Construct a dual-path instruction fine-tuning mechanism for the semantic understanding path and the semantic generation path, and implement graph instruction adjustment and graph-to-text reconstruction based on the projection of the target subgraph in the large language model space.
[0015] In this technical solution, the question text and text graph raised by the user are obtained, the question text and the text of the text graph node are encoded using a language model, and the query semantic embedding representation and the semantic embedding representation are obtained. The similarity between the query semantic embedding representation and the node semantic embedding representation is calculated, and the nodes with the largest similarity are selected to form an initial candidate node set. The initial candidate node set is optimized, and the subgraph is expanded to obtain the target subgraph using the nodes of the initial optimized candidate node set as the starting point. The topological relationship of the target subgraph is modeled to generate a node-level embedding representation. The node-level embedding representation is projected into the text semantic space of the large language model, and a dual-path instruction fine-tuning mechanism for the semantic understanding path and the semantic generation path is constructed to achieve graph instruction adjustment and graph-to-text reconstruction. This solves the contradiction between node-level local structure perception and global graph reasoning, and improves generalization ability.
[0016] Preferably, the expression of the text graph is:
[0017]
[0018] in, and ε represent the node set and edge set of the text graph respectively, Represents the natural language description features of the node;
[0019] The process of calculating the similarity between the query semantic embedding representation and the node embedding representation is as follows:
[0020] Use language model to encode text graph node text into low-dimensional semantic embedding vector z i ;
[0021] The language model is used to encode the question text raised by the user or the node text that needs to be classified into the query semantic embedding z q ;
[0022] Use cosine similarity to calculate the problem embedding and the set of all node embeddings The similarity s i , the expression is:
[0023]
[0024] in, is the node set of the text graph.
[0025] Preferably, the similarity s i Arrange them in descending order, select the nodes corresponding to the similarity of the first k arrangements, and form the initial candidate node set The expression is:
[0026]
[0027] Among them, top-k() is a filtering operator, which means that from v i Filter the k nodes with the highest similarity, v i represents the i-th candidate node.
[0028] Preferably, the optimizing the initial candidate node set includes:
[0029] Use the weighted optimization algorithm to perform structured pruning on the initial candidate node set to obtain the initial optimized candidate node set after structured pruning The expression is:
[0030]
[0031] Among them, PCST represents the optimization algorithm used.
[0032] Preferably, the process of starting with the nodes of the initial optimized candidate node set, performing topological retrieval, dynamically evaluating the marginal benefits of the nodes, optimizing the original text graph, and obtaining the target subgraph is as follows:
[0033] Define the i-th candidate node v i and the optimized candidate node set The average embedding similarity of the nodes in The expression is:
[0034]
[0035] Define the i-th candidate node and the optimized candidate node set Semantic consistency The expression is:
[0036]
[0037] Define the i-th candidate node v i and the optimized candidate node set The gain function of the nodes in The expression is:
[0038]
[0039] Among them, α and β are parameters that adjust the tension between semantic focus strength and information diversity;
[0040] Arrange the gains of neighboring nodes of the boundary nodes in the optimized candidate node set from large to small, select the candidate nodes whose gains exceed the threshold δ and the candidate nodes corresponding to the gains of the first n arrangements, and add them to the candidate node set In the process, a new set of candidate nodes is formed
[0041] Prune the original text graph and only keep the original text graph structure where both endpoints are located in the new candidate node set , construct the target subgraph
[0042] Preferably, the topological relationship of the target subgraph is modeled to generate a node-level embedding representation H graph , the process is:
[0043] For the target subgraph Initialize the node features in;
[0044] Generate and update information on the initialized subgraph to obtain node embedding H graph , the expression is:
[0045]
[0046] Among them, φ1 is the parameter of the graph neural network GNN, R represents the set of real numbers;
[0047] Embed H into the node graph Perform normalization to obtain the final node embedding H graph .
[0048] Preferably, the node-level embedding representation is projected into the large language model space to obtain the projection h of the target subgraph in the large language model space. graph , the expression is:
[0049] h graph =mean(MLP φ2 (H graph ))
[0050] Among them, φ2 is the parameter of the parameterized multilayer perceptron MLP, mean() represents the average operation, H graph represents the node-level embedding representation, LLM stands for Large Language Model;
[0051] The projection h of the target subgraph in the large language model space graph Keep consistent with the text semantic space of the large language model.
[0052] Preferably, the semantic understanding path includes: constructing a description input of the target subgraph, encoding the description input into a text semantic embedding H using a large language model text , the process is:
[0053] Dynamically construct subgraph descriptions x based on different task characteristics desc , where, for the node classification task, the target node text attribute x is used q As a description of the subgraph, the expression is:
[0054] x desc =x q
[0055] For knowledge question answering tasks, the subgraph elements are parsed into a set of triples, expressed as:
[0056]
[0057] Among them, v i Represents the head node, e i Indicates tail type, x i Represents the tail node;
[0058] The query x q With the description of the subgraph x desc Splice and get the subgraph input x input , the expression is:
[0059] x input =concat(x q ,x desc )
[0060] Among them, concat() represents the splicing operation;
[0061] Encode the subgraph input x using a large language model input , generate text semantic embedding H text , the expression is:
[0062] H text =LLMEmbed(x input ).
[0063] Preferably, the semantic understanding path further includes: using the optimization target function Target GIT , based on text semantic embedding H text and the projection h of the target subgraph in the large language model space graph Adjust the graph instruction, the expression is:
[0064]
[0065] Among them, y label represents the supervision label related to the task type, φ1 is the parameter of the graph neural network, and φ2 is the parameter of the adaptive readout model.
[0066] Preferably, the semantic generation path includes: using text semantic embedding H text The reconstruction from graph to text is realized with the question text raised by the user. The process is as follows:
[0067] The user’s question text or the task instruction text is encoded using a large language model into instruction embedding H instruction ;
[0068] Using the optimization objective function Target GTR Based on instruction embedding H instruction and the projection h of the target subgraph in the large language model space graph , generate the original text description y of the target subgraph node desc , the expression is:
[0069] Target GTR =logp(y desc |h graph ,H instruction )
[0070] The original text description of the target subgraph node is used as the answer to the question.
[0071] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0072] The present invention proposes a graph question answering method based on text graph subgraph enhancement. The question text and text graph node text are encoded using a language model to obtain query semantic embedding representation and semantic embedding representation, thereby improving the accuracy of locating the target subgraph structure. The similarity between the query semantic embedding representation and the node semantic embedding representation is calculated, and nodes with larger similarity are selected to form an initial candidate node set. The initial candidate node set is optimized to further eliminate noise and strengthen the constraints on the target subgraph structure. With the nodes of the initial optimized candidate node set as the starting point, topological retrieval is performed, and the marginal benefits of the nodes are dynamically evaluated to obtain the target subgraph, effectively suppressing the branching of the graph structure. Explosion and noise accumulation ensure information completeness, computational efficiency and noise robustness; model the topological relationship of the target subgraph, generate node-level embedding representation, project the node-level embedding representation into the text semantic space of the large language model, and construct a dual-path instruction fine-tuning mechanism of the semantic understanding path and the semantic generation path to realize graph instruction adjustment and graph-to-text reconstruction; utilize the deep semantic understanding ability of LLM, and establish a graph-text cross-modal dynamic adaptation mechanism through the parameterized mapping layer to achieve the optimal balance between information compression and semantic fidelity; this method not only takes into account node-level local structure perception and global graph reasoning, but also improves generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 A schematic diagram showing a flow chart of a graph question answering method based on text graph subgraph enhancement proposed in Example 1 of the present invention;
[0074] Figure 2 This represents the process from obtaining the question text and text graph raised by the user to calculating the initial optimized candidate node set proposed in Example 2 of the present invention;
[0075] Figure 3 It represents the process of constructing the target subgraph proposed in embodiment 2 of the present invention;
[0076] Figure 4 It shows the process of the graphic instruction adjustment task and the graphic-to-text reconstruction adjustment task proposed in the second embodiment of the present invention. DETAILED DESCRIPTION
[0077] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;
[0078] In order to better illustrate this embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the actual size;
[0079] It is understandable to those skilled in the art that descriptions of certain well-known contents may be omitted in the drawings.
[0080] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0081] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting this patent;
[0082] Example 1
[0083] This embodiment proposes a graph question answering method based on text graph subgraph enhancement. The flowchart of this method is shown in Figure 1 , including the following steps:
[0084] S1. Obtain the question text and text graph raised by the user, wherein the text graph includes nodes and edges;
[0085] S2. Encode the question text or node text using a language model to obtain a query semantic embedding representation. Input the node text of the text graph into the language model for encoding to obtain a semantic embedding representation. Calculate the similarity between the query semantic embedding representation and the node embedding representation.
[0086] S3. Arrange the similarities in descending order, select the nodes corresponding to the first K arranged similarities, form an initial candidate node set, optimize the initial candidate node set, and output the initial optimized candidate node set;
[0087] S4. Starting from the nodes of the initial optimized candidate node set, perform topological retrieval, dynamically evaluate the marginal benefits of the nodes, optimize the original text graph, and obtain the target subgraph;
[0088] S5. Model the topological relationship of the target subgraph, generate node-level embedding representations, and project the node-level embedding representations into the text semantic space of the large language model to obtain the projection of the target subgraph in the large language model space.
[0089] S6. Construct a dual-path instruction fine-tuning mechanism for the semantic understanding path and the semantic generation path, and implement graph instruction adjustment and graph-to-text reconstruction based on the projection of the target subgraph in the large language model space.
[0090] The method proposed in this embodiment preserves local node details while leveraging global structures for multi-hop reasoning. At the same time, it ensures that key information can be extracted from both small-scale local subgraphs and large-scale global graphs, thereby improving cross-task generalization and reasoning capabilities. The innovative method proposed in this embodiment is based on large-scale language models, graph neural networks, and cross-modal alignment technology. By theoretically considering the upper bound of mutual information and the lower bound of projection error, it provides a new approach to resolving the contradiction between node-level local structure perception and global graph reasoning in traditional methods.
[0091] In this embodiment, the question text and the text of the text graph node are encoded using a language model to obtain the query semantic embedding representation and the semantic embedding representation, thereby improving the accuracy of locating the target subgraph structure; the similarity between the query semantic embedding representation and the node semantic embedding representation is calculated, and the nodes with greater similarity are selected to form an initial candidate node set, which is then optimized to further eliminate noise and strengthen the constraints on the target subgraph structure; starting from the nodes of the initial optimized candidate node set, a topological search is performed, and the marginal benefits of the nodes are dynamically evaluated to obtain the target subgraph, thereby effectively suppressing the branch explosion and noise accumulation of the graph structure. It ensures information completeness, computational efficiency and noise robustness; models the topological relationship of the target subgraph, generates a node-level embedding representation, projects the node-level embedding representation into the text semantic space of the large language model, and constructs a dual-path instruction fine-tuning mechanism for the semantic understanding path and the semantic generation path to achieve graph instruction adjustment and graph-to-text reconstruction; it not only utilizes the deep semantic understanding capability of LLM, but also establishes a graph-text cross-modal dynamic adaptation mechanism through a parameterized mapping layer to achieve an optimal balance between information compression and semantic fidelity; this method not only takes into account node-level local structure perception and global graph reasoning, but also improves generalization ability.
[0092] Example 2
[0093] In this embodiment, the expression of the text graph is:
[0094]
[0095] in, and Represent the node set and edge set of the text graph respectively, Represents the natural language description features of the node;
[0096] The process of calculating the similarity between the query semantic embedding representation and the node embedding representation is as follows:
[0097] Use language model to encode text graph node text into low-dimensional semantic embedding vector z i ;
[0098] The language model is used to encode the question text raised by the user or the node text that needs to be classified into the query semantic embedding z q ;
[0099] Specifically, the language model used is the SentenceBERT model;
[0100] Use cosine similarity to calculate the problem embedding and the set of all node embeddings The similarity s i , the expression is:
[0101]
[0102] in, is the node set of the text graph.
[0103] In this embodiment, the similarity s i Arrange them in descending order, select the nodes corresponding to the similarity of the first k arrangements, and form the initial candidate node set The expression is:
[0104]
[0105] Among them, top-k() is a filtering operator, which means that from v i Filter the k nodes with the highest similarity, v i represents the i-th candidate node.
[0106] In this embodiment, the optimization of the initial candidate node set includes:
[0107] Use the weighted optimization algorithm to perform structured pruning on the initial candidate node set to obtain the initial optimized candidate node set after structured pruning The expression is:
[0108]
[0109] Among them, PCST represents the optimization algorithm used.
[0110] Specifically, from obtaining the question text and text graph raised by the user to calculating the initial optimization candidate node set The process diagram is as follows Figure 2 shown.
[0111] In this embodiment, the process of performing topological retrieval starting from the nodes of the initial optimized candidate node set and dynamically evaluating the marginal benefits of the nodes to expand the subgraph to obtain the target subgraph is as follows:
[0112] Define the i-th candidate node v i and the optimized candidate node set The average embedding similarity of the nodes in The expression is:
[0113]
[0114] Define the i-th candidate node and the optimized candidate node set Semantic consistency The expression is:
[0115]
[0116] Define the i-th candidate node vi and the optimized candidate node set The gain function of the nodes in The expression is:
[0117]
[0118] Among them, α and β are parameters that adjust the tension between semantic focus strength and information diversity;
[0119] Arrange the gains of neighboring nodes of the boundary nodes in the optimized candidate node set from large to small, select the candidate nodes whose gains exceed the threshold δ and the candidate nodes corresponding to the gains of the first n arrangements, and add them to the candidate node set In the process, a new set of candidate nodes is formed
[0120] Prune the original text graph so that only the two endpoints in the original text graph structure are located in the new candidate node set , construct the target subgraph The process of constructing the target subgraph is as follows Figure 3 shown.
[0121] Specifically, the dual constraints of gain threshold and quantity truncation are used to achieve the progressive growth of subgraphs. Only high-value nodes are expanded in each iteration, effectively suppressing the branch explosion and noise accumulation of the graph structure. Finally, by keeping all endpoints located at The edges in the graph construct a compact and semantically consistent target subgraph.
[0122] In this embodiment, the topological relationship of the target subgraph is modeled to generate a node-level embedding representation H graph , the process is:
[0123] For the target subgraph Initialize the node features in;
[0124] Generate and update information on the initialized subgraph to obtain node embedding H graph , the expression is:
[0125]
[0126] Among them, φ1 is the parameter of the graph neural network GNN, R represents the set of real numbers;
[0127] Embed H into the node graph Perform normalization to obtain the final node embedding H graph .
[0128] In this embodiment, the node-level embedding representation is projected into the large language model space to obtain the projection h of the target subgraph in the large language model space. graph , the expression is:
[0129] h graph =mean(MLP φ2 (H graph ))
[0130] Among them, φ2 is the parameter of the parameterized multilayer perceptron MLP, mean() represents the average operation, H graph represents the node-level embedding representation, LLM stands for Large Language Model;
[0131] The projection h of the target subgraph in the large language model space graph Keep consistent with the text semantic space of the large language model.
[0132] In this embodiment, the semantic understanding path includes: constructing a description input of the target subgraph, encoding the description input into a text semantic embedding H using a large language model text , the process is:
[0133] Dynamically construct subgraph descriptions x based on different task characteristics desc , where, for the node classification task, the target node text attribute x is used q As a description of the subgraph, the expression is:
[0134] x desc =x q
[0135] For knowledge question answering tasks, the subgraph elements are parsed into a set of triples, expressed as:
[0136]
[0137] Among them, v i Represents the head node, e i Indicates tail type, x i Represents the tail node;
[0138] The query x q With the description of the subgraph x desc Splice and get the subgraph input x input , the expression is:
[0139] x input =concat(x q ,x desc )
[0140] Among them, concat() represents the splicing operation;
[0141] Encode the subgraph input x using a large language model input , generate text semantic embedding H text , the expression is:
[0142] H text =LLMEmbed(x input ).
[0143] In this embodiment, the semantic understanding path also includes: using the optimization target function Target GIT , based on text semantic embedding H text and the projection h of the target subgraph in the large language model space graph Adjust the graph instruction, the expression is:
[0144]
[0145] Among them, y label represents the supervision label related to the task type, φ1 is the parameter of the graph neural network, and φ2 is the parameter of the adaptive readout model.
[0146] Specifically, the method proposed in this invention encapsulates the graph task into a query q and generates the result y=f(q) through a unified model framework. In the text graph, we locate the most relevant subgraph according to a specific query q. Its information is incorporated into the parameterized model to enhance the generation process. and ε represent the node set and edge set of the text graph respectively, and The natural language description features of the storage node. Ultimately, the output of the model can be expanded to:
[0147] Specifically, the process of the graph instruction adjustment task of the semantic understanding path and the graph-to-text reconstruction adjustment task of the semantic generation path is as follows: Figure 4 shown.
[0148] Example 3
[0149] In this embodiment, experiments are conducted on a graph question answering method based on text graph subgraph enhancement proposed in the present invention to verify the effectiveness of the method proposed in the present invention under different settings. By applying the model to unseen datasets or tasks without additional fine-tuning, its zero-sample adaptability is evaluated to understand the generalization performance of the model; the effectiveness of the subgraph retrieval strategy in extracting key structural information and assisting LLM reasoning is verified through experiments to evaluate its role in improving model performance; the effectiveness of these two mechanisms in converting graph structural information into a representation understandable to LLM is analyzed through experiments to verify its effectiveness in utilizing structural information. The method of the present invention is trained and evaluated on two benchmark datasets to ensure its effectiveness and generalization ability in different scenarios.
[0150] Specifically, the GraphQA benchmark is used to compare the performance gap between the present invention and the prior art. The comparison results are shown in Table 1.
[0151] Table 1
[0152]
[0153] The GraphQA benchmark is a benchmark that is closely aligned with the mainstream usage patterns of large language models (LLMs), where users typically initiate requests in the form of questions. It combines knowledge graphs and generative commonsense reasoning to effectively evaluate the performance of LLMs in these scenarios. The benchmark integrates three existing datasets: ExlaGraphs, SceneGraphs, and WebQSP. ExlaGraphs is a dataset designed for generative commonsense reasoning that focuses on creating explanation graphs for position prediction in debates. WebQSP is a large-scale multi-hop knowledge graph question answering dataset that contains many questions that require multi-hop reasoning to answer. SceneGraphs is a scene graph dataset used to evaluate the performance of models in scene understanding tasks.
[0154] For the WebQSP dataset, we used F1 score, Hit@1, and recall as evaluation metrics to comprehensively measure model performance. For the ExlaGraphs dataset, which focuses on commonsense reasoning, we used accuracy as the primary evaluation metric. To ensure a fair comparison, we selected the Llama-2-7b1 base model as the baseline. Furthermore, we chose Sentence-BERT as the text encoder and GraphTransformer as the graph encoder.
[0155] In this example, experiments were conducted on multiple benchmark datasets to evaluate the performance of the proposed graph question answering method based on text graph subgraph enhancement in different graph tasks such as node classification and graph question answering, in order to understand its performance in diverse task scenarios. The comparison results are shown in Table 2:
[0156] Table 2
[0157]
[0158] This paper selects four datasets: Cora, Citeseer, WikiCS, and Instagram. These datasets cover a variety of real-world scenarios, such as citation networks, social media interactions, and encyclopedia content.
[0159] In this embodiment, the zero-shot performance of the proposed method is compared with that of the large language model (LLM) under various settings. The comparison results are shown in Table 3:
[0160] Table 3
[0161]
[0162] The results show that the proposed method outperforms the fine-tuned LLM under all conditions. In particular, in cross-task scenarios, the fine-tuned LLM performs poorly when answering domain-specific questions, while the proposed method maintains strong zero-shot performance. This demonstrates that the proposed method can effectively leverage structured information from different domains. In this example, the performance improvements achieved by different strategies during LLM reasoning without any fine-tuning are compared. The comparison results are shown in Table 4:
[0163] Table 4
[0164]
[0165] In Table 4, the proposed method improves the F1 score by 20% over the baseline model. This is particularly important in question answering scenarios because it can provide users with more correct candidate entities to choose from.
[0166] In this embodiment, the performance of the existing large model ChatGPT and the present invention in processing problems involving multiple entities in a label is compared. The comparison results are shown in Table 5:
[0167] Table 5
[0168]
[0169] In this example, we remove the graph-adaptive readout mechanism and graph-to-text reconstruction (GTR) and design ablation experiments to evaluate the contribution of these two mechanisms to model performance on the WebQSP and ExlaGraphs datasets. The evaluation results are shown in Table 6. Experiments were conducted under two different fine-tuning approaches: prompt tuning and fine-tuning using LoRa. Under the prompt fine-tuning setting, the full model (G-Tuning) demonstrates strong performance and is effective for tasks involving structured data. However, removing either the graph-adaptive readout mechanism or the graph-to-text reconstruction significantly degrades performance, highlighting their critical role in capturing key information about the graph structure. Switching to the LoRa fine-tuning setting, the full model achieves even higher performance, demonstrating greater potential for improvement. Notably, even under this more robust fine-tuning approach, removing any one module still leads to performance degradation. Both the graph-adaptive readout mechanism and the graph-to-text reconstruction significantly improve performance, whether using lightweight prompt fine-tuning or more in-depth fine-tuning.
[0170] Table 6
[0171]
[0172] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A graph question answering method based on text graph subgraph enhancement, characterized in that: The following steps are involved: Obtaining a question text and a text graph raised by a user, wherein the text graph includes nodes and edges; Encode the question text or node text using the language model to obtain the query semantic embedding representation, input the node text of the text graph into the language model for encoding, obtain the semantic embedding representation, and calculate the similarity between the query semantic embedding representation and the node embedding representation; Arrange the similarities in descending order, select the nodes corresponding to the first K arranged similarities to form the initial candidate node set, optimize the initial candidate node set, and output the initial optimized candidate node set; Starting from the nodes of the initial optimized candidate node set, perform topological retrieval, dynamically evaluate the marginal benefits of nodes, optimize the original text graph, and obtain the target subgraph; Model the topological relationship of the target subgraph, generate node-level embedding representation, project the node-level embedding representation into the text semantic space of the large language model, and obtain the projection of the target subgraph in the large language model space; A dual-path instruction fine-tuning mechanism is constructed for the semantic understanding path and the semantic generation path, and graph instruction adjustment and graph-to-text reconstruction are achieved based on the projection of the target subgraph in the large language model space.
2. A graph question answering method based on text graph subgraph enhancement according to claim 1, characterized in that: The expression of the text graph is: in, and ε represent the node set and edge set of the text graph respectively, Represents the natural language description features of the node; The process of calculating the similarity between the query semantic embedding representation and the node embedding representation is as follows: Use language model to encode text graph node text into low-dimensional semantic embedding vector z i ; The language model is used to encode the question text raised by the user or the node text that needs to be classified into the query semantic embedding z q ; Use cosine similarity to calculate the problem embedding and the set of all node embeddings The similarity s i , the expression is: in, is the node set of the text graph.
3. The graph question answering method based on text graph subgraph enhancement according to claim 2 is characterized in that: The similarity s i Arrange them in descending order, select the nodes corresponding to the similarity of the first k arrangements, and form the initial candidate node set The expression is: Among them, top-k() is a filtering operator, which means that from v i Filter the k nodes with the highest similarity, v i represents the i-th candidate node.
4. A graph question answering method based on text graph subgraph enhancement according to claim 3, characterized in that: The optimizing of the initial candidate node set includes: Use the weighted optimization algorithm to perform structured pruning on the initial candidate node set to obtain the initial optimized candidate node set after structured pruning The expression is: Among them, PCST represents the optimization algorithm used.
5. The graph question answering method based on text graph subgraph enhancement according to claim 4 is characterized in that: The process of starting with the nodes of the initial optimized candidate node set, performing topological retrieval, dynamically evaluating the marginal benefits of the nodes, optimizing the original text graph, and obtaining the target subgraph is as follows: Define the i-th candidate node v i and the optimized candidate node set The average embedding similarity of the nodes in The expression is: Define the i-th candidate node and the optimized candidate node set Semantic consistency The expression is: Define the i-th candidate node v i and the optimized candidate node set The gain function of the nodes in The expression is: Among them, α and β are parameters that adjust the tension between semantic focus strength and information diversity; Arrange the gains of neighboring nodes of the boundary nodes in the optimized candidate node set from large to small, select the candidate nodes whose gains exceed the threshold δ and the candidate nodes corresponding to the gains of the first n arrangements, and add them to the candidate node set In the process, a new set of candidate nodes is formed Prune the original text graph and only keep the original text graph structure where both endpoints are located in the new candidate node set , construct the target subgraph 6. The graph question answering method based on text graph subgraph enhancement according to claim 5 is characterized in that: Model the topological relationship of the target subgraph and generate a node-level embedding representation H graph , the process is: For the target subgraph Initialize the node features in; Generate and update information on the initialized subgraph to obtain node embedding H graph , the expression is: Among them, φ1 is the parameter of the graph neural network GNN, R represents the set of real numbers; Embed H into the node graph Perform normalization to obtain the final node embedding H graph .
7. The graph question answering method based on text graph subgraph enhancement according to claim 6 is characterized in that: The node-level embedding representation is projected into the large language model space to obtain the projection h of the target subgraph in the large language model space graph , the expression is: Among them, φ2 is the parameter of the parameterized multilayer perceptron MLP, mean() represents the average operation, H graph represents the node-level embedding representation, LLM stands for Large Language Model; The projection h of the target subgraph in the large language model space graph Keep consistent with the text semantic space of the large language model.
8. The graph question answering method based on text graph subgraph enhancement according to claim 7 is characterized in that: The semantic understanding path includes: constructing a description input of the target subgraph, encoding the description input into a text semantic embedding H using a large language model text , the process is: Dynamically construct subgraph descriptions x based on different task characteristics desc , where, for the node classification task, the target node text attribute x is used q As a description of the subgraph, the expression is: x desc =x q For knowledge question answering tasks, the subgraph elements are parsed into a set of triples, expressed as: Among them, v i Represents the head node, e i Indicates tail type, x i Represents the tail node; The query x q With the description of the subgraph x desc Splice and get the subgraph input x input , the expression is: x input =concat(x q ,x desc ) Among them, concat() represents the splicing operation; Encode the subgraph input x using a large language model input , generate text semantic embedding H text , the expression is: H text =LLMEmbed(x input )。 9. The graph question answering method based on text graph subgraph enhancement according to claim 8, characterized in that: The semantic understanding path also includes: using the optimization target function Target GIT , based on text semantic embedding H text and the projection h of the target subgraph in the large language model space graph Adjust the graph instruction, the expression is: Among them, y label represents the supervision label related to the task type, φ1 is the parameter of the graph neural network, and φ2 is the parameter of the adaptive readout model.
10. The graph question answering method based on text graph subgraph enhancement according to claim 9, characterized in that: The semantic generation path includes: using text semantic embedding H text The reconstruction from graph to text is realized with the question text raised by the user. The process is as follows: The user’s question text or the task instruction text is encoded using a large language model into instruction embedding H instruction ; Using the optimization objective function Target GTR Based on instruction embedding H instruction and the projection h of the target subgraph in the large language model space graph , generate the original text description y of the target subgraph node desc , the expression is: Target GTR =logp(y desc |h graph ,H instruction ) The original text description of the target subgraph node is used as the answer to the question.
Citation Information
Cited By
Node embedding method based on knowledge graph, electronic equipment and storage medium
CN120875000A
Large model general graph data mining method based on compact graph description and instruction tuning
CN121210723A