Graph reasoning big language model construction method based on graph tool instruction learning

By designing graph information extraction and task information extraction processes in large language models, building diversified task generation rules and high-quality data sets, and fine-tuning them using low-rank adaptation methods, the problems of insufficient tool call accuracy and limited task generalization capabilities in graph inference tasks are solved, and efficient and reliable graph inference task reasoning are achieved.

CN120106136APending Publication Date: 2025-06-06UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166998.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing large language model based on tool learning faces the problems of insufficient tool call accuracy and limited task generalization capabilities in graph inference tasks, resulting in inefficient tool call efficiency, unreliable inference results, and difficult to adapt to complex graph inference tasks.

Method used

By designing a clear process of graph information extraction and task information extraction, combining regularized processing methods, standardized tool call command generation process, optimize the model's understanding of graph topological information in natural language. At the same time, a diverse graph inference task generation rules and high-quality data sets GTools were constructed, and the model was fine-tuned using the parameter-efficient low-rank adaptation method (LoRA) to develop an open source graph inference model GraphForge.

Benefits of technology

It significantly improves the accuracy of tool call and task generalization capabilities of the model in graph inference tasks, and realizes efficient and reliable complex graph inference task inference, and maintains low cost while performing performance consistent with high-cost GPT-4o-FC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005272750010000041
    Figure BDA0005272750010000041
  • Figure BDA0005272750010000051
    Figure BDA0005272750010000051
  • Figure BDA0005272750010000055
    Figure BDA0005272750010000055
Patent Text Reader

Abstract

The invention provides a graph reasoning large language model construction method based on graph tool instruction learning, and aims to comprehensively improve the tool calling accuracy and task generalization ability of a model in a graph reasoning task. According to the method, firstly, by designing a clear graph information extraction and task information extraction process and combining a regularization processing method, the understanding ability of the model on graph topology information in a natural language is optimized, so that tool calling failures and error results are remarkably reduced, and the tool calling efficiency is improved; the execution efficiency and the result accuracy of the model in the graph reasoning task can be effectively improved. And 2, by constructing diversified graph reasoning task generation rules and a high-quality graph reasoning task data set GTools, various task types and different scales of graph structures are covered, and the adaptability of the model to complex graph problems is enhanced. Meanwhile, a low-rank adaptation method with efficient parameters is adopted to carry out fine adjustment on the model, so that the model can adapt to various graph reasoning tasks and reasoning problems containing more complex graph information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large language models, and in particular to a method for constructing a graph reasoning large language model based on graph tool instruction learning. Background Art

[0002] In recent years, the rapid development of large language models (LLMs) has had a profound impact on the field of natural language processing. Through large-scale data training, these models have demonstrated powerful language generation and reasoning capabilities. However, with the complexity of application scenarios, the reasoning ability of the generative model itself can no longer meet the needs of many professional fields, especially in tasks involving complex structured data processing. To solve this problem, researchers have proposed an innovative solution - tool learning. The core idea of ​​tool learning is to decompose complex tasks into the execution process of specific tools by calling external tools or algorithms, so as to make up for the lack of ability of the generative model in specific tasks. Based on this concept, Toolformer, as a pioneering research in the field of tool learning, significantly improves the ability of language models in fields such as numerical computing and data analysis by generating tool call commands. This method provides new ideas for the functional expansion of large language models and opens up new directions for processing complex structured data (such as graph data).

[0003] Graph reasoning tasks have become an important application scenario for tool learning methods due to their unique characteristics. Graph data not only has high connectivity and non-Euclidean characteristics, but also contains rich combinatorial structures. These characteristics make graph reasoning tasks particularly dependent on algorithms. For example, classic graph algorithms such as Dijkstra's algorithm and maximum flow algorithm can efficiently solve problems such as shortest path and flow optimization, while these complex calculation processes are often difficult to achieve through direct reasoning of generative models. Therefore, tool learning methods provide an efficient solution for graph reasoning tasks. In this field, current research mainly focuses on two types of methods: tool instruction fine-tuning methods represented by Graph-ToolFormer, and function calling methods represented by GPT-4o-Function Calling. The advantages and disadvantages of these two methods are analyzed below.

[0004] The core idea of ​​the tool instruction fine-tuning method is to fine-tune a large language model so that it can generate call instructions for specific tools to solve complex tasks. Graph-ToolFormer is a typical representative of this direction. It integrates the generation of tool call commands into the model training process, enabling the model to better understand the requirements of graph reasoning tasks. For example, Graph-ToolFormer can generate instructions to call classic graph algorithms (such as Dijkstra's algorithm) to efficiently solve the single-source shortest path problem.

[0005] The advantages of this method are: strong targeting, through fine-tuning specific tasks, Graph-ToolFormer can better adapt to specific graph reasoning scenarios. For example, in tasks such as the shortest path or ring detection, the model shows high tool call accuracy; the task understanding ability is enhanced, by introducing the training of task descriptions, the model shows strong understanding ability when dealing with specific tasks, and can generate more accurate tool call instructions. However, it still has many limitations: insufficient generalization ability, because the tool instruction fine-tuning method relies on the training data of specific tasks, it performs poorly on unseen tasks (such as generalizing from the shortest path problem to the topological sorting problem). This task dependency leads to limited adaptability of the model to the diversity of graph tasks, and it is difficult to effectively cope with the generalization ability of graph tasks; the flexibility of calling is limited, and the tool call of Graph-ToolFormer often relies on fixed tool names and calling formats. When the task description or tool environment changes, the model may not be able to adjust flexibly, resulting in call failure or poor execution effect.

[0006] The function call method is a new tool integration method. Its core idea is to embed the tool description directly into the input prompt of the large language model, so that the model can dynamically generate call instructions and directly control the execution of the tool. The closed-source large model represented by GPT-4o-FunctionCalling (hereinafter referred to as GPT-4o-FC) has significantly improved the flexibility and adaptability of tool calls through this method.

[0007] Its advantages are: strong flexibility, the function call method can call external tools without additional fine-tuning of the model. This flexibility enables the model to adapt to more diverse graph reasoning tasks, especially showing strong generalization ability on unseen tasks. Adapting to diverse tools, since the function call method allows the model to dynamically generate call instructions, it does not rely on fixed tool names and formats, and can adapt to a wider range of tool environments and task requirements. However, its limitations are also prominent: limited graph data processing capabilities, graph reasoning tasks usually involve complex graph structures and multiple relationships, which places extremely high demands on the graph data understanding ability of large language models. However, the generation process of the function call method mainly depends on the model's understanding and generation capabilities of natural language, and it is difficult to directly capture the non-Euclidean characteristics, high connectivity and rich combination structures of graph data. This defect may cause the model to extract incomplete information or misunderstand when processing complex graph reasoning tasks, thereby affecting the accuracy of tool calls and the final execution effect of the task; it has high requirements on the reasoning ability of open source models. The application of the function call method in the field of graph reasoning requires the model to have the ability to accurately understand multi-level information, including task descriptions, graph structure information, and comprehensive analysis of tool documents. This capability is more prominent in closed-source large models (such as GPT-4o-FC), but for open-source models, due to the limitations of their training scale and optimization strategies, it is often difficult to meet the corresponding performance requirements. In particular, when dealing with complex graph reasoning tasks, open-source models may not be able to accurately generate tool call instructions, or there may be misunderstandings during the process of multi-information fusion, resulting in call failures or inaccurate results. This high requirement for model capabilities limits the widespread application of function call methods in the open-source ecosystem.

[0008] In summary, existing large language models based on tool learning face two major problems in graph reasoning tasks: 1) Insufficient accuracy of graph tool calls, which is manifested in the difficulty in accurately understanding and extracting graph topology information in natural language when generating tool call commands, resulting in call failures or incorrect results; 2) Limited generalization capabilities of graph tasks. Since the generation and execution of tool calls often rely on the training of the model on specific tasks, the model is difficult to adapt to unseen graph reasoning tasks or more complex graph problems. These problems significantly limit the performance of large language models in graph reasoning tasks, resulting in inefficient tool calls, unreliable reasoning results, and reducing the applicability and solution efficiency of the model in complex graph reasoning tasks.

[0009] In order to solve the above problems, a solution that can improve the performance of graph reasoning tasks is proposed by combining the respective advantages of the tool instruction fine-tuning method and the function call method. On the one hand, the accuracy of the tool instruction fine-tuning method on specific tasks and its adaptability to complex graph tasks provide an important reference for improving the model's understanding of tasks; on the other hand, the flexibility and task generalization ability of the function call method in the form of a toolset provide an effective way to solve unseen tasks. By combining the advantages of the above two methods, the goal of the present invention application is to solve the key problems of current large language models in graph reasoning tasks, improve the accuracy of tool calls, and enhance the generalization ability of the model on diversified tasks, thereby achieving efficient and highly reliable reasoning for complex graph reasoning tasks. Summary of the invention

[0010] Existing large language models face two key problems in graph reasoning tasks that need to be solved urgently. The first is the lack of accuracy in graph tool calls, which is specifically manifested in the difficulty of the model to accurately understand and extract graph topology information in natural language when generating tool call commands. This deficiency often leads to errors or non-expectations in the generated tool call commands, which in turn leads to call failures or erroneous results. The root cause of this problem is that the model has limited ability to understand natural language descriptions of complex graph structures, and there is a lack of standardized processing mechanism for the generation of tool call commands. The direct consequence is low tool call efficiency and unreliable reasoning results, which seriously affects the performance of the model in actual graph reasoning tasks. The second is the limited generalization ability of graph tasks. The training and reasoning process of existing models usually depends on data of specific tasks and lacks the ability to adapt to unseen tasks or more complex graph problems. This significantly reduces the performance of the model on new tasks and makes it difficult to generalize to unseen graph reasoning tasks. At the same time, the reasoning ability of the model is also insufficient for tasks containing complex topological structures or diverse graph information. These problems significantly limit the practical application of large language models in the field of graph reasoning and reduce their applicability and solution efficiency in complex scenarios.

[0011] In order to solve the above technical problems, the present invention proposes a method for constructing a large language model for graph reasoning based on graph tool instruction learning, aiming to comprehensively improve the tool call accuracy and task generalization ability of the model in graph reasoning tasks. The objectives of the present invention are mainly reflected in two aspects: First, by designing clear graph information extraction and task information extraction processes, combined with regularization processing methods, standardizing the tool call command generation process, and optimizing the model's understanding of graph topology information in natural language, thereby significantly reducing the occurrence of tool call failures and erroneous results. The realization of this goal can effectively improve the execution efficiency and result accuracy of the model in graph reasoning tasks. Second, by constructing diversified graph reasoning task generation rules and high-quality graph reasoning task datasets GTools, covering a variety of task types and graph structures of different scales, the model's adaptability to complex graph problems is enhanced. At the same time, a parameter-efficient low-rank adaptation method (LoRA) is used to fine-tune the model so that the model can adapt to various graph reasoning tasks and reasoning problems containing more complex graph information. To achieve the above goals, the core technologies of the present invention include the following three aspects:

[0012] 1) Graph Task Generation Module: By designing rules such as graph scale diversity, task description diversity, and answer uniqueness, it generates a wide range of graph reasoning tasks and divides complex graph reasoning task information into graph information extraction and task information extraction tasks to improve the model's task understanding ability.

[0013] 2) Tool call module: Through the design of clear tool call instructions and the standardization of model outputs using regular expressions, the model can accurately generate tool call commands and use the tool execution results to generate high-quality task labels.

[0014] 3) Model fine-tuning module: Based on the constructed high-quality dataset GTools, the parameter-efficient low-rank adaptation method (LoRA) is used to fine-tune the Llama3-8B base model, and the open source graph reasoning model GraphForge is developed, which achieves performance consistent with the function-call-based GPT-4o-FC while maintaining low costs.

[0015] The specific technical solution of the method for constructing a large language model for graph reasoning based on graph tool instruction learning of the present invention is as follows:

[0016] A method for constructing a large language model for graph reasoning based on graph tool instruction learning includes the following steps:

[0017] Step 1: Rule-based graph task generation and decomposition: Generate graph reasoning tasks based on four rules: graph scale diversity, multiple descriptions of graph reasoning tasks, balanced answers to graph reasoning tasks, and unique answers to graph reasoning tasks. Then divide graph reasoning task information into graph information extraction and task information extraction.

[0018] Step 2: The graph reasoning task obtains the output of the large language model through tool call instructions, uses regularization processing on its output, and uses the tool execution results to generate task labels;

[0019] Step 3: Define the original dataset based on the graph reasoning task, tool call instructions, and the output of the large language model. Obtain the dataset GTools by filtering through the matching function. Use the parameter-efficient low-rank adaptation method LoRA to fine-tune the Llama3-8B base model on the dataset GTools to obtain the open source graph reasoning large model GraphForge.

[0020] The four rules of graph scale diversity, multiple descriptions of graph reasoning tasks, balanced answers to graph reasoning tasks, and unique answers to graph reasoning tasks are as follows:

[0021] Graph size diversity: Graph information is organized in natural language form, and the size of the graph is reasonably constrained. The number of nodes in the graph is controlled between 2 and 40, and the number of edges is at most 300.

[0022] Multiple descriptions of graph reasoning tasks: Design more than one natural language description for each graph reasoning task;

[0023] Balanced answers for graph reasoning tasks: In tasks that require judging true or false, ensure that the answers in the generated graph data are evenly distributed;

[0024] Unique answer for graph reasoning tasks: For topological sorting tasks that may have multiple valid solutions, the uniqueness of the answer is ensured during the graph generation process.

[0025] The output of the large language model is formalized as follows:

[0026]

[0027] Among them, I (s) is the tool calling instruction, x (s) For graph reasoning tasks, including graph information extraction and task information extraction Use collection express;

[0028] The specific formalization of generating task labels using tool execution results is as follows:

[0029]

[0030] in is the output task label, Tool represents the tool, Output results for the large model after regularization.

[0031] The construction process of the dataset GTools is as follows:

[0032] Based on the graph reasoning task x (s) Its instruction I (s) And the output g of the large language model (s) Defined as the original dataset

[0033]

[0034] Decide whether to keep the output g based on the matching function M (s) :

[0035]

[0036] The matching function M is specifically:

[0037]

[0038]

[0039] where g *(s) is the standard label of each subtask, g * is the answer label, ∧ represents the union of the various information output by the large model and the standard label.

[0040] Through the above filtering process, we get the data set

[0041]

[0042] The dataset Organized in Alpaca format, the dataset GTools is obtained.

[0043] The parameter-efficient low-rank adaptation method LoRA is as follows:

[0044]

[0045] in is the actual output of the large language model, It is a graph reasoning task and instruction I (s) Combination of α controls the original weight matrix The update range, and is a low-rank matrix, r is the selected rank, less than min(d m ,d n ),in represents the field of real numbers, d m Represents the dimension of the output vector, d nRepresents the dimension of the input vector; during the learning process, only matrices A and B are updated.

[0046] The loss function of the open source graph reasoning model GraphForge fine-tuning for:

[0047]

[0048] in, represents the average loss on the entire dataset GTools, It represents the number of samples in the training data set, and cross-entropy is the cross entropy.

[0049] The cross entropy is defined as:

[0050]

[0051] Among them, j is the index of the output vector, indicating the jth category or dimension in the output vector. is the true label of the jth category, is the probability value of the jth category predicted by the model.

[0052] The beneficial effects of the present invention are as follows:

[0053] 1. This paper proposes a set of generation and evaluation rules for graph reasoning tasks to ensure the unique results of answer labels for reasoning tasks and the uniform distribution of different task sizes.

[0054] 2. Based on the graph generation rules and screening, a high-quality large model fine-tuning dataset (GTools) was constructed, and the parameter-efficient low-rank adaptation method (LoRA) method was used to fine-tune the open source large model Llama3-8B to obtain the large model GraphForge for graph reasoning tasks, which improved the graph reasoning ability of the model.

[0055] 3. The GraphForge model fine-tuned based on Llama3-8B performs well in graph reasoning tasks, which is consistent with the performance of the high-cost GPT-4o-FC, but with extremely low cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 The figure is a schematic diagram of the process of constructing the method of the present invention. DETAILED DESCRIPTION

[0057] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of a method for constructing a graph reasoning large language model based on graph tool instruction learning in conjunction with the accompanying drawings.

[0058] This embodiment proposes a method for constructing a large language model for graph reasoning based on graph tool instruction learning, aiming to improve the tool call accuracy and task generalization ability of the model in graph reasoning tasks. This method covers twenty graph-related tasks and graph structures of different types and sizes by constructing diverse graph reasoning task generation rules and high-quality data sets GTools, providing rich training and evaluation resources for the model. At the same time, the parameter-efficient low-rank adaptation method (LoRA) is used to fine-tune the Llama3-8B model, optimizing the performance of the model when processing complex graph reasoning tasks. The main contributions of the present invention include: 1) a set of graph reasoning task generation rules supporting multiple task types is proposed; 2) a high-quality graph reasoning task fine-tuning data set GTools is constructed; 3) an open source graph reasoning large model GraphForge is developed based on the parameter-efficient low-rank adaptation method (LoRA) fine-tuning technology. Experimental results show that GraphForge surpasses the existing mainstream large language models in the performance of graph reasoning tasks, and achieves performance consistent with GPT-4o-FC at a very low cost.

[0059] like Figure 1 As shown, this method is divided into three basic modules: 1) Graph task generation module: By designing rules such as graph scale diversity, task description diversity, and answer uniqueness, a wide range of graph reasoning tasks are generated, and complex tasks are decomposed into two sub-tasks: graph information extraction and task information extraction, so as to improve the model's task understanding ability. 2) Tool call module: Through the design of clear tool call instructions and the standardization of model outputs in combination with regular expressions, the model can accurately generate tool call commands and use the tool execution results to generate high-quality task labels. 3) Model fine-tuning module: Based on the constructed high-quality dataset GTools, the parameter-efficient low-rank adaptation method (LoRA) is used to fine-tune the Llama3-8B basic model, and the open source graph reasoning large model GraphForge is developed, which achieves performance consistent with GPT4o while maintaining low cost. Next, a method for constructing a large graph reasoning language model GraphForge based on graph tool instruction learning is further explained:

[0060] Step 1: Rule-based graph task generation and decomposition;

[0061] First, based on the problem that the existing graph reasoning task datasets have few categories and small scale, a graph task generation module is constructed. The goal of this module is to generate diverse and high-quality graph reasoning tasks and decompose complex tasks into subtasks to improve the model's task understanding ability. Specifically, the following steps are included:

[0062] Step 1.1: Graph rule formulation: In order to achieve diverse and high-coverage task generation, the following rules are designed:

[0063] A. Diversity of graph size: Graph information is organized in natural language. To ensure that its length does not exceed the general processing capacity limit of large language models (4096 word length), the size of the graph is reasonably constrained, that is, the number of nodes in the graph is controlled between 2 and 40, and the number of edges is at most 300. This design can not only test the processing ability of large language models for graph information, but also ensure the diversity and parsability of tasks. In the natural language description of the graph, nodes are represented by numbers, and the connection relationship between nodes is expressed in brackets. For example, a simple undirected graph can be described as: "The graph contains 3 nodes and 2 edges. The nodes are numbered 1, 2, and 3. The connections between the nodes are as follows: (1,2) and (2,3)." For a directed graph, the connection relationship will be clearly marked with the direction, for example: "The graph contains 4 nodes and 3 directed edges. The nodes are numbered 1, 2, 3, and 4. The connections between the nodes are as follows: (1,2), (2,3), and (3,4)." This natural language organization method can effectively express the structural information of the nodes and edges of the graph, while ensuring the flexibility and readability of the task description, which is convenient for parsing and processing by large language models.

[0064] B. Multiple descriptions of graph reasoning tasks: In order to evaluate the recognition ability of large language models for graph reasoning tasks, we designed five different descriptions for each graph reasoning task. These descriptions test the model's depth of understanding and generalization ability of the task by changing the way the problem is expressed. For example, for the shortest path problem, you can ask questions in different ways: "What is the shortest path from node 1 to node 2?" or "What is the fastest way to get from node 1 to node 2?" For another example, for the ring detection task, you can ask questions: "Is there a ring in the graph?" or "Does this graph contain a path from a node back to that node?" Through this diverse task description, we can comprehensively examine whether the model can accurately understand different forms of task descriptions and generate correct tool call instructions, thereby verifying its task recognition ability and robustness.

[0065] C. Balance of answers in graph reasoning tasks: In tasks that require true or false judgments (such as the output of a ring detection task is whether a ring exists or not), ensure that the distribution of answers in the generated graph data is balanced to avoid the model being biased towards a certain result, thereby ensuring the objectivity of the test. The graph generation process first needs to determine the number of nodes, generate a node set according to the preset graph scale (for example, the number of nodes is between 2 and 40), and then randomly determine the connection relationship between nodes by setting the edge generation probability (such as 0.1, 0.3, etc.), and randomly generate connection relationships between nodes to construct the graph structure. On this basis, in order to achieve the balance of answers, we adjust the generated graph data. For example, in the ring detection task, ensure that the number of graphs containing rings and those not containing rings is equal; in the path existence task, ensure that the number of graphs with paths and those without paths is balanced. Through this probability-based generation method, the randomness and diversity of the graph data are guaranteed, and the distribution of answers is effectively controlled, thereby improving the scientific nature of the model test.

[0066] D. Unique answer for graph reasoning tasks: For topological sorting tasks that may have multiple valid solutions, the uniqueness of the answer is ensured during the graph generation process. This helps compare the results of large language models with the standard answers, making the evaluation more accurate.

[0067] Step 1.2: Graph task generation:

[0068] Based on the application of the above rules, the graph task generation module can generate graph reasoning tasks with high diversity and balance. On this basis, we selected eleven different graph reasoning tasks, and further divided each task based on its directedness and undirectedness. In total, we constructed a graph reasoning dataset GTools containing twenty tasks, of which the maximum triangle sum task is only for undirected graphs, and the topological sorting task is only for directed graphs, as shown in Table 1:

[0069] Table 1. Twenty graph reasoning tasks included in the GTools dataset

[0070]

[0071]

[0072] Step 2: Tool call generation: Since the graph and task are determined, accurate answer labels can be obtained through algorithmic procedures. Using Llama3-8B as the base model, the tool call module is used to solve various graph reasoning tasks. The clear steps of reasoning are as follows Figure 1 As shown in the second part. For each task, two objectives are defined: graph information extraction and task information extraction Use collection The output g of each subtask is represented by(s) The large language model is based on the corresponding text instructions I (s) and task x (s) generate.

[0073]

[0074] Among them, the text instruction I (s) It refers to the natural language input prompts designed for each task, which are used to guide the large language model to generate task-related outputs. Specifically, I (s) is a text instruction that clearly describes the goals and requirements of the current subtask. For example, for graph information extraction Mission I (s) The instruction may be: "Please extract the information of the nodes and edges of the graph from the following text description. Nodes are represented by numbers and edges are represented by brackets." Mission I (s) The instruction may be: "Please extract the task type and related parameters, such as the start and end point numbers, based on the description and task requirements of the following figure." The input data x (s) It consists of a natural language description of the graph and the specific requirements of the task, for example: "The graph contains 5 nodes and 6 edges. The nodes are numbered 1, 2, 3, 4, and 5. The connection relationships of the edges are (1,2), (2,3), (3,4), (4,5), (5,1), and (2,4). The task is to determine whether there is a path from node 1 to node 4."

[0075] To output the text (s) Convert to a format that can be used by the tool Using regular expressions R (s) Process the output:

[0076]

[0077] This will allow you to obtain image information and task information Next, get the output label through the corresponding tool

[0078]

[0079] The corresponding tool Tool refers to an algorithm program or tool module that can perform calculations or reasoning based on the graph structure and task type. Specifically, different tasks will call different tools. For example, for the shortest path task, the tool can be a program module based on the Dijkstra algorithm, with the input being the node and edge information of the graph and the task information of the starting point and the end point, and the output being the shortest path length from the starting point to the end point or the path itself. For another example, for the ring detection task, the tool can be an algorithm module based on depth-first search (DFS), with the input being the node and edge information of the graph, and the output being a Boolean value (True or False) of whether there is a ring in the graph.

[0080] Step 3: Fine-tune the large model based on parameter-efficient low-rank adaptation methods;

[0081] Step 3.1: Further screening of the GTools dataset:

[0082] In order to ensure the high quality and reliability of the dataset, the output g (s) Strict conditions are introduced. First, a matching function M is defined to check whether all relevant tags match:

[0083]

[0084]

[0085] where g *(s) is the standard label of each subtask, g * is the answer label, and ∧ represents the union of the various information output by the large model and the standard label. (s) and its corresponding instruction I (s) And the output g of the large language model (s) Defined as the original dataset

[0086]

[0087] Decide whether to keep the output g based on the matching function M (s) :

[0088]

[0089] Through the above filtering process, we get the data set :

[0090]

[0091] Step 3.2: Fine-tuning based on parameter-efficient low-rank adaptation methods

[0092] After the comparison is completed, the data set Organized in Alpaca format, which is commonly used to fine-tune Llama to improve instruction following capabilities, a dataset GTools containing 40,000 instances was finally constructed, with 2,000 instances per task. The open source graph reasoning model GraphForge was obtained by fine-tuning the Llama3-8B base model on the GTools dataset.

[0093] In the process of fine-tuning large language models, updating model parameters is a core issue. However, since large language models usually have billions or even hundreds of billions of parameters, directly updating the model parameters will bring huge computing and storage costs. This high resource demand usually places high demands on hardware devices, limiting the application of large-scale models on ordinary computing devices.

[0094] Taking the open source large model Llama3-8B as an example, its parameters reach billions of levels, and complete fine-tuning needs to be run on a high-performance GPU, which takes up a lot of video memory and computing resources. For example, for full parameter fine-tuning, the model needs to load all weights and optimize them, which not only requires large-capacity video memory support, but also significantly increases training time. In addition, comprehensive updates of model parameters are prone to overfitting problems, especially when the scale of training data is limited.

[0095] In order to overcome the above problems, this embodiment adopts a parameter-efficient low-rank adaptation method (Low-Rank Adaptation, LoRA) in the fine-tuning process. The parameter-efficient low-rank adaptation method (LoRA) is an efficient fine-tuning method designed specifically for large models. Its core idea is to introduce only a small number of trainable parameters for update while freezing the original model parameters, thereby significantly reducing the computational and storage costs while maintaining the generalization ability of the model. Specifically, the parameter-efficient low-rank adaptation method decomposes the parameter matrix to be updated into two low-rank matrices. This decomposition method greatly reduces the amount of training parameters and reduces the GPU memory occupancy and computational burden. At the hardware level, parameter-efficient low-rank adaptation methods exhibit the following significant advantages: First, low-rank decomposition significantly reduces the demand for video memory, making it possible to efficiently fine-tune large models even in environments with limited hardware resources, such as a single consumer-grade GPU (such as the NVIDIA RTX series); second, since most of the parameters of the original model are frozen, the gradient calculation and update during training only involve the newly added low-rank parameters, thereby reducing the amount of calculation and speeding up training; finally, parameter-efficient low-rank adaptation methods only need to save the newly added low-rank parameters without re-storing the entire large model, which makes the fine-tuned model more lightweight and easy to deploy on resource-constrained devices (such as edge devices or mobile devices).

[0096] The GTools dataset used for training in the fine-tuning process can be represented as the following triples after the above filtering:

[0097]

[0098] Among them I (s) is the instruction of subtask s, For graph reasoning tasks, For I (s) Instructions for input The expected output.

[0099] Formally, for a linear layer, we use Indicates that It is a graph reasoning task and instruction I (s) A combination of is the weight matrix, where represents the field of real numbers, d m Represents the dimension of the output vector, d n Representing the dimension of the input vector, the parameter-efficient low-rank adaptation method introduces the following low-rank update:

[0100]

[0101] in is the actual output of the large language model, and is a low-rank matrix, r is the selected rank, significantly less than min(d m ,d n ), and α controls the update amplitude of the original weight matrix W. During the learning process, only matrices A and B are updated.

[0102] Defining the loss function for GraphForge fine-tuning for:

[0103]

[0104] in, Represents the entire dataset The average loss on It is composed of three The training data set consists of Represents the number of samples in the training data set. The specific definition of the cross-entropy loss function is:

[0105]

[0106] Among them, j is the index of the output vector, indicating the jth category (or dimension) in the output vector, is the true label of the jth category (usually 0 or 1), is the probability value of the jth category predicted by the model. The cross entropy measures the output distribution With the true distribution The difference between them provides the optimization direction.

[0107] During training, a learning rate of 1e-5 and a warm-up ratio of 0.1 were used. The batch size was set to 4 and a cosine learning rate scheduler was used to periodically adjust the learning rate. GraphForge was fine-tuned over 3 epochs on an 80GB NVIDIA A800 GPU.

[0108] In order to fully verify the effectiveness of the method for building a large language model for graph reasoning tasks based on tool call optimization proposed in this paper, a series of experiments were designed and implemented. These experiments were carried out from the dimensions of model performance evaluation and cross-domain task performance verification, aiming to evaluate the accuracy and generalization ability of the GraphForge model in graph reasoning tasks. The experimental results show that GraphForge has reached the most advanced level in open source models, surpassed closed source models in multiple tasks, and greatly improved the overall performance of graph reasoning tasks.

[0109] To verify the performance of the model, experiments were conducted on the following two datasets:

[0110] GTools test set: Based on the generation rules of GTools, 500 additional test instances are created for each task, covering 20 graph reasoning tasks, including loop detection, path existence, maximum flow, shortest path, etc. These test data have no overlap with the training data and are designed to evaluate the reasoning ability of the model on known tasks.

[0111] NLGraph dataset: To further verify the generalization ability of the model, we selected the following tasks from the public dataset NLGraph:

[0112] Conventional tasks: path existence, cycle detection, topological sorting, maximum flow, and shortest path. These tasks have certain similarities with GTools tasks and can verify the consistent performance of the model across data sets.

[0113] Cross-domain tasks: bipartite graph matching and GNN tasks. These tasks are beyond the scope of GTools and are designed to test the model's adaptability to unseen tasks.

[0114] The experiment uses answer accuracy as the only evaluation indicator, which is defined as the percentage of questions correctly predicted by the model in the test set to the total number of questions. This indicator can intuitively reflect the reasoning performance of the model on various tasks.

[0115] In order to verify the performance advantages of GraphForge, the following baseline models were selected for comparison:

[0116] Graph-ToolFormer, a large model based on tool instructions: Since Graph-ToolFormer has only released a fine-tuned model based on GPT-J-6B, for fair comparison, Llama3-8B is fine-tuned on the GTools dataset based on its method as the Graph-ToolFormer baseline model.

[0117] Function call-based large models GLM4-0520-FC, GPT-3.5-FC, GPT-4o-FC: These models represent the current large language models based on function call methods, covering different parameter sizes and performance levels.

[0118] Table 2. GraphForge performance on GTools

[0119]

[0120] The performance of GraphForge compared with tool instructions and function call methods is shown in Table 2. Compared with the tool instruction-based Graph-ToolFormer, GraphForge consistently outperforms by 30%. Compared with the function call-based GPT-4o-FC, GraphForge's reasoning performance is comparable to the high-cost GPT-4o-FC. Compared with GLM4-0520-FC and GPT-3.5-FC, GraphForge leads in accuracy by more than 20%.

[0121] Table 3. Performance of GraphForge on NLGraph

[0122] Task GPT-4-Turbo GLM4-0520-FC GPT-4o-FC GraphForge Ring Detection 66.75 100 100 99.73 Path exists 84.57 99.46 99.14 99.76 Shortest Path 51.51 86.84 99.73 99.47 Topological sort 23.25 98.14 99.38 99.38 Maximum flow 6.57 84.86 98.28 99.14 Bipartite Graph Detection 33.92 98.03 100 99.6 Graph Neural Networks 61.53 97.08 100 99.58 average value 46.87 94.92 99.5 99.52

[0123] In addition, Table 3 shows the performance of GraphForge on the NLGraph dataset. The results show that GraphForge also shows strong performance on the public dataset NLGraph. For cross-domain tasks, GraphForge also shows strong generalization capabilities.

[0124] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A method for constructing a large language model for graph reasoning based on graph tool instruction learning, characterized in that: The following steps are involved: Step 1: Rule-based graph task generation and decomposition: Generate graph reasoning tasks based on four rules: graph scale diversity, multiple descriptions of graph reasoning tasks, balanced answers to graph reasoning tasks, and unique answers to graph reasoning tasks. Then divide graph reasoning task information into graph information extraction and task information extraction. Step 2: The graph reasoning task obtains the output of the large language model through tool call instructions, uses regularization processing on its output, and uses the tool execution results to generate task labels; Step 3: Define the original dataset based on the graph reasoning task, tool call instructions, and the output of the large language model. Obtain the dataset GTools by filtering through the matching function. Use the parameter-efficient low-rank adaptation method LoRA to fine-tune the Llama3-8B base model on the dataset GTools to obtain the open source graph reasoning large model GraphForge.

2. According to the method for constructing a large language model based on graph reasoning and graph tool instruction learning according to claim 1, it is characterized in that: The four rules of graph scale diversity, multiple descriptions of graph reasoning tasks, balanced answers to graph reasoning tasks, and unique answers to graph reasoning tasks are as follows: Graph size diversity: Graph information is organized in natural language form, and the size of the graph is reasonably constrained. The number of nodes in the graph is controlled between 2 and 40, and the number of edges is at most 300. Multiple descriptions of graph reasoning tasks: Design more than one natural language description for each graph reasoning task; Balanced answers for graph reasoning tasks: In tasks that require judging true or false, ensure that the answers in the generated graph data are evenly distributed; Unique answer for graph reasoning tasks: For topological sorting tasks that may have multiple valid solutions, the uniqueness of the answer is ensured during the graph generation process.

3. According to claim 2, a method for constructing a large language model based on graph reasoning and graph tool instruction learning is characterized in that: The output of the large language model is formalized as follows: Among them, I (s) is the tool calling instruction, x (s) For graph reasoning tasks, including graph information extraction and task information extraction Use collection express; The specific formalization of generating task labels using tool execution results is as follows: in is the output task label, Tool represents the tool, Output results for the large model after regularization.

4. According to claim 3, a method for constructing a large language model based on graph reasoning and graph tool instruction learning is characterized in that: The construction process of the dataset GTools is as follows: Based on the graph reasoning task x (s) Its instruction I (s) And the output g of the large language model (s) Defined as the original dataset Decide whether to keep the output g based on the matching function M (s) : The matching function M is specifically: where g *(s) is the standard label of each subtask, g * is the answer label, ∧ represents the union of the various information output by the large model and the standard label. Through the above filtering process, we get the data set The dataset Organized in Alpaca format, the dataset GTools is obtained.

5. According to claim 4, a method for constructing a large language model based on graph reasoning and graph tool instruction learning is characterized in that: The parameter-efficient low-rank adaptation method LoRA is as follows: in is the actual output of the large language model, It is a graph reasoning task and instruction i (s) Combination of α controls the original weight matrix The update range, and is a low-rank matrix, r is the selected rank, less than min(d m ,d n ),in represents the field of real numbers, d m Represents the dimension of the output vector, d n Represents the dimension of the input vector; during the learning process, only matrices A and B are updated.

6. A method for constructing a large language model based on graph reasoning and graph tool instruction learning according to claim 5, characterized in that: The loss function of the open source graph reasoning model GraphForge fine-tuning for: in, represents the average loss on the entire dataset GTools, It represents the number of samples in the training data set, and cross-entropy is the cross entropy.

7. A method for constructing a large language model based on graph reasoning and graph tool instruction learning according to claim 6, characterized in that: The cross entropy is defined as Among them, j is the index of the output vector, indicating the jth category or dimension in the output vector. is the true label of the jth category, is the probability value of the jth category predicted by the model.