A tool chain planning method based on semantic perception hypergraph and large language model fusion

CN122819210APending Publication Date: 2026-09-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610958952.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供了一种基于语义感知超图与大语言模型融合的工具链规划方法,以通过语义兼容性判定、超图高阶协同建模与时序防越界修正,解决接口匹配僵化、高阶拓扑缺失及自回归幻觉问题,提升工具链规划的泛化能力和鲁棒性

Benefits of technology

[0007]本申请实施例提供了一种基于语义感知超图与大语言模型融合的工具链规划方法,首先,基于包含多个不同功能接口的工具库中各工具的输入输出类型信息,采用字符串完全匹配与语义空间相似度计算相结合的分段赋值策略,构建反映工具间接口兼容性程度的语义兼容性矩阵,进而,从历史成功执行的工具调用轨迹中提取高频共现的工具组合作为超边,并基于超边构建表示工具间高阶协同关系的初始超图关联矩阵;进一步的,利用语义兼容性矩阵对初始超图关联矩阵进行加权修正,基于修正后的加权超图结构对工具的初始特征和用户请求文本的文本特征进行图信息聚合,输出融合了接口兼容性与高阶协同拓扑的增强图特征;进而,将增强图特征与用户请求文本分别编码为提示词张量后拼接输入至预设推理模型进行自回归推理,生成包含任务节点以及任务节点之间拓扑连接关系的初始结构化工具链结果;从而,对初始结构化工具链结果进行解析,检测任务节点参数中对后方未生成节点的非法前向引用,将检测到的非法前向引用修正为指向前置相邻节点的合法引用,输出符合时序因果约束的可执行工具链。本申请的技术方案,通过构建语义兼容性矩阵,实现了对语义一致但命名各异的异构接口的灵活兼容性判定,解决了字符串精确匹配泛化能力差的问题;通过从历史轨迹中提取高频共现工具组合作为超边并构建超图关联矩阵,实现了对多工具高阶共现模式的有效表征,解决了普通图无法捕获高阶拓扑关系的问题;通过对初始结构化工具链结果中非法前向引用的检测与修正,实现了对自回归生成时序逻辑错误的矫正,解决了模型生成复杂图结构时易产生越界幻觉的问题,最终输出了符合时序因果约束的可执行工具链,提升了工具链规划的鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819210A_ABST
    Figure CN122819210A_ABST
Patent Text Reader

Abstract

The application provides a tool chain planning method based on semantic perception hypergraph and large language model fusion, constructs a semantic compatibility matrix reflecting the interface compatibility degree between tools based on the input and output type information of each tool; constructs an initial hypergraph association matrix from the hyperedges extracted from the tool call track of historical successful execution; the initial hypergraph association matrix is weighted and corrected by using the semantic compatibility matrix, the initial features of the tools and the text features of the user request text are aggregated based on the weighted hypergraph structure after correction, and enhanced graph features are output; based on the enhanced graph features and the user request text, an initial structured tool chain result is generated; the initial structured tool chain result is analyzed, the detected illegal forward reference is corrected to a legal reference pointing to the front adjacent node, and an executable tool chain is output, solving the problems of interface matching rigidity, high-order topological loss and self-recurrence illusion, and improving the generalization ability and robustness of tool chain planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the fields of artificial intelligence and natural language processing technology, and in particular to a toolchain planning method based on the fusion of semantically perceptive hypergraphs and large language models. Background Technology

[0002] In recent years, the reasoning ability of large-scale language models has been significantly improved. Using them for external tool invocation and complex task planning has become a core research direction, aiming to break through the limitation of models relying solely on internal knowledge and solve multi-step engineering problems by calling external interfaces.

[0003] Currently, a graph neural network-based assisted planning method is commonly used. Specifically, a general graph is constructed using tools as nodes and complete string matching of input and output types as the basis for edge connections. The graph features are then encoded using a graph neural network and input into a large-scale language model, which generates the toolchain planning results through autoregression. This approach assists the model's decision-making by introducing topological relationships.

[0004] The main drawbacks of the above technical solutions are as follows: First, type compatibility judgment relies on exact string matching, which cannot handle heterogeneous interfaces with the same semantics but different names, resulting in poor generalization ability; Second, ordinary graphs can only represent pairwise connections, which cannot capture high-order co-occurrence patterns used by multiple tools in collaboration, and lack high-order topological prior guidance; Third, when the model autoregressively generates complex graph structures, it is prone to generating illegal forward references to nodes that have not yet been generated, leading to planning failure. Summary of the Invention

[0005] This application provides a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models. By using semantic compatibility determination, high-order topology missing and autoregressive illusion problems, the method can solve the problems of interface matching rigidity, high-order topology missing and autoregressive illusion, thereby improving the generalization ability and robustness of toolchain planning.

[0006] This application provides a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models. The method includes: Based on the input and output type information of each tool in a tool library containing multiple different functional interfaces, a segmented assignment strategy combining string full matching and semantic space similarity calculation is adopted to construct a semantic compatibility matrix that reflects the degree of interface compatibility between tools. High-frequency co-occurring tool combinations are extracted from the historical successful tool call trajectories as hyperedges, and an initial hypergraph association matrix representing high-order collaborative relationships between tools is constructed based on the hyperedges; The initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix. Based on the corrected weighted hypergraph structure, graph information is aggregated for the initial features of the tool and the textual features of the user request text, and an enhanced graph feature that integrates interface compatibility and high-order collaborative topology is output. The enhanced graph features and the user request text are encoded into prompt word tensors, which are then concatenated and input into a preset inference model for autoregressive inference, generating an initial structured toolchain result containing task nodes and the topological connection relationships between task nodes. The initial structured toolchain result is parsed, illegal forward references to subsequent nodes that have not yet been generated are detected in the task node parameters, the detected illegal forward references are corrected to legal references pointing to the preceding adjacent nodes, and an executable toolchain that conforms to the temporal causal constraints is output.

[0007] This application provides a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models. First, based on the input / output type information of each tool in a tool library containing multiple different functional interfaces, a segmented assignment strategy combining string full matching and semantic space similarity calculation is used to construct a semantic compatibility matrix reflecting the degree of interface compatibility between tools. Then, high-frequency co-occurring tool combinations are extracted from historically successfully executed tool call trajectories as hyperedges, and an initial hypergraph association matrix representing high-order collaborative relationships between tools is constructed based on these hyperedges. Further, the initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix, and the corrected weighted hypergraph is then used to construct the final hypergraph association matrix. The structure aggregates graph information from the initial features of the tool and the textual features of the user request text, outputting enhanced graph features that integrate interface compatibility and high-order collaborative topology. Then, the enhanced graph features and the user request text are encoded into prompt word tensors and concatenated before being input into a pre-defined inference model for autoregressive inference, generating an initial structured toolchain result containing task nodes and the topological connections between them. The initial structured toolchain result is then parsed, detecting illegal forward references in the task node parameters to nodes that have not yet been generated. These illegal forward references are corrected to legal references to preceding adjacent nodes, outputting an executable toolchain that conforms to temporal causal constraints. The technical solution of this application, by constructing a semantic compatibility matrix, achieves flexible compatibility determination for heterogeneous interfaces with consistent semantics but different names, solving the problem of poor generalization ability of string precise matching; by extracting high-frequency co-occurrence tool combinations from historical trajectories as hyperedges and constructing a hypergraph association matrix, it achieves effective representation of high-order co-occurrence patterns of multiple tools, solving the problem that ordinary graphs cannot capture high-order topological relationships; by detecting and correcting illegal forward references in the initial structured toolchain results, it achieves correction of time-series logic errors in autoregressive generation, solving the problem of out-of-bounds illusions when the model generates complex graph structures, and finally outputs an executable toolchain that conforms to time-series causal constraints, improving the robustness of toolchain planning. Attached Figure Description

[0008] To more clearly illustrate the technical solutions of the exemplary embodiments of this application, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the embodiments to be described in this application, and not all of them. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.

[0009] Figure 1 This is a flowchart illustrating a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models provided in an embodiment of this application. Figure 2A flowchart illustrating another toolchain planning method based on the fusion of semantically aware hypergraph and large language model provided in this application embodiment; Figure 3 This is a schematic diagram illustrating the implementation process of determining a semantic compatibility matrix according to an embodiment of this application; Figure 4 A flowchart illustrating another toolchain planning method based on the fusion of semantically aware hypergraph and large language model provided in this application embodiment; Figure 5 This is a schematic diagram illustrating the implementation process of determining enhancement map features according to an embodiment of this application. Detailed Implementation

[0010] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present application, not the entire structure.

[0011] Before introducing the technical solutions provided in the embodiments of this application, the application scenarios of the solutions can be described first. This embodiment is applicable to various scenarios that require planning heterogeneous interfaces in a large-scale tool library. Currently, although interface compatibility judgment methods based on string exact matching and tool relationship modeling methods based on ordinary graphs are widely used in toolchain planning, traditional methods have obvious limitations. In practical applications, since the tool library contains multiple different functional interfaces, the input and output types of each interface have heterogeneous characteristics with semantically consistent but different expressions. Moreover, in the historical successful execution trajectory, there are generally high-order co-occurrence patterns where multiple tools must be used in concert. At the same time, when large language models generate complex directed acyclic graph structures, they are prone to generating the illusion of illegal forward references to nodes that have not been generated. Traditional methods are difficult to achieve global optimization at the levels of interface compatibility generalization, high-order collaborative topology representation, and consistency of time-series causal constraints, which can easily lead to problems such as tool connection failure, low planning search efficiency, and unexecutable generated results. Therefore, there is an urgent need for a planning method that can integrate semantic-aware compatibility judgment, high-order relation mining of hypergraph structures, and temporal boundary protection post-processing correction to improve the robustness, generalization ability, and executability of toolchain completion and planning.

[0012] Example 1 Figure 1 This is a flowchart illustrating a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models provided in this application embodiment. This embodiment can be applied to various situations that require planning heterogeneous interfaces in a large-scale tool library.

[0013] like Figure 1As shown, the toolchain planning method based on the fusion of semantically aware hypergraphs and large language models provided in this embodiment of the invention includes the following steps: S110. Based on the input and output type information of each tool in a tool library containing multiple different functional interfaces, a segmented assignment strategy combining string full matching and semantic space similarity calculation is adopted to construct a semantic compatibility matrix that reflects the degree of interface compatibility between tools.

[0014] The tool library with different functional interfaces refers to a collection of software components or service modules containing multiple modules with different operation entry points and calling specifications. Input / output type information refers to the data category descriptions of the data received and the data category descriptions of the data generated by each tool.

[0015] Among these, string matching refers to a method of determining whether two text strings are completely identical character by character at the character sequence level. Semantic space similarity refers to a quantitative indicator that measures the degree of semantic similarity by the geometric distance between vectors after mapping text information to a continuous vector space. Segmented assignment strategy refers to a weight allocation method that sets different value rules according to different condition intervals.

[0016] The semantic compatibility matrix is ​​a two-dimensional structured data organization method used to record the compatibility weights calculated based on the degree of interface type matching between any two tools in the tool library.

[0017] Specifically, the first step is to obtain the data categories that each tool can receive and generate, which will serve as the basis for subsequent judgments. When evaluating the compatibility of any two tools, a precise string-level comparison is performed first. If the output type information of one tool is completely identical to the input type information of another tool in terms of text characters, it is directly determined to be highly compatible. When a complete string match does not occur, semantic space similarity calculation is performed, which maps type information to a semantic space to measure the closeness of their meanings. The strength of compatibility is determined based on the calculated similarity. Based on the above two different levels of judgment results, a segmented assignment strategy is adopted, that is, different intervals of compatibility weight values ​​are assigned to complete string matches, highly semantically similar cases, and other general cases. Finally, these weight values ​​are organized to form a semantic compatibility matrix that can comprehensively reflect the degree of interface compatibility between the tools in the tool library.

[0018] S120. Extract frequently co-occurring tool combinations from the historical successful tool call trajectories as hyperedges, and construct an initial hypergraph association matrix based on the hyperedges to represent high-order collaborative relationships between tools.

[0019] Here, a tool call trajectory refers to a series of consecutive, successfully completed tool calls recorded in chronological order during the execution process. A tool portfolio refers to a set of tools that are used together during the execution of a task. A hyperedge is a connection structure used to connect two or more tool nodes, indicating a collaborative working relationship between these tools.

[0020] The initial hypergraph association matrix is ​​a two-dimensional data structure used to record the attribution relationship between each tool in the tool library and each hyperedge.

[0021] Specifically, to capture complex collaborative patterns between tools that transcend pairwise relationships, the first step is to collect historical tool call trajectories that have successfully completed tasks as a data source. Within these trajectories, statistical analysis identifies frequently co-occurring tool combinations—high-frequency co-occurrence tool combinations—representing statistically significant and validated collaborative patterns in practice. Subsequently, each identified high-frequency co-occurrence tool combination is defined as a hyperedge. The unique characteristic of a hyperedge is its ability to connect multiple tools simultaneously, thus expressing the higher-order relationships of these tools working collaboratively as a whole. Finally, using all tools as rows and all identified hyperedges as columns, a corresponding attribution label is assigned based on whether each tool belongs to a particular hyperedge. The data structure organized according to this rule constitutes the initial hypergraph association matrix representing the higher-order collaborative relationships between tools.

[0022] S130. The initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix. Based on the corrected weighted hypergraph structure, the initial features of the tool and the textual features of the user request text are aggregated to output an enhanced graph feature that integrates interface compatibility and high-order collaborative topology.

[0023] Here, the initial features of a tool refer to a set of predefined primitive attribute vectors for each tool in the tool library, used to characterize the basic functional semantics of that tool. User request text refers to the raw input statement expressed in natural language, describing the current task objective or requirement. Text features refer to the vectorized representation extracted from the user request text after encoding, capable of representing its semantic connotation.

[0024] Among them, enhanced graph features refer to the feature representation output by inputting the initial features of the tools and text features into the graph structure propagation mechanism, which integrates interface compatibility information and high-order collaborative topological relationships between tools.

[0025] Specifically, during the graph information propagation process, the initial hypergraph association matrix is ​​first weighted and corrected using a semantic compatibility matrix. The semantic compatibility matrix records the interface compatibility weights between any two tools, determined by the degree of matching between input and output types, while the initial hypergraph association matrix expresses the attribution relationship between tools and hyperedges. The specific way to combine these two is to use the interface compatibility constraints reflected by the semantic compatibility matrix as weight coefficients to scale the full connectivity relationship between tools obtained by clique expansion of the initial hypergraph association matrix position by position. This results in the semantic-level physical suppression or enhancement of the connection relationship that originally only depended on the topological structure, thus obtaining a corrected weighted hypergraph structure. Based on this, the weighted hypergraph structure is used as the skeleton of information propagation. Graph information aggregation is performed on the initial features of the tools and the textual features of the user request text. The initial features of the tools provide the basic semantic information of each tool, while the textual features of the user request text introduce the goal-oriented information of the current task. By repeatedly transmitting and fusing these features along the weighted hypergraph structure between interconnected tool nodes, an enhanced graph feature that integrates interface compatibility and high-order collaborative topology is finally output. This feature includes both the high-order collaboration patterns formed by historical collaboration experience between tools and the connection feasibility strength information determined by interface type compatibility.

[0026] S140. Encode the enhanced graph features and the user request text into prompt word tensors, then concatenate them and input them into the preset inference model for autoregressive inference to generate an initial structured toolchain result containing task nodes and the topological connection relationships between task nodes.

[0027] Among them, the cue word tensor refers to the encoding of graph features or text information into a fixed-length numerical vector sequence that can be directly processed by the model.

[0028] The pre-built inference model refers to a deep learning network that has been pre-constructed and parameter-tuned to generate output sequences word-by-word based on input information. In this embodiment, optionally, considering the high memory consumption of complex toolchain planning tasks and the efficiency requirements of fine-tuning large-size language models in specific domains, the pre-built inference model is a large language model that has undergone efficient parameter fine-tuning. A large language model refers to a large pre-trained language model with hundreds of millions of parameters. The efficient parameter fine-tuning specifically includes: quantizing the backbone weights of the large language model into low-bit floating-point data types using double quantization technology; reducing the memory usage of intermediate activation values ​​during training using gradient checkpointing technology; freezing the backbone weights of the large language model using low-rank adaptive fine-tuning technology and injecting low-rank bypasses only into the query matrix, key matrix, value matrix, and output matrix of the attention mechanism; and performing gradient accumulation backpropagation based on a paging optimizer and cross-entropy loss function to obtain the fine-tuned large language model. In the low-rank adaptive fine-tuning technique, the rank of the low-rank bypass is set to 16, the scaling factor is set to 32, and the dropout ratio is set to 0.05; the double quantization technique uses a 4-bit floating-point data type and the calculation data type is set to a 16-bit floating-point number; the cumulative steps of gradient accumulation are set to 8.

[0029] The initial structured toolchain result refers to an organized data description containing task nodes and their interconnections, which is directly generated by a pre-defined inference model and has not yet undergone temporal validity verification.

[0030] Specifically, to integrate graph structure information with user requirements for toolchain generation, the enhanced graph features are first encoded and mapped into graph cue tensors of fixed length with dimensions matching the internal representation of the pre-defined inference model. Simultaneously, the user request text is encoded into a text cue tensor using the pre-defined inference model's word embedding mechanism. The graph cue tensor is then concatenated with the text cue tensor at the beginning, forming a complete input context. This allows the pre-defined inference model to preferentially perceive the prior topological knowledge inherent in the graph structure during subsequent generation. The concatenated complete input context is then fed into the pre-defined inference model, triggering an autoregressive inference process. The model progressively predicts the next output unit based on the already generated content, adding the newly generated unit to the context at each step to continue generating subsequent content. In this way, the pre-defined inference model progressively outputs an organized data description containing task nodes and the topological connections between them. Since this result has not undergone temporal validity verification, it is referred to as the initial structured toolchain result.

[0031] S150. Parse the initial structured toolchain result, detect illegal forward references to subsequent nodes that have not been generated in the task node parameters, correct the detected illegal forward references to legal references pointing to the preceding adjacent nodes, and output an executable toolchain that conforms to the temporal causal constraints.

[0032] Among them, task node parameters refer to the field content carried by each task node, which specifies the output results or configuration information of other nodes that it depends on during runtime.

[0033] Here, "ungenerated subsequent nodes" refers to subsequent task nodes that are after the currently being processed task node but have not yet been generated by the model. An illegal forward reference occurs when the parameters of the current task node incorrectly point to the index of a node in the sequence that does not yet exist; such a reference violates the temporal causal relationship.

[0034] Among them, a valid reference pointing to the preceding adjacent node refers to correcting an originally invalid reference to point to the task node that precedes the current task node, which must have already been generated and exists.

[0035] Among them, the executable toolchain refers to a sequence of task nodes that can be directly called and executed after time sequence legality verification and out-of-bounds reference correction, where the dependencies between all nodes conform to the causal constraints of a directed acyclic graph.

[0036] In this embodiment, the executable toolchain that conforms to the temporal causal constraint is a structured representation containing a sequence of task nodes and topological dependencies between nodes; wherein, each task node contains a node index identifier, tool call parameters and a preceding node reference field, and the structured representation is used to directly drive the automatic invocation and execution of the corresponding functional interfaces in the tool library.

[0037] This can be understood as follows: the executable toolchain, after undergoing temporal validity verification and conforming to temporal causality constraints, is not simply a list of tools in terms of data organization. Instead, it is a structured representation containing a sequence of task nodes and topological dependencies between nodes. The task node sequence refers to an ordered set of multiple task nodes arranged in the order of execution. The topological dependencies between nodes indicate the other task nodes that each task node depends on; that is, the current node can only start running after the dependent node has completed its execution. Specifically, the internal structure of each task node contains three core components: a node index identifier used to uniquely distinguish and locate each task node in the sequence; tool call parameters indicating which functional interface in the tool library the node needs to call and the specific configuration information passed in; and a predecessor node reference field recording the index identifiers of the predecessor task nodes that the current node depends on, through which the validity of the dependency relationship can be traced and verified. Ultimately, this structured representation containing complete node information and dependencies does not rely on additional interpretation or transformation and can be directly used to drive the automatic invocation and execution of the corresponding functional interfaces in the tool library, that is, triggering the tool calls specified by each task node in the order determined by the topological dependencies between nodes.

[0038] Specifically, because large language models are prone to generating logical errors that violate causal temporal order when generating sequences with topological dependencies, post-processing corrections are needed for the initial structured toolchain results directly output by the model. Specifically, the initial structured toolchain results are first parsed, decomposed into a sequence of sequentially arranged task nodes, and each task node's parameters are checked for references to other nodes. During the check, special attention is paid to illegal forward references, i.e., parameters of the current task node pointing to an ungenerated node located after it but not yet generated. Such references violate the basic temporal constraints of directed acyclic graphs. Once an illegal forward reference is detected, it is forcibly corrected to a legal reference pointing to a preceding neighbor node. That is, the index originally pointing to an unknown subsequent node is rewritten to the index of the task node preceding the current task node. Since this preceding neighbor node already exists and must have been defined before the current node, the correction satisfies causal temporal order logic. For the starting node or boundary cases that occur during the correction process, if the preceding neighbor node does not exist, the original parameters are retained or the node is treated as an independent node. After the above node-by-node detection and correction, the final output is an executable toolchain in which all reference relationships conform to the temporal causal constraint.

[0039] In this embodiment, optionally, the specific implementation of generating an executable toolchain that conforms to temporal causality constraints may include the following steps: S1501. Parse the initial structured toolchain result to obtain the task node sequence, and traverse each task node in the task node sequence.

[0040] Here, the task node sequence refers to an ordered set of multiple task nodes arranged in the order of generation, obtained after parsing the initial structured toolchain result. A task node is a single unit in the sequence that represents an independent tool call operation.

[0041] Specifically, parsing refers to breaking down and identifying the original output text generated by the model according to predefined format specifications, extracting individual task nodes, and organizing them into a task node sequence based on their order of appearance in the original output. After obtaining the task node sequence, each task node in the sequence is traversed sequentially from beginning to end, treating each task node as the current processing object and performing subsequent reference detection and correction operations on that task node until all task nodes in the sequence have been processed. The combination of parsing and traversal ensures that the reference relationships of each task node in the sequence can be examined one by one, laying the foundation for the subsequent detection and correction of illegal forward references.

[0042] S1502. For the task node at the current index position, extract the reference index from the parameter field of the task node based on regular expressions.

[0043] Here, the current index position refers to the position number of the task node being processed in the sequence during traversal. The parameter field is a specific data area within the task node used to store runtime configuration information or dependency references. The reference index is a value extracted from the parameter field that points to the position numbers of other task nodes in the sequence.

[0044] Specifically, each task node contains a parameter field that stores the index identifiers of other nodes it depends on during runtime. To accurately extract this index information from the raw text content of the parameter field, regular expressions are used as a matching tool. Regular expressions are a text retrieval technique based on specific pattern matching rules, capable of locating and extracting substrings that match a predefined syntax pattern from a string. The content of the parameter field of the current task node is compared with the regular expression used to match the node reference format, extracting one or more reference indices. These reference indices represent the position numbers of other task nodes that the current task node depends on in the sequence. Through this extraction operation, the originally unstructured parameter field text is converted into structured reference values ​​that can be used for subsequent logical judgment.

[0045] S1503. If the reference index is greater than or equal to the current index position and the value of the current index position minus one is greater than or equal to zero, then the reference index is rewritten to the value of the current index position minus one to obtain a corrected valid reference; if the reference index is less than the current index position, then the original reference index is retained unchanged; if the reference index is greater than or equal to the current index position and the value of the current index position minus one is less than zero, then the current task node is treated as an independent node and its reference field is set to null.

[0046] Independent node processing refers to the operation method where, when a task node cannot find a valid predecessor dependency node, the node is regarded as a starting node that does not depend on any other node and is not referenced or rewritten.

[0047] Specifically, after extracting the reference index from the parameter field of the current task node, the validity of the reference index needs to be determined and corresponding correction operations need to be performed. First, compare the size relationship between the reference index and the current index position: if the reference index is greater than or equal to the current index position, and the value of the current index position minus one is greater than or equal to 0, it means that the reference index points to the current task node itself or a task node in the sequence that is located after the current node but has not yet been generated. This reference violates the temporal causality constraint, so it needs to be rewritten to the value of the current index position minus one, that is, the dependency relationship needs to be forcibly corrected to point to the preceding adjacent node of the current task node. The value obtained after this rewriting is called the corrected valid reference.

[0048] If the reference index is less than the current index position, it means that the reference index points to a node that already exists in the sequence and is located before the current task node. This kind of reference conforms to the temporal causality constraint, so the original reference index is kept unchanged without any modification.

[0049] In addition, a boundary case needs to be handled: if the reference index is greater than or equal to the current index position and the value of the current index position minus one is less than zero, this means that the current index position is already the starting position of the entire sequence, and there is no previous node to reference. Therefore, this task node does not have a valid forward reference condition, and is thus treated as an independent node, that is, its reference field is set to null, indicating that this node does not depend on any previous node. Through the processing of the above three conditional branches, it is ensured that the reference relationship of each task node can ultimately satisfy the temporal legality requirement.

[0050] S1504. Based on the correction and retention results of the reference index, output an executable toolchain that conforms to the temporal causal constraint.

[0051] In this embodiment, after the validity of the reference index of each task node in the task node sequence is determined and the corresponding correction operation is performed, the reference indexes in the parameter fields of all task nodes have been processed: the originally invalid reference indexes are rewritten as corrected valid references, the originally valid reference indexes are retained unchanged, and boundary cases are properly resolved by retaining the original parameters or handling them independently. Based on the above correction and retention results performed on each task node, these processed task nodes are reorganized according to their original sequence order to form a complete task node sequence in which the dependencies between all nodes no longer contain invalid forward references. Each task node in this sequence satisfies the temporal constraint that the node pointed to by its reference index must be located before itself. Finally, this corrected and reorganized complete task node sequence is taken as the output result, which is the executable toolchain that conforms to the temporal causality constraint.

[0052] This application provides a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models. First, based on the input / output type information of each tool in a tool library containing multiple different functional interfaces, a segmented assignment strategy combining string full matching and semantic space similarity calculation is used to construct a semantic compatibility matrix reflecting the degree of interface compatibility between tools. Then, high-frequency co-occurring tool combinations are extracted from historically successfully executed tool call trajectories as hyperedges, and an initial hypergraph association matrix representing high-order collaborative relationships between tools is constructed based on these hyperedges. Further, the initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix, and the corrected weighted hypergraph is then used to construct the final hypergraph association matrix. The structure aggregates graph information from the initial features of the tool and the textual features of the user request text, outputting enhanced graph features that integrate interface compatibility and high-order collaborative topology. Then, the enhanced graph features and the user request text are encoded into prompt word tensors and concatenated before being input into a pre-defined inference model for autoregressive inference, generating an initial structured toolchain result containing task nodes and the topological connections between them. The initial structured toolchain result is then parsed, detecting illegal forward references in the task node parameters to nodes that have not yet been generated. These illegal forward references are corrected to legal references to preceding adjacent nodes, outputting an executable toolchain that conforms to temporal causal constraints. The technical solution of this application, by constructing a semantic compatibility matrix, achieves flexible compatibility determination for heterogeneous interfaces with consistent semantics but different names, solving the problem of poor generalization ability of string precise matching; by extracting high-frequency co-occurrence tool combinations from historical trajectories as hyperedges and constructing a hypergraph association matrix, it achieves effective representation of high-order co-occurrence patterns of multiple tools, solving the problem that ordinary graphs cannot capture high-order topological relationships; by detecting and correcting illegal forward references in the initial structured toolchain results, it achieves correction of time-series logic errors in autoregressive generation, solving the problem of out-of-bounds illusions when the model generates complex graph structures, and finally outputs an executable toolchain that conforms to time-series causal constraints, improving the robustness of toolchain planning.

[0053] Example 2 Figure 2 This is a schematic diagram of a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models provided in this application embodiment. Based on the foregoing embodiments, this embodiment will provide a detailed description of S110 and S120, and the specific implementation can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0054] like Figure 2 As shown, the method specifically includes the following steps: To incorporate the originally discrete and heterogeneous interface type information into a unified continuous semantic space for computation, before constructing a semantic compatibility matrix reflecting the degree of interface compatibility between tools, it is necessary to convert the input and output type information of each tool in the tool library into vectorized representations that can be used for semantic similarity measurement. Specific implementation methods may include the following steps: S1. Convert the input type set and output type set of each tool in the tool library into text strings respectively.

[0055] The input type set refers to the summative combination of all data categories that each tool can receive. The output type set refers to the summative combination of all data categories that each tool can generate. The text string refers to the transformation of the originally structured type set information into a continuous, character-based natural language sequence.

[0056] Specifically, the input type set is a data set containing multiple type names, and the output type set is similar. The elements within these sets are typically organized in a structured way, such as in lists or sets. However, the language model expects to receive input in the form of continuous character sequences. Therefore, it is necessary to concatenate the type names in the input type set of each tool into a coherent text string according to certain delimiting rules. The same operation is applied to the output type set, concatenating the type names into another independent text string. After this transformation, the originally discrete, structured input and output type sets become two continuous text strings that can be directly read and encoded by the language model. Each text string fully preserves the semantic information of all type names in the original sets.

[0057] S2. Input the text string into a lightweight pre-trained language model for feature encoding.

[0058] Lightweight pre-trained language models refer to deep learning text encoding models with fewer parameters, lower computational overhead, and pre-trained on large-scale corpora.

[0059] Specifically, a lightweight pre-trained language model is a deep learning network pre-trained on a large-scale text corpus. Its core capability lies in its ability to map natural language text of arbitrary length into a fixed-dimensional continuous vector, which effectively preserves the semantic information of the original text. When a text string is fed into the lightweight pre-trained language model, the model internally uses a multi-layered neural network to encode and extract features from each character and its context, ultimately producing a dense floating-point vector at the model's output layer. This process is feature encoding, which essentially transforms the text string from a raw character sequence into a compressed, semantically rich numerical vector representation. This vector will subsequently be used as a semantic vector in similarity calculations.

[0060] S3. Based on the feature vectors output by the lightweight pre-trained language model, determine the input semantic vectors corresponding to the input type set and the output semantic vectors corresponding to the output type set.

[0061] The input semantic vector is a dense numerical representation of the semantic meaning of the text strings corresponding to the input type set, obtained by feature encoding. The output semantic vector is a dense numerical representation of the semantic meaning of the text strings corresponding to the output type set, obtained by feature encoding.

[0062] Specifically, after the lightweight pre-trained language model completes feature encoding of the input text string, it generates a corresponding feature vector at its output. For each tool, there are two independent text strings: one derived from the input type set, and the other from the output type set. These two text strings are fed into the lightweight pre-trained language model for feature encoding, resulting in two independent feature vectors. The feature vector generated from the encoded text string corresponding to the input type set is determined as the input semantic vector, which carries the comprehensive semantic information of all data categories that the tool can receive. Similarly, the feature vector generated from the encoded text string corresponding to the output type set is determined as the output semantic vector, which carries the comprehensive semantic information of all data categories that the tool can generate. Through this process, the originally discrete input and output type sets are mapped to continuous, computable input and output semantic vectors, respectively.

[0063] S4. Determine the semantic similarity between tools based on the cosine similarity between the input semantic vector and the output semantic vector.

[0064] Specifically, cosine similarity is a numerical metric that measures the degree of directional consistency between two vectors in a high-dimensional space; a higher value indicates a closer semantic relationship between the two vectors. When evaluating the compatibility between a source tool and a target tool, the output semantic vector of the source tool and the input semantic vector of the target tool are used as the computational objects. These two vectors are substituted into the cosine similarity calculation framework to examine the degree of matching between the output semantic meaning of the source tool and the input semantic meaning of the target tool in the semantic space. The calculated cosine similarity value directly reflects the semantic similarity between this pair of tools. A higher value indicates that the output type of the source tool and the input type of the target tool are semantically closer, even if they are not completely identical in string form, a potential compatibility relationship can still be captured through semantic similarity. In this way, the semantic similarity between each pair of source and target tools is determined, providing a quantitative basis for the soft rule judgment in the subsequent segmentation assignment strategy.

[0065] S210. If the output type set of the source tool and the input type set of the target tool intersect at the string level, the compatibility weight between the source tool and the target tool is assigned a preset maximum value; if the output type set and the input type set do not intersect at the string level and the semantic similarity is greater than a preset threshold, the compatibility weight is assigned a continuous floating-point value of the semantic similarity; if the output type set and the input type set do not intersect at the string level and the semantic similarity is less than or equal to a preset threshold, the compatibility weight is assigned a preset minimum positive value.

[0066] In this context, the source tool refers to the tool that serves as the starting point for connections, and its output will be used to match the inputs of other tools. The target tool refers to the tool that serves as the ending point for connections, and its input will be used to receive the outputs of the source tool. The compatibility weight is a numerical value that quantifies the degree of interface matching between the source tool and the target tool.

[0067] The preset maximum value refers to the highest value assigned to the compatibility weight when the strings are completely matched. A continuous floating-point value refers to any real number between the minimum and maximum values ​​assigned to the compatibility weight when the strings are semantically similar but not completely matched. The preset minimum positive value refers to a positive number close to zero but not zero assigned to the compatibility weight when the strings do not match and the semantic similarity is low.

[0068] Specifically, when assessing the possibility of a valid interface connection between any two tools, it is necessary to consider the matching of their type information at both the string form and semantic meaning levels. First, check whether the output type set of the source tool overlaps with the input type set of the target tool; that is, whether all data categories that the source tool can generate overlap with all data categories that the target tool can receive. If there is an intersection at the string level—that is, at least one data type has the same name text—it indicates that the two tools have a strictly accurate interface compatibility relationship. In this case, the compatibility weight between the source and target tools is assigned a preset maximum value to reflect the absolute reliability of the hard match.

[0069] When the output type set of the source tool and the input type set of the target tool do not have any completely identical data type names at the string level, the judgment is then transferred to the semantic level: the semantic similarity between the semantics of the output type of the source tool and the semantics of the input type of the target tool is calculated. If the semantic similarity is greater than a preset threshold, it means that although the type names of the two are not completely identical in characters, their semantic meanings are close enough to have a practically usable compatibility relationship. At this time, the compatibility weight is assigned a continuous floating-point value of the semantic similarity itself, so that the weight can accurately reflect the difference in the degree of semantic closeness.

[0070] If there is no intersection at the string level and the calculated semantic similarity is less than or equal to a preset threshold, it indicates that the source tool and the target tool lack effective compatibility in terms of both literal form and semantic meaning. In this case, a preset minimum positive value is assigned to the compatibility weight. This minimum positive value can both express the fact that the two are almost incompatible and maintain the weak connectivity of the graph structure in numerical computation to prevent the gradient vanishing problem in subsequent information propagation. Through the above three conditional branch assignment strategies, the compatibility weight between each tool pair is reasonably determined.

[0071] S220. Based on the assignment of compatibility weights between tools, construct a semantic compatibility matrix that reflects the degree of interface compatibility between tools.

[0072] Specifically, the semantic compatibility matrix is ​​a two-dimensional data organization method, where rows and columns represent all tools in the tool library, and the order of rows and columns is consistent. The element value at the i-th row and j-th column position in the matrix is ​​the compatibility weight assigned between the i-th tool as the source tool and the j-th tool as the target tool. Through this structured organization of a two-dimensional matrix, the interface compatibility degree between any two tools can be directly obtained by querying the element value at the corresponding position in the matrix. The constructed semantic compatibility matrix completely retains the assigned compatibility weights between all tool pairs, including the preset maximum value assigned for a complete string match, the continuous floating-point value assigned when the semantic similarity is greater than a preset threshold, and the preset minimum positive value assigned when the semantic similarity is less than or equal to the preset threshold. This provides a structured weight input for subsequent weighted correction of the hypergraph structure using this matrix.

[0073] Figure 3 This is a schematic diagram illustrating the implementation process for determining a semantic compatibility matrix. Specifically, it first parses the description files (e.g., a predefined JSON format tool description dictionary) of all nodes in the tool library. For any two tools in the tool library, they are defined as the source tools. and target tools .extract The set of "output-type" attributes, and The set of "input-type" attributes.

[0074] Next, a lightweight pre-trained language model can be loaded to convert the extracted type list into a text string and feed it into the model for feature encoding to obtain a fixed-dimensional (e.g., The dense floating-point semantic feature vectors (dimension 1) are denoted as the output semantic vectors. With input semantic vector .

[0075] After obtaining the semantic vectors, the system calculates the cosine similarity between the two in the semantic space. The calculation formula is as follows: To balance the absolute accuracy of existing hard rules with the generalization ability of soft rules, this embodiment employs a segmented assignment strategy to construct a global semantic compatibility matrix. : (1) Hard rule absolute matching: if Output type set and If the set of input types has a non-empty intersection at the string level, it indicates strict interface compatibility, and therefore the maximum weight is assigned. ; (2) Soft rule semantic generalization: If the strings do not intersect, but the calculated semantic similarity is greater than the set empirical threshold, then continuous floating-point weights are assigned. ; (3) Weak exploration penalty: For the remaining cases that do not meet the above two conditions, a minimum value (e.g., 0.01) is assigned. The purpose of setting this minimum value is to prevent gradient vanishing during subsequent graph neural network training and to maintain the weak connectivity exploration capability of the network structure.

[0076] S230. Filter out invalid nodes that are not in the current tool library from the historical successfully executed tool call traces to obtain valid tool call traces.

[0077] Invalid nodes refer to tool nodes that exist in the historical call history but are not in the current tool library. Valid tool call history refers to the sequence of tools remaining after filtering out invalid nodes from the original historical call history.

[0078] Specifically, each tool node can be iterated through from the historical successfully executed tool call trajectories. Each tool node is checked to see if it exists in the current tool library. Tool nodes not in the current tool library are identified as invalid and removed from the trajectories. Meanwhile, all valid nodes in the current tool library and their original order are preserved. After this filtering operation, the original trajectory, which might have contained invalid nodes, is transformed into a new sequence consisting only of valid nodes from the current tool library. This new sequence is called the valid tool call trajectory. It serves as input data for subsequent frequent itemset mining, ensuring that each tool combination in the mining results can directly find a corresponding tool instance in the current tool library.

[0079] S240. Based on the preset minimum support threshold, the frequent itemset mining algorithm is used to extract high-frequency co-occurring tool combinations from the effective tool call trajectory, and each high-frequency co-occurring tool combination is determined as a hyperedge.

[0080] The minimum support threshold refers to the lowest possible frequency of an item set to be considered a high-frequency co-occurrence item set. Frequent itemset mining algorithms are data mining methods used to identify all item sets in a dataset whose frequency exceeds the minimum support threshold.

[0081] In this embodiment, after obtaining the effective tool call trajectories, it is necessary to extract tool combinations that frequently appear together historically. These combinations represent collaborative work patterns that have been repeatedly verified in practice. The specific extraction process relies on a preset minimum support threshold, which is a proportional value between zero and one, used to determine whether a tool combination has statistically significant high-frequency characteristics. A frequent itemset mining algorithm is used to scan all trajectories in the effective tool call trajectories. This algorithm can count the frequency of each possible tool combination appearing in all trajectories, and tool combinations whose frequency is greater than or equal to the preset minimum support threshold are identified as high-frequency co-occurring tool combinations. For each identified high-frequency co-occurring tool combination, the multiple tools within it are considered as a whole collaborative unit, and the tool combination as a whole is defined as a hyperedge. The core characteristic of a hyperedge is that it can connect two or more tool nodes simultaneously, thereby expressing a high-order collaborative relationship between these tools as a whole to complete the task, rather than being limited to simple pairwise connections. Through this process, each extracted high-frequency co-occurring tool combination corresponds to a hyperedge.

[0082] S250. Based on all hyperedges, construct an initial hypergraph association matrix with a dimension equal to the total number of tools multiplied by the total number of hyperedges.

[0083] Specifically, when the tool belongs to a hyperedge, the matrix elements take a first preset value; when the tool does not belong to a hyperedge, the matrix elements take a second preset value. The first preset value refers to the fixed value filled into the matrix when the tool belongs to a hyperedge; the second preset value refers to the fixed value filled into the matrix when the tool does not belong to a hyperedge.

[0084] Specifically, after extracting all high-frequency co-occurring tool combinations from the effective tool call trajectories and identifying them as hyperedges, these hyperedges need to be organized into a mathematical form that facilitates subsequent graph structure calculations, namely, the initial hypergraph association matrix. The dimensions of this matrix are determined by two parameters: the first dimension equals the total number of tools in the tool library, representing all possible tool nodes; the second dimension equals the total number of mined hyperedges, representing all identified high-frequency tool combinations. During construction, each row of the matrix corresponds to a tool, and each column corresponds to a hyperedge. Each element in the matrix indicates whether the tool corresponding to that row belongs to the hyperedge corresponding to that column. Specifically, when a tool belongs to a hyperedge, i.e., the tool appears in the high-frequency co-occurring tool combination corresponding to that hyperedge, the element at the corresponding position in the matrix takes a first preset value, which is usually used to indicate that the attribution relationship is valid; when a tool does not belong to a hyperedge, i.e., the tool does not appear in the high-frequency co-occurring tool combination corresponding to that hyperedge, the element at the corresponding position in the matrix takes a second preset value, which is usually used to indicate that the attribution relationship is invalid. The initial hypergraph association matrix constructed in this way records the complete attribution relationships between all tools and all hyperedges in the form of a two-dimensional table, providing a structured topological input for subsequent hypergraph convolution and information propagation using this matrix.

[0085] S260. The initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix. Based on the corrected weighted hypergraph structure, the initial features of the tool and the textual features of the user request text are aggregated to output an enhanced graph feature that integrates interface compatibility and high-order collaborative topology.

[0086] S270. Encode the enhanced graph features and the user request text into prompt word tensors respectively, then concatenate them and input them into the preset inference model for autoregressive inference to generate an initial structured toolchain result containing task nodes and the topological connection relationships between task nodes.

[0087] S280. Parse the initial structured toolchain result, detect illegal forward references to subsequent nodes that have not been generated in the task node parameters, correct the detected illegal forward references to legal references pointing to the preceding adjacent nodes, and output an executable toolchain that conforms to the temporal causal constraints.

[0088] The technical solution of this application constructs a semantic compatibility matrix that reflects the degree of interface compatibility between tools. This is achieved through a segmented assignment strategy that combines hard rules of complete string matching with soft rules of semantic space similarity calculation. On one hand, complete string matching ensures the highest compatibility weight is assigned when the interface types are strictly consistent, guaranteeing the deterministic connection between explicitly compatible tool pairs. On the other hand, semantic similarity assigns continuous floating-point weights to tool pairs with mismatched strings but similar semantics, breaking through the rigid constraint of requiring complete string matching in traditional methods. This effectively solves the problem of difficulty in aligning heterogeneous interfaces due to naming differences, significantly expanding the coverage of effective connections between tools. Simultaneously, a preset minimum positive value is assigned to completely incompatible tool pairs, clarifying the incompatibility relationship while maintaining the weak connectivity of the graph structure to prevent gradient vanishing. This provides information-rich and numerically stable compatibility weight input for subsequent graph information aggregation.

[0089] The technical solution of this application, when constructing an initial hypergraph association matrix representing high-order collaborative relationships between tools based on hyperedges, ensures that the constructed hypergraph structure is entirely based on the currently available set of tools by filtering out invalid nodes not in the current tool library, thus avoiding interference from invalid nodes on subsequent collaborative relationship modeling. Secondly, by employing a frequent itemset mining algorithm and setting a minimum support threshold, statistically significant high-frequency co-occurring tool combinations can be automatically selected from historical trajectories, effectively eliminating noise interference caused by accidental co-occurrence and ensuring that the extracted hyperedges reflect real and stable collaborative working patterns. Finally, by constructing an initial hypergraph association matrix multiplied by the total number of tools and the total number of hyperedges, each high-frequency co-occurring tool combination is encoded as a hyperedge that can simultaneously connect multiple tools, thereby breaking through the limitation that ordinary graphs can only represent pairwise relationships, realizing a structured expression of high-order collaborative relationships between tools, and providing a topological input rich in multi-tool collaboration patterns for subsequent graph information aggregation.

[0090] Example 3 Figure 4 This is a schematic diagram of a toolchain planning method based on the fusion of semantically aware hypergraphs and large language models provided in this application embodiment. Based on the foregoing embodiments, this embodiment will provide a detailed description of S130-S140, and the specific implementation method can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0091] like Figure 4 As shown, the method specifically includes the following steps: S310. Based on the input and output type information of each tool in a tool library containing multiple different functional interfaces, a segmented assignment strategy combining string full matching and semantic space similarity calculation is adopted to construct a semantic compatibility matrix that reflects the degree of interface compatibility between tools.

[0092] S320. Extract frequently co-occurring tool combinations from the historical successful tool call trajectories as hyperedges, and construct an initial hypergraph association matrix based on the hyperedges to represent high-order collaborative relationships between tools.

[0093] S330. Calculate the clique expansion matrix based on the initial hypergraph association matrix, and multiply the clique expansion matrix element-wise with the semantic compatibility matrix to obtain the weighted adjacency matrix.

[0094] The clique expansion matrix is ​​a square matrix used to represent the pairwise connections between tool nodes, obtained by converting the hypergraph structure described by the initial hypergraph association matrix into a regular graph structure through the hyperedge clique expansion operation. The weighted adjacency matrix is ​​a comprehensive weight matrix obtained by multiplying the clique expansion matrix and the semantic compatibility matrix element-wise, where each element reflects both the original hypergraph topological connections between tools and the strength of interface semantic compatibility.

[0095] In this embodiment, the clique expansion matrix is ​​calculated based on the initial hypergraph association matrix. The core idea of ​​clique expansion is to transform each hyperedge connecting multiple tool nodes into a fully connected clique subgraph. That is, any two tool nodes belonging to the same hyperedge are connected by a pairwise edge. After performing clique expansion on all hyperedges and merging all connected edges, the resulting clique expansion matrix is ​​a square matrix, where its rows and columns correspond to the tools in the tool library. Each element in the matrix indicates whether there is at least one common hyperedge between two corresponding tools. However, the clique expansion matrix only reflects the topological connection relationship between tools in historical collaboration patterns and does not consider whether these connections are truly compatible at the interface type level. To simultaneously inject interface compatibility constraints, the clique expansion matrix and the semantic compatibility matrix are multiplied element-wise, that is, multiplication is performed on the elements at the same position in the two matrices. Since the semantic compatibility matrix records the compatibility weights between each pair of tools based on the degree of matching between input and output types, after element-wise multiplication, the values ​​of the elements in the clique expansion matrix that were originally connection relationships are multiplied by their corresponding compatibility weights. This significantly reduces the connection weights between tool pairs that are topologically connected but have incompatible interfaces, while preserving or enhancing the connection weights between tool pairs that are topologically connected and have compatible interfaces. The new matrix obtained after this element-wise multiplication operation is the weighted adjacency matrix, where each element encodes both the original hypergraph topological connectivity and the strength of interface semantic compatibility.

[0096] S340. Based on the initial hypergraph association matrix and hyperedge degree matrix, determine the first information transmission result from the node to the hyperedge.

[0097] The first information transmission result refers to the hyperedge feature representation obtained after aggregating information from the tool node along the hyperedge direction during the hypergraph convolution process.

[0098] Specifically, the attribution relationship between each tool node and each hyperedge can be extracted based on the initial hypergraph association matrix. Combined with the number of tool nodes connected by each hyperedge recorded in the hyperedge degree matrix, the initial features of all tool nodes belonging to the same hyperedge are weighted and aggregated according to their attribution relationship, thereby obtaining the aggregated feature representation corresponding to each hyperedge. This aggregation result is the first information transmission result from the node to the hyperedge.

[0099] In this embodiment, optionally, the specific implementation method for determining the first information transmission result from a node to a hyperedge based on the initial hypergraph association matrix and the hyperedge degree matrix may include: determining the hyperedge feature matrix based on the inverse matrix of the hyperedge degree matrix, the transpose of the initial hypergraph association matrix, and the initial embedding feature matrix of the tool; and using the hyperedge feature matrix as the first information transmission result.

[0100] Specifically, the first step is to obtain the inverse of the hyperedge degree matrix. The hyperedge degree matrix records the number of tool nodes connected to each hyperedge, and its inverse is used to normalize the aggregation results of each hyperedge to avoid the difference in numerical magnitude between hyperedges with different numbers of connected nodes affecting the balance of information transmission. Simultaneously, the transpose of the initial hypergraph association matrix needs to be obtained. The initial hypergraph association matrix records the affiliation relationship between each tool and each hyperedge. After transposing it, rows correspond to hyperedges and columns correspond to tools, allowing each row to easily collect information from all tool nodes belonging to that hyperedge. With these two matrices, the inverse of the hyperedge degree matrix, the transpose of the initial hypergraph association matrix, and the initial embedding feature matrix of the tools are multiplied sequentially: first, the transpose of the initial hypergraph association matrix is ​​multiplied by the initial embedding feature matrix of the tools to collect information from tool nodes to hyperedges, i.e., each hyperedge obtains the initial features of all tool nodes belonging to it; then, the multiplication is multiplied by the inverse of the hyperedge degree matrix to normalize and scale the aggregated features collected for each hyperedge. After this series of matrix multiplication operations, the resulting matrix is ​​the hyperedge feature matrix. Each row of this matrix corresponds to a hyperedge, and the value in each row represents the comprehensive feature representation of the hyperedge aggregated from all the tool nodes belonging to it. This hyperedge feature matrix is ​​directly used as the first information transmission result from the node to the hyperedge.

[0101] S350. Based on the weighted adjacency matrix and the node degree matrix, determine the second information transmission result from the hyperedge to the node.

[0102] The second information transmission result refers to the node feature representation obtained after transmitting information back to the tool node from the hyperedge in the reverse direction during the hypergraph convolution process.

[0103] Specifically, based on the connection strength between each pair of tool nodes recorded in the weighted adjacency matrix after weighting by interface compatibility, and combined with the number of connection edges of each tool node recorded in the node degree matrix, the hyperedge features obtained after the first information transmission are back propagated back to each tool node according to the connection strength in the weighted adjacency matrix. The information propagated to the same tool node is then weighted and aggregated to obtain the updated feature representation of each tool node. This aggregation result is the second information transmission result from the hyperedge to the node.

[0104] In this embodiment, optionally, the specific implementation method for determining the second information transmission result from the hyperedge to the node based on the weighted adjacency matrix and the node degree matrix may include: determining the node features after aggregation by topological and semantic dual constraints based on the negative 1 / 2 power of the node degree matrix, the weighted adjacency matrix, the negative 1 / 2 power of the node degree matrix, and the hyperedge feature matrix; and using the node features as the second information transmission result.

[0105] Specifically, the negative 1 / 2 power of the node degree matrix is ​​first obtained. The node degree matrix records the number of edges connected to each tool node in the weighted adjacency matrix. Its negative 1 / 2 power is used to perform symmetric normalization on the received results of each node during the propagation process, so as to balance the difference in information magnitude between nodes with different degrees. Next, the four components—the negative half-power of the node degree matrix, the weighted adjacency matrix, the negative half-power of the node degree matrix, and the hyperedge feature matrix—are multiplied in a specific order: First, the weighted adjacency matrix is ​​multiplied on the left and right by the negative half-power of the node degree matrix to achieve bilateral symmetric normalization of the weighted adjacency matrix, ensuring that both connection strength and node degree are considered during information propagation. Then, this bilaterally normalized result is multiplied by the hyperedge feature matrix, so that the aggregated features carried in each hyperedge are propagated back to each tool node according to the topological and semantic constraints defined in the weighted adjacency matrix. That is, tool nodes obtain more information contributions from the hyperedges of adjacent nodes with which they have connections and high interface compatibility. After this series of multiplication operations, each row of the resulting matrix corresponds to a tool node. The value in each row represents the comprehensive feature representation of the tool node, which is aggregated from all relevant hyperedges and integrates high-order collaborative topology and interface semantic compatibility constraints. This result is called the node feature after aggregation by both topological and semantic constraints, and it is directly used as the second information transmission result from hyperedge to node.

[0106] S360. Based on the first and second information transmission results, the updated enhanced graph features are determined through nonlinear transformation and learnable parameter transformation.

[0107] Specifically, the first information transmission result is the hyperedge feature matrix obtained after aggregating from nodes to hyperedges. The second information transmission result is the node feature matrix obtained after returning from the hyperedges to the nodes. Since subsequent processing requires node-level feature representations, the second information transmission result already provides the aggregated features at the node level. Based on this, a nonlinear transformation is applied to the node features, i.e., it is input into a nonlinear activation function. The nonlinear mapping of the activation function allows the features to fit more complex functional relationships. Subsequently, the nonlinearly transformed node features are transformed with a learnable parameter matrix. This learnable parameter matrix is ​​a set of parameters that is continuously updated during training. Its function is to perform a linear projection on the features, mapping the feature space to a specified dimension and adjusting the numerical distribution of the features to adapt to the needs of subsequent tasks.

[0108] After two steps of nonlinear transformation and learnable parameter transformation, the original node features are transformed into new feature representations. These feature representations not only integrate the high-order collaborative topology information inherited from the initial hypergraph association matrix, but also embed the interface compatibility constraints carried by the semantic compatibility matrix through the weighted adjacency matrix. They possess both nonlinear expressive power and learnable parameter adaptation capabilities. This final output is the updated enhanced graph feature.

[0109] Figure 5 This is a schematic diagram illustrating the implementation process of determining augmented graph features. Specifically, firstly, a pre-trained language model is used to encode and pool the textual metadata such as the names and function descriptions of each tool in the tool library. After linear projection, a vector representation of each tool is obtained. All tool vectors are then stacked vertically to obtain the initial tool embedding. After obtaining the dimension... Initial tool embedding feature matrix After (among them) (Consistent with the hidden layer dimensions of subsequent large language models), this invention designs a network layer for graph information aggregation. Traditional hypergraph convolution relies solely on topological structure; this embodiment utilizes the matrix constructed in step S1. Multiplicative injection at the graph structure level was performed.

[0110] First, define the hyperedge degree matrix. Compute the information transfer from the node to the hyperedge (Node -> Hyperedge): Subsequently, the node degree matrix is ​​defined. Calculate the information transfer from hyperedge to node (Hyperedge -> Node). Since the clique expansion matrix of the hypergraph is equivalent to... In this embodiment, the continuous semantic weight matrix Perform the Hadamard product (i.e., element-wise multiplication, denoted as) with it. This means that even if semantically incompatible tools are within the same hyperedge, their information exchange will be physically suppressed: Finally, non-linear updates of node features are performed using single-layer or multi-layer graph convolution: in, Let be the learnable parameter matrix of the graph network layer, and ReLU be the activation function. The graph features output by the network fully integrate knowledge of high-order cooperative topology and interface semantic compatibility.

[0111] S370. Encode the enhanced graph features and the user request text into prompt word tensors respectively, then concatenate them and input them into the preset inference model for autoregressive inference to generate an initial structured toolchain result containing task nodes and the topological connection relationships between task nodes.

[0112] In this embodiment, optionally, the specific implementation of generating the initial structured toolchain result containing task nodes and the topological connection relationships between task nodes may include the following steps: S3701. Input the enhanced map features into a multilayer perceptron projection module containing linear layers and activation functions.

[0113] Among them, the multilayer perceptron projection module refers to a feedforward neural network component composed of stacked linear layers and nonlinear activation functions, which is used to map input features to the target dimension space.

[0114] Specifically, after obtaining the enhanced graph features, they need to be converted into a vector form that matches the word embedding dimension of the preset inference model. This conversion is accomplished by the multilayer perceptron projection module. The multilayer perceptron projection module is a neural network submodule composed of linear layers and activation functions connected in series. The role of the linear layers is to perform linear transformations on the input features, that is, to map the input features from the original dimension to the target dimension through weight matrices and bias terms. The activation function follows the linear layers, applying a non-linear mapping to the output of the linear transformation, enabling the module to express more complex functional relationships. After the enhanced graph features are fed into the multilayer perceptron projection module, the features first undergo linear transformations of the linear layers to adjust their dimensions, and then undergo non-linear processing by the activation function. Finally, a new representation with dimensional transformation and feature enhancement is obtained at the output of the module. The core purpose of this process is to compress or expand the semantic information of the enhanced graph features to the vector dimension expected by the preset inference model, while improving the expressive power of the features through non-linear transformations, thus preparing features for the subsequent generation of graph cue word tensors with the same word embedding dimension.

[0115] S3702. Based on the projection feature vector output by the multilayer perceptron projection module, determine the graph hint word tensor with a fixed length and a dimension consistent with the word embedding dimension of the preset inference model.

[0116] Here, the projected feature vector refers to the intermediate vector representation output after the enhanced graph features are input into the projection module of the multilayer perceptron and processed by linear transformation and activation function. The graph cue word tensor refers to the numerical tensor obtained by further adjusting the projected feature vector to a fixed length and making it completely consistent with the word embedding dimension of the preset inference model, which serves as the encoding input for graph structure information.

[0117] Specifically, after the multilayer perceptron projection module performs linear transformations and activation function processing on the input augmented graph features, it generates a projection feature vector at its output. This projection feature vector already contains information about the augmented graph features after dimensional mapping and nonlinear transformation. However, the preset inference model requires the input to be a tensor of fixed length with dimensions consistent with its word embedding dimensions. Therefore, it is necessary to further determine the graph cue word tensor that meets this format requirement based on the projection feature vector. Specifically, the projection feature vector itself may be a one-dimensional vector or a sequence containing multiple feature segments. It needs to be organized into a sequence structure of a specific length, which is a pre-defined fixed value. For example, the projection feature vector may be split into several consecutive segments or repeated to form a specified number of tokens. At the same time, it is necessary to ensure that the vector dimension of each token is exactly equal to the word embedding dimension of the preset inference model. This way, each token in the graph cue word tensor can be dimensionally aligned with the word embedding vector inside the model, thus allowing it to be directly received and processed by the model. The graph cue tensor obtained after this determination process is essentially a serialized tensor composed of a fixed number of vectors, each of which has the same dimension as the word embedding dimension of the preset inference model. Its function is to serve as the encoding carrier of graph structure information at the model input level.

[0118] S3703. Input the user request text into the word embedding layer of the preset inference model, and determine the text prompt word tensor based on the word embedding vector sequence output by the word embedding layer.

[0119] Among them, the text prompt word tensor refers to the serialized tensor formed by converting each word in the text into a corresponding vector and organizing these vectors in sequence after the user request text input is embedded in the word embedding layer of the preset inference model, and using it as the encoded input of the text information.

[0120] Specifically, in addition to graph structure information, the task objective description carried in the user request text also needs to be encoded into an input format that the model can process. This encoding process is accomplished through the word embedding layer built into the predefined inference model. The user request text is input into the word embedding layer of the predefined inference model as a raw natural language string. The word embedding layer is the lowest-level component of the predefined inference model, and it maintains a lookup table that maps discrete words to a continuous vector space. The word embedding layer first segments the user request text into multiple independent words according to the model's predefined word segmentation rules. Then, for each word, it retrieves its corresponding dense vector representation from the lookup table and arranges these vectors in the order in which the words appear in the original text. Finally, it forms a sequence structure composed of multiple vectors connected end to end. This sequence structure is the word embedding vector sequence output by the word embedding layer. Based on this word embedding vector sequence, it is directly used as a text prompt word tensor. The text prompt word tensor is a tensor that exists in the form of a vector sequence. Its length is equal to the number of word units after the user request text is segmented. The dimension of each vector is equal to the word embedding dimension of the preset inference model, which completely preserves the semantic information of each word unit in the user request text and its order relationship in the text.

[0121] S3704. Concatenate the image prompt tensor to the beginning of the text prompt tensor to synthesize a complete input context.

[0122] Specifically, after obtaining the graph cue tensor and the text cue tensor separately, they need to be merged into a unified input sequence for processing by the pre-defined inference model. The merging method involves concatenating the graph cue tensor to the beginning of the text cue tensor. Concatenating to the beginning means placing all vector elements of the graph cue tensor in their original order within the tensor, followed by the first vector element of the text cue tensor immediately after the last vector element of the graph cue tensor, and so on until all vector elements of the text cue tensor are placed. This concatenation order places the graph structure information carried by the graph cue tensor at the very beginning of the entire input sequence, while the task semantic information carried by the text cue tensor follows immediately after. After this concatenation operation, the two originally independent tensors are merged into a new serialized tensor, which sequentially contains all vectors from both the graph cue tensor and the text cue tensor. This merged whole is called the complete input context. The complete input context includes both prior knowledge of the graph structure and semantic information of the user request. Furthermore, since the graph cue tensor is located at the forefront, the pre-defined inference model can prioritize the guidance of graph structure information during the autoregressive generation process.

[0123] S3705. Input the complete input context into the preset inference model, and determine the initial structured toolchain result containing task nodes and the topological connection relationships between task nodes based on the autoregressive decoding output of the preset inference model.

[0124] Specifically, after synthesizing the complete input context, it is fed into the pre-defined inference model for processing. The pre-defined inference model uses autoregressive decoding to generate output content. Autoregressive decoding means that when generating the output sequence, the model generates only one output unit at each time step and appends the newly generated output unit to the end of the existing output sequence as additional context information for generating the next output unit. This process is repeated until a termination marker is generated. Specifically, the pre-defined inference model first receives the complete input context as initial context information, and then generates the first output unit based on this context information in the first decoding time step. In each subsequent decoding time step, the model merges all previously generated output units with the original complete input context to form a new context, and generates the next output unit accordingly. As the decoding process progresses, the model gradually outputs a series of structured content, which is organized according to a pre-defined format. This content includes multiple task nodes and their related attribute information, as well as topological connections such as dependencies or calling order between these task nodes. Since this output is directly generated by the preset inference model and has not undergone any post-processing verification, it is called the initial structured toolchain result. It serves as the input for the subsequent anti-boundary parsing step and awaits further temporal validity correction.

[0125] The technical solution of this application, when generating an initial structured toolchain result containing task nodes and the topological connections between task nodes, projects enhanced graph features into a graph hint tensor with the same embedding dimension as the preset inference model, and concatenates the graph hint tensor to the front of the text hint tensor to synthesize a complete input context. On the one hand, the multilayer perceptron projection module maps the graph structure information into the same semantic space as the text word embedding, so that the graph hint tensor and the text hint tensor can be processed uniformly by the same model at the representation level, avoiding the problem of dimension mismatch. On the other hand, placing the graph hint tensor at the front of the text hint tensor allows the preset inference model to prioritize the perception of the prior topological knowledge contained in the graph structure during the autoregressive decoding process, thereby guiding the task node sequence generated by the model to better conform to the collaborative dependency relationship between tools, effectively improving the quality and rationality of the initial structured toolchain result.

[0126] S380. Parse the initial structured toolchain result, detect illegal forward references to subsequent nodes that have not been generated in the task node parameters, correct the detected illegal forward references to legal references pointing to the preceding adjacent nodes, and output an executable toolchain that conforms to the temporal causal constraints.

[0127] The technical solution of this application, when outputting enhanced graph features that integrate interface compatibility and high-order collaborative topology, obtains a weighted adjacency matrix by multiplying the initial hypergraph association matrix through clique expansion and the semantic compatibility matrix element-wise. Then, it sequentially performs the first information transfer from node to hyperedge and the second information transfer from hyperedge to node. Finally, it determines the enhanced graph features through nonlinear transformation and learnable parameter transformation. The technical effects are as follows: First, by using the semantic compatibility matrix to weight and correct the clique expansion matrix, the connection strength between incompatible tool pairs during information propagation is effectively suppressed, thereby directly injecting prior knowledge of the type system into the graph structure. Second, through the bidirectional information transfer mechanism from node to hyperedge and then to node, the full flow and integration of high-order collaborative topology information in the hypergraph structure is achieved. Third, by introducing nonlinear transformation and learnable parameter transformation, the model can adaptively adjust feature representation and capture more complex functional relationships. The final output enhanced graph features simultaneously encode interface compatibility constraints between tools, the high-order topology of historical collaboration patterns, and the textual semantics of the current task, providing rich and structured feature inputs for subsequent toolchain generation.

[0128] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

Claims

1. A toolchain planning method based on the fusion of semantically aware hypergraphs and large language models, characterized in that, The method includes: Based on the input and output type information of each tool in a tool library containing multiple different functional interfaces, a segmented assignment strategy combining string full matching and semantic space similarity calculation is adopted to construct a semantic compatibility matrix that reflects the degree of interface compatibility between tools. High-frequency co-occurring tool combinations are extracted from the historical successful tool call trajectories as hyperedges, and an initial hypergraph association matrix representing high-order collaborative relationships between tools is constructed based on the hyperedges; The initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix. Based on the corrected weighted hypergraph structure, graph information is aggregated for the initial features of the tool and the textual features of the user request text, and an enhanced graph feature that integrates interface compatibility and high-order collaborative topology is output. The enhanced graph features and the user request text are encoded into prompt word tensors, which are then concatenated and input into a preset inference model for autoregressive inference, generating an initial structured toolchain result containing task nodes and the topological connection relationships between task nodes. The initial structured toolchain result is parsed, illegal forward references to subsequent nodes that have not yet been generated are detected in the task node parameters, the detected illegal forward references are corrected to legal references pointing to the preceding adjacent nodes, and an executable toolchain that conforms to the temporal causal constraints is output.

2. The method according to claim 1, characterized in that, Before constructing the semantic compatibility matrix that reflects the degree of interface compatibility between tools, the following steps are also included: Convert the input type set and output type set of each tool in the tool library into text strings respectively; The text string is input into a lightweight pre-trained language model for feature encoding; Based on the feature vectors output by the lightweight pre-trained language model, determine the input semantic vectors corresponding to the input type set and the output semantic vectors corresponding to the output type set; The semantic similarity between the tools is determined based on the cosine similarity between the input semantic vector and the output semantic vector.

3. The method according to claim 2, characterized in that, Based on the input and output type information of each tool in the tool library containing multiple different functional interfaces, a segmented assignment strategy combining complete string matching and semantic space similarity calculation is used to construct a semantic compatibility matrix reflecting the degree of interface compatibility between tools, including: If the output type set of the source tool and the input type set of the target tool have an intersection at the string level, then the compatibility weight between the source tool and the target tool is assigned a preset maximum value. If the output type set and the input type set do not intersect at the string level, and the semantic similarity is greater than a preset threshold, then the compatibility weight is assigned a continuous floating-point value of the semantic similarity. If the output type set and the input type set do not intersect at the string level, and the semantic similarity is less than or equal to the preset threshold, then the compatibility weight is assigned a preset minimum positive value. Based on the assigned compatibility weights among the tools, a semantic compatibility matrix reflecting the degree of interface compatibility between the tools is constructed.

4. The method according to claim 1, characterized in that, The process of extracting frequently co-occurring tool combinations from historically successful tool call trajectories as hyperedges, and constructing an initial hypergraph association matrix representing high-order collaborative relationships between tools based on these hyperedges, includes: Filter out invalid nodes that are not in the current tool library from the historical successfully executed tool call traces to obtain valid tool call traces; Based on a preset minimum support threshold, a frequent itemset mining algorithm is used to extract high-frequency co-occurring tool combinations from the effective tool call trajectories, and each high-frequency co-occurring tool combination is determined as a hyperedge. Based on all hyperedges, construct an initial hypergraph association matrix with a dimension of the total number of tools multiplied by the total number of hyperedges; Specifically, when the tool belongs to the hyperedge, the matrix element takes the first preset value, and when the tool does not belong to the hyperedge, the matrix element takes the second preset value.

5. The method according to claim 1, characterized in that, The initial hypergraph association matrix is ​​weighted and corrected using the semantic compatibility matrix. Based on the corrected weighted hypergraph structure, graph information is aggregated between the initial features of the tool and the textual features of the user request text. The result is an enhanced graph feature that integrates interface compatibility and higher-order collaborative topology, including: The clique expansion matrix is ​​calculated based on the initial hypergraph association matrix. The clique expansion matrix is ​​then multiplied element-wise with the semantic compatibility matrix to obtain the weighted adjacency matrix. Based on the initial hypergraph association matrix and hyperedge degree matrix, determine the first information transmission result from the node to the hyperedge; Based on the weighted adjacency matrix and node degree matrix, the second information transmission result from the hyperedge to the node is determined; Based on the first information transmission result and the second information transmission result, the updated enhanced graph features are determined through nonlinear transformation and learnable parameter transformation.

6. The method according to claim 5, characterized in that, The determination of the first information transmission result from a node to a hyperedge based on the initial hypergraph association matrix and hyperedge degree matrix includes: The hyperedge feature matrix is ​​determined based on the inverse of the hyperedge degree matrix, the transpose of the initial hypergraph association matrix, and the initial embedding feature matrix of the tool. The hyperedge feature matrix is ​​used as the result of the first information transmission.

7. The method according to claim 5, characterized in that, The determination of the second information transmission result from the hyperedge to the node based on the weighted adjacency matrix and the node degree matrix includes: Based on the negative 1 / 2 power of the node degree matrix, the weighted adjacency matrix, the negative 1 / 2 power of the node degree matrix, and the hyperedge feature matrix, the node features after aggregation under both topological and semantic constraints are determined. The node features are used as the result of the second information transmission.

8. The method according to claim 1, characterized in that, The process involves encoding the enhanced graph features and the user request text into prompt word tensors, concatenating them, and inputting them into a preset inference model for autoregressive inference. This generates an initial structured toolchain result containing task nodes and the topological connections between task nodes, including: The enhanced map features are input into a multilayer perceptron projection module containing linear layers and activation functions; Based on the projection feature vector output by the multilayer perceptron projection module, a graph cue word tensor with a fixed length and a dimension consistent with the word embedding dimension of the preset inference model is determined. The user request text is input into the word embedding layer of the preset inference model. Based on the word embedding vector sequence output by the word embedding layer, the text prompt word tensor is determined. The graph prompt tensor is concatenated to the beginning of the text prompt tensor to synthesize a complete input context. The complete input context is input into the preset inference model, and based on the autoregressive decoding output of the preset inference model, an initial structured toolchain result containing task nodes and the topological connection relationships between task nodes is determined.

9. The method according to claim 1, characterized in that, The process of parsing the initial structured toolchain result, detecting illegal forward references to subsequent nodes that have not yet been generated in the task node parameters, correcting the detected illegal forward references to legal references to the preceding adjacent nodes, and outputting an executable toolchain that conforms to temporal causality constraints includes: The initial structured toolchain result is parsed to obtain a task node sequence, and each task node in the task node sequence is traversed. For the task node at the current index position, extract the reference index from the parameter field of the task node based on regular expressions; If the reference index is greater than or equal to the current index position and the value of the current index position minus one is greater than or equal to zero, then the reference index is rewritten to the value of the current index position minus one, resulting in a corrected valid reference; If the reference index is less than the current index position, the original reference index is retained unchanged; If the reference index is greater than or equal to the current index position and the value of the current index position minus one is less than zero, then the current task node will be treated as an independent node and its reference field will be set to null. Based on the correction and retention results of the reference index, an executable toolchain that conforms to the temporal causal constraint is output.

10. The method according to any one of claims 1 to 9, characterized in that, The executable toolchain that conforms to the temporal causal constraints is a structured representation that includes a sequence of task nodes and topological dependencies between nodes; Each task node includes a node index identifier, tool call parameters, and a preceding node reference field. This structured representation is used to directly drive the automatic invocation and execution of the corresponding functional interfaces in the tool library.