Multi-modal scene RAG tool dynamic enhancement agent building method and system

By constructing a tool vector space and dependency graph, and combining context awareness and four-dimensional scoring, tools are dynamically selected and updated, solving the problem of low efficiency in RAG tool selection and realizing efficient tool recommendation and use by agents in multimodal environments.

CN120821756AActive Publication Date: 2025-10-21ZHEJIANG SHUXIN NETWORK CO LTD

Patent Information

Application Number
CN202511318260.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-21
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

The existing RAG tool selection mechanism lacks context awareness and cannot dynamically adjust the tool selection strategy according to the specific scenario and environmental constraints of the task. This results in low tool selection efficiency, failure to maximize resource value, and lack of tool usage feedback mechanism, making it difficult to learn and optimize from historical experience.

Method used

By generating composite vectors for tools and constructing a tool vector space, we use context-aware retrieval vectors to perform approximate nearest neighbor search. Combining tool dependency graphs and rule filtering, we select target tools using four-dimensional scoring weights and perform incremental updates using statistical vectors through the tools, and establish a standardized calling interface.

Benefits of technology

It enables efficient tool recommendation and use by intelligent agents in multimodal environments, improves the accuracy and comprehensiveness of tool retrieval, and enhances the adaptability and task processing performance of intelligent agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821756A_ABST
    Figure CN120821756A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal scene RAG tool dynamic enhancement agent construction method and system, and relates to the technical field of scene construction, and the method comprises the steps: generating a tool composite vector, constructing a context awareness retrieval vector to carry out tool retrieval, selecting an optimal tool in combination with a tool dependence graph and a four-dimensional scoring mechanism, and constructing a multi-modal scene RAG tool dynamic enhancement agent. And finally, updating the tool vector based on execution result feedback. According to the method, the optimal tool combination can be adaptively selected, the execution efficiency and accuracy of the intelligent agent in a dynamic task environment are improved, and resource consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to scene construction technology, and in particular to a method and system for dynamically enhancing intelligent body construction using RAG tools in multimodal scenes. Background Art

[0002] With the rapid development of artificial intelligence (AI), Retrieval-Augmented Generation (RAG) has become a key core technology in intelligent agent systems. In multimodal scenarios, agents must process diverse information types, such as text, images, and audio, and use appropriate tools to complete complex tasks. Existing RAG tool systems primarily rely on predefined tool sets, with agents selecting the appropriate tool to perform tasks based on user requests.

[0003] Traditional tool selection methods typically rely on keyword matching or simple similarity calculations. These single-minded tool vector representations fail to fully capture the diverse characteristics of tools. Furthermore, dependencies and synergies between tools are often overlooked in traditional systems, making it difficult to develop effective tool combination strategies.

[0004] The existing RAG tool selection mechanism lacks contextual awareness and is unable to dynamically adjust the tool selection strategy according to the specific scenarios and environmental constraints of the task, resulting in inefficient tool selection in complex and changing scenarios and failure to accurately meet the actual needs of users.

[0005] The existing technology has a single dimension for tool evaluation, and mainly relies on functional relevance for tool selection, ignoring multi-dimensional evaluation indicators such as tool performance consumption, execution efficiency, and value return, and is unable to maximize the value of tool use under limited resources.

[0006] The existing RAG system lacks a tool usage feedback mechanism and is unable to dynamically update tool representations based on historical execution data. This makes it difficult for the system to learn and optimize from historical experience, and is unable to continuously improve the accuracy and efficiency of tool selection during usage. Summary of the Invention

[0007] The embodiments of the present invention provide a method and system for building a RAG tool dynamically enhanced intelligent agent in a multimodal scenario, which can solve the problems in the prior art.

[0008] A first aspect of an embodiment of the present invention provides a method for dynamically enhancing an intelligent agent using a RAG tool in a multimodal scenario, comprising: Generate a tool composite vector and store it in the tool vector space for subsequent tool retrieval, receive the agent's task request, and construct a context-aware retrieval vector based on the task request; performing an approximate nearest neighbor search in the tool vector space using the context-aware retrieval vector to obtain an initial tool of a candidate tool set, performing a graph traversal in a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, performing rule filtering on the initial tool and the associated tools based on current environmental constraints, and updating the candidate tool set; Determine the four-dimensional scoring weights based on the scenario type of the current task, collect the real-time operating indicators of each tool in the candidate tool set and calculate the scores of each dimension, perform a weighted sum of the scores of each dimension and the four-dimensional scoring weights to obtain the multidimensional utility score of each tool, and select the tool with the highest multidimensional utility score as the target tool; The task request is converted into a standardized call request according to the input interface mode of the target tool, and the target tool is called using the standardized call request to obtain the execution result; the performance value consumption index of the target tool during the execution process is collected, and the tool usage statistical vector of the target tool is updated based on the execution result and the performance value consumption index, and the updated tool usage statistical vector is used to generate a new tool composite vector.

[0009] Performing an approximate nearest neighbor search in the tool vector space using the context-aware retrieval vector to obtain initial tools of a candidate tool set includes: Constructing an initial retrieval feature based on a context-aware retrieval vector, mapping the context-aware retrieval vector to a tool vector space to establish a tool knowledge graph, and performing knowledge enhancement on the context-aware retrieval vector based on the tool knowledge graph to obtain an enhanced retrieval vector; constructing an approximate nearest neighbor retrieval index in the tool vector space, performing an approximate nearest neighbor search on the enhanced retrieval vector based on the approximate nearest neighbor retrieval index, and calculating a similarity score between the enhanced retrieval vector and the tool vector; The tool vectors are sorted according to the similarity scores, and tools corresponding to a preset number of tool vectors with the highest similarity scores are selected as initial tools of the candidate tool set.

[0010] Mapping the context-aware retrieval vector to a tool vector space to establish a tool knowledge graph, and performing knowledge enhancement on the context-aware retrieval vector based on the tool knowledge graph to obtain an enhanced retrieval vector includes: Constructing a dynamic mapping matrix, wherein the dynamic mapping matrix maps the context-aware retrieval vector to a tool vector space to obtain a tool vector, obtaining a mapping difference between a reference tool vector and the tool vector to generate mapping update information, and superimposing the mapping update information on the dynamic mapping matrix to update the dynamic mapping matrix; Constructing a tool knowledge graph in the tool vector space, performing mapping conversion on adjacent tool nodes in the tool knowledge graph using the dynamic mapping matrix, and multiplying the conversion result by the association information of the adjacent tool nodes to construct an inter-node association strength value in the tool knowledge graph; The context-aware retrieval vector is subjected to knowledge enhancement, the tool vector is used as a starting node, the associated node vector is searched in the tool knowledge graph using the dynamic mapping matrix, and the association strength value is used as a weight coefficient, the weight coefficient is changed with the association strength value, and the weighted associated node vector is fused with the context-aware retrieval vector to obtain an enhanced retrieval vector.

[0011] Traversing a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, and filtering the initial tool and the associated tools based on current environment constraints include: Using the initial tool as the starting node to traverse the tool dependency graph, calculating the node transition probability of each traversal step, traversing along the edges of the tool dependency graph based on the node transition probability, and grouping the tool nodes encountered during the traversal into an associated tool set; Establishing a rule template library, storing a plurality of rule templates in the rule template library, each rule template including rule structure information and combination pattern information, inputting the combination of the initial tool and the tools in the associated tool set in historical execution records into each rule template, calculating the impact degree and complexity score of each rule template on the tool combination, and taking the weighted sum of the impact degree and the complexity score as the rule importance weight; The rule templates in the rule template library are sorted and screened according to the rule importance weights, and the rule templates with the highest rule importance weights are selected to form a rule set. The initial tool and each tool combination in the associated tool set are input into the rule set for rule verification, and the current environmental constraints are extracted to filter the verification results.

[0012] Converting the task request into a standardized call request according to the input interface mode of the target tool, and using the standardized call request to call the target tool and obtain the execution result includes: Constructing a multimodal parsing engine, using the multimodal parsing engine to perform feature extraction on the task request to generate an initial feature vector, and inputting the initial feature vector into the multimodal parsing engine again for deep feature extraction to obtain a request feature vector with a hierarchical semantic representation; Performing semantic feature parsing on the request feature vector based on the input interface mode of the target tool, mapping the hierarchical semantic representation in the request feature vector to a standard semantic space according to the multimodal parsing engine, and generating standardized semantic features containing normalized semantic information; Performing semantic consistency verification on the normalized semantic information in the standardized semantic feature, calculating the semantic similarity between the request feature vector and the standardized semantic feature, and converting the verified standardized semantic feature into a standardized call request according to the input interface mode of the target tool when the semantic similarity is greater than a preset similarity threshold; The standardized call request is transmitted to the target tool for calling, and the call execution result of the target tool is collected.

[0013] Updating the tool usage statistics vector of the target tool based on the execution result and the performance value consumption indicator, and using the updated tool usage statistics vector to generate a new tool composite vector includes: Based on the execution results and the performance value consumption index, the tool usage statistics vector of the target tool is progressively updated. The progressive update is achieved by weighted combination of the original statistics vector and the current execution statistics information. The weight of the weighted combination is determined by the historical information retention coefficient. The historical information retention coefficient dynamically decays with the number of executions. An update vector is generated based on the tool usage statistical vector after progressive update. The update vector is obtained by progressively fusing the tool usage statistical vector with environmental context information. The progressive fusion adopts the same weight coefficient as the progressive update to generate a new tool composite vector based on the update vector.

[0014] A second aspect of an embodiment of the present invention provides a RAG tool dynamically enhanced agent building system for multimodal scenarios, including: The first unit is configured to generate a tool composite vector and store it in a tool vector space for subsequent tool retrieval, receive a task request from an agent, and construct a context-aware retrieval vector based on the task request; a second unit configured to use the context-aware retrieval vector to perform an approximate nearest neighbor search in the tool vector space to obtain an initial tool of a candidate tool set, perform a graph traversal in a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, perform rule filtering on the initial tool and the associated tools based on current environment constraints, and update the candidate tool set; The third unit is configured to determine a four-dimensional scoring weight based on the scenario type of the current task, collect real-time operating indicators of each tool in the candidate tool set and calculate scores in each dimension, perform a weighted sum of the scores in each dimension and the four-dimensional scoring weight to obtain a multidimensional utility score for each tool, and select the tool with the highest multidimensional utility score as the target tool; The fourth unit is used to convert the task request into a standardized call request according to the input interface mode of the target tool, use the standardized call request to call the target tool and obtain the execution result; collect the performance value consumption index of the target tool during the execution process, update the tool usage statistical vector of the target tool based on the execution result and the performance value consumption index, and use the updated tool usage statistical vector to generate a new tool composite vector.

[0015] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0016] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0017] The beneficial effects of this application are as follows: The present invention dynamically enhances the intelligent agent building method through RAG tools in multimodal scenarios, realizes intelligent tool recommendation and use, can dynamically select the most suitable tool according to the task scenario, and effectively improves the task execution efficiency of the intelligent agent in complex multimodal environments.

[0018] The present invention constructs a tool composite vector and tool dependency graph, which fuses and represents multi-dimensional information such as the tool's semantic description, interface definition, usage statistics, and technical characteristics. Combined with context-aware retrieval and graph traversal technology, it can not only accurately match the tools required by the current task, but also discover potential related tools, significantly enhancing the accuracy and comprehensiveness of tool retrieval.

[0019] This invention designs a tool utility evaluation mechanism based on four-dimensional scoring, dynamically adjusts the scoring weights to adapt to the needs of different scenarios, and continuously optimizes the tool selection strategy by collecting tool operation indicators in real time. At the same time, it establishes a standardized tool calling interface, so that the intelligent agent can seamlessly connect with various tools, greatly improving the adaptability and task processing effect of the intelligent agent in a multimodal environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the process of dynamically enhancing the intelligent agent building method of the RAG tool in a multimodal scenario according to an embodiment of the present invention; Figure 2 This is a flowchart of tool rule verification and filtering according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0022] The technical solution of the present invention is described in detail below with reference to specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0023] Figure 1 Schematic diagram of the process of dynamically enhancing the intelligent agent building method of the RAG tool in the multimodal scene of the embodiment of the present invention, such as Figure 1 As shown, the method includes: Generate a tool composite vector and store it in the tool vector space for subsequent tool retrieval, receive the agent's task request, and construct a context-aware retrieval vector based on the task request; performing an approximate nearest neighbor search in the tool vector space using the context-aware retrieval vector to obtain an initial tool of a candidate tool set, performing a graph traversal in a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, performing rule filtering on the initial tool and the associated tools based on current environmental constraints, and updating the candidate tool set; Determine the four-dimensional scoring weights based on the scenario type of the current task, collect the real-time operating indicators of each tool in the candidate tool set and calculate the scores of each dimension, perform a weighted sum of the scores of each dimension and the four-dimensional scoring weights to obtain the multidimensional utility score of each tool, and select the tool with the highest multidimensional utility score as the target tool; The task request is converted into a standardized call request according to the input interface mode of the target tool, and the target tool is called using the standardized call request to obtain the execution result; the performance value consumption index of the target tool during the execution process is collected, and the tool usage statistical vector of the target tool is updated based on the execution result and the performance value consumption index, and the updated tool usage statistical vector is used to generate a new tool composite vector.

[0024] In an optional embodiment, using the context-aware retrieval vector to perform an approximate nearest neighbor search in the tool vector space to obtain initial tools of the candidate tool set includes: Constructing an initial retrieval feature based on a context-aware retrieval vector, mapping the context-aware retrieval vector to a tool vector space to establish a tool knowledge graph, and performing knowledge enhancement on the context-aware retrieval vector based on the tool knowledge graph to obtain an enhanced retrieval vector; constructing an approximate nearest neighbor retrieval index in the tool vector space, performing an approximate nearest neighbor search on the enhanced retrieval vector based on the approximate nearest neighbor retrieval index, and calculating a similarity score between the enhanced retrieval vector and the tool vector; The tool vectors are sorted according to the similarity scores, and tools corresponding to a preset number of tool vectors with the highest similarity scores are selected as initial tools of the candidate tool set.

[0025] When constructing initial retrieval features based on the context-aware retrieval vector, the system receives a user query request, such as "I want to analyze recent sales data and generate a chart report." This text input is encoded using a pretrained language model, such as an encoder with a 12-layer transformer architecture, an input sequence length of 128, and a hidden layer dimension of 768. The encoding process considers the semantic representation of the key terms "analysis," "sales data," "chart," and "report" in the query text, generating an initial vector representation. This vector representation is then mapped through a fully connected layer into a 512-dimensional initial retrieval feature vector for subsequent retrieval operations in the tool vector space.

[0026] The system employs dimensional alignment to map context-aware search vectors to the tool vector space to create a tool knowledge graph. Assuming the tool vector space is a 384-dimensional space, the system uses a projection matrix to convert the 512-dimensional context-aware search vectors into a 384-dimensional representation. This projection matrix is ​​obtained through supervised learning on 20,000 query-tool matching samples. The mapped vectors serve as nodes in the graph, establishing connections with tool nodes in the tool vector space. The system assigns attributes to each tool node, including functional descriptions, parameter information, and usage scenarios. For example, a data analysis tool node includes attributes such as "supports pivot tables" and "visualization chart types." Edges between nodes are established based on functional similarity and call dependencies. Edge weights reflect the strength of the relationship and range from 0 to 1.

[0027] When enhancing context-aware search vectors based on the tool knowledge graph, the system utilizes a graph attention network mechanism. Starting from the initial search vector, the system explores relevant nodes within two hops along the knowledge graph and collects their representation vectors. For the query example above, the system identifies three core tool nodes: "data analysis tools," "visualization components," and "report generator," and obtains their representation vectors. These relevant vectors are then weighted and aggregated, with weights calculated based on their semantic relevance to the initial search vector. For example, "data analysis tools" has a weight of 0.6, "visualization components" has a weight of 0.3, and "report generator" has a weight of 0.1. By combining the initial search vector with the weighted relevant node vectors, a richer enhanced search vector is generated, while also maintaining a 384-dimensional representation.

[0028] To construct an approximate nearest neighbor search index in the tool vector space, a hierarchical navigation graph index structure was used. This index, constructed for 10,000 tool vectors, consists of three layers. The top layer contains approximately 50 anchor nodes, the middle layer contains approximately 500 nodes, and the bottom layer contains all tool nodes. By precalculating inter-node distances and establishing adjacency relationships, the index supports efficient retrieval of large vector collections. Once constructed, the index supports sublinear time complexity, with an average query time of approximately 5 milliseconds on the 384-dimensional vector space.

[0029] When performing an approximate nearest neighbor search on an enhanced search vector based on the approximate nearest neighbor search index, the system employs a hierarchical navigation strategy. Starting at the top layer, the distance between the query vector and the anchor nodes is calculated, and the five closest anchor nodes are selected to proceed to the next layer of search. In the middle layer, the distance to the query vector is further calculated for each incoming node, and the 15 closest nodes are selected to proceed to the bottom layer of search. At the bottom layer, the system calculates the cosine similarity between the query vector and the candidate tool nodes, with values ​​ranging from -1 to 1, where larger values ​​indicate greater similarity.

[0030] For the preceding query, the system finds several potential matching tools, such as "PivotTable Analyzer" (similarity 0.92), "Business Intelligence Dashboard" (similarity 0.89), "Data Visualization Builder" (similarity 0.85), and "Report Automation Tool" (similarity 0.83). The system calculates the similarity scores between all tool vectors and the enhanced search vector and records them in the result list.

[0031] When sorting tool vectors by similarity score, the scores are sorted in descending order, and the top N tools with the highest similarity scores are selected as the initial tools in the candidate tool set. In this example, N is set to 5, and the system selects "Pivot Analyzer," "Business Intelligence Dashboard," "Data Visualization Builder," "Report Automation Tool," and "Spreadsheet Analysis Plugin" (with a similarity of 0.81) as the initial tools in the candidate tool set. These tools serve as the basis for subsequent tool selection and combination to meet users' data analysis and report generation needs.

[0032] Through the implementation of the above technical means, the system can intelligently recommend the most suitable tool set based on the user's natural language query, improving the accuracy and relevance of tool recommendations, enabling the artificial intelligence assistant to respond to user needs more accurately.

[0033] In an optional embodiment, mapping the context-aware retrieval vector to a tool vector space to establish a tool knowledge graph, and performing knowledge enhancement on the context-aware retrieval vector based on the tool knowledge graph to obtain an enhanced retrieval vector includes: Constructing a dynamic mapping matrix, wherein the dynamic mapping matrix maps the context-aware retrieval vector to a tool vector space to obtain a tool vector, obtaining a mapping difference between a reference tool vector and the tool vector to generate mapping update information, and superimposing the mapping update information on the dynamic mapping matrix to update the dynamic mapping matrix; Constructing a tool knowledge graph in the tool vector space, performing mapping conversion on adjacent tool nodes in the tool knowledge graph using the dynamic mapping matrix, and multiplying the conversion result by the association information of the adjacent tool nodes to construct an inter-node association strength value in the tool knowledge graph; The context-aware retrieval vector is subjected to knowledge enhancement, the tool vector is used as a starting node, the associated node vector is searched in the tool knowledge graph using the dynamic mapping matrix, and the association strength value is used as a weight coefficient, the weight coefficient is changed with the association strength value, and the weighted associated node vector is fused with the context-aware retrieval vector to obtain an enhanced retrieval vector.

[0034] A dynamic mapping matrix is ​​constructed to map the context-aware retrieval vector to the tool vector space. For example, consider a real-world application scenario where a context-aware retrieval vector represents a user's search for "how to handle attachments in emails." This context-aware retrieval vector can be represented as a multidimensional vector, for example, a 256-dimensional vector. The tool vector space is a vector representation of various tool functions, such as "download attachment," "preview attachment," and "forward attachment." The initial values ​​of the dynamic mapping matrix can be obtained through pre-training, for example, by training with a large amount of user query and tool usage data.

[0035] The user's current context-aware retrieval vector is obtained and mapped to the tool vector space through a dynamic mapping matrix to obtain a tool vector. At the same time, the system obtains a reference tool vector, which can be the optimal tool vector determined based on the user's historical behavior data. The system calculates the mapping difference between the tool vector and the reference tool vector and generates mapping update information. For example, if the cosine similarity between the mapped tool vector and the reference tool vector is lower than a preset threshold (such as 0.8), it is considered that there is a mapping difference and the dynamic mapping matrix needs to be updated. The mapping update information can be the difference vector between the two vectors. This difference vector is superimposed on the dynamic mapping matrix according to a certain weight (such as 0.1) to achieve the update of the dynamic mapping matrix. This update method can enable the mapping matrix to dynamically adjust according to the user's usage, thereby improving the accuracy of the mapping.

[0036] A tool knowledge graph is constructed in the tool vector space. The tool knowledge graph contains multiple tool nodes, each representing a tool function, and there are associations between the nodes. For example, the two tool nodes "Download Attachment" and "Preview Attachment" are associated because they are both related to attachment processing. The system uses a dynamic mapping matrix to map and transform adjacent tool nodes in the tool knowledge graph, multiplying the transformation result by the association information of the adjacent tool nodes to obtain the strength of the association between the nodes. The association information can be a predefined association type and strength. For example, the association type between "Download Attachment" and "Preview Attachment" is "Attachment Processing", and the initial association strength is 0.7. After the mapping transformation is performed using the dynamic mapping matrix, if the conversion result has a high similarity with the vector representation of the target node, the association strength is increased; if the similarity is low, the association strength is weakened.

[0037] For example, suppose the tool vector for the "Download Attachment" node is V1, and the tool vector for the "Preview Attachment" node is V2. The initial correlation strength between them is 0.7. V1 is mapped and transformed using the dynamic mapping matrix M, resulting in the mapping result M(V1). The cosine similarity between M(V1) and V2 is calculated. Assuming the result is 0.85, the new correlation strength value can be updated by taking the weighted average of 0.7 and 0.85, for example, to 0.75. In this way, through the dynamic mapping matrix, the node correlation strength in the tool knowledge graph is dynamically adjusted based on actual usage.

[0038] Knowledge enhancement is performed on the context-aware retrieval vector, using the tool vector as the starting node and searching for associated node vectors in the tool knowledge graph. Continuing with the previous example, after the user query "How to handle attachments in emails" is mapped to a tool vector, assuming the best matching tool node is "Download attachment," the system uses the "Download attachment" node as the starting node and searches for other associated nodes, such as "Preview attachment" and "Forward attachment." The system uses a dynamic mapping matrix to search for associated node vectors in the tool knowledge graph, using the association strength as a weighting factor. For example, if the association strength between "Download attachment" and "Preview attachment" is 0.75, then 0.75 will be used as the weighting factor when fusing the "Preview attachment" node vector.

[0039] All associated node vectors are weighted and summed according to their association strength to produce a composite vector. For example, assuming that the associated nodes for "download attachment" are "preview attachment" (association strength 0.75) and "forward attachment" (association strength 0.6), the composite vector can be expressed as: 0.75 × V2 + 0.6 × V3 (where V2 is the vector representation of "preview attachment" and V3 is the vector representation of "forward attachment"). The system then fuses this composite vector with the original context-aware retrieval vector to produce an enhanced retrieval vector. This fusion can be performed using simple vector addition, weighted averaging, or more complex methods such as attention mechanisms.

[0040] Through the above steps, the context-aware retrieval vector is mapped to the tool vector space and enriched with knowledge based on the tool knowledge graph, resulting in an enhanced retrieval vector. This enhanced retrieval vector incorporates both the original query information and relevant tool knowledge and can be used in subsequent retrieval tasks to improve retrieval precision and relevance. In practical applications, this approach can help the system better understand user queries, recommend more appropriate tools and operations, and enhance the user experience.

[0041] In an optional embodiment, traversing a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, and filtering the initial tool and the associated tools based on current environment constraints includes: Using the initial tool as the starting node to traverse the tool dependency graph, calculating the node transition probability of each traversal step, traversing along the edges of the tool dependency graph based on the node transition probability, and grouping the tool nodes encountered during the traversal into an associated tool set; Establishing a rule template library, storing a plurality of rule templates in the rule template library, each rule template including rule structure information and combination pattern information, inputting the combination of the initial tool and the tools in the associated tool set in historical execution records into each rule template, calculating the impact degree and complexity score of each rule template on the tool combination, and taking the weighted sum of the impact degree and the complexity score as the rule importance weight; The rule templates in the rule template library are sorted and screened according to the rule importance weights, and the rule templates with the highest rule importance weights are selected to form a rule set. The initial tool and each tool combination in the associated tool set are input into the rule set for rule verification, and the current environmental constraints are extracted to filter the verification results.

[0042] like Figure 2 As shown, the method includes: When traversing the tool dependency graph using the initial tool as the starting node, a random walk algorithm is used to calculate node transition probabilities. For tool node A, the transition probability of its adjacent node B is calculated as: the weight of the edge from A to B divided by the sum of the weights of all edges originating from A. For example, if tool A is connected to tools B, C, and D, with edge weights of 0.5, 0.3, and 0.2, respectively, then the transition probability from A to B is 0.5 / (0.5+0.3+0.2)=0.5, the transition probability from A to C is 0.3 / 1.0=0.3, and the transition probability from A to D is 0.2 / 1.0=0.2.

[0043] During the actual traversal process, starting from the initial tool node, the next tool node to be visited is selected based on the node transition probability. To avoid local loops, a random restart probability of 0.15 is set, meaning there is a 0.15 probability of returning to the initial tool node and restarting the traversal. The traversal depth is set to 5, meaning a maximum of 5 jumps. Each tool node encountered during the traversal is added to the associated tool set. To increase the diversity of the results, 10 random walks are performed, all encountered tool nodes are merged, and sorted by access frequency. The top 15 most frequently accessed tools are selected to form the final associated tool set.

[0044] Taking language translation tools as an initial tool as an example, the associated tools obtained through graph traversal include: speech recognition, text editing, grammar checking, vocabulary expansion, content generation, semantic analysis, image recognition, multimedia processing, document format conversion and other tools.

[0045] When creating a rule template library, the stored rule templates contain rule structure information and combination mode information. Rule structure information describes the logical structure of the rule, such as "If tool A is used, tool B must / must not be used" or "When tool A and tool B are used together, parameter X must be set to Y." Combination mode information describes the usage mode of the tool combination, such as "serial execution," "parallel execution," and "conditional triggering."

[0046] Each rule template contains at least the following fields: rule ID, rule description, rule condition, rule result, applicable tool type, combination mode, and priority. For example, a rule template might be: {Rule ID: "R001", Rule Description: "When using a text processing tool and a language model tool together, the text length must be less than the system limit", Rule Condition: "Tool A Type = Text Processing AND Tool B Type = Language Model", Rule Result: "Check that the input text length is less than or equal to the system limit", Applicable Tool Types: ["Text Processing", "Language Model"], Combination Mode: "Serial Execution", Priority: "High"}.

[0047] To calculate the impact of a rule template on a tool combination, we analyze the historical execution records for the combination of the initial tool and associated tools. Specifically, we extract metrics such as the success rate, user satisfaction, and execution efficiency of the tool combination from the historical records. We then calculate the similarity between the current tool combination and the historical records, and then weight the impact based on this similarity. For example, if the initial tool is "Language Translation" and the associated tool is "Grammar Check," and historical records show a 95% success rate, a 4.8 / 5 user satisfaction rating, and a 30% improvement in execution efficiency, the impact is 0.95 × 0.4 + 4.8 / 5 × 0.3 + 0.3 × 0.3 = 0.766.

[0048] The complexity score calculation takes into account factors such as the complexity of the rule structure, the number of tools involved, and the number of parameters. The higher the complexity, the more difficult it is to apply in practice, and rules with lower complexity are preferred. For example, if a rule involves two tools, three parameters, and a simple conditional structure, the complexity score is 0.3; if a rule involves five tools, 10 parameters, and a structure with multiple levels of nested conditions, the complexity score is 0.8.

[0049] The rule importance weight is calculated by weighting the impact and complexity scores: Rule importance weight = Impact × 0.7 - Complexity score × 0.3. For example, if the impact is 0.766 and the complexity score is 0.3, the rule importance weight is 0.766 × 0.7 - 0.3 × 0.3 = 0.446.

[0050] Rule templates are sorted and filtered based on their importance weights, and the top five rule templates with the highest weights are selected to form a rule set. Each tool combination in the initial tool and associated tool set is entered into the rule set for rule validation to check whether the tool combination meets the rule conditions and meets the rule result requirements.

[0051] When filtering validation results based on current environmental constraints, consider environmental factors such as system resource limitations, user permissions, network conditions, and security requirements. For example, in a low-bandwidth network environment, filter out tool combinations that require large amounts of data transfer; on resource-constrained devices, filter out computationally intensive tool combinations; and for users with limited permissions, filter out tool combinations that require advanced permissions.

[0052] For example, if the current environment is a mobile device with average network conditions, rule verification results indicate that this combination's performance degrades by 50% under high network latency. The environmental constraints include "network latency < 100ms" and "device memory > 2GB." Based on the current network latency (120ms) and device memory (3GB), this tool combination is filtered out because it does not meet the network latency requirements.

[0053] Through the above rule filtering process, the most appropriate tool combination that meets the current environment constraints can be screened out from the associated tool set of the initial tool, thereby improving the applicability and execution efficiency of the tool combination.

[0054] In an optional embodiment, converting the task request into a standardized call request according to the input interface mode of the target tool, and using the standardized call request to call the target tool and obtain the execution result includes: Constructing a multimodal parsing engine, using the multimodal parsing engine to perform feature extraction on the task request to generate an initial feature vector, and inputting the initial feature vector into the multimodal parsing engine again for deep feature extraction to obtain a request feature vector with a hierarchical semantic representation; Performing semantic feature parsing on the request feature vector based on the input interface mode of the target tool, mapping the hierarchical semantic representation in the request feature vector to a standard semantic space according to the multimodal parsing engine, and generating standardized semantic features containing normalized semantic information; Performing semantic consistency verification on the normalized semantic information in the standardized semantic feature, calculating the semantic similarity between the request feature vector and the standardized semantic feature, and converting the verified standardized semantic feature into a standardized call request according to the input interface mode of the target tool when the semantic similarity is greater than a preset similarity threshold; The standardized call request is transmitted to the target tool for calling, and the call execution result of the target tool is collected.

[0055] A multimodal parsing engine was constructed, utilizing a deep neural network architecture consisting of a feature extraction layer and a semantic understanding layer. The feature extraction layer, comprised of multiple convolutional neural networks, processes input information from various modalities, such as text, images, and audio. The semantic understanding layer, using a transformer architecture, achieves cross-modal semantic fusion and understanding. When the system receives a task request from a user, the multimodal parsing engine extracts features from the request content and generates an initial feature vector. For example, when a user submits a request to "query sales figures for the past three months and generate a trend chart," the system first performs word segmentation on the text, extracting keyword elements such as "query," "past three months," "sales figures," "generate," and "trend chart," and converts these elements into a 512-dimensional initial feature vector.

[0056] To obtain a deeper semantic representation, the initial feature vector is re-entered into the multimodal parsing engine for deeper feature extraction. During this process, the parsing engine uses an attention mechanism to capture the correlations between different feature elements and, combined with a pre-trained language model, performs semantic enhancement, ultimately generating a 1024-dimensional request feature vector. This feature vector provides a hierarchical semantic representation, encompassing information such as the task request's intent (e.g., "query" operation), object (e.g., "sales" data), and constraint (e.g., "last three months" timeframe). In the above example, the system identifies the user's intent as data query and visualization, requiring the use of database query tools and chart generation tools.

[0057] The request feature vector is semantically parsed based on the input interface schema of the target tool. The system maintains a tool library containing input interface descriptions, parameter requirements, and functional specifications for various tools. For the sales data query tool, its input interface requirements include data type parameters, time range parameters, and query condition parameters. The system utilizes a pre-built interface mapping model to map the hierarchical semantic representation in the request feature vector to a standard semantic space. This mapping process is implemented using a semantic matching algorithm that calculates the cosine similarity between the request feature vector and the tool interface description features and selects the interface schema with the highest similarity as the mapping target. After mapping, the system generates standardized semantic features containing standardized semantic information, such as operation type = "query", data type = "sales", time range = "last three months", and output format = "trend chart".

[0058] The semantic consistency of the normalized semantic information in the standardized semantic features is verified, and the semantic similarity between the request feature vector and the standardized semantic features is calculated. The similarity calculation adopts a measurement method based on vector space. In actual applications, when the semantic similarity is greater than 0.85 (the preset similarity threshold), the system believes that the standardization process retains the core semantics of the user's original request. For example, the original request "query the sales of the past three months and generate a trend chart" becomes "get the sales data of the past 90 days and generate a line chart display" after standardization. The semantic similarity between the two is calculated to be 0.92, which exceeds the threshold, so the verification is passed. After the verification is passed, the system converts the standardized semantic features into standardized call requests according to the input interface mode of the target tool. For the sales data query tool, the standardized call request is formatted in JSON format: {"action":"query", "data type ":"sales", "time range ":"90days", "output":"trend chart "}.

[0059] The standardized call request is passed to the target tool for calling, and the request is sent to the corresponding tool service through the API gateway. For the sales data query tool, the system establishes a secure connection channel, transmits the standardized call request, and sets the request timeout to 3 seconds. After receiving the request, the target tool performs the corresponding operation to complete the retrieval of sales data and the generation of trend charts. The system collects the call execution results of the target tool, including the operation status code, return data, and error information. During the collection process, the system uses an asynchronous listening mechanism to continuously detect the execution status. When the execution completion signal is detected, the execution result is obtained and standardized processing is performed. For the sales data query example, the execution results obtained by the system include the average daily sales data of the past 90 days and the generated trend chart URL address.

[0060] In terms of error handling, when semantic similarity verification fails, the standardized mapping strategy is readjusted and a second conversion attempt is made using an alternative mapping model. If three consecutive conversion attempts fail verification, the system will provide the user with a request ambiguity prompt, requiring the user to clarify the request intent. In addition, the system has a built-in request conversion logging function that records the intermediate results and similarity calculation values ​​during each conversion process to facilitate subsequent optimization of the conversion algorithm. Through the above technical implementation, the system can efficiently and accurately convert the user's multimodal task requests into standardized call requests that can be recognized by the target tool, realizing intelligent and automated tool calls.

[0061] In an optional embodiment, updating the tool usage statistics vector of the target tool based on the execution result and the performance value consumption indicator, and using the updated tool usage statistics vector to generate a new tool composite vector includes: Based on the execution results and the performance value consumption index, the tool usage statistics vector of the target tool is progressively updated. The progressive update is achieved by weighted combination of the original statistics vector and the current execution statistics information. The weight of the weighted combination is determined by the historical information retention coefficient. The historical information retention coefficient dynamically decays with the number of executions. An update vector is generated based on the tool usage statistical vector after progressive update. The update vector is obtained by progressively fusing the tool usage statistical vector with environmental context information. The progressive fusion adopts the same weight coefficient as the progressive update to generate a new tool composite vector based on the update vector.

[0062] After receiving a task request from a user, the system selects an appropriate target tool from the tool library to execute the task. Upon completion, the system records the execution results and corresponding performance cost metrics. The execution results can be Boolean values ​​(success or failure), numerical scores (such as an effectiveness score between 0 and 1), or composite metrics (such as multi-dimensional evaluation results). Performance cost metrics include quantifiable cost metrics such as execution time, computing resource consumption, and the number of API calls.

[0063] Based on this information, the tool usage statistics vector for the target tool is progressively updated. The tool usage statistics vector includes multiple dimensions, such as success rate, average execution time, and average resource consumption. The update process uses a weighted combination approach: the original statistics vector is merged with the current execution statistics. For example, suppose the original tool statistics vector is [0.85, 3.2, 0.45] (representing an 85% success rate, an average execution time of 3.2 seconds, and an average resource consumption of 0.45 units), the current execution result is [1.0, 2.8, 0.4], and the historical information retention coefficient is 0.9. The updated statistics vector is [0.9×0.85+0.1×1.0, 0.9×3.2+0.1×2.8, 0.9×0.45+0.1×0.4], or [0.865, 3.16, 0.445].

[0064] The historical information retention coefficient is a key parameter that determines the system's reliance on historical data. This coefficient decays dynamically with the number of executions, allowing the system to gradually adjust its emphasis on new data. The specific decay method can be set as follows: the initial historical information retention coefficient is 0.95, and after every 10 executions, the coefficient is multiplied by 0.99. This way, the system retains more historical data in the early stages of the tool to build a stable statistical foundation, while gradually increasing its sensitivity to new data as the tool is used more frequently.

[0065] After the progressive update is complete, an update vector is generated based on the updated tool usage statistics vector. This update vector is then fused with environmental context information. Environmental context information includes user preferences, current task type, system load status, and more. For example, if the system identifies the current task as time-sensitive, it will increase the weight of the execution time metric during the fusion process; if it identifies the task as precision-sensitive, it will increase the weight of the accuracy metric.

[0066] The specific fusion process also adopts a weighted combination method. Assuming that the environment context vector is [0.7, 0.3, 0.9] (representing that the current environment attaches 0.7 to the success rate, 0.3 to the execution time, and 0.9 to the resource consumption), using the same weight coefficient 0.9 as the progressive update, the update vector is calculated as [0.9×0.865+0.1×0.7,0.9×3.16+0.1×0.3, 0.9×0.445+0.1×0.9], that is, [0.8485, 2.874, 0.4905].

[0067] Based on the resulting update vectors, a new tool composite vector is generated. This composite vector comprehensively expresses the current state and usefulness of the tool, directly influencing subsequent tool selection and task allocation decisions. During the generation process, the system normalizes the update vectors to make the data across dimensions comparable. For example, execution time typically has a larger value, while resource consumption typically has a smaller value. Normalization allows comparisons between these dimensions to fall within the 0-1 range.

[0068] Normalization can be performed using the minimum-maximum normalization method, calculating the corresponding metrics for all similar tools in the tool library. Assuming the execution time range for similar tools in the tool library is [1.5, 8.0], the execution time 2.874 in the above update vector will be normalized to (2.874 - 1.5) / (8.0 - 1.5) = 0.212. After performing similar operations on all dimensions, the resulting composite tool vector will be [0.78, 0.212, 0.35].

[0069] The tool composite vector is globally updated periodically (e.g., every 100 calls or every 24 hours) to reflect long-term usage trends and performance changes. In practice, if a tool exhibits persistent performance degradation for a specific task type (e.g., slow API response), the system will gradually reduce the tool's selection probability for the corresponding task through a progressive update mechanism and trigger an alert to prompt system administrators to check the tool's status.

[0070] The above mechanism ensures that the system can dynamically adjust tool selection strategies based on actual tool usage, improving task execution efficiency and resource utilization. By dynamically adjusting the historical information retention coefficient, the system not only retains historical usage experience but also responds promptly to changes in tool performance, achieving intelligent and adaptive tool management.

[0071] A second aspect of an embodiment of the present invention provides a RAG tool dynamically enhanced agent building system for multimodal scenarios, including: The first unit is configured to generate a tool composite vector and store it in a tool vector space for subsequent tool retrieval, receive a task request from an agent, and construct a context-aware retrieval vector based on the task request; a second unit configured to use the context-aware retrieval vector to perform an approximate nearest neighbor search in the tool vector space to obtain an initial tool of a candidate tool set, perform a graph traversal in a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, perform rule filtering on the initial tool and the associated tools based on current environment constraints, and update the candidate tool set; The third unit is configured to determine a four-dimensional scoring weight based on the scenario type of the current task, collect real-time operating indicators of each tool in the candidate tool set and calculate scores in each dimension, perform a weighted sum of the scores in each dimension and the four-dimensional scoring weight to obtain a multidimensional utility score for each tool, and select the tool with the highest multidimensional utility score as the target tool; The fourth unit is configured to convert the task request into a standardized call request based on the input interface mode of the target tool, use the standardized call request to call the target tool and obtain the execution result; collect the performance value consumption index of the target tool during execution; update the tool usage statistics vector of the target tool based on the execution result and the performance value consumption index; and use the updated tool usage statistics vector to generate a new tool composite vector. According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0072] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0073] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A RAG tool dynamically enhances the intelligent agent building method for multimodal scenarios, characterized in that: include: Generate a tool composite vector and store it in the tool vector space for subsequent tool retrieval, receive the agent's task request, and construct a context-aware retrieval vector based on the task request; performing an approximate nearest neighbor search in the tool vector space using the context-aware retrieval vector to obtain an initial tool of a candidate tool set, performing a graph traversal in a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, performing rule filtering on the initial tool and the associated tools based on current environmental constraints, and updating the candidate tool set; Determine the four-dimensional scoring weights based on the scenario type of the current task, collect the real-time operating indicators of each tool in the candidate tool set and calculate the scores of each dimension, perform a weighted sum of the scores of each dimension and the four-dimensional scoring weights to obtain the multidimensional utility score of each tool, and select the tool with the highest multidimensional utility score as the target tool; The task request is converted into a standardized call request according to the input interface mode of the target tool, and the target tool is called using the standardized call request to obtain the execution result; the performance value consumption index of the target tool during the execution process is collected, and the tool usage statistical vector of the target tool is updated based on the execution result and the performance value consumption index, and the updated tool usage statistical vector is used to generate a new tool composite vector.

2. The method according to claim 1, characterized in that Performing an approximate nearest neighbor search in the tool vector space using the context-aware retrieval vector to obtain initial tools of a candidate tool set includes: Constructing an initial retrieval feature based on a context-aware retrieval vector, mapping the context-aware retrieval vector to a tool vector space to establish a tool knowledge graph, and performing knowledge enhancement on the context-aware retrieval vector based on the tool knowledge graph to obtain an enhanced retrieval vector; constructing an approximate nearest neighbor retrieval index in the tool vector space, performing an approximate nearest neighbor search on the enhanced retrieval vector based on the approximate nearest neighbor retrieval index, and calculating a similarity score between the enhanced retrieval vector and the tool vector; The tool vectors are sorted according to the similarity scores, and tools corresponding to a preset number of tool vectors with the highest similarity scores are selected as initial tools of the candidate tool set.

3. The method according to claim 2, characterized in that Mapping the context-aware retrieval vector to a tool vector space to establish a tool knowledge graph, and performing knowledge enhancement on the context-aware retrieval vector based on the tool knowledge graph to obtain an enhanced retrieval vector includes: Constructing a dynamic mapping matrix, wherein the dynamic mapping matrix maps the context-aware retrieval vector to a tool vector space to obtain a tool vector, obtaining a mapping difference between a reference tool vector and the tool vector to generate mapping update information, and superimposing the mapping update information on the dynamic mapping matrix to update the dynamic mapping matrix; Constructing a tool knowledge graph in the tool vector space, performing mapping conversion on adjacent tool nodes in the tool knowledge graph using the dynamic mapping matrix, and multiplying the conversion result by the association information of the adjacent tool nodes to construct an inter-node association strength value in the tool knowledge graph; The context-aware retrieval vector is subjected to knowledge enhancement, the tool vector is used as a starting node, the associated node vector is searched in the tool knowledge graph using the dynamic mapping matrix, and the association strength value is used as a weight coefficient, the weight coefficient is changed with the association strength value, and the weighted associated node vector is fused with the context-aware retrieval vector to obtain an enhanced retrieval vector.

4. The method according to claim 1, wherein Traversing a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, and filtering the initial tool and the associated tools based on current environment constraints include: Using the initial tool as the starting node to traverse the tool dependency graph, calculating the node transition probability of each traversal step, traversing along the edges of the tool dependency graph based on the node transition probability, and grouping the tool nodes encountered during the traversal into an associated tool set; Establishing a rule template library, storing a plurality of rule templates in the rule template library, each rule template including rule structure information and combination pattern information, inputting the combination of the initial tool and the tools in the associated tool set in historical execution records into each rule template, calculating the impact degree and complexity score of each rule template on the tool combination, and taking the weighted sum of the impact degree and the complexity score as the rule importance weight; The rule templates in the rule template library are sorted and screened according to the rule importance weights, and the rule templates with the highest rule importance weights are selected to form a rule set. The initial tool and each tool combination in the associated tool set are input into the rule set for rule verification, and the current environmental constraints are extracted to filter the verification results.

5. The method according to claim 1, wherein Converting the task request into a standardized call request according to the input interface mode of the target tool, and using the standardized call request to call the target tool and obtain the execution result includes: Constructing a multimodal parsing engine, using the multimodal parsing engine to perform feature extraction on the task request to generate an initial feature vector, and inputting the initial feature vector into the multimodal parsing engine again for deep feature extraction to obtain a request feature vector with a hierarchical semantic representation; Performing semantic feature parsing on the request feature vector based on the input interface mode of the target tool, mapping the hierarchical semantic representation in the request feature vector to a standard semantic space according to the multimodal parsing engine, and generating standardized semantic features containing normalized semantic information; Performing semantic consistency verification on the normalized semantic information in the standardized semantic feature, calculating the semantic similarity between the request feature vector and the standardized semantic feature, and converting the verified standardized semantic feature into a standardized call request according to the input interface mode of the target tool when the semantic similarity is greater than a preset similarity threshold; The standardized call request is transmitted to the target tool for calling, and the call execution result of the target tool is collected.

6. The method according to claim 1, characterized in that Updating the tool usage statistics vector of the target tool based on the execution result and the performance value consumption indicator, and using the updated tool usage statistics vector to generate a new tool composite vector includes: Based on the execution results and the performance value consumption index, the tool usage statistics vector of the target tool is progressively updated. The progressive update is achieved by weighted combination of the original statistics vector and the current execution statistics information. The weight of the weighted combination is determined by the historical information retention coefficient. The historical information retention coefficient dynamically decays with the number of executions. An update vector is generated based on the tool usage statistical vector after progressive update. The update vector is obtained by progressively fusing the tool usage statistical vector with environmental context information. The progressive fusion adopts the same weight coefficient as the progressive update to generate a new tool composite vector based on the update vector.

7. A RAG tool for multimodal scenarios dynamically enhanced intelligent agent building system, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is configured to generate a tool composite vector and store it in a tool vector space for subsequent tool retrieval, receive a task request from an agent, and construct a context-aware retrieval vector based on the task request; a second unit configured to use the context-aware retrieval vector to perform an approximate nearest neighbor search in the tool vector space to obtain an initial tool of a candidate tool set, perform a graph traversal in a tool dependency graph based on the initial tool to obtain associated tools of the initial tool, perform rule filtering on the initial tool and the associated tools based on current environment constraints, and update the candidate tool set; The third unit is configured to determine a four-dimensional scoring weight based on the scenario type of the current task, collect real-time operating indicators of each tool in the candidate tool set and calculate scores in each dimension, perform a weighted sum of the scores in each dimension and the four-dimensional scoring weight to obtain a multidimensional utility score for each tool, and select the tool with the highest multidimensional utility score as the target tool; The fourth unit is used to convert the task request into a standardized call request according to the input interface mode of the target tool, use the standardized call request to call the target tool and obtain the execution result; collect the performance value consumption index of the target tool during the execution process, update the tool usage statistical vector of the target tool based on the execution result and the performance value consumption index, and use the updated tool usage statistical vector to generate a new tool composite vector.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Task processing method and device based on intelligent agent and electronic equipment

    CN119166834A

  • Reasoning optimization method and device for large model and medium

    CN119168067A

  • Large language model tool matching method and system based on knowledge graph

    CN119294493A

  • Tool recommendation method and device and electronic equipment

    CN119829753A

  • Agent-based tool combination and task processing method and device, equipment and medium

    CN120276875A

Cited By

  • Intelligent agent tool calling decision-making method based on problem scene

    CN121542307A