A learning training method and device for a tool invocation capability, and a storage medium

By constructing a multi-dimensional quantitative evaluation and fine-grained reward mechanism to train the language model agent, the problems of weak generalization ability, lack of dynamic decision-making and insufficient semantic understanding in the existing technology are solved, and the reliability and adaptability of tool invocation of enterprise-level agents are improved.

CN121638310BActive Publication Date: 2026-06-23DIGITAL CHINA CHINA CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DIGITAL CHINA CHINA CO LTD
Filing Date
2026-02-04
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies suffer from weak generalization ability, lack of dynamic decision-making and error correction capabilities, insufficient semantic understanding, and improper handling of capability boundaries when building enterprise-level intelligent agents. This leads to performance degradation of the model when faced with novel task combinations or error scenarios, and makes it unable to effectively handle user requests that exceed the capabilities of the toolset, thus affecting system reliability and user experience.

Method used

By constructing an initial context containing user questions, tool sets, and historical interaction information, tool invocation instructions are generated and multi-dimensional quantitative evaluations are performed. Combined with a fine-grained reward mechanism, the language model agent is iteratively trained to enhance its tool invocation capabilities, including optimization of format, semantics, and dynamic decision-making.

Benefits of technology

It enhances the generalization and dynamic decision-making capabilities of language model agents, reduces semantic error rates, improves adaptability and reliability for complex tasks, ensures that the model can identify boundaries and autonomously adjust strategies when it exceeds its capabilities, and reduces ineffective operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121638310B_ABST
    Figure CN121638310B_ABST
Patent Text Reader

Abstract

The application provides a tool calling capability learning training method and device and a storage medium. The method comprises the following steps: constructing initial context information to be trained; inputting the initial context information into a language model agent to generate a response result; performing multi-dimensional quantitative evaluation and processing on a tool calling instruction to generate fine-grained reward information; training the initial context information based on the response result, an execution result and correctness reward information to obtain target context information; and iteratively training the language model agent according to trajectory data generated by historical interaction information, the fine-grained reward information and the target context information. The application constructs initial context information containing a user question, a tool set and historical interaction information, generates a response result containing a tool calling instruction and performs multi-dimensional quantitative evaluation, and iteratively trains the language model agent in combination with a fine-grained reward mechanism, thereby improving the tool calling reliability and adaptability of the language model agent.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • CN120598075A

  • CN120654815A