Function tool calling method, device and system and nonvolatile storage medium
By presenting the problem text and obtaining user responses during the call of the function tool of the large language model, the problem of the large language model illusion is solved, and the function tool is accurately called based on the output results of the large language model, improving the accuracy and sustainability of the calling process.
Patent Information
- Application Number
- CN202510279135.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, large language models cannot be interrupted during the inference process of function tool calling, resulting in the inability to supplement the input instructions when the user input instructions are not clear, and the problem of hallucination of large language models is prone to occur, and the call of function tools cannot be realized based on the results output by the large language model.
The input text of the target large language model is generated based on the initial text input by the target object, the prompt word template and the historical session data of the target object. During the execution of the function tool call, it is determined whether the target object needs to input reference information, display the problem text generated by the large model and obtain the reply text of the target object, so as to correct the output result and continue to execute the call of the function tool.
It realizes the interruption of the big model inference process, eliminates the illusion of the big model caused by unclear input content, and thus solves the problem of not being able to realize the function tool call based on the output results of the large language model, and improves the accuracy and sustainability of the function tool call process.
Smart Images

Figure CN120216619A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic digital data processing, and in particular, to a function tool call method, device, system, and non-volatile storage medium. Background Art
[0002] In related technologies, when implementing function tool call inference through a large language model, since the inference process of the model cannot be interrupted, it is impossible to supplement and explain the input instruction when the user's input instruction is unclear, resulting in the problem of large language model hallucinations, and it is impossible to implement the function tool call according to the result output by the large language model.
[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present application provide a function tool call method, device, system, and non-volatile storage medium, so as to at least solve the technical problem that the function tool call cannot be implemented according to the result output by the large language model due to the inability to solve the large language model hallucination problem in related technologies.
[0005] According to one aspect of the embodiments of the present application, a function tool call method is provided, including: generating an input text of a target large language model based on an initial text input by a target object, a prompt template, and historical session data of the target object; inputting the input text into the target large language model, and obtaining a first output result output by the target large language model, where the first output result includes information related to the called function tool, and the first output result is an output result that passes the verification; performing a call process of the function tool corresponding to the relevant information according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, obtaining a question text generated by the target large language model for indicating the reference information and presenting it to the target object; after obtaining a reply text input by the target object for replying to the question text, correcting the first output result according to the reply text and the question text to obtain a second output result, and continuing to perform the function tool call process according to the second output result.
[0006] Optionally, input the input text into the target large language model and obtain the first output result output by the target large language model, including: obtaining the third output result generated by the target large language model based on the input text; performing format verification on the third output result, where the format verification includes verifying whether the function tool name and input parameter format in the third output result are in a preset format; in the case where the third output result passes the format verification, performing accuracy verification on the third output result, where the accuracy verification includes verifying whether the function tool name is in the preset tool name list and verifying whether the input parameters conform to the preset rules; in the case where the third output result passes the accuracy verification, determining the third output result as the first output result.
[0007] Optionally, after performing format verification on the third output result, the method further includes: in the case where the third output result fails to pass the format verification, generating a first correction instruction based on the verification result of the format verification; storing the first correction instruction and the second output result as additional content in the cache; updating the input text based on the additional content in the cache and obtaining the fourth output result generated by the large language model based on the updated input text; each time the fourth output result is obtained, performing format verification on the fourth output result until the fourth output result passes the format verification, and then using the fourth output result that has passed the format verification as the third output result.
[0008] Optionally, the method further includes: in the case where the third output result passes the accuracy verification, clearing the cache.
[0009] Optionally, after performing accuracy verification on the second output result, the method further includes: in the case where the third output result fails to pass the accuracy verification, generating a second correction instruction based on the verification result of the accuracy verification; storing the second correction instruction and the second output result as additional content in the cache; updating the input text based on the additional content in the cache and obtaining the fifth output result generated by the large language model based on the updated input text; each time the fifth output result is obtained, performing accuracy verification on the fifth output result until the fifth output result passes the accuracy verification, and then using the fifth output result that has passed the accuracy verification as the first output result.
[0010] Optionally, the method further includes: in the case where the fifth output result passes the accuracy verification, clearing the cache.
[0011] Optionally, the process of invoking the function tool corresponding to the relevant information based on the first output result includes: determining the function tool name in the first output result; in the case where the function tool name is the preset function tool name, obtaining the question text and the corresponding reply text, and storing the question text and the reply text as additional content in the decision cache.
[0012] Optionally, after obtaining the response text input by the target object for answering the question text, correcting the first output result according to the response text and the question text to obtain the second output result includes: correcting the input text according to the additional content in the decision cache, and calling the large language model to process the corrected input text to obtain the second output result.
[0013] Optionally, the prompt template includes a system instruction prompt template and a user instruction prompt template; generating the input text of the target large language model according to the initial text input by the target object, the prompt template and the historical conversation data of the target object includes: constructing the input text of the target large language model according to the initial text input by the target object, the system instruction prompt template, the user instruction prompt template and the historical conversation data of the target object, wherein the system instruction prompt template is used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt template is used to reflect the problem description information indicated by the initial text.
[0014] According to another aspect of the embodiments of the present application, there is also provided a function tool call device, including: a first processing module, configured to generate the input text of the target large language model according to the initial text input by the target object, the prompt template and the historical conversation data of the target object; a second processing module, configured to input the input text into the target large language model and obtain the first output result output by the target large language model, where the first output result includes the relevant information of the function tool called, and the first output result is the output result that passes the verification; a third processing module, configured to execute the call process of the function tool corresponding to the relevant information according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, obtain the question text generated by the target large language model for indicating the reference information and display it to the target object; a fourth processing module, configured to, after obtaining the response text input by the target object for answering the question text, correct the first output result according to the response text and the question text to obtain the second output result, and continue to execute the call process of the function tool according to the second output result.
[0015] According to another aspect of the embodiments of the present application, there is also provided a function tool call system applicable to the function tool call method, including: a prompt word management unit for managing prompt word templates, where the prompt word templates include a system instruction prompt word template and a user instruction prompt word template. The system instruction prompt word template is used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt word template is used to embody the problem description information indicated by the initial text; a large language model unit for scheduling the large language model and configuring the standardized interface of the large language model; a model output format parsing unit for storing various text structuring information parsing rules and verification rules of the output results of the large language model; a tool parsing and verification unit for storing and managing the parsing and verification rules of function tools, where the parsing and verification rules include a preset unified pre-tool parsing and verification rule and a custom post-tool parsing and verification rule configured by the target object; a tool execution unit for parsing and verifying the output results of the large language model according to the parsing and verification rules and scheduling the function tools according to the parsing and verification results; a memory storage unit for storing persistent historical session data and temporarily memorized cache data.
[0016] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium storing a program, where when the program runs, it controls the device where the non-volatile storage medium is located to execute the function tool call method.
[0017] According to another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor, and the processor is used to run the program stored in the memory, where when the program runs, it executes the function tool call method.
[0018] According to another aspect of the embodiments of the present application, there is also provided a computer program product including a computer program, and when the computer program is executed by a processor, it implements the function tool call method.
[0019] In the embodiments of the present application, an input text for a target large language model is generated based on the initial text input by the target object, a prompt template, and the historical conversation data of the target object; the input text is input into the target large language model, and a first output result output by the target large language model is obtained, where the first output result includes information related to the function tool called, and the first output result is an output result that passes the verification; the calling process of the function tool corresponding to the relevant information is executed according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, the question text generated by the target large language model for indicating the reference information is obtained and displayed to the target object; after obtaining the reply text input by the target object for replying to the question text, the first output result is corrected according to the reply text and the question text to obtain a second output result, and the calling process of the function tool is continued according to the second output result. By judging whether the target object needs to input reference information to display the question text generated by the large model and obtain the reply text input by the target object, the purpose of interrupting the inference process of the large model is achieved, thereby realizing the technical effect of eliminating the large model hallucination caused by unclear input content, and further solving the technical problem that the function tool call cannot be realized according to the result output by the large language model due to the inability to solve the large language model hallucination problem in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0021] Figure 1 is a schematic structural diagram of a computer terminal (portable terminal) provided according to an embodiment of the present application;
[0022] Figure 2 is a schematic flowchart of a function tool call method provided according to an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of an in-loop verification process provided according to an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of the verification effect of the in-loop verification process provided according to an embodiment of the present application;
[0025] Figure 5 is a schematic flowchart of a function tool call process without interruption in the middle provided according to an embodiment of the present application;
[0026] Figure 6 is a schematic diagram of the cached content and execution result in a function tool call process without interruption in the middle provided according to an embodiment of the present application;
[0027] Figure 7 is a schematic diagram of the pre - interruption process of a function tool call process with mid - interruption provided according to an embodiment of the present application;
[0028] Figure 8 is a schematic diagram of the cached content and execution result of the pre - interruption process of a function tool call process with mid - interruption provided according to an embodiment of the present application;
[0029] Figure 9 is a schematic diagram of the post - interruption process of a function tool call process with mid - interruption provided according to an embodiment of the present application;
[0030] Figure 10 is a schematic diagram of the cached content and execution result of the post - interruption process of a function tool call process with mid - interruption provided according to an embodiment of the present application;
[0031] Figure 11 is a schematic diagram of a reply process provided according to an embodiment of the present application;
[0032] Figure 12 is a schematic diagram of a function tool call process provided according to an embodiment of the present application;
[0033] Figure 13 is a framework schematic diagram of a function tool call method provided according to an embodiment of the present application;
[0034] Figure 14 is a structural schematic diagram of a function tool call device provided according to an embodiment of the present application;
[0035] Figure 15 is a structural schematic diagram of a function tool call system provided according to an embodiment of the present application. Detailed implementation manners
[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0038] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained as follows:
[0039] LLM (Large Language Model): A large language model is a deep learning model trained based on a vast amount of text data. It can not only generate natural language text but also deeply understand the meaning of the text and handle various natural language tasks such as text summarization, question answering, and translation.
[0040] Prompt Engineering: It refers to a technology that guides a generative AI model to generate a specific output through carefully designed prompts. In the field of natural language processing, prompt engineering is usually used to guide a large language model to generate text output that meets specific requirements. For example, by designing a prompt, a large language model can be made to generate an article describing a certain topic or a dialogue response that conforms to a specific format. Therefore, designing and adjusting a set of good prompt templates is a key technology to improve the output quality of a generative AI model. A good prompt template can make a large language model generate text output that meets specific requirements according to the specified format.
[0041] CoT (Chain of Thought): The chain of thought is an improved prompting strategy used to improve the performance of LLM in complex reasoning tasks such as arithmetic reasoning, common sense reasoning, and symbolic reasoning. The chain of thought does not simply construct a prompt with input-output pairs but combines intermediate reasoning steps that can lead to the final output in the prompt. Simply put, the chain of thought is a discrete prompting learning. More specifically, it is context learning under a large model (that is, without training, adding examples to the front of the current sample input and letting the model input these texts at once to complete the task). Compared with the previous traditional context learning, the chain of thought has additional intermediate derivation prompts.
[0042] ReAct (Reasoning and Acting): The reasoning and acting framework means that the LLM can, based on logical reasoning (Reason), construct a complete series of actions (Action), and then, according to the execution results of the actions (Observation), conduct the next round of logical reasoning and construct the next action until the desired goal (Final Answer) is achieved. The LLM model has shown excellent performance in logical reasoning. Therefore, there is reason to believe that the LLM model can also perform logical reasoning, learn knowledge, make decisions, and execute like humans. In actual use, the LLM may have hallucinations and misjudgments. This is because the knowledge the LLM is exposed to during training is limited. Therefore, when conducting logical analysis on data beyond that used in the training process, the LLM will start fabricating some reasons. To solve this problem, the method adopted in this application is to provide the necessary knowledge to the LLM before it makes an analysis and decision. The role of the ReAct approach is to coordinate the LLM model and external information acquisition, and interact with other functions (function tools) to supplement the knowledge and capabilities that the LLM does not possess.
[0043] SFT (Supervised Fine-Tuning): Supervised fine-tuning is a method based on transfer learning. By using training data in a specific domain or task to conduct supervised fine-tuning training on a pre-trained model, the model can better adapt to new domain tasks. This method has shown excellent performance in many fields, such as text classification, image recognition, and bioinformatics.
[0044] With the continuous iterative development of large language model technology in recent years, the technology of large language models is no longer limited to conversational chat and simple prompt engineering applications. Nowadays, more and more large language models already have the ability to call function tools. The large language model conducts thinking deduction based on system instructions and the prompt words input by the user, can judge the user's intention, select the corresponding function tool from the tool library, parse the input parameters of the function from the context of the conversation, and can achieve the call of the function tool by combining engineering code technology, and return the function execution result to the large model for answering, completing the closed-loop of the tool scheduling task based on the large language model.
[0045] However, since the large language model itself is pre-trained based on massive amounts of data, this means that there is always some knowledge information that is not in the training data of the large language model. For knowledge areas that the large language model is "unfamiliar" with, it may generate fabricated content when responding to user instructions, resulting in unavoidable hallucinations. The hallucination problem of the large language model will bring instability to tasks such as the large language model calling function tools that require structured information processing.
[0046] At present, there are two main technical routes for realizing function tool call reasoning with large language models in related technologies. One is to activate the CoT (Chain of Thought) reasoning ability of the large language model based on prompt word templates such as the ReAct (Reasoning and Acting) framework to realize the parsing of function tool names and input parameters; the other is to introduce special instruction role symbols (such as <|Tool|>, <|Function|>, <|Observation|>, etc.) to arrange the function tool related instruction training data during the SFT (Supervised Fine-Tuning) instruction fine-tuning training, and then perform instruction fine-tuning training of the large language model to achieve the purpose of natively supporting function tool names and input parameter parsing. However, no matter which technical route is used, the parsing call of the function tool depends entirely on the semantic understanding and logical reasoning ability of the large language model itself. For large language models with large model parameters (such as commercial models with hundreds of billions of parameters), the parsing accuracy of the function tool call is higher, but for models with small parameters (such as open source native models or quantitative models with billions of parameters), the parsing accuracy of the function tool call will be relatively low. Especially when the user input instructions are unclear, the hallucination problem of the large language model is prominent, and it is easy to make up nonsense.
[0047] In order to solve this problem, an embodiment of the present application provides an enhanced function tool parsing and calling function tool calling method based on the ReAct framework prompt word engineering technology route, which can effectively reduce the erroneous judgment interference of the model hallucination phenomenon, improve the accuracy of function tool calling of large language models, and especially improve the sustainability and accuracy of function tool calling in multiple rounds of dialogue with large language models, which is described in detail below.
[0048] According to an embodiment of the present application, a method embodiment of a function tool calling method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0049] The method embodiments provided by the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a function tool call method. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b,..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0050] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any other combination. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is used for a processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0051] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the function tool call method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned function tool call method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0052] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0053] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0054] Under the above operating environment, an embodiment of the present application provides a function tool call method, as Figure 2 shown, the method includes the following steps:
[0055] Step S202, generating an input text for the target large language model according to the initial text input by the target object, the prompt word template, and the historical conversation data of the target object;
[0056] In the technical solution provided in step S202, the prompt word template includes a system instruction prompt word template and a user instruction prompt word template; the step of generating an input text for the target large language model according to the initial text input by the target object, the prompt word template, and the historical conversation data of the target object includes: constructing an input text for the target large language model according to the initial text input by the target object, the system instruction prompt word template, the user instruction prompt word template, and the historical conversation data of the target object, where the system instruction prompt word template is used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt word template is used to reflect the problem description information indicated by the initial text.
[0057] As an alternative implementation manner, the above target object includes a user or a terminal device used by the user, and the above system instruction prompt word template includes a ReAct framework system instruction prompt word template. The input text of the large language model can be constructed by combining the two prompt word templates of the ReAct framework system instruction and the user instruction provided by the embodiment of the present application with the historical conversation information between the user and the large language model.
[0058] In currently common large language conversation models, three types of role instructions are usually designed, namely System role instruction, User role instruction, and AI role instruction (Assistant). In the following statements, System instruction refers to the System role instruction, User instruction refers to the User role instruction, and Assistant instruction refers to the AI role instruction.
[0059] The System instruction is a global instruction requirement, and its scope of action covers the entire conversation process between the user and the large language model; the User instruction is a task instruction for the user to interact with the large language model. After receiving the current User instruction, the large language model will respond to the user's instruction requirements in the subsequent reply and generate relevant content for the reply through the Assistant instruction.
[0060] In some embodiments of the present application, in order to enable the large language model to call external function tools and interact with the outside through prompt engineering technology, the embodiments of the present application adopt a prompt template design strategy under the ReAct framework to guide the inference learning of the internal thought chain (CoT) of the large language model, so that the large language model can accurately respond to complex prompt requirements. Optionally, in the embodiments of the present application, the prompt templates of the ReAct framework are divided into system instruction templates and user instruction templates. The following is a specific description of the prompt templates:
[0061] (1) System instruction prompt template of the ReAct framework
[0062] To respond to the user's request as accurately as possible, you can use the following provided tools:
[0063] {tools}
[0064] The user may have multiple requests, and different requests may require the use of different tools. Continue writing in the following format:
[0065]
[0066]
[0067] Precautions:
[0068] The annotation information " / / ..." in the above JSON structure is only to help you understand the content requirements. The JSON structure you output should never carry the " / / ..." annotation information.
[0069] (2) User instruction prompt template of the ReAct framework
[0070] Please continue the following content based on the context information. Remember that you always need to output "Thought:" first, and then "Action:" to determine the actions to be taken:
[0071] Question:{question}{agent_scratchpad}
[0072] It can be seen that the above two prompt templates contain four variables. Among them, the variables {tools} and {tool_names} in the system instruction prompt template are variable parameters related to function tools and need to specify the values of the variables in advance; the variables {question} and {agent_scratchpad} in the user instruction prompt template are text descriptions in the ReAct instruction format. After the user inputs the instruction, the operation framework proposed in the embodiments of the present application will automatically read relevant information from the cache.
[0073] {tools} is the text description of the usage conditions of the function tool, which needs to include at least the name (Name), usage scenario description (Description), and input parameter description (Arguments) of each function tool. The input parameter description needs to include the parameter name (Parameter Name), parameter description (Parameter Description), and parameter value type (Parameter Type).
[0074] In some embodiments of the present application, the following JSON format text can be used to describe the usage conditions of the function tool to help the large language model understand in what scenarios to call the tool:
[0075]
[0076]
[0077]
[0078] It should be noted that the above is only one optional implementation method, and the present application embodiments do not limit which statement method is specifically used to describe the usage conditions of the function tool.
[0079] It should be noted that different from the implementation method of ReAct prompting engineering in related technologies, the ReAct prompting engineering framework provided in the embodiments of the present application has an interrupted ReAct process in the middle, allowing the large language model to interact with the user during the ReAct process. Therefore, there is a fixed function tool with a built-in implementation in the tools of the embodiments of the present application. The name of this function tool is "user_input". Its input parameter "query" is the question that the model needs to ask the user, and the output result after the tool is executed is equal to the value of its input parameter "query", that is, the output is the question that the model needs to ask the user. Exemplarily, the Python code for the function implementation method of "user_input" is as follows:
[0080] def user_input(query:str)->str:
[0081] return query
[0082] Exemplarily, the description text of the usage conditions of its function tool is as follows:
[0083]
[0084]
[0085] The {tool_names} variable is all the function tool names that appear in the description of the usage conditions of the function tools in the {tools} variable. Generally, it is described in a list format. As an embodiment, the list text description format can be used: tool_names = ["function tool name 1", "function tool name 2",...]. It should be noted that this is only one of the implementation methods, and the specific statement method of which function tool name is sampled is not specifically limited in the embodiments of the present application.
[0086] The {question} variable is the original question or request input by the user in the description of "Question: the original input question that you must answer" before the start of the ReAct process. It will be automatically read from the original question cache in the memory storage unit, that is, question = original question cache.
[0087] The {agent_scratchpad} variable is the text information of the "Thought: / Action: / Observation:" thought chain process after the ReAct process starts. Among them, the text contents of "Thought" and "Action" are automatically inferred by the large language model, and the text content of "Observation" is passed in from the outside. It can include the return result after the successful execution of the function tool, the return result after the failure of the function tool execution, and the correction instruction for the function tool verification and parsing implemented in the embodiments of the present application. In the following statements, the embodiments of the present application will only describe the implementation method of the correction instruction for the function tool verification and parsing, while the return results of the successful or failed execution of the function tool are customized by the user according to specific scenario requirements, and the embodiments of the present application do not make specific restrictions. {agent_scratchpad} will be automatically read from the model tool decision cache and the function tool parsing verification correction instruction cache of the ReAct process in the memory storage unit, that is, agent_scratchpad = the model tool decision cache of the ReAct process + the function tool parsing verification correction instruction cache.
[0088] In some embodiments of the present application, since the user will have multiple rounds of conversations during the process of interacting with the large language model, it is necessary to save these historical conversation records. In the embodiments of the present application, they are stored in the historical conversation module of the memory storage unit. The historical conversation module stores the previous conversation records between the user and the large language model after the ReAct process ends, while the conversation records between the user and the large language model generated during the ReAct process will be temporarily stored in the conversation cache during the ReAct process of the memory storage unit. After the ReAct process ends, the conversation cache during the ReAct process of the memory storage unit will be updated and saved to the historical conversation module to become the long-term memory between the user and the large language model.
[0089] As an optional implementation manner, when the user first interacts with the large language model, the historical conversation, the model tool decision cache of the ReAct process, the function tool parsing verification correction instruction cache, and the conversation cache during the ReAct process in the memory storage unit are all initialized with null values, and the original problem cache will save the original problem or request input by the user.
[0090] Therefore, when the user inputs a question or a request, it is necessary to combine the ReAct framework system instructions, the historical conversation, and the ReAct framework user instructions into the input text of the large language model. Assuming that the special symbols corresponding to the System instruction, User instruction, and Assistant instruction roles of the large language model are <|system|>, <|user|>, and <|assistant|> respectively, exemplarily, the specific format of the large language model input text is as follows:
[0091] <|system|>
[0092] The text content of the ReAct framework system instructions (functions and tools have been predefined here, and descriptions of tools and tool_names are passed in, including the built-in "user_input" function tool)
[0093] <|user|>
[0094] ... (historical conversation)
[0095] <|assistant|>
[0096] ... (historical conversation)
[0097] <|user|>
[0098] The text content of the ReAct framework user instructions (here, the question will be read from the original question cache of the memory storage unit, and the agent_scratchpad will be read from the model tool decision cache and the function tool parsing, verification, and correction instruction cache of the ReAct process in the memory storage unit. The text content of this part will be automatically adjusted cyclically as the ReAct process progresses)
[0099] It should be noted that the format described above is exemplary. The specific instruction role symbols and the splicing method of the symbols and the text are determined by the settings of each large language model manufacturer, and the embodiments of this application do not make specific restrictions.
[0100] Step S204, input the input text into the target large language model and obtain the first output result output by the target large language model. Among them, the first output result includes the relevant information of the called function tool, and the first output result is the output result that passes the verification;
[0101] It should be noted that most large language models support the function of specifying stop words. For example, when using the official API interfaces of some large language models, the parameter for the stop word function can be "stop". As long as stop = "stop word" is set, when the model generates each word, once the generated word is equal to the stop word, the generation process will terminate, and the text content generated before that will be directly returned as the output result. It should be noted that the stop word function is a common function for generative AI. Whether it is the API interface of a commercial model or self-deploying using an open-source model, the stop word function can be implemented. In the embodiments of this application, the specific implementation of the stop word is not limited.
[0102] As an alternative implementation, after preparing the input text of the large language model, set the stop word of the large language model to "Observation:". This is because under the system instructions of the ReAct framework, the large language model will reason and generate relevant content in the order of "Thought: / Action: / Observation:", and since the content of "Observation:" is passed in from the outside and not generated by the large language model, setting the stop word to "Observation:", the large language model will only output the text content of the "Thought: / Action:" part. That is to say, the text format of the output result of a large language model that meets the instruction requirements should be as follows:
[0103] Thought: Logical reasoning and thinking statements of the large language model itself
[0104] Action:
[0105] ```json
[0106] {{
[0107] "action":"Name of the function tool"
[0108] "action_input":{"Parameter name 1":Parameter value,"Parameter name 2":Parameter value,...}
[0109] }}
[0110] In the technical solution provided in step S204, the steps of inputting the input text into the target large language model and obtaining the first output result output by the target large language model include: obtaining a third output result generated by the target large language model based on the input text; performing format verification on the third output result, where the format verification includes verifying whether the function tool name and the format of the input parameters in the third output result are in a preset format; in the case where the third output result passes the format verification, performing accuracy verification on the third output result, where the accuracy verification includes verifying whether the function tool name is in a preset tool name list and verifying whether the input parameters conform to preset rules; in the case where the third output result passes the accuracy verification, determining the third output result as the first output result.
[0111] As an optional implementation manner, after performing format verification on the third output result, in the case where the third output result fails to pass the format verification, a first correction instruction may be generated according to the verification result of the format verification; the first correction instruction and the second output result are stored in the cache as additional content; the input text is updated according to the additional content in the cache, and a fourth output result generated by the large language model based on the updated input text is obtained; after each fourth output result is obtained, format verification is performed on the fourth output result until the fourth output result passes the format verification, and then the fourth output result that passes the format verification is used as the third output result.
[0112] As an optional implementation manner, after performing accuracy verification on the second output result, in the case where the third output result fails to pass the accuracy verification, a second correction instruction may be generated according to the verification result of the accuracy verification; the second correction instruction and the second output result are stored in the cache as additional content; the input text is updated according to the additional content in the cache, and a fifth output result generated by the large language model based on the updated input text is obtained; after each fifth output result is obtained, accuracy verification is performed on the fifth output result until the fifth output result passes the accuracy verification, and then the fifth output result that passes the accuracy verification is used as the first output result.
[0113] In some embodiments of the present application, the cache may also be cleared in the case where the fifth output result passes the accuracy verification. It should be noted that the cleared cache includes the first correction instruction for format verification and the second correction instruction for accuracy verification.
[0114] As an optional implementation manner, for the text content in the ReAct framework format output by the large language model, an in-loop accuracy verification parsing of whether the function tool name and the input parameters conform to the JSON format can be performed through a model output format parser, so that the large language model outputs a function tool name and an input parameter parsing text that meet the JSON format requirements.
[0115] Due to the inevitable hallucination problem of large language models, that is, for knowledge content that large language models have never seen during the pre-training stage, they will fabricate it randomly, or due to the model itself generating biased text content during the process of calculating and predicting the next word, this text content may deviate from factual statements. For example, it does not generate content according to the specified format requirements, or when the user does not provide relevant information, the large language model makes assumptions or guesses on its own and generates information content that does not conform to the user's facts.
[0116] Therefore, the text content generated by large language models under the ReAct framework instructions may not necessarily strictly conform to the format requirements. For models with a larger parameter scale, the hallucination problem will be less, but for models with a smaller parameter scale, the hallucination problem will be more prominent.
[0117] To solve this problem, the embodiments of this application provide an Figure 3 inner loop verification process as shown. As can be seen from Figure 3 , in the embodiments of this application, a model output format parser (Output Parser) is used to verify the text content generated by the large language model in JSON format. If the text content generated by the large language model does not conform to the JSON format requirements of the ReAct framework, corresponding function tool verification and parsing correction instructions will be generated. And the observed value of its misaligned "Observation", together with the text content generated by the large language model, will be stored in the function tool parsing verification correction instruction cache of the memory storage unit in the form of appended content. Then enter the inner loop process, read the original question cache of the memory storage unit, the model tool decision cache of the ReAct process, and the function tool parsing verification correction instruction cache by the ReAct framework user instruction, and update the value of the {question} variable in the ReAct framework user instruction template with the text content of the original question cache, and update the value of the {agent_scratchpad} variable in the ReAct framework user instruction template with the text content after splicing the model tool decision cache of the ReAct process and the function tool parsing verification correction instruction cache, and recombine them into a new large language model input text and pass it into the large language model to perform a new round of model text output in the ReAct framework format. Thus, an inner loop verification process for whether the function tool name and input parameters conform to the JSON format is formed.
[0118] In some embodiments of this application, the verification effect obtained after executing the Figure 3 verification process shown is as Figure 4 shown.
[0119] As an alternative implementation, the specific parsing and verification algorithm process of the model output format parser (Output Parser) in the inner loop during the current round of parsing is as follows:
[0120] In the first step, use regular expressions to define the output text format that conforms to the ReAct framework instruction format. Here, there are two types of format parsing. The first is the verification and parsing of the "Thought: / Action:" structure, and its regular expression is:
[0121] output_pattern = r"Thought:(.*?)Action:(.*)"
[0122] The second is the verification and parsing of the JSON text format of the function tool name and input parameters inside the "Action" content, and its regular expression is:
[0123] action_pattern = r"```[a-zA-Z]*\n(.*?)```"
[0124] In the second step, assume that text is the text content generated by the large language model. Perform format parsing through the regular expression output_pattern. If text does not meet the format requirements of output_pattern, the verification fails, and the corrective instruction is:
[0125] Observation:Error: Please continue writing according to the "Thought: / Action: / Observation:" format!
[0126] In the third step, if text meets the format requirements of output_pattern, assume that thought and action are the contents of "Thought:" and "Action:" respectively extracted from the model output text. If the text length of thought is zero, the verification fails, and the corrective instruction is:
[0127] Observation:Error: The content of "Thought:" cannot be empty. Please try again!
[0128] In the fourth step, if the text length of thought is greater than zero, then determine whether the content of action meets the format requirements of action_pattern. If not, the verification fails, and the corrective instruction is:
[0129] Observation: Error: The content of "Action:" should start with "```json\n" and end with "\n```", and the middle part should be JSON structure parameters without " / / ..." content. Please try again!
[0130] Step 5: If the action meets the format requirements of action_pattern, assume action_string is the function tool name and input parameters in JSON text format extracted from the action. Execute eval(action_string) through the eval() function in Python. If the execution fails, it means that action_string is not in a legal JSON text format, and the verification fails. The corrected instruction is still:
[0131] Observation: Error: The content of "Action:" should start with "```json\n" and end with "\n```", and the middle part should be JSON structure parameters without " / / ..." content. Please try again!
[0132] Step 6: If action_string is in a legal JSON text format, at this time action_string is in a dictionary format. Determine whether the "action" key name is included in action_string. If not, the verification fails, and the corrected instruction is:
[0133] Observation: Error: The JSON structure in the "Action:" content is missing the parameter "action". Please re - output!
[0134] Step 7: If the dictionary format of action_string contains the "action" key name, then determine whether the "action_input" key name is included. If not, the verification fails, and the corrected instruction is:
[0135] Observation: Error: The JSON structure in the "Action:" content is missing the parameter "action_input". Please re - output!
[0136] Step 8: If the dictionary format of action_string contains both the "action" and "action_input" parameters, since the value of the "action" parameter should be the name of the function tool and must be in string format, if it does not meet this requirement, the verification fails, and the corrected instruction is:
[0137] Observation: Error: The value type of the JSON structure parameter "action" in "Action:" content is String. Please re - output!
[0138] Step 9, if the value of the "action" parameter in the dictionary format of action_string conforms to the string format, since the value of the "action_input" parameter is the input parameter of the function tool and must be in dictionary format, if it does not meet the requirement, the verification fails, and the corrective instruction is:
[0139] Observation: Error: The value type of the JSON structure parameter "action_input" in "Action:" content is Dict. Please re - output!
[0140] Step 10, if all of the above steps 1 - 9 pass the verification, the output text of the large - language model fully conforms to the output format requirements of the ReAct framework instructions and can pass the verification of the model output format parser in the inner loop.
[0141] As an optional implementation, when the output text of the large - language model passes the verification of the model output format parser in the inner loop, the model output format parser will output the function tool name and input parameters parsed, which is in a dictionary format, in the form of:
[0142] {
[0143] "action": "function tool name"
[0144] "action_input": {"parameter name 1": parameter value, "parameter name 2": parameter value,...}
[0145] }
[0146] In some embodiments of the present application, after the format verification passes, the function tool name and input parameter parsing text that conform to the JSON format can be parsed, and the inner - loop accuracy verification of whether the function tool name is a valid tool name and whether the input parameters are compliant can be performed through the tool parsing validator, and the function tool name and input parameter parsing text that pass the compliance verification are output, and the inner - loop process ends.
[0147] Optionally, after the output text of the large language model passes the verification of the model output format parser, it can only be confirmed that the format of the model output text conforms to the format requirements of the ReAct framework instructions. Due to the existence of the hallucination problem of the large language model, it is possible that the model fabricates a non-existent function tool name. Or even if the function tool name is correct, its input parameters are fabricated by the large language model itself. Or the large language model treats the parameters of other function tools as the parameters of another function tool. Therefore, further accuracy verification of the function tool name and input parameters needs to be performed through the ToolValidator during the inner loop process.
[0148] As Figure 3 shown, if the function tool name and input parameters in the output result after format verification cannot pass the verification of the Tool Validator, corresponding correction instructions for function tool verification and parsing will also be generated. And the correction instructions will be used as the observation value of "Observation", together with the text content generated by the large language model, and stored in the function tool parsing verification correction instruction cache of the memory storage unit in the form of appended content, and then enter the inner loop process again.
[0149] As an optional implementation manner, the Tool Validator includes a unified pre-tool verification rule (Pre Validator) and a custom post-tool verification rule (Post Validator). The custom post-tool verification rule can be set by the user according to actual needs and is not limited in the embodiments of this application. The verification effect of the ReAct inner loop through the Tool Validator is as Figure 4 shown.
[0150] Optionally, assume that T0 is the name of the function tool parsed from the data result passing the format verification, T = [T1, T2,...] is the list of all set function tool names, but T does not include the "final_answer" function name, S0 = {arg1, agr2,...} is the set composed of the input parameter names parsed corresponding to T0, and S = [S1, S2,...] is the list of sets composed of the input parameter names of all set function tools T = [T1, T2,...]. The specific algorithm flow of the unified pre-tool verification rule (PreValidator) of the Tool Validator in the current round of verification during the inner loop is as follows:
[0151] In the first step, when T0 ≠ "final_answer" and is not included in the list of T, the verification fails, and the correction instruction is:
[0152] Observation: Error: T0 is not a valid tool. You can only select one from [T1, T2,...]. Please select the correct tool and re-enter!
[0153] Step 2: When T0 ≠ "final_answer" and T0 is included in the list of T, it means T0 is a valid function tool. Calculate the complement of set S0 with respect to each set S in the set list S, where i = 1, 2,... i SD
[0154] SD i = symmetric_difference(S0, S i )
[0155] Then calculate the number of elements in each complement SD i to obtain the corresponding list of the number of elements in the complement sets N = [N1, N2,...]. Assume that N min is the smallest number of elements in the complement set, and the corresponding pre-set function tool is T min , and the set of its input parameter names is S min . The smallest number of elements in the complement set indicates that the set of input parameter names S min corresponding to the pre-set function tool T min is closer to the set of input parameter names S0 corresponding to the parsed function tool T0. Therefore, T min is more likely to be the correct tool to be called. If T0 ≠ T min at this time, the verification fails, and the correction instruction is:
[0156] Observation: Error: The currently input parameter S0 is closer to the input parameter S min required by the tool T min . Please correct the value of the "action" parameter in the JSON structure parameter of the Action: content to T min , and correct the value of the "action_input" parameter to the dictionary format input parameter S min required by the tool T min .
[0157] Step 3: When T0 ≠ "final_answer", if T0 = T min , it means the parsed function tool is correct. Next, it is necessary to verify whether the set of input parameter names S0 corresponding to T0 is correct. Assume that the correct set of input parameter names originally required by T0 is S c . Calculate the difference set of S c with respect to S0, that is, the elements that belong to Sc but not the set that does not belong to S0, and the difference set of S0 with respect to S c , that is, the set that belongs to S0 but not to S c :
[0158] D c = S c - S0
[0159] D0 = S0 - S c
[0160] Step 4, when T0 ≠ "final_answer", if D c is not empty, but D0 is empty, the verification fails, and the correction instruction is:
[0161] Observation: Error: The input parameter required by T0 is S c , but the above input parameter is missing D c , and D needs to be added c parameter!
[0162] Step 5, when T0 ≠ "final_answer", if D0 is not empty, but D c is empty, the verification fails, and the correction instruction is:
[0163] Observation: Error: The input parameter required by T0 is S c , but the above input parameter has an extra D0, and the D0 parameter needs to be removed!
[0164] Step 6, when T0 ≠ "final_answer", if D c and D0 are both not empty, the verification fails, and the correction instruction is:
[0165] Observation: Error: The input parameter required by T0 is S c , but the above input parameter is missing D c , has an extra D0, and D needs to be added c parameter and remove the D0 parameter!
[0166] Step 7, when T0 = "final_answer", and the number of elements in the set S0 corresponding to the parsed input parameter names is not equal to 1 or the element value of the set S0 is not equal to "answer", the verification fails, and the correction instruction is:
[0167] Observation: Error: The value format of "action_input" of "final_answer" is {"answer": String}, please correct and re - output!
[0168] Step 8: If the above steps 1 to 7 all pass the verification, it is determined that the function tool name and input parameters in the output result have passed the unified pre-validator rule. Next, the verification process of the custom post-validator rule will be entered.
[0169] As an optional implementation, since the custom post-validator rule involves the specific rule requirements of each tool for its input parameters. Exemplarily, assume that T0 is a function tool that allows users to provide a mobile phone number for traffic query. The function tool name of T0 is "business_search", and its input parameters are "phone_nbr" and "search_type", and the value range of "search_type" is ["traffic", "phone bill", "call duration"]. Then the algorithm flow of its custom post-validator rule can be designed as follows:
[0170] Step 1: Determine whether the string value of "phone_nbr" only contains Arabic numerals. If not, the verification fails, and the correction instruction is:
[0171] Observation: Error: The value of the parameter "phone_nbr" of "business_search" is the mobile phone number provided by the user. If the user does not provide it, ask the user.
[0172] Step 2: If the string value of "phone_nbr" only contains Arabic numerals, determine whether it is an 11-digit number. If not, the verification fails, and the correction instruction is:
[0173] Observation: Error: The value of the parameter "phone_nbr" of "business_search" is not the mobile phone number provided by the user. The number of digits is 11. Ask the user.
[0174] Step 3: If the string value of "phone_nbr" meets the requirement of 11-digit Arabic numerals, determine whether the first three digits of the string are within the mobile phone number segments of telecom operators, such as 143, 189, etc. If not, the verification fails, and the correction instruction is:
[0175] Observation: Error: The value of the parameter "phone_nbr" of "business_search" is not a valid mobile phone number. If the user does not provide it, ask the user.
[0176] Step 4: Check whether the string value of "search_type" is within the range of ["traffic", "phone charges", "call duration"]. If not, the check fails. The correct instruction is:
[0177] Observation:Error:The value of the parameter "search_type" of "business_search" can only be selected from ["traffic", "phone charges", "call duration"]. Please reselect the correct value.
[0178] Step 5: If the above steps 1 to 4 are verified, it is determined that the function tool name and input parameters in the output result have passed the verification process of the custom post-tool verification rule (Post Validator).
[0179] In some embodiments of the present application, when the function tool name and input parameters in the output result pass the unified pre-tool validation rule (Pre Validator) and the custom post-tool validation rule (Post Validator) in turn, the inner loop process can be ended, and the function tool parsing verification correction instruction cache of the memory storage unit is cleared. At this point, the parsing and verification of the function tool is completed, and the final correct function tool name and input parameters are output, and the output text of the large language model is output at the same time. The model output text at this time is the output text of the successful parsing of the function tool.
[0180] It should be noted that when multiple function tools are called, each function tool has its corresponding inner loop verification process. When multiple function tools need to be called, for each function tool to be called, the large language model will output the output results for each function tool in turn, and then execute the above inner loop verification process. When the verification passes, the function tool execution process will be called, and after the execution result is returned to the large language model, the large language model will then determine the next function tool to be called, and then execute the inner loop verification process for the next function tool, and so on, until all the function tools that need to be called are called.
[0181] Step S206, executing the calling process of the function tool corresponding to the relevant information according to the first output result, and if it is determined during the execution that the target object needs to input reference information, obtaining the question text for indicating the reference information generated by the target large language model and displaying it to the target object;
[0182] In the technical solution provided in step S206, the process of invoking the function tool corresponding to the relevant information according to the first output result includes: determining the function tool name in the first output result; when the function tool name is a preset function tool name, obtaining the question text and the response text corresponding to the question text, and storing the question text and the response text as additional content in the decision cache.
[0183] Step S208, after obtaining the response text input by the target object for answering the question text, correct the first output result according to the response text and the question text to obtain a second output result, and continue to execute the function tool invocation process according to the second output result.
[0184] As an optional implementation manner, the step of correcting the first output result to obtain the second output result according to the response text and the question text after obtaining the response text input by the target object for answering the question text includes: correcting the input text according to the additional content in the decision cache, and invoking the large language model to process the corrected input text to obtain the second output result.
[0185] In some embodiments of the present application, after the end of the inner loop tool verification and parsing process, the function tool names and input parameters that pass the compliance verification will be obtained. Then, the tool executor can be called to execute the tool, and the model output result of the successful parsing of the function tool and the execution result of the tool are stored in the cache. After the user instruction template of the ReAct framework reads and splices the cache information, it enters the outer loop process of the large language model scheduling. The large language model will perform the next round of decision-making reasoning and function tool selection and invocation.
[0186] In the related art, the ReAct prompt engineering framework does not allow interruption during the process of multi-tool serial scheduling. However, the ReAct framework multi-tool serial scheduling process provided by the embodiments of the present application includes two scenarios. The first is the ReAct outer loop multi-tool serial scheduling process without interruption in the middle, and the second is the ReAct outer loop multi-tool serial scheduling process with interruption in the middle.
[0187] Optionally, as Figure 5 shown, the ReAct outer loop multi-tool serial scheduling process without interruption in the middle is as follows:
[0188] The first step is to call the tool executor (ToolExecutor) to execute the function tool for the function tool names and input parameters that pass the compliance verification.
[0189] Step 2: If the name of the function tool to be executed is not the built-in "user_input", then use the execution result of the tool as the observation value of "Observation", and store it together with the text content generated by the large language model, that is, the model output text with successful function tool parsing, in the model tool decision cache of the ReAct process in the memory storage unit in the form of appended content. The model output text with successful function tool parsing refers to the text contained in the output result where the corresponding function tool name and input parameters pass the inner loop verification. The format of the appended text content is as follows:
[0190] Thought: Logical reasoning and thinking statements of the large language model itself
[0191]
[0192] Step 3: The ReAct framework user instruction reads the original question cache in the memory storage unit, the model tool decision cache in the ReAct process, and the function tool parsing verification and correction instruction cache (which has been emptied at this time due to the end of the inner loop process). Then, use the text content of the original question cache to update the value of the {question} variable in the ReAct framework user instruction template, and use the text content obtained by splicing the model tool decision cache in the ReAct process and the function tool parsing verification and correction instruction cache to update the value of the {agent_scratchpad} variable in the ReAct framework user instruction template. Then, recombine them into a new large language model input text, and then pass it into the large language model for a new round of model text output in the ReAct framework format. The large language model will perform the next round of decision-making reasoning and function tool selection and call, and enter the outer loop process.
[0193] In the function tool call process without interruption in the middle, the content cached in the cache and the final execution result are as Figure 6 shown. As can be seen from Figure 6 , after the ReAct framework user instruction reads the cache in the memory storage unit, it enters the outer loop for multi-tool serial scheduling and returns the final answer to the user.
[0194] As an alternative implementation, as shown in Figure 7 and Figure 9 , the interrupted ReAct outer loop multi-tool serial scheduling process is divided into two stages: before and after the interruption. The process of the pre-interruption stage is as shown in Figure 7 and includes the following steps:
[0195] Step 1: For the function tool name and input parameters that pass the compliance verification, call the tool executor (ToolExecutor) to execute the function tool.
[0196] Step 2: If the name of the function tool to be executed is the built-in "user_input", the observation value of "Observation" is first set to null, and together with the text content generated by the large language model, that is, the model output text with successful function tool parsing, it is stored in the model tool decision cache of the ReAct process in the memory storage unit in the form of appended content. The format of the appended text content is as follows:
[0197] Thought: The logical reasoning and thinking statements of the large language model itself
[0198]
[0199] Step 3: During the entire interaction between the user and the large language model, the conversations between the user and the large language model generated during the ReAct process will be temporarily stored in the session cache during the ReAct process in the memory storage unit. Since the large language model selects the "user_input" tool to ask questions to the user at this time, an interaction with the user is generated. Therefore, the value of the input parameter of the "user_input" function tool, that is, the question asked by the model to the user, needs to be stored in the session cache during the ReAct process in the memory storage unit. The content of the session cache during the ReAct process is retained until the end of the ReAct process, and then the question asked by the model to the user is returned to the user to complete an interrupted interaction. This is the first half of the process.
[0200] As an optional implementation, the cache content and execution results in the first half of the interrupted ReAct outer loop multi-tool serial scheduling process are as Figure 8 shown. As can be seen from Figure 8 , after the ReAct framework reads the cache of the memory storage unit for the user instruction, it enters the next round of inner loop function tool scheduling decision reasoning. In some exemplary embodiments, such as in the user query traffic scenario, since the user's original question is to query traffic but does not provide a mobile phone number, the large language model independently reasons that it needs to call the built-in "user_input" tool to ask the user for the mobile phone number, thus interrupting the ReAct outer loop process.
[0201] In some embodiments of the present application, the process after interruption is as Figure 9 shown, including the following steps:
[0202] In the first step, the question information returned by the first half of the process is presented to the user. After the user receives the question information returned by the first half of the ReAct outer loop process, the user will input the content of the user's answer to the model. At this time, since the original question cache in the memory storage unit stores the user's original question at the beginning, the input of the user in the second half of the process will not update the information in the original question cache, but will be used as the result after the execution of the "user_input" function tool in the first half of the process, that is, as the observation value of "Observation", and stored in the model tool decision cache of the ReAct process in the memory storage unit. At this time, together with the content already stored in the model tool decision cache of the ReAct process in the memory storage unit in the first half, it is appended to the text content format of the model tool decision cache of the ReAct process as follows:
[0203] Thought: The logical reasoning and thinking statements of the large language model itself
[0204]
[0205] In the second step, the ReAct framework user instruction reads the original question cache in the memory storage unit, the model tool decision cache of the ReAct process (at this time, it combines the output text that has been successfully parsed by the function tool and the result of the tool execution stored in the previous and subsequent two stages), and the function tool parsing verification and correction instruction cache (which has been emptied at this time due to the end of the inner loop process), and updates the value of the {question} variable in the ReAct framework user instruction template with the text content of the original question cache, and updates the value of the {agent_scratchpad} variable in the ReAct framework user instruction template with the text content after splicing the model tool decision cache of the ReAct process and the function tool parsing verification and correction instruction cache, and returns to step S101 to recombine into a new large language model input text, and then passes it into the large language model for a new round of model text output in the ReAct framework format. The large language model will perform the next round of decision-making reasoning and function tool selection and call, and enter the outer loop process.
[0206] In some embodiments of the present application, Figure 10Shows the cache content and execution results of the second half of the interrupted ReAct outer loop multi-tool serial scheduling process. In the traffic query scenario, the large language model interacts with the user midway and asks the user for the mobile phone number. The answer input by the user from the outside will be used as the execution result of the function tool "user_input" output in the first half, and will be stored as the observation value of "Observation" in the model tool decision cache of the ReAct process in the memory storage unit. At this time, the ReAct framework resumes the interrupted ReAct outer loop process after reading the cache in the memory storage unit. Since the user has supplemented the number information, the large language model outputs the correct function tool parsing text in the next outer loop and successfully obtains the execution result of the tool. Then, in the next outer loop, since the large language model has obtained the final answer to the user's original question, it calls the built-in "final_answer" tool to give the final answer.
[0207] When the large language model infers on its own that it needs to use the built-in "final_answer" function tool, it means that the large language model has obtained the final answer and can return the final answer to the user. At this time, the outer loop process ends.
[0208] In some embodiments of the present application, when the large language model determines that it can call the final answer, it ends the outer loop process, saves the conversation record generated during the ReAct process, and returns the final answer to the user.
[0209] As Figure 11 shown, after the large language model obtains the final answer, it means that the entire ReAct process has ended. At this time, the final answer can be returned to the user, and at the same time, the final answer of the model is stored in the conversation cache during the ReAct process in the memory storage unit. Then, the temporary conversation record generated during the ReAct process is stored in the historical conversation module of the memory storage unit as long-term memory. After the historical conversation record is updated, since the ReAct process has ended, the original question cache in the memory storage unit, the model tool decision cache of the ReAct process, the function tool parsing verification and correction instruction cache, and the conversation cache during the ReAct process are all cleared and return to the initial state to accurately enter the next round of user Q&A.
[0210] According to the embodiments of the present application, there is also provided a function tool call process as Figure 12 shown, including the following steps:
[0211] S1202: Construct the input text of the large language model through two prompt templates of the specially designed ReAct framework system instruction and the user instruction, in combination with the historical conversation between the user and the large language model;
[0212] S1204: Input the input text of the constructed large language model into the large language model. The large language model makes decision reasoning and selects and invokes function tools based on its own logical reasoning ability, and outputs a function tool call decision text in the ReAct framework format;
[0213] S1206: For the text content in the ReAct framework format output by the model, perform an in-loop accuracy verification and parsing on whether the function tool name and input parameters conform to the JSON format through a model output format parser, and output a function tool name and input parameter parsing text that meets the JSON format requirements;
[0214] S1208: For the function tool name and input parameter parsing text that conform to the JSON format, perform an in-loop accuracy verification on whether the function tool name is a valid tool name and whether the input parameters are compliant through a tool parsing validator, output a function tool name and input parameter parsing text that pass the compliance verification, and end the in-loop process;
[0215] S1210: When the in-loop tool verification and parsing process ends, for the function tool name and input parameters that pass the compliance verification obtained, call the tool executor to execute the tool, and store the model output result of successful function tool parsing and the execution result of the tool in the cache. After the ReAct framework's user instruction template reads and splices the cache information, it enters the out-of-loop process of large language model scheduling. The large language model will perform the next round of decision reasoning and function tool selection and invocation;
[0216] S1212: When the large language model determines that the final answer can be invoked, end the out-of-loop process, save the session record generated during the ReAct process at the same time, and return the final answer to the user.
[0217] Such as Figure 13As shown, by generating the input text of the target large language model based on the initial text input by the target object, the prompt template, and the historical conversation data of the target object; inputting the input text into the target large language model, and obtaining the first output result output by the target large language model, where the first output result includes the relevant information of the function tool called, and the first output result is the output result that passes the verification; according to the first output result, execute the call process of the function tool corresponding to the relevant information, and when it is determined that the target object needs to input reference information during the execution process, obtain the question text generated by the target large language model for indicating the reference information and display it to the target object; after obtaining the response text input by the target object for answering the question text, correct the first output result according to the response text and the question text to obtain the second output result, and continue to execute the call process of the function tool according to the second output result. By judging whether the target object needs to input reference information to display the question text generated by the large model and obtain the response text input by the target object, the purpose of interrupting the inference process of the large model is achieved, thereby realizing the technical effect of eliminating the large model hallucination caused by unclear input content, and further solving the technical problem that the function tool call cannot be realized according to the result output by the large language model due to the inability to solve the large language model hallucination problem in the related technology.
[0218] In addition, the function tool call method provided by the embodiments of the present application can automatically parse and accurately verify the deviated text content generated by the large language model during the inference process of calling the function tool due to the hallucination problem, and generate an error correction instruction through a triple parsing and verification algorithm process, so that the large language model can complete self-thinking and automatic correction under the internal and external double-loop ReAct inference framework, thereby greatly improving the accuracy and stability of the large language model in calling the function tool. Especially in the multi-round dialogue scenario between the user and the large language model, the sustainability, stability, and accuracy of the function tool call of the large language model have been greatly improved, thus improving the practicality of the large language model in the actual production environment.
[0219] Specifically, due to the inevitable hallucination problem of large language models, when responding to the system and the user's structured information output instructions, it cannot ensure that each output strictly conforms to the structured format requirements of the instructions. That is to say, the output of the model is not stable, and the hallucination phenomenon will cause the model to fabricate out of thin air and generate error information that does not meet expectations. In view of the hallucination problem of large language models, in the process of parsing the text content output by them, this application embodiment successively passes through a triple parsing and verification algorithm process of an Output Parser, a Pre Validator for unified pre-tool verification rules, and a Post Validator for custom post-tool verification rules, and realizes the self-thinking and automatic correction mechanism of large language models through the ReAct framework prompt instructions. For the deviated text content format generated by large language models due to hallucination problems, the triple parsing and verification algorithm greatly reduces the risk of large language models outputting error information to users, and enables large language models to automatically correct errors in the inner loop, improving the accuracy and practicality of large language models in calling function tools.
[0220] In addition, due to two common problems in current large language models, one problem is the accuracy of understanding long texts. As the length of the input context gradually increases, the semantic understanding and logical reasoning abilities of large language models gradually decline. Especially for large language models with a relatively small parameter scale, the longer the input text content, the more unstable the output text quality. Another problem is that large language models are also quite sensitive to the quality of the input text. Input texts with chaotic logic or messy arrangements will affect the internal semantic understanding and logical reasoning of large language models. If the input text contains a lot of error information and error correction instructions, there is a high probability that the large language model will crash in the subsequent multi-round conversations, including function tool call errors, a decline in self-correction ability, and the phenomenon of repeating the same content continuously like a repeater. The ReAct framework implemented in the embodiments of this application uses an internal and external double-loop method. In the internal loop, function tool parsing and verification are implemented. For the incorrect format text output by the large language model and the error correction instructions generated during the verification process, they will only be stored in a temporary memory function tool parsing verification correction instruction cache in the internal loop. When the large language model completes self-correction and successfully outputs the correct function tool parsing text content, as the internal loop ends, the function tool parsing verification correction instruction cache in the memory storage unit will be automatically cleared, and the error information and error correction instructions will not be carried into the next round of function tool call parsing in the external loop. Incorrect or messy texts will only appear in the internal loop, ensuring that the function tool parsing text content stored in the model tool decision cache of the ReAct process in the memory storage unit during the external loop is correct content. On the one hand, this method can greatly compress the length of the text input to the large language model during multi-round conversations. On the other hand, it avoids the influence of previous error information and error correction instructions on the logical reasoning of the large language model in the next round of function tool call decisions. This internal and external double-loop framework greatly improves the sustainability and accuracy of function tool calls during the multi-round conversations between users and large language models.
[0221] It should be noted that during the multi-tool serial scheduling process of the ReAct prompting engineering framework in the related technology, interruption is not allowed midway. This will cause the large language model to fabricate information that the user did not provide due to the existence of hallucination problems when the user's input question or request is unclear. Although the large language model can be prohibited from fabricating information through system instructions or user instructions, for models with a relatively small parameter scale, the influence of the prompt words is often unstable and cannot effectively reduce the hallucination problem of the large language model. The multi-tool serial scheduling process of the ReAct framework implemented in the embodiments of this application includes two scenarios. The first is the ReAct outer loop multi-tool serial scheduling process without interruption midway, which is the conventional ReAct framework loop process; the second is the ReAct outer loop multi-tool serial scheduling process that can be interrupted midway. This method realizes the midway interruption of the ReAct process by introducing the built-in function tool "user_input". When the large language model responds through self-inference or corrective instructions parsed and verified by the function tool, the large language model can achieve midway interaction with the user through the function tool "user_input". When the user's input question or request is unclear or the user does not provide the input parameter information required by the function tool, by asking the user questions, the user's next answer is used to timely intervene and supplement information to the large language model, greatly reducing the phenomenon that the large language model fabricates information due to incomplete input request information in the early stage.
[0222] The embodiments of this application provide a function tool call device. Figure 14 It is the structural schematic diagram of the device. As can be seen Figure 14 from it, the device includes: a first processing module 140, configured to generate the input text of the target large language model according to the initial text input by the target object, the prompt word template, and the historical conversation data of the target object; a second processing module 142, configured to input the input text into the target large language model and obtain the first output result output by the target large language model, where the first output result includes the relevant information of the called function tool, and the first output result is the output result that passes the verification; a third processing module 144, configured to execute the call process of the function tool corresponding to the relevant information according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, obtain the question text generated by the target large language model for indicating the reference information and display it to the target object; a fourth processing module 146, configured to, after obtaining the reply text input by the target object for replying to the question text, correct the first output result according to the reply text and the question text to obtain the second output result, and continue to execute the call process of the function tool according to the second output result.
[0223] In some embodiments of the present application, the prompt template includes a system instruction prompt template and a user instruction prompt template; the steps for the first processing module 140 to generate the input text of the target large language model based on the initial text input by the target object, the prompt template, and the historical conversation data of the target object include: constructing the input text of the target large language model based on the initial text input by the target object, the system instruction prompt template, the user instruction prompt template, and the historical conversation data of the target object, where the system instruction prompt template is used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt template is used to reflect the problem description information indicated by the initial text.
[0224] In some embodiments of the present application, the steps for the second processing module 142 to input the input text into the target large language model and obtain the first output result output by the target large language model include: obtaining the third output result generated by the target large language model based on the input text; performing format verification on the third output result, where the format verification includes verifying whether the function tool name and the format of the input parameters in the third output result are in a preset format; in the case where the third output result passes the format verification, performing accuracy verification on the third output result, where the accuracy verification includes verifying whether the function tool name is in the preset tool name list and verifying whether the input parameters conform to the preset rules; in the case where the third output result passes the accuracy verification, determining the third output result as the first output result.
[0225] In some embodiments of the present application, after performing format verification on the third output result, the second processing module 142 is further configured to: in the case where the third output result fails to pass the format verification, generate a first correction instruction based on the verification result of the format verification; store the first correction instruction and the second output result as additional content in the cache; update the input text based on the additional content in the cache and obtain the fourth output result generated by the large language model based on the updated input text; each time the fourth output result is obtained, perform format verification on the fourth output result until the fourth output result passes the format verification, and then use the fourth output result that has passed the format verification as the third output result.
[0226] In some embodiments of the present application, after verifying the accuracy of the second output result, the second processing module 142 is further configured to: in the case where the third output result fails the accuracy verification, generate a second correction instruction according to the verification result of the accuracy verification; store the second correction instruction and the second output result in the cache as additional content; update the input text according to the additional content in the cache, and obtain a fifth output result generated by the large language model according to the updated input text; after obtaining the fifth output result each time, verify the accuracy of the fifth output result until the fifth output result passes the accuracy verification, and use the fifth output result that passes the accuracy verification as the first output result.
[0227] In some embodiments of the present application, the second processing module 142 is further configured to: in the case where the fifth output result passes the accuracy verification, clear the cache.
[0228] In some embodiments of the present application, the steps for the third processing module 144 to execute the calling process of the function tool corresponding to the relevant information according to the first output result include: determining the function tool name in the first output result; in the case where the function tool name is a preset function tool name, obtaining the question text and the reply text corresponding to the question text, and storing the question text and the reply text in the decision cache as additional content.
[0229] In some embodiments of the present application, after obtaining the reply text input by the target object for replying to the question text, the steps for the fourth processing module 146 to correct the first output result according to the reply text and the question text to obtain the second output result include: correcting the input text according to the additional content in the decision cache, and calling the large language model to process the corrected input text to obtain the second output result.
[0230] It should be noted that each module in the above function tool calling device may be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it may be presented in the following forms, but not limited to: each of the above modules is presented as a processor, or the functions of each of the above modules are implemented by a processor.
[0231] According to an embodiment of the present application, there is also provided a function tool calling system as Figure 15 shown. From Figure 15As can be seen, the system includes: a prompt word management unit 150 for managing prompt word templates, where the prompt word templates include a system instruction prompt word template and a user instruction prompt word template. The system instruction prompt word template is used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt word template is used to reflect the problem description information indicated by the initial text; a large language model unit 152 for scheduling the large language model and configuring a standardized interface of the large language model; a model output format parsing unit 154 for storing various text structured information parsing rules and verification rules of the output results of the large language model; a tool parsing and verification unit 156 for storing and managing the parsing and verification rules of function tools, where the parsing and verification rules include a preset unified pre-tool parsing and verification rule and a custom post-tool parsing and verification rule configured by the target object; a tool execution unit 158 for parsing and verifying the output results of the large language model according to the parsing and verification rules and scheduling the function tools according to the parsing and verification results; a memory storage unit 1510 for storing persistently stored historical session data and temporarily memorized cache data.
[0232] In some embodiments of the present application, the above function tool calling method is also used to execute Figure 2 the function tool calling method shown in
[0233] In some embodiments of the present application, the prompt word management unit 150 is used to store and manage the system instruction prompt word template of the ReAct framework, the user instruction prompt word template of the ReAct framework, and other prompt word templates used for task instruction interaction with the large language model in the entire prompt word engineering technology.
[0234] In some embodiments of the present application, the large language model unit 152 is used to manage the large language model and provides a unified standardized interface to dock various prompt word templates of the upstream prompt word management unit and various output format parsers of the downstream model output format parsing unit.
[0235] In some embodiments of the present application, the model output format parsing unit 154 is used to store and manage various structured information parsing rules and verification rules of the structured output text of the large language model, including the output text structured parser in the ReAct framework format and other output text structured parsers such as JSON structure, dictionary structure, list structure, and string structure.
[0236] In some embodiments of the present application, the tool parsing and verification unit 156 is used to store and manage the parsing and verification rules of function tools, including unified pre-tool verification rules and custom post-tool verification rules. The unified pre-tool verification rules implement unified interface scheduling, while the custom post-tool verification rules provide a standardized interface paradigm, allowing users to combine different function tools to achieve unified interface scheduling using the standardized interface paradigm;
[0237] In some embodiments of the present application, the tool execution unit 158 is used to implement unified function scheduling according to the parsed function tool name and its input parameters, uniformly encapsulate the results of the function tool execution, use the execution result of the tool as the observation value of "Observation" in the ReAct framework prompt template, and splice the text content of the successfully parsed function tool output by the large language model and the execution result of the tool, and store them in the model tool decision cache of the ReAct process in the memory storage unit;
[0238] In some embodiments of the present application, the memory storage unit 1510 is used to store and manage historical session information that needs to be remembered for a long time, as well as the original question cache for temporary memory, the model tool decision cache of the ReAct process, the function tool parsing and verification correction instruction cache, and the session cache during the ReAct process. The content of the historical session is saved for a long time, while the cache information is automatically cleared after each ReAct process ends.
[0239] According to the embodiments of the present application, a non-volatile storage medium is also provided. A program is stored in the non-volatile storage medium. When the program runs, it controls the device where the non-volatile storage medium is located to execute the following function tool calling method: generate the input text of the target large language model based on the initial text input by the target object, the prompt template, and the historical session data of the target object; input the input text into the target large language model, and obtain the first output result output by the target large language model, where the first output result includes the relevant information of the called function tool, and the first output result is the output result that has passed the verification; execute the calling process of the function tool corresponding to the relevant information according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, obtain the question text generated by the target large language model for indicating the reference information and display it to the target object; after obtaining the reply text input by the target object for replying to the question text, correct the first output result according to the reply text and the question text to obtain the second output result, and continue to execute the function tool calling process according to the second output result.
[0240] According to an embodiment of the present application, an electronic device is further provided, including a memory and a processor. The processor is configured to run a program stored in the memory. When the program runs, it executes the following function tool call method: generating an input text for a target large language model based on the initial text input by the target object, a prompt template, and the historical conversation data of the target object; inputting the input text into the target large language model and obtaining a first output result output by the target large language model, where the first output result includes relevant information of the function tool to be called, and the first output result is an output result that passes the verification; executing the call process of the function tool corresponding to the relevant information according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, obtaining a question text generated by the target large language model for indicating the reference information and presenting it to the target object; after obtaining the reply text input by the target object for replying to the question text, correcting the first output result according to the reply text and the question text to obtain a second output result, and continuing to execute the function tool call process according to the second output result.
[0241] According to an embodiment of the present application, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, it implements the following function tool call method: generating an input text for a target large language model based on the initial text input by the target object, a prompt template, and the historical conversation data of the target object; inputting the input text into the target large language model and obtaining a first output result output by the target large language model, where the first output result includes relevant information of the function tool to be called, and the first output result is an output result that passes the verification; executing the call process of the function tool corresponding to the relevant information according to the first output result, and when it is determined that the target object needs to input reference information during the execution process, obtaining a question text generated by the target large language model for indicating the reference information and presenting it to the target object; after obtaining the reply text input by the target object for replying to the question text, correcting the first output result according to the reply text and the question text to obtain a second output result, and continuing to execute the function tool call process according to the second output result.
[0242] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0243] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0244] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0245] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0246] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0247] The above is only the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A function tool calling method, characterized in that: include: Generating input text of a target large language model according to an initial text input by a target object, a prompt word template, and historical conversation data of the target object; Inputting the input text into the target large language model, and obtaining a first output result output by the target large language model, wherein the first output result includes relevant information of the called function tool, and the first output result is an output result that has passed verification; Executing a calling process of a function tool corresponding to the relevant information according to the first output result, and in the case where it is determined during the execution that the target object needs to input reference information, obtaining a question text generated by the target large language model for indicating the reference information and displaying it to the target object; After obtaining the reply text input by the target object to reply to the question text, the first output result is corrected according to the reply text and the question text to obtain the second output result, and the calling process of the function tool is continued according to the second output result.
2. The function tool calling method according to claim 1, characterized in that: Inputting the input text into the target large language model and obtaining a first output result output by the target large language model includes: Obtaining a third output result generated by the target large language model according to the input text; Performing format verification on the third output result, wherein the format verification includes verifying whether the format of the function tool name and the input parameter in the third output result is a preset format; In the case that the third output result passes the format check, performing accuracy check on the third output result, wherein the accuracy check includes checking whether the function tool name is in a preset tool name list, and checking whether the input parameter complies with preset rules; In a case where the third output result passes the accuracy check, the third output result is determined to be the first output result.
3. The function tool calling method according to claim 2, characterized in that: After format checking the third output result, the method further includes: If the third output result fails the format check, generating a first correction instruction according to the check result of the format check; storing the first correction instruction and the second output result as additional content in a cache; updating the input text according to the additional content in the cache, and obtaining a fourth output result generated by the large language model according to the updated input text; After the fourth output result is obtained each time, the format check is performed on the fourth output result until the fourth output result passes the format check, and the fourth output result that passes the format check is used as the third output result.
4. The function tool calling method according to claim 2, characterized in that: After checking the accuracy of the second output result, the method further includes: If the third output result fails the accuracy check, generating a second correction instruction according to the check result of the accuracy check; storing the second correction instruction and the second output result as additional content in a cache; updating the input text according to the additional content in the cache, and obtaining a fifth output result generated by the large language model according to the updated input text; After the fifth output result is obtained each time, the accuracy check is performed on the fifth output result until the fifth output result passes the accuracy check, and the fifth output result that passes the accuracy check is used as the first output result.
5. The function tool calling method according to claim 4, characterized in that: The method further comprises: When the fifth output result passes the accuracy check, the cache is cleared.
6. The function tool calling method according to claim 1, characterized in that: The process of calling the function tool corresponding to the relevant information according to the first output result includes: Determine the function tool name in the first output result; In the case where the function tool name is a preset function tool name, the question text and the answer text corresponding to the question text are obtained, and the question text and the answer text are stored in the decision cache as additional content.
7. The function tool calling method according to claim 6, characterized in that: After obtaining the reply text input by the target object to reply to the question text, correcting the first output result according to the reply text and the question text to obtain the second output result includes: The input text is corrected according to the additional content in the decision cache, and the large language model is called to process the corrected input text to obtain the second output result.
8. The function tool calling method according to claim 1, characterized in that: The prompt word templates include system instruction prompt word templates and user instruction prompt word templates; Generating the input text of the target large language model according to the initial text input by the target object, the prompt word template and the historical conversation data of the target object includes: The input text of the target large language model is constructed based on the initial text input by the target object, the system instruction prompt word template, the user instruction prompt word template and the historical conversation data of the target object, wherein the system instruction prompt word template is used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt word template is used to reflect the problem description information indicated by the initial text.
9. A function tool calling device, characterized in that: include: A first processing module, configured to generate input text of a target large language model according to an initial text input by a target object, a prompt word template and historical conversation data of the target object; a second processing module, configured to input the input text into the target large language model, and obtain a first output result output by the target large language model, wherein the first output result includes relevant information of the called function tool, and the first output result is an output result that has passed verification; A third processing module is used to execute the calling process of the function tool corresponding to the relevant information according to the first output result, and when it is determined during the execution process that the target object needs to input reference information, obtain the question text generated by the target large language model for indicating the reference information and display it to the target object; The fourth processing module is used to obtain the reply text input by the target object to reply to the question text, correct the first output result according to the reply text and the question text to obtain the second output result, and continue to execute the calling process of the function tool according to the second output result.
10. A function tool calling system, characterized in that: The function tool calling method applicable to any one of claims 1 to 8 comprises: a prompt word management unit, used to manage the prompt word templates, wherein the prompt word templates include system instruction prompt word templates and user instruction prompt word templates, the system instruction prompt word templates are used to indicate the generation rule information and output format information followed by the target large language model, and the user instruction prompt word templates are used to reflect the problem description information indicated by the initial text; A large language model unit, used for scheduling the large language model and configuring a standardized interface of the large language model; A model output format parsing unit, used to store various text structured information parsing rules and verification rules of the output results of the large language model; A tool parsing and verification unit, used for storing and managing parsing and verification rules of function tools, wherein the parsing and verification rules include preset unified pre-tool parsing and verification rules and customized post-tool parsing and verification rules configured by the target object; A tool execution unit, configured to perform parsing verification on the output result of the large language model according to the parsing verification rule, and schedule the function tool according to the parsing verification result; The memory storage unit is used to store the historical session data in persistent storage and the cache data in temporary memory.
11. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the function tool calling method described in any one of claims 1 to 8.
12. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the function tool calling method described in any one of claims 1 to 8 when running.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Task execution method, device and system, electronic device and storage medium
CN120723411A