Large-model multi-type tool collaborative reasoning method, system and equipment
Through an adaptive exploration and reinforcement learning strategy based on token entropy, the uncertainty in the reasoning process of large language models is quantified and diversified reasoning paths are generated. This solves the problems of single reasoning paths and rigid decisions in large language models, and improves the accuracy and robustness of the final answer.
Patent Information
- Application Number
- CN202511261103.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
When large language models call external reasoning tools, the reasoning path is single and the decision-making is rigid. In particular, suboptimal paths are easily selected at decision nodes, resulting in uncertainty and insufficient robustness in the reasoning process.
An adaptive exploration and reinforcement learning strategy based on token entropy is adopted. By real-time monitoring of token entropy changes during the reasoning process, the uncertainty of decision nodes is quantified, and the exploration mechanism is triggered at high token entropy nodes to generate diverse reasoning paths. The GRPO algorithm is used to train and optimize the collaborative reasoning capabilities of multiple types of tools in large models.
It improves the intelligence of the reasoning tool selection strategy for large models under uncertain situations, and improves the accuracy and robustness of the final answer.
Smart Images

Figure CN120745849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and natural language processing technology, and in particular to a large-model multi-type tool collaborative reasoning method, system and equipment. Background Art
[0002] Currently, to solve open and complex reasoning tasks, large language models (LLMs) are deeply integrated with external reasoning tools. However, when large language models call external reasoning tools to perform complex reasoning tasks, there are problems such as a single reasoning path and rigid decision-making, especially the tendency to choose suboptimal paths at decision nodes.
[0003] Agent-based reinforcement learning has become a key training paradigm for enabling dynamic interaction between large language models and the environment. Mainstream approaches often use inference path-based optimization algorithms, such as GRPO (Group Relative Policy Optimization). These algorithms sample inference paths invoked by the LLM inference tool and reward participants based on the final results.
[0004] However, this model has a significant limitation: it treats reasoning as a whole, ignoring the fine-grained decisions made by the large model at each inference tool invocation step. Previous research has shown that each tool response introduces a high degree of uncertainty, manifesting as high token entropy. Existing methods, due to their excessive focus on comparing the entire path, underexplore these key nodes, making it difficult to develop diverse reasoning tool usage strategies. This, in turn, limits the optimization potential of the model's reasoning path and the robustness of the resulting solution. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a large-model multi-type tool collaborative reasoning method, system and equipment, designs an adaptive exploration and reinforcement learning strategy based on token entropy, realizes the intelligent improvement of the reasoning tool selection strategy of large models under uncertain conditions, and improves the reasoning accuracy of the final answer.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a large-model multi-type tool collaborative reasoning method, comprising: For the original reasoning path generated by the large model for the problem, calculate the preceding reasoning nodes generated by the large model after each invocation of the reasoning tool. The entropy mean of the token fragments is used as the entropy value of the current inference node; According to the entropy values of the current inference node, the previous inference node, and the initial inference node, the entropy change of the current inference node is determined. According to the comparison between the entropy change and the set entropy change threshold, the high token entropy node is determined. When the next inference tool is called at a high token entropy node, a different inference tool from the original inference path is randomly called and the inference process is continued to generate a new inference path. The entropy value of each inference node in the new inference path is continued to be calculated until there is no high token entropy node. The original reasoning path and all the generated new reasoning paths are used as training sets to train the large model based on the reinforcement learning algorithm; The trained large model is used to generate the final answer to the problem being processed.
[0007] As an alternative implementation, , the process of collaborative reasoning based on multi-type tools using a large model is modeled as follows: ; Among them, the factor Represents the reasoning process of multiple types of tool calls, Representing a chain of thought The number of tokens in For location token, Indicates location All previous tokens; Indicates the introduction of the reasoning toolset Model instructions; Indicates location Feedback from all previous history calls to the inference tool; factor represents the answer generation process, Indicate the answer The number of tokens, For location The model generation results are For location Previous model history generation results.
[0008] As an alternative implementation, the reasoning toolset Including local corpus retrieval tools, web online retrieval tools and code tools.
[0009] As an optional implementation method, during the reasoning process, the thinking process content is filled in <think>Identify the internal; fill in the operation of calling multiple types of tools <tool>The inside of the label contains the name of the reasoning tool and the parameters related to the reasoning tool; the return results of calling multiple types of tools are filled in<tool_result> The mark is used as context information for subsequent reasoning steps; until the maximum number of times the reasoning tool is called is reached or the large model automatically ends reasoning, the reasoning result is filled in <answer>Identify the interior as the final answer.
[0010] As an optional implementation, the entropy value of each token fragment for: ; in, is the dictionary size, For location Token probability; For location The probability of the jth token in the corresponding dictionary; Indicates A large model with learnable parameters.
[0011] As an optional implementation, the entropy change of the current inference node for: ; in, is the entropy value of the current inference node; is the entropy value of the previous inference node; is the entropy value of the initial inference node; is the dictionary size.
[0012] As an optional implementation, the process of comparing the entropy change amount with the set entropy change threshold includes: If the entropy change of the current reasoning node is less than the entropy change threshold, the multi-type tool calling strategy of the original reasoning path is maintained; If the entropy change of the current inference node is greater than or equal to the entropy change threshold, the current inference node is determined to be a high token entropy node.
[0013] As an optional implementation, the process of training a large model based on a reinforcement learning algorithm is modeled as follows: ; in, is the mathematical expectation function, is the number of groups, Generate results for the query, To limit the range of variation, It is the clip range parameter; Relative advantage within the group; represents the KL regularization term; is the importance sampling ratio; are the learnable parameters of the model.
[0014] In a second aspect, the present invention provides a large-model multi-type tool collaborative reasoning system, comprising: The entropy calculation module is configured to calculate the original reasoning path generated by the large model for the problem, and calculate the forward and reverse of the reasoning nodes generated by the large model after each inference tool call. The entropy mean of the token fragments is used as the entropy value of the current inference node; The entropy change judgment module is configured to determine the entropy change of the current inference node based on the entropy values of the current inference node, the previous inference node, and the initial inference node, and determine the high token entropy node based on the comparison of the entropy change with the set entropy change threshold; The adaptive exploration module is configured to randomly call a different reasoning tool from the original reasoning path and continue the reasoning process when the next reasoning tool is called at a high token entropy node, generating a new reasoning path and continuing to calculate the entropy value of each reasoning node on the new reasoning path until there are no more high token entropy nodes; The training module is configured to use the original reasoning path and all the generated new reasoning paths as training sets to train the large model based on the reinforcement learning algorithm; The reasoning module is configured to use the trained large model to generate the final answer to the problem to be processed.
[0015] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0016] Compared with the prior art, the present invention has the following beneficial effects: When calling tools to perform complex reasoning tasks based on large language models, there are problems such as a single reasoning path and rigid decision-making, especially the tendency to choose suboptimal paths at decision nodes. The present invention designs an adaptive exploration and reinforcement learning strategy based on token entropy, which is applied to collaborative reasoning tasks of multiple types of tools based on large models. By monitoring the changes in token entropy during the reasoning process in real time, the uncertainty of decision nodes in the reasoning process is quantified, and an exploration mechanism is triggered at high token entropy nodes to generate diversified reasoning paths. At the same time, based on the generated diversified reasoning paths, the GRPO algorithm is used to train and optimize the collaborative reasoning capabilities of multiple types of tools of the large model, realizing the intelligent improvement of the reasoning tool selection strategy of the large model under uncertain conditions and improving the reasoning accuracy of the final answer.
[0017] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0019] Figure 1 Flowchart of the large-model multi-type tool collaborative reasoning method provided in Example 1 of the present invention; Figure 2 This is an example diagram of collaborative reasoning of multiple types of tools using a large model, as provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0021] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0022] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that the terms "comprise" and "include" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0024] Example 1 This embodiment provides a large model multi-type tool collaborative reasoning method, such as Figure 1 As shown, including: For the original reasoning path generated by the large model for the problem, calculate the preceding reasoning nodes generated by the large model after each invocation of the reasoning tool. The entropy mean of the token fragments is used as the entropy value of the current inference node; According to the entropy values of the current inference node, the previous inference node, and the initial inference node, the entropy change of the current inference node is determined. According to the comparison between the entropy change and the set entropy change threshold, the high token entropy node is determined. When the next inference tool is called at a high token entropy node, a different inference tool from the original inference path is randomly called and the inference process is continued to generate a new inference path. The entropy value of each inference node in the new inference path is continued to be calculated until there is no high token entropy node. The original reasoning path and all the generated new reasoning paths are used as training sets to train the large model based on the reinforcement learning algorithm; The trained large model is used to generate the final answer to the problem being processed.
[0025] This embodiment proposes a large-model multi-type tool collaborative reasoning method based on token entropy, which aims to improve the ability of large language models to solve complex problems. This method integrates and schedules three external reasoning tools: local corpus retrieval, Web online retrieval, and code execution. It guides the large model to perform iterative reasoning of "thinking-calling-feedback" through structured thinking chain prompts. Its core innovation lies in the introduction of an adaptive exploration and reinforcement learning strategy based on token entropy. By real-time monitoring of token entropy changes during the reasoning process, the uncertainty of decision nodes in the reasoning process is quantified. When the uncertainty exceeds the set threshold, the exploration mechanism is triggered to generate a new reasoning path, and the GRPO algorithm is used to train and optimize multiple paths, significantly enhancing the reasoning tool selection of the large model under uncertain conditions, thereby greatly improving the accuracy and robustness of the final answer.
[0026] The method of this embodiment is described in detail below.
[0027] 1. Setting the type of reasoning tool.
[0028] In this embodiment, the reasoning tools used include local corpus search tools, Web online search tools and code tools; Among them, the local corpus retrieval tool: performs retrieval-augmented generation (RAG) technology on the local corpus (including privatized knowledge and related business data) to retrieve relevant information.
[0029] Web online search tool: For open, real-time questions, it calls the network search API (Application Programming Interface) to extract relevant content and summarize key information to parse web search results in response to queries.
[0030] Code tool: For complex logical operation problems, LLM generates Python code snippets and executes the code in a sandbox environment, returning execution results or error messages based on the correctness of the code.
[0031] 2. Large model reasoning process.
[0032] The reasoning format of LLM is controlled by the prompt word design. In the reasoning process of LLM, the content of the thinking process is required to be filled in. <think> …< / think> Identify internal; operations that call multiple types of tools are required to fill in <tool> …< / tool> The tag contains the name of the reasoning tool (such as local corpus search tool, web online search tool and code tool) and the parameters related to the reasoning tool (such as local corpus search query, web online search query and code function description); the return results of calling multiple types of tools are required to be filled in<tool_result> …< / tool_result> The mark is used as context information for subsequent steps to continue the LLM reasoning process. The above process is iterated until the maximum number of times the reasoning tool is called is reached, or the LLM automatically ends the reasoning. At this time, the reasoning result is filled in <answer> …< / answer> Mark inside as the final answer.
[0033] 3. Definition of the collaborative reasoning process of multiple types of tools.
[0034] Given a natural language query question and external reasoning toolset ,In the process of LLM generating reasoning paths, the reasoning tool is called autonomously, and real-time feedback is linked to the thinking chain to promote the accuracy of the model reasoning process.
[0035] The process is modeled as: ; Among them, the factor Represents the reasoning process of multiple types of tool calls, Representing a chain of thought The number of tokens in For location token, Indicates location All previous tokens; Indicates the introduction of the reasoning toolset Model instructions; Indicates location Feedback from all previous history calls to the inference tool; factor represents the answer generation process, Indicate the answer The number of tokens, For location The model generation results are For location Previous model history generation results.
[0036] 4. Reinforcement learning method based on token entropy.
[0037] We use token entropy to study uncertainty in LLM reasoning. Low-entropy tokens reflect deterministic token positions, while high-entropy tokens reflect high uncertainty (such as ambiguity or additional exploration space in multi-step reasoning). These tokens typically correspond to decision nodes in model reasoning. Previous research has shown that maintaining high uncertainty and exploration at key node tokens in LLM reasoning improves reasoning quality.
[0038] So, first, calculate the entropy of a single token fragment : ; in, Indicates the dictionary size, Indicates location Token probability; For location The probability of the jth token in the corresponding dictionary; Indicates is an LLM model with learnable parameters.
[0039] Secondly, we continue to pay attention to the changes in token entropy of the inference nodes generated by the large model during the LLM inference process after each inference tool call, adaptively select high token entropy nodes, and infer the uncertainty of the model's prediction at the current stage.
[0040] Specifically: For each reasoning path generated by the large model for the problem, In the second inference tool call phase, the front of the inference node generated by the large model is calculated. The mean entropy of token segments , and use this as the entropy value of the current reasoning node.
[0041] Also consider the entropy value of the current inference node , relative to the entropy value of the previous inference node and the entropy value of the initial inference node The gain situation, define the entropy change : ; in, Indicates the first answer of the large model The mean entropy of token segments.
[0042] Then, in order to encourage LLM to adaptively explore the reasoning tool selection path in the high token entropy stage, this embodiment designs the entropy change threshold , pay attention to the change of token entropy after each inference tool call, and control the distribution probability of the GRPO algorithm when sampling the inference path.
[0043] Specifically: when Less than the entropy change threshold When , the multi-type tool calling strategy of the original reasoning path is maintained.
[0044] when Greater than or equal to the entropy change threshold When , the current inference node is determined to be a high token entropy node; When the next inference tool call is made at a high token entropy node, a reasoning tool different from the original reasoning path is randomly called to perform exploration. A new reasoning path is generated by inserting the inference tool call prompt into the original reasoning path and continuing the reasoning process. The entropy value of each reasoning node on the new reasoning path is continuously calculated until there is no high token entropy node.
[0045] For example, in the original reasoning path, if the next call at the current high token entropy node is a local corpus retrieval tool, then a web online retrieval tool or code tool is randomly called at this time to enhance the LLM's exploration ability under uncertainty.
[0046] Thus, LLM generates For each original reasoning path, a new reasoning path for additional exploration is generated according to the change in entropy, and finally a new reasoning path is constructed by the original reasoning path and all the generated new reasoning paths. The inference paths are used as training sets to implement GRPO reinforcement learning training.
[0047] The GRPO reinforcement learning algorithm is designed as follows: ; in, is the mathematical expectation function, Indicates the number of groups, Indicates the query generation result, Indicates the range of limited changes, It is the clip range parameter; Indicates the relative advantage within the group. is the reward function, which includes two parts: answer accuracy and format reward. is the average reward within the group, Standard deviation of rewards within the group; Represents the KL regularization term, which is used to limit the model difference from being too large to prevent training collapse; represents the importance sampling ratio, represents the reference strategy, Indicates the target strategy; are the model learnable parameters, Indicates location The query generates results; Indicates location The previous query generates results.
[0048] like Figure 2 An example of the above scheme is shown. Figure 2 Take the above solution as an example to explain in detail.
[0049] Specifically: Question q generates two original reasoning paths by LLM.
[0050] (1) For the first original reasoning path, fill in the first <think>After the identification, the first Think reasoning node is generated. Then, after the local corpus retrieval tool is called for the first time, reasoning continues to generate the second Think reasoning node. And so on. After the local corpus retrieval tool is called for the second time, the third Think reasoning node is generated. After the Web online retrieval tool is called for the third time, the fourth Think reasoning node is generated.
[0051] It should be noted here that in the process of generating the original reasoning path, the LLM autonomously calls the reasoning tool and links the real-time feedback to the thinking chain.
[0052] For each inference node, calculate the previous The entropy value of each token fragment in the token fragment , then calculate the front The entropy mean of the token fragments is used as the entropy value of the current inference node ; Among them, the entropy mean of the first Think reasoning node is .
[0053] According to the entropy value of the current inference node , the entropy value of the previous inference node and the entropy value of the initial inference node , determine the entropy change of the current inference node , according to the entropy change Comparison with the set entropy change threshold , determine the high token entropy nodes; Depend on Figure 2 It can be seen that the second Think reasoning node generated by the first call to the local corpus retrieval tool is a high token entropy node. Then, when the next reasoning tool is called here, a reasoning tool different from the original reasoning path is randomly called, that is, a random selection is made from the Web online retrieval tool and the code tool. Figure 2 The code selection tool is displayed in the dialog box, and then the reasoning process continues to generate a new Think reasoning node; After the new Think reasoning node, LLM will automatically call the reasoning tool to execute the reasoning process, and finally generate the first new reasoning path; and continue to calculate the entropy value of each reasoning node for this new reasoning path until there is no high token entropy node. Figure 2 This process is not shown.
[0054] (2) Similarly, for the second original reasoning path, fill in the first <think>After the identification, the first Think reasoning node is generated. Then, after the local corpus retrieval tool is called for the first time, reasoning continues to generate the second Think reasoning node. Similarly, after the Web online retrieval tool is called for the second time, the third Think reasoning node is generated. After the code tool is called for the third time, the fourth Think reasoning node is generated.
[0055] For each inference node, calculate the previous The entropy value of each token fragment in the token fragment , then calculate the front The entropy mean of the token fragments is used as the entropy value of the current inference node ; Among them, the entropy mean of the first Think reasoning node is .
[0056] According to the entropy value of the current inference node , the entropy value of the previous inference node and the entropy value of the initial inference node , determine the entropy change of the current inference node , according to the entropy change Comparison with the set entropy change threshold , determine the high token entropy nodes; Depend on Figure 2 It can be seen that the third Think reasoning node generated by the second call to the Web online search tool is a high token entropy node. Then, when the next reasoning tool is called here, a reasoning tool different from the original reasoning path is randomly called, that is, a random selection is made from the local corpus search tool and the code tool. Figure 2 In the figure, select the local corpus retrieval tool, and then continue the reasoning process to generate a new Think reasoning node; After the new Think reasoning node, LLM will automatically call the reasoning tool to execute the reasoning process, and finally generate a second new reasoning path; and continue to calculate the entropy value of each reasoning node for this new reasoning path until there is no high token entropy node. Figure 2 This process is not shown.
[0057] Therefore, the training set is composed of two original reasoning paths and two new reasoning paths. The large model is trained based on the reinforcement learning algorithm, and finally the trained large model is used to generate the final answer.
[0058] It is understandable that the implementation of this embodiment is not limited to Figure 2 The process shown, Figure 2 This is just an example.
[0059] In this embodiment, combined with the above implementation process, the implementation of the large-model multi-type tool collaborative reasoning method specifically includes the following steps: Step 1: Initialize three types of external reasoning tools, including local corpus retrieval, web online retrieval, and code execution, to prepare for LLM to generate highly accurate reasoning paths.
[0060] Step 2: Generate structured thinking. Follow the preset prompt word format and <think>Generate a structured thinking process within the identification. This step will clarify the query intent and plan the initial reasoning path and reasoning tool selection for solving the problem.
[0061] Step 3: Calling the reasoning tool. According to the thinking results, the big model <tool>The most appropriate reasoning tool is selected and called within the identifier. The LLM encapsulates the reasoning tool name and required reasoning tool parameters, such as the need to retrieve the problem or code function description.
[0062] Step 4: Integrate the feedback from the reasoning tool. The big model receives the results returned by the reasoning tool and fills them into<tool_result> This feedback serves as key contextual information to guide and correct subsequent reasoning steps.
[0063] Step 5: LLM iterative reasoning and answer generation. The large model loops through the iterative process of "thinking-calling-feedback" until the termination condition is met. <answer>The final answer that integrates all the information is generated and output within the logo.
[0064] Step 6: Quantify inference uncertainty. By calculating the entropy of key token segments, we can monitor the uncertainty of the model's predictions in real time. High entropy values typically correspond to key nodes that require decision-making or exploration.
[0065] Step 7: Triggering Adaptive Exploration. For the original reasoning path generated by the LLM, the entropy change after each invocation of the reasoning tool is calculated and compared with a preset threshold. This mechanism is used to determine whether exploratory reasoning should be initiated during a period of high uncertainty.
[0066] Step 8: Perform exploratory reasoning. When the entropy gain exceeds a threshold, the model is forced to randomly select a new tool for exploration. This behavior aims to generate a new reasoning path to enhance the model's ability to explore under uncertainty.
[0067] Step 9: Reinforcement Learning Strategy Optimization. The original reasoning path generated by LLM and the new reasoning path generated by exploration are used as training samples. The GRPO algorithm is applied for reinforcement learning. This aims to optimize the reasoning tool selection strategy of the large model at different decision points and improve the overall reasoning quality.
[0068] When calling tools to perform complex reasoning tasks for large language models, there are problems such as a single reasoning path and rigid decision-making, especially the tendency to choose suboptimal paths at decision nodes. This embodiment designs an adaptive exploration and reinforcement learning strategy based on token entropy, which is applied to multi-type tool collaborative reasoning tasks based on large models. By monitoring the token entropy changes in the reasoning process in real time, the uncertainty of the decision nodes in the reasoning process is quantified, and the exploration mechanism is triggered at high token entropy nodes to generate diversified reasoning paths. At the same time, based on the generated diversified reasoning paths, the GRPO algorithm is used to train and optimize the collaborative reasoning capabilities of the multi-type tools of the large model, realizing the intelligent improvement of the reasoning tool selection strategy of the large model under uncertain conditions and improving the reasoning accuracy of the final answer.
[0069] It should be noted that all data is obtained in compliance with laws and regulations and with user consent, and the data is used legally.
[0070] Example 2 This embodiment provides a large-model multi-type tool collaborative reasoning system, including: The entropy calculation module is configured to calculate the original reasoning path generated by the large model for the problem, and calculate the forward and reverse of the reasoning nodes generated by the large model after each inference tool call. The entropy mean of the token fragments is used as the entropy value of the current inference node; The entropy change judgment module is configured to determine the entropy change of the current inference node based on the entropy values of the current inference node, the previous inference node, and the initial inference node, and determine the high token entropy node based on the comparison of the entropy change with the set entropy change threshold; The adaptive exploration module is configured to randomly call a different reasoning tool from the original reasoning path and continue the reasoning process when the next reasoning tool is called at a high token entropy node, generating a new reasoning path and continuing to calculate the entropy value of each reasoning node on the new reasoning path until there are no more high token entropy nodes; The training module is configured to use the original reasoning path and all the generated new reasoning paths as training sets to train the large model based on the reinforcement learning algorithm; The reasoning module is configured to use the trained large model to generate the final answer to the problem to be processed.
[0071] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0072] In further embodiments, there is also provided: An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed by the processor, wherein when the computer instructions are executed by the processor, the method described in Example 1 is performed. For the sake of brevity, no further details are given here.
[0073] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0074] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0075] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method described in Example 1 is performed.
[0076] The method in Example 1 can be directly implemented as a hardware processor, or can be implemented using a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, it will not be described in detail here.
[0077] A computer program product includes a computer program, which implements the method described in embodiment 1 when executed by a processor.
[0078] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions contained in program modules, which are executed in a device on a real or virtual processor of a target to perform the process / method described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided between program modules as needed. The machine-executable instructions for the program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.
[0079] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0080] In the context of the present invention, computer program code or related data can be carried by any appropriate carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, and the like.
[0081] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0082] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.< / answer> < / tool> < / think> < / think> < / think> < / answer> < / tool> < / think>
Claims
1. A large-model multi-type tool collaborative reasoning method, characterized by: include: For the original reasoning path generated by the large model for the problem, calculate the preceding reasoning nodes generated by the large model after each invocation of the reasoning tool. The entropy mean of the token fragments is used as the entropy value of the current inference node; According to the entropy values of the current inference node, the previous inference node, and the initial inference node, the entropy change of the current inference node is determined. According to the comparison between the entropy change and the set entropy change threshold, the high token entropy node is determined. When the next inference tool is called at a high token entropy node, a different inference tool from the original inference path is randomly called and the inference process is continued to generate a new inference path. The entropy value of each inference node in the new inference path is continued to be calculated until there is no high token entropy node. The original reasoning path and all the generated new reasoning paths are used as training sets to train the large model based on the reinforcement learning algorithm; The trained large model is used to generate the final answer to the problem being processed.
2. A large-model multi-type tool collaborative reasoning method according to claim 1, characterized in that: On the issue , the process of collaborative reasoning based on multi-type tools using a large model is modeled as follows: ; Among them, the factor Represents the reasoning process of multiple types of tool calls, Representing a chain of thought The number of tokens in For location token, Indicates location All previous tokens; Indicates the introduction of the reasoning toolset Model instructions; Indicates location Feedback from all previous history calls to the inference tool; factor represents the answer generation process, Indicate the answer The number of tokens, For location The model generation results are For location Previous model history generation results.
3. A large-model multi-type tool collaborative reasoning method as claimed in claim 2, characterized in that: Reasoning toolset Including local corpus retrieval tools, web online retrieval tools and code tools.
4. A large-model multi-type tool collaborative reasoning method according to claim 2, characterized in that: In the process of reasoning, the content of the thinking process is filled in <think>Identify the internal; fill in the operation of calling multiple types of tools <tool>The inside of the label contains the inference tool name and inference tool parameters; the return result of calling multiple types of tools is filled in<tool_result> The mark is used as context information for subsequent reasoning steps; until the maximum number of times the reasoning tool is called is reached or the large model automatically ends reasoning, the reasoning result is filled in <answer> Identify the interior as the final answer.< / answer> < / tool> < / think> 5. The method for collaborative reasoning of large models and multiple types of tools according to claim 1, characterized in that: Entropy value of each token fragment for: ; in, is the dictionary size, For location Token probability; For location The probability of the jth token in the corresponding dictionary; Indicates A large model with learnable parameters.
6. A large-model multi-type tool collaborative reasoning method according to claim 1, characterized in that: Entropy change of the current inference node for: ; in, is the entropy value of the current inference node; is the entropy value of the previous inference node; is the entropy value of the initial inference node; is the dictionary size.
7. The method for collaborative reasoning of large models and multiple types of tools according to claim 1, characterized in that: The process of comparing the entropy change amount with the set entropy change threshold includes: If the entropy change of the current reasoning node is less than the entropy change threshold, the multi-type tool calling strategy of the original reasoning path is maintained; If the entropy change of the current inference node is greater than or equal to the entropy change threshold, the current inference node is determined to be a high token entropy node.
8. The large-model multi-type tool collaborative reasoning method according to claim 1, characterized in that: The process of training a large model based on the reinforcement learning algorithm is modeled as follows: ; in, is the mathematical expectation function, is the number of groups, Generate results for the query, To limit the range of variation, It is the clip range parameter; Relative advantage within the group; represents the KL regularization term; is the importance sampling ratio; are the learnable parameters of the model.
9. A large-model multi-type tool collaborative reasoning system, characterized by: include: The entropy calculation module is configured to calculate the original reasoning path generated by the large model for the problem, and calculate the forward and reverse of the reasoning nodes generated by the large model after each inference tool call. The entropy mean of the token fragments is used as the entropy value of the current inference node; The entropy change judgment module is configured to determine the entropy change of the current inference node based on the entropy values of the current inference node, the previous inference node, and the initial inference node, and determine the high token entropy node based on the comparison of the entropy change with the set entropy change threshold; The adaptive exploration module is configured to randomly call a different reasoning tool from the original reasoning path and continue the reasoning process when the next reasoning tool is called at a high token entropy node, generating a new reasoning path and continuing to calculate the entropy value of each reasoning node on the new reasoning path until there are no more high token entropy nodes; The training module is configured to use the original reasoning path and all the generated new reasoning paths as training sets to train the large model based on the reinforcement learning algorithm; The reasoning module is configured to use the trained large model to generate the final answer to the problem to be processed.
10. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 8 is completed.
Citation Information
Patent Citations
Interactive clinical decision support system and method based on large model knowledge enhancement
CN118335360A
Multi-agent cooperative scheduling method and system based on maximum entropy reinforcement learning
CN119623933A
Training method and device based on online tree search, equipment and medium
CN120338059A
A multi-choice question answering system that generates evidence-based answers
KR102696474B1
Knowledge graph optimized prompt for open-domain common sense reasoning decision making with artificial intelligence
US20240160955A1
Cited By
Inference model training method and device, electronic equipment, medium and product
CN121257757A
Structural semantic flow modeling-based interpretable text question and answer method and system
CN121365742A