Dynamic retrieval decision scheme determination method and system based on Monte Carlo tree search
By formalizing the inference search process into a Markov decision-making process, and using the Monte Carlo tree search and inference path preference learning mechanism, dynamically optimizing the model inference path, the problems of low retrieval accuracy and redundant retrieval in traditional RAG methods are solved, and efficient and accurate complex problem handling is achieved.
Patent Information
- Application Number
- CN202510502738.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the fields of intelligent question-and-answer and knowledge retrieval, the traditional search augmented generation (RAG) method has problems such as low accuracy in knowledge retrieval, incomplete information capture, fixed search timing and frequency, and the lack of multi-step inference dynamic programming for complex problems.
The dynamic search decision-making scheme based on Monte Carlo tree search is adopted, and the model inference path is dynamically optimized by formalizing the inference search process into the Markov decision-making process, flexible action combination is supported, and the inference path preference learning mechanism is used to achieve in-depth coordination between retrieval and reasoning, and the model inference path dynamically optimize the model inference path.
It reduces invalid search, improves the computing efficiency of the model and the accuracy of answering complex questions, realizes independent decision-making when to call external search, and reduces the computational cost.
Smart Images

Figure CN120011413A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and natural language processing, and in particular relates to a method and system for determining a dynamic retrieval decision solution based on Monte Carlo tree search. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The application of traditional retrieval-augmented generation (RAG) methods in fields such as intelligent question answering and knowledge retrieval mainly relies on the static combination of large language models and external knowledge bases. Typical retrieval schemes include: Solution 1: Serial retrieval generation solution, which forces the retrieval module to be called in each round of reasoning of the large language model. This results in a strong correlation between the number of retrievals and the complexity of the problem, and a high computational cost. Solution 2: Rule-driven retrieval generation solution, which calls the retrieval module based on keyword matching or fixed trigger conditions (such as confidence thresholds), but it is still difficult to deal with fuzzy queries or requires multi-step reasoning; Solution 3: End-to-end joint training solution, which jointly optimizes the large language model and the retrieval module, but lacks control over the intermediate reasoning path, and is prone to answer failure due to the accumulation of incorrect intermediate answers in multi-hop reasoning.
[0004] In summary, the traditional Retrieval-augmented Generation (RAG) method has the disadvantages of low knowledge retrieval accuracy and incomplete information capture, fixed retrieval timing and frequency, which easily leads to redundant retrieval, and lacks multi-step reasoning dynamic programming for complex problems. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a method and system for determining a dynamic retrieval decision solution based on Monte Carlo tree search, which can formalize the reasoning retrieval process into a Markov decision process and support flexible action combinations, and realize deep coordination of retrieval and reasoning through Monte Carlo tree search and reasoning path preference learning mechanism, reduce invalid retrieval, and improve the computational efficiency of the model and the accuracy of answering complex questions.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: A first aspect of the present invention provides a method for determining a dynamic retrieval decision solution based on Monte Carlo tree search.
[0007] In one or more embodiments, a method for determining a dynamic retrieval decision solution based on Monte Carlo tree search is provided, comprising: An initial query question is received and used as a starting point of the reasoning process to trigger the invocation of a large language model operation; wherein the reasoning process is constructed in advance as a Markov decision process; Call the large language model to perform subquery generation and action selection operations, iterate and explore the reasoning actions of the subquery problem based on Monte Carlo tree search, and dynamically filter the reasoning path by considering the confidence score and honesty score of the large language model to obtain a complete reasoning trajectory set; Based on the complete set of reasoning trajectories, preference pairs are constructed for each sub-query question, and then the reasoning path is optimized to enable the large language model to autonomously judge the best retrieval time.
[0008] As an implementation method, the process of dynamically screening the reasoning path includes: For the subquery problems iteratively generated by the large language model, the confidence score of the large language model is generated accordingly, and then the corresponding reasoning action is matched and selected to generate the honesty score of the large language model accordingly, and the corresponding reward function is obtained; In the process of iteratively exploring the reasoning actions of the subquery problem based on the Monte Carlo tree search, the score of each reasoning path is updated through the reward function. After the iteration, the reasoning path with a score higher than the preset threshold is screened.
[0009] As an implementation manner, the reasoning action is any one of a direct answer, a retrieved answer, a query conversion, and a final answer.
[0010] As an implementation method, if the confidence score of the large language model exceeds a preset confidence score threshold, a direct answer or final answer action is allowed to be selected; otherwise, a retrieval answer or query conversion action is triggered.
[0011] As an implementation method, when the inference action is a direct answer or a final answer, the honesty score of the large language model ranges from ; When the inference action is to retrieve answers or query conversion, the honesty score of the language model is 1.
[0012] As an implementation mode, the expression of the reward function is obtained by summing the upper confidence interval and the question-answering utility factor of the large language model; the question-answering utility factor of the large language model is obtained by multiplying the confidence score and the honesty score of the large language model by a constant factor.
[0013] A second aspect of the present invention provides a dynamic retrieval decision solution determination system based on Monte Carlo tree search.
[0014] In one or more embodiments, a dynamic retrieval decision solution determination system based on Monte Carlo tree search includes: A reasoning process starting point acquisition module, which is used to receive an initial query question and use it as the starting point of the reasoning process to trigger the call of the large language model operation; wherein the reasoning process is pre-constructed as a Markov decision process; The inference path dynamic optimization module is used to call the large language model to perform subquery generation and action selection operations, iterates the inference actions of the subquery problem based on the Monte Carlo tree search, and dynamically screens the inference path by considering the confidence score and honesty score of the large language model to obtain a complete set of inference trajectories; The module for autonomously judging the best retrieval timing is used to construct preference pairs for each sub-query question based on the complete set of reasoning trajectories, and then optimize the reasoning path to enable the large language model to autonomously judge the best retrieval timing.
[0015] A third aspect of the present invention provides a computer-readable storage medium.
[0016] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search as described above.
[0017] A fourth aspect of the present invention provides a computer program product.
[0018] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search as described above.
[0019] A fifth aspect of the present invention provides an electronic device.
[0020] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search as described above are implemented.
[0021] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention constructs the reasoning process as a Markov decision process, quantifies the retrieval decision path, and then combines it with an evaluation strategy based on Monte Carlo tree search and introduces a combination of confidence and honesty indicators to dynamically optimize the model reasoning path, breaking through the traditional fixed retrieval mode and realizing autonomous decision-making when to call external retrieval, reducing redundant retrieval and lowering computing costs.
[0022] (2) The present invention includes four types of reasoning actions: direct answer, retrieval answer, query conversion and final answer, covering the scenario of collaboration between large language model parameterized knowledge and external knowledge, and supports question decomposition and query rewriting, which helps to improve the accuracy of retrieval of complex questions.
[0023] (3) The present invention uses Monte Carlo tree search to effectively explore potential reasoning trajectories. By iteratively constructing and evaluating reasoning paths, it can prioritize paths at a lower retrieval cost, balancing accuracy and computational efficiency. In addition, it constructs preference pairs based on each sub-query question to optimize the reasoning path, significantly improving the dynamic decision-making and reasoning capabilities of enhanced generation based on large language model retrieval, giving the large language model human-like reasoning flexibility and resource utilization efficiency, and is suitable for knowledge-intensive complex question-answering scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0025] Figure 1 It is a flow chart of a method for determining a dynamic retrieval decision solution based on Monte Carlo tree search according to an embodiment of the present invention; Figure 2 It is a process diagram of a dynamic retrieval decision solution based on Monte Carlo tree search in an embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of a dynamic retrieval decision solution determination system based on Monte Carlo tree search according to an embodiment of the present invention; Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0027] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0029] Figure 1 FIG. 1 is a flow chart of a method for determining a dynamic retrieval decision solution based on Monte Carlo tree search in an embodiment of the present invention. Figure 1 The method for determining a dynamic retrieval decision solution based on Monte Carlo tree search in the embodiment shown may include the following steps S101 to S103.
[0030] S101, receiving the initial query question and using it as the starting point of the reasoning process to trigger the call of the large language model operation; wherein the reasoning process is pre-constructed as a Markov decision process (MDP), thus achieving the quantification of the retrieval decision path.
[0031] The large language model of the embodiment of the present invention refers to a deep learning model trained with a large amount of text data, so that the model can generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics by training on huge data sets. The core idea is to learn the patterns and structures of natural language through large-scale unsupervised training, and to simulate the human language cognition and generation process to a certain extent. For example, large language models such as GPT model, LLaMA model or DeepSeek model.
[0032] Specifically, the reasoning retrieval process of the large language model is defined as a four-tuple MDP ( ): Represents the state space, each state Describes the reasoning process of a large language model, where For the initial query, Respectively represent Subquery questions and intermediate answers generated by step reasoning; .
[0033] Representing the action space, the embodiment of the present invention defines four types of reasoning actions, including: direct answer, retrieval answer, query conversion and final answer, which bridges the gap between large language model reasoning and human cognition in the RAG (Retrieval-augmented Generation) scenario.
[0034] Specifically, direct answering: using the parameter knowledge of a large language model to directly answer questions without relying on any external knowledge; Retrieve answers: Retrieve relevant knowledge from external knowledge bases to support subsequent reasoning; Query transformation: transforming questions, such as rephrasing, back-off prompts, and question decomposition; Final answer: Combine the historical reasoning process, intermediate answers, and initial questions to generate the final answer and stop model reasoning.
[0035] Indicates state transition, based on state Performing inference actions After the operation, the environment will update the status When the reasoning action For direct answers, the intermediate answer The model itself generates an intermediate answer through analysis; when the inference action When searching for answers, the intermediate answers The model retrieves documents and generates intermediate answers; when the inference action Intermediate answer for query conversion The model divides the subproblems into Rewrite and use as an intermediate answer; when reasoning For the final answer, the intermediate answer The final answer is generated by the model combined with information from the historical reasoning process.
[0036] Represents the reward function. In this embodiment, the expression of the reward function is obtained by summing the upper confidence interval and the question-answering utility factor of the large language model; the question-answering utility factor of the large language model is the product of the confidence score and the honesty score of the large language model and a constant factor. The specific expression is as follows: ; The above reward function consists of two parts; represents the Upper Confidence Bound Apply to Tree (UCT), which is used to balance the exploration and exploitation of Monte Carlo tree search; Represents the accumulated rewards of the child node, Indicates the number of visits to a child node. Indicates the number of visits to the parent node; is a constant that controls the trade-off between exploration and exploitation; Represents the logarithmic function with base e.
[0037] represents the question-answering utility factor of the large language model, Indicates the use of historical status Assess the current problem The model confidence score ranges from ; Indicates the evaluation of the intermediate answer The model honesty score of .
[0038] When the reasoning action is a direct answer or a final answer, The value range is ; When the inference action is to retrieve answers or query transformations, The value of is 1; Represents a constant factor used to balance the confidence and honesty of the large language model.
[0039] Compared with the existing upper confidence interval algorithm, the reward function proposed in the embodiment of the present invention is not only evaluated based on the correctness of the final answer and the cost of the number of retrievals, but also focuses on the confidence and honesty of the large language model in answering the question, which helps to reduce the hallucination of the large language model while balancing model reasoning and retrieval.
[0040] The embodiment of the present invention transforms the dynamic retrieval decision process into a sequence decision problem that can be quantified and optimized. The method breaks through the limitations of fixed retrieval modes in traditional retrieval enhancement generation technology and realizes the coordinated optimization of retrieval timing, frequency and reasoning path.
[0041] S102, calling the large language model to perform subquery generation and action selection operations, iteratively exploring the reasoning actions of the subquery problem based on the Monte Carlo tree search, and dynamically screening the reasoning path by considering the confidence score and honesty score of the large language model to obtain a complete reasoning trajectory set.
[0042] Specifically, the confidence score and honesty score of the large language model are both in the range of 0 to 1.
[0043] Each complete reasoning trace in the complete reasoning trace set consists of a subquery question and its corresponding intermediate answer and reasoning action.
[0044] Monte Carlo tree search is a heuristic search algorithm based on random simulation. It is mainly used for optimal strategy selection in complex decision-making scenarios. Its core idea is to gradually build an asymmetric search tree by repeatedly simulating decision paths, reduce invalid searches by dynamically focusing on high-potential branches, and finally find the solution with the greatest benefit. The core of Monte Carlo tree search generation lies in the action space, which defines the scope of tree exploration.
[0045] In step S102, the process of dynamically screening the reasoning path includes: S1021, for the subquery problem iteratively generated by the large language model, a confidence score of the large language model is generated accordingly, and then a corresponding reasoning action is selected and matched and an honesty score of the large language model is generated accordingly to obtain a corresponding reward function; S1022, in the process of iteratively exploring the reasoning action of the subquery problem based on the Monte Carlo tree search, the score of each reasoning path is updated through the reward function, and after the iteration is completed, the reasoning path with a score higher than the preset threshold is screened.
[0046] The four types of reasoning actions in the embodiment of the present invention define a highly diverse action space . The model reasoning retrieval path is constructed through Monte Carlo tree search, and different solution strategies are integrated for each sub-query in the reasoning process. Different from the existing methods, the embodiment of the present invention introduces the concepts of model confidence and honesty in the process of Monte Carlo tree search to reduce the illusion of large language models and improve the efficiency of reasoning path search. Specifically, for a certain input question, the large language model generates a series of sub-questions and explores four reasoning actions; in each round of reasoning action exploration, it consists of the following three parts: First, the large language model uses the historical state Assess the current problem Confidence.
[0047] Secondly, when the confidence score is greater than the threshold (its value can be set according to actual conditions), it means that the large language model is confident in reasoning about the current question, and the reasoning actions can be direct answer and final answer. Otherwise, it means that the large language model needs additional information or in-depth analysis for reasoning about the current question, and the reasoning actions can be retrieval answer and query conversion.
[0048] Finally, for the direct answer and the final answer, the large language model evaluates the honesty of the answer to the current question, with a value range of ; For retrieval answering and query conversion, the current question answer honesty is set to 1, assuming that the large language model fully trusts the information obtained through retrieval and query conversion.
[0049] Through each round of iteration, the size of the Monte Carlo tree continues to increase, all nodes continue to update their scores according to the reward function, and different reasoning paths are explored.
[0050] S103, constructing a preference pair for each sub-query question based on the complete reasoning trajectory set, and then optimizing the reasoning path to enable the large language model to autonomously judge the best retrieval time.
[0051] The embodiment of the present invention utilizes the Monte Carlo tree search method to effectively explore potential reasoning trajectories. By iteratively constructing and evaluating reasoning paths, the paths can be prioritized at a lower retrieval cost, balancing accuracy and computational efficiency. The above process search is performed on a training data set to obtain an adaptive reasoning retrieval process for a specific problem, which includes the optimal answer strategy for each sub-query question to determine whether retrieval is required. Based on the reasoning retrieval process data, the present invention constructs a preference pair for each sub-query to indicate the preferred action selection.
[0052] In some optional embodiments, the existing Group Relative Policy Optimization (GRPO) algorithm can be used to optimize the reasoning path, that is, to obtain the reasoning retrieval strategy, so as to enable the large language model to autonomously judge the best retrieval time. Among them, the GRPO algorithm is a reinforcement learning (RL) algorithm specifically used to enhance the reasoning ability in large language models (LLMs). The GRPO algorithm optimizes the model by evaluating groups of responses that are related to each other. This method can improve training efficiency and make the GRPO algorithm an ideal choice for reasoning tasks that require complex problem solving and long-chain thinking.
[0053] It is understandable here that, in other embodiments, preference pairs may be constructed based on each sub-query question, and other existing optimization algorithms may be used to optimize the reasoning path to enable the large language model to autonomously determine the best retrieval time.
[0054] This embodiment enables the large language model to learn when to retrieve external information through a preference learning process based on reinforcement learning, thereby maximizing the model's reasoning retrieval capability while reducing unnecessary retrieval.
[0055] like Figure 3 As shown, the dynamic retrieval decision solution determination system based on Monte Carlo tree search provided by the embodiment of the present invention can be implemented in a software manner. The dynamic retrieval decision solution determination system based on Monte Carlo tree search includes the following software modules: The inference process starting point acquisition module 301 is used to receive the initial query question and use it as the starting point of the inference process to trigger the invocation of the large language model operation; wherein the inference process is constructed as a Markov decision process in advance; The inference path dynamic optimization module 302 is used to call the large language model to perform subquery generation and action selection operations, iteratively explore the inference actions of the subquery problem based on the Monte Carlo tree search, and dynamically screen the inference path by considering the confidence score and honesty score of the large language model to obtain a complete inference trajectory set; The best search timing autonomous judgment module 303 is used to construct a preference pair for each sub-query question based on the complete reasoning trajectory set, and then optimize the reasoning path to realize the autonomous judgment of the best search timing of the large language model.
[0056] It should be noted here that Figure 3 The dynamic retrieval decision scheme based on Monte Carlo tree search determines the various modules in the system and Figure 1 The various steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search correspond one to one, and the specific implementation process is the same, which will not be repeated here.
[0057] Figure 4 The schematic diagram of the structure of an electronic device provided by an embodiment of the present invention can be understood as follows: Figure 4 Only an exemplary structure of an electronic device is shown, not all structures, and can be implemented as needed. Figure 4 Partial or complete structure shown.
[0058] An electronic device provided by an embodiment of the present invention includes: at least one processor 401, a memory 402, a user interface 403 and at least one network interface 404. The components in the dynamic retrieval decision scheme based on Monte Carlo tree search are coupled together through a bus system 405. It can be understood that the bus system 405 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 405 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, in Figure 4 In the figure, various buses are labeled as bus system 405. The user interface 403 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a touch pad or a touch screen.
[0059] It is understood that the memory 402 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The memory 402 in the embodiment of the present invention can store the following: Figure 1 The computer program corresponding to each step in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search shown in . Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks. The application program can include various application programs.
[0060] In other embodiments, the dynamic retrieval decision scheme determination system based on Monte Carlo tree search is an example of an implementation using a combination of software and hardware. The dynamic retrieval decision scheme determination system based on Monte Carlo tree search can also be directly embodied as a combination of software modules executed by the processor 401. The software module can be located in a storage medium, and the storage medium is located in the memory 402. The processor 401 reads the executable instructions included in the software module in the memory 402, and combines with the necessary hardware (for example, including the processor 401 and other components connected to the bus 405) to complete the dynamic retrieval decision scheme determination method based on Monte Carlo tree search provided in an embodiment of the present invention.
[0061] As an example, the processor 401 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0062] As an example of hardware implementation of the dynamic retrieval decision scheme determination system based on Monte Carlo tree search provided in an embodiment of the present invention, the device provided in an embodiment of the present invention can be directly executed by a processor 401 in the form of a hardware decoding processor, for example, one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components to implement the dynamic retrieval decision scheme determination method based on Monte Carlo tree search provided in an embodiment of the present invention.
[0063] The memory 402 in the embodiment of the present invention is used to store various types of data to support the operation of the dynamic retrieval decision solution determination system based on Monte Carlo tree search. Examples of these data include: any executable instructions for operating on the dynamic retrieval decision solution determination system based on Monte Carlo tree search, such as executable instructions, and the program for implementing the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search in the embodiment of the present invention can be included in the executable instructions.
[0064] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, the computer program including a computer program for executing Figure 1 In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions defined in the device of the present application are executed.
[0065] in, Figure 1 The computer program instructions corresponding to the method shown may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0066] The above description is only the preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for determining a dynamic retrieval decision solution based on Monte Carlo tree search, characterized in that: include: An initial query question is received and used as a starting point of the reasoning process to trigger the invocation of a large language model operation; wherein the reasoning process is constructed in advance as a Markov decision process; Call the large language model to perform subquery generation and action selection operations, iterate and explore the reasoning actions of the subquery problem based on Monte Carlo tree search, and dynamically filter the reasoning path by considering the confidence score and honesty score of the large language model to obtain a complete reasoning trajectory set; Based on the complete set of reasoning trajectories, preference pairs are constructed for each sub-query question, and then the reasoning path is optimized to enable the large language model to autonomously judge the best retrieval time.
2. The method for determining a dynamic retrieval decision solution based on Monte Carlo tree search according to claim 1, characterized in that: The process of dynamically screening the reasoning path includes: For the subquery problems iteratively generated by the large language model, the confidence score of the large language model is generated accordingly, and then the corresponding reasoning action is matched and selected to generate the honesty score of the large language model accordingly, and the corresponding reward function is obtained; In the process of iteratively exploring the reasoning actions of the subquery problem based on the Monte Carlo tree search, the score of each reasoning path is updated through the reward function. After the iteration, the reasoning path with a score higher than the preset threshold is screened.
3. The method for determining a dynamic retrieval decision solution based on Monte Carlo tree search according to claim 1, characterized in that: The reasoning action is any one of a direct answer, a retrieved answer, a query conversion, and a final answer.
4. The method for determining a dynamic retrieval decision solution based on Monte Carlo tree search according to claim 3, characterized in that: If the confidence score of the large language model exceeds the preset confidence score threshold, a direct answer or final answer action is allowed; otherwise, a retrieval answer or query conversion action is triggered.
5. The method for determining a dynamic retrieval decision solution based on Monte Carlo tree search according to claim 3, characterized in that: When the inference action is a direct answer or a final answer, the honesty score of the large language model ranges from ; When the inference action is to retrieve answers or query conversion, the honesty score of the language model is 1.
6. The method for determining a dynamic retrieval decision solution based on Monte Carlo tree search according to claim 2, characterized in that: The expression of the reward function is obtained by summing the upper confidence interval and the question-answering utility factor of the large language model; the question-answering utility factor of the large language model is the product of the confidence score and the honesty score of the large language model and a constant factor.
7. A dynamic retrieval decision solution determination system based on Monte Carlo tree search, characterized in that: include: A reasoning process starting point acquisition module, which is used to receive an initial query question and use it as the starting point of the reasoning process to trigger the call of the large language model operation; wherein the reasoning process is pre-constructed as a Markov decision process; The inference path dynamic optimization module is used to call the large language model to perform subquery generation and action selection operations, iterates the inference actions of the subquery problem based on the Monte Carlo tree search, and dynamically screens the inference path by considering the confidence score and honesty score of the large language model to obtain a complete set of inference trajectories; The module for autonomously judging the best retrieval timing is used to construct preference pairs for each sub-query question based on the complete set of reasoning trajectories, and then optimize the reasoning path to enable the large language model to autonomously judge the best retrieval timing.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search as described in any one of claims 1 to 6 are implemented.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search as described in any one of claims 1 to 6 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for determining a dynamic retrieval decision solution based on Monte Carlo tree search as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method, device and equipment for training large language model
CN118153624A
Customer service question and answer method and device
CN118673122A
Large model illusion relieving method and device, equipment and storage medium
CN118964583A
Large model agent interactive question and answer task decision-making method, device and equipment and medium
CN119166778A
Intelligent agent language ability automatic testing method based on multi-language model self-evolution
CN119357069A
Cited By
Search task processing method and target scoring model training method
CN120218173A
Multi-source heterogeneous data multi-mode mixed retrieval method and system based on large model reasoning
CN120509496A
Multi-source heterogeneous data multi-mode hybrid retrieval method and system based on large model reasoning
CN120509496B
Multi-path legal inference engine based on MCP protocol
CN120764667A
Industrial production safety monitoring method, equipment and medium
CN121071756A