Enhanced large language model processing method, system and platform based on Monte Carlo tree search and storage medium

By introducing the Monte Carlo tree search algorithm into the large language model, dynamically adjusting the cluster width and backtracking update confidence data, the problem of low performance of traditional large language models in complex inference tasks is solved, and more efficient and accurate inference processing is achieved.

CN120011510APending Publication Date: 2025-05-16GUANGZHOU ZHONGKE YIDE TECH CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510091179.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional large language models have low performance when dealing with complex inference tasks, especially in applications such as mathematical problem-solving, chain thinking reasoning, context-related question-and-answer reasoning, and the reasoning process is inefficient and poor accuracy.

Method used

Using an enhanced large language model processing method based on Monte Carlo tree search, a first model corresponding to the large language model is created, a data set corresponding to the logical reasoning process is generated and obtained, and the first model is fine-tuned according to the data set. Then, based on task requirements, the search tree is divided into multiple initial paths, each path allocates multiple nodes for independent subtree expansion, and candidate output is generated in each subtree, and the cluster width is dynamically adjusted in combination with confidence data. The expansion path of each subtree is independently simulated, backtracking and updating the confidence data of the node, and generating optimized inference response data.

Benefits of technology

It significantly improves the inference processing efficiency and accuracy of large language models in complex inference tasks, and can more effectively deal with complex tasks such as mathematical problem-solving, chain thinking reasoning, and context-related question-and-answer reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011510A_ABST
    Figure CN120011510A_ABST
Patent Text Reader

Abstract

The invention discloses an enhanced large language model processing method, system and platform based on Monte Carlo tree search and a storage medium. The method comprises the following steps: creating a first model corresponding to a big language, generating and obtaining a data set corresponding to a logical reasoning process, and performing fine tuning processing on the first model according to the data set; dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent sub-tree expansion, generating at least two candidate outputs based on a first model in each sub-tree, and dynamically adjusting and processing a corresponding cluster width in combination with confidence data; independently simulating and processing the expansion path of each sub-tree, generating corresponding first data, backtracking, updating and processing confidence data of each node according to the first data, and generating corresponding second data at the same time; wherein the second data is the reasoning response data after enhancement optimization processing, and the corresponding system, platform and storage medium can improve the reasoning processing efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence processing technology, and specifically relates to an enhanced large language model processing method, system, platform and storage medium based on Monte Carlo tree search. Background Art

[0002] At present, with the development of artificial intelligence technology, large language models (LLMs) have made significant progress in the field of natural language processing. However, traditional LLMs have limitations in dealing with complex reasoning tasks, especially in logical reasoning and decision making, that is, LLMs have poor performance in complex reasoning tasks, especially in applications such as mathematical problem solving, chain thinking reasoning, and context-related question-answering reasoning, and there are low efficiency and poor accuracy in the reasoning process.

[0003] Therefore, in view of the technical problems and defects of the above LLM in processing complex reasoning tasks, it is urgent to design and develop an enhanced large language model processing method, system, platform and storage medium based on Monte Carlo tree search. Summary of the invention

[0004] In order to overcome the shortcomings and difficulties of the above-mentioned prior art, the purpose of the present invention is to provide an enhanced large language model processing method, system, platform and storage medium based on Monte Carlo tree search to improve the efficiency and accuracy of reasoning processing.

[0005] The first object of the present invention is to provide an enhanced large language model processing method based on Monte Carlo tree search; the second object of the present invention is to provide an enhanced large language model processing system based on Monte Carlo tree search; the third object of the present invention is to provide an enhanced large language model processing platform based on Monte Carlo tree search; the fourth object of the present invention is to provide a computer-readable storage medium.

[0006] The first object of the present invention is achieved in that the method comprises the following steps:

[0007] Creating a first model corresponding to the large language, generating and acquiring a data set corresponding to the logical reasoning process, and fine-tuning the first model according to the data set; wherein the first model is a large language model;

[0008] Based on the task requirements, the search tree is divided into at least two initial paths, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding beam width is dynamically adjusted in combination with the confidence data;

[0009] Independently simulate and process the expansion path of each subtree and generate corresponding first data. Backtrack and update the confidence data of each node based on the first data and generate corresponding second data at the same time; wherein the first data is the simulation result data; the second data is the inference response data after enhanced optimization processing.

[0010] Furthermore, the data sets include math problem solving, chain thinking reasoning data sets, context-related question-answering reasoning data sets, long-span logical reasoning data sets, and domain-specific knowledge question-answering reasoning data sets.

[0011] Further, the search tree is divided into at least two initial paths based on task requirements, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding beam width is dynamically adjusted in combination with confidence data, and also includes:

[0012] Calculate and generate confidence data corresponding to the path nodes; the calculation formula is as follows:

[0013]

[0014] In the formula, R n is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients;

[0015] Based on the search tree, the paths with high confidence are preferentially expanded, and the beam width is dynamically adjusted in combination with the score difference of the process reward model; the adjustment calculation formula is as follows:

[0016]

[0017] Where M min is the minimum beam width, M max is the maximum beam width, and variance(PRM_scores) is the variance of the current process reward model score.

[0018] Further, the independent simulation processes the expansion path of each subtree and generates corresponding first data, and backtracks and updates the confidence data of each node according to the first data to generate corresponding second data, and also includes:

[0019] Creating a second model corresponding to the subtree expansion path; wherein the second model is a process reward model;

[0020] Based on the second model and in combination with the ε-greedy strategy, corresponding third data are generated; wherein the third data are path quality evaluation data; wherein the calculation formula of the exploration ratio ε is:

[0021]

[0022] In the formula, ∈ min is the minimum value of the exploration probability, ∈0 is the initial exploration probability, and t is the current number of iterations;

[0023] Based on the path cumulative reward, combined with the process reward and path length penalty, the corresponding path selection is automatically adjusted during the reasoning process to balance the breadth and depth of the search; the cumulative reward value calculation formula is:

[0024]

[0025] Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty.

[0026] Further, the independent simulation processes the expansion path of each subtree and generates corresponding first data, and backtracks and updates the confidence data of each node according to the first data to generate corresponding second data, and also includes:

[0027] After the path simulation is completed, the node information is updated and processed based on the simulation results, and subsequent decisions are optimized; the node's cumulative reward is updated in real time based on the current simulation reward; the update formula is:

[0028]

[0029] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It is a simulated reward;

[0030] After the backtracking path update simulation is completed, backtrack from the leaf node to the root node to update the visit count and historical rewards of all visited nodes.

[0031] The second object of the present invention is achieved in this way: the system is used to implement the enhanced large language model processing method based on Monte Carlo tree search; the system comprises:

[0032] A first data processing unit, configured to create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model;

[0033] A second data processing unit is used to divide the search tree into at least two initial paths based on task requirements, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding beam width is dynamically adjusted in combination with confidence data;

[0034] The third data processing unit is used to independently simulate and process the expansion path of each subtree, and generate corresponding first data, backtrack and update the confidence data of each node according to the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; and the second data is the inference response data after enhanced optimization processing.

[0035] Furthermore, the data set includes math problem solving, chain thinking reasoning data set, context-related question-answering reasoning data set, long-span logical reasoning data set, and domain-specific knowledge question-answering reasoning data set;

[0036] The second data processing unit further includes:

[0037] The first calculation module is used to calculate and generate confidence data corresponding to the path node; wherein the calculation formula is as follows:

[0038]

[0039] In the formula, R n is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients;

[0040] The first processing module is used to preferentially expand and process high-confidence paths based on the search tree, and dynamically adjust the bundle width in combination with the score difference of the process reward model; wherein the adjustment calculation formula is as follows:

[0041]

[0042] Where M min is the minimum beam width, M max is the maximum value of the beam width, variance(PRM_scores) is the variance of the current process reward model score;

[0043] And / or, the third data processing unit further includes:

[0044] A first creation module is used to create a second model corresponding to the subtree expansion path; wherein the second model is a process reward model;

[0045] The first generating module is used to generate corresponding third data based on the second model and in combination with the ε-greedy strategy; wherein the third data are path quality evaluation data; wherein the calculation formula of the exploration ratio ε is:

[0046]

[0047] In the formula, ∈ min is the minimum value of the exploration probability, ∈0 is the initial exploration probability, and t is the current number of iterations;

[0048] The second processing module is used to automatically adjust the corresponding path selection during the reasoning process based on the path cumulative reward, combined with the process reward and the path length penalty item, to balance the breadth and depth of the search; wherein the cumulative reward value calculation formula is:

[0049]

[0050] Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty.

[0051] Furthermore, the third data generating unit further includes:

[0052] The third processing module is used to update the node information according to the simulation results after the path simulation is completed, and optimize the subsequent decision-making; the accumulated rewards of the node are updated in real time according to the current simulation rewards; the update formula is:

[0053]

[0054] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It is a simulated reward;

[0055] The fourth processing module is used to trace back from the leaf node to the root node after the backtracking path update simulation is completed, and update the number of visits and historical rewards of all visited nodes.

[0056] The third object of the present invention is achieved as follows: it includes a processor, a memory and an enhanced large language model processing platform control program based on Monte Carlo tree search; wherein the enhanced large language model processing platform control program based on Monte Carlo tree search is executed by the processor, the enhanced large language model processing platform control program based on Monte Carlo tree search is stored in the memory, and the enhanced large language model processing platform control program based on Monte Carlo tree search implements the enhanced large language model processing method based on Monte Carlo tree search.

[0057] The fourth object of the present invention is achieved in this way: the computer-readable storage medium stores an enhanced large language model processing platform control program based on Monte Carlo tree search, and the enhanced large language model processing platform control program based on Monte Carlo tree search implements the enhanced large language model processing method based on Monte Carlo tree search.

[0058] The present invention creates a first model corresponding to a large language through a method, generates and obtains a data set corresponding to a logical reasoning process, and fine-tunes the first model according to the data set; wherein the first model is a large language model; based on task requirements, the search tree is divided into at least two initial paths, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding cluster width is dynamically adjusted in combination with confidence data; the expansion path of each subtree is independently simulated and processed, and the corresponding first data is generated, and the confidence data of each node is backtracked and updated according to the first data, and the corresponding second data is generated at the same time; wherein the first data is the simulation result data; the second data is the reasoning response data after enhanced optimization processing, as well as the system, platform and storage medium corresponding to the method, so as to improve the efficiency and accuracy of reasoning processing.

[0059] That is to say, the present invention proposes an enhanced large language model push processing method based on Monte Carlo tree search, which optimizes the reasoning efficiency and accuracy through reasonable path expansion, dynamic adjustment strategy and simulation feedback mechanism. The scheme of the present invention combines the fine-tuning technology of the large language model and the path selection algorithm of Monte Carlo tree search, and can dynamically select and optimize the reasoning path according to the requirements of different reasoning tasks, thereby improving the performance of the large language model in complex reasoning tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0061] Figure 1 A schematic flow chart of an enhanced large language model processing method based on Monte Carlo tree search according to the present invention;

[0062] Figure 2 It is a flow chart of an embodiment of an enhanced large language model processing method based on Monte Carlo tree search of the present invention;

[0063] Figure 3This is a schematic diagram of an architecture of an enhanced large language model processing system based on Monte Carlo tree search according to the present invention;

[0064] Figure 4 It is a schematic diagram of the architecture of an embodiment of an enhanced large language model processing system based on Monte Carlo tree search of the present invention;

[0065] Figure 5 This is a schematic diagram of an architecture of an enhanced large language model processing platform based on Monte Carlo tree search according to the present invention;

[0066] Figure 6 A schematic diagram of a computer-readable storage medium architecture in an embodiment of the present invention. DETAILED DESCRIPTION

[0067] In order to better understand the purpose, technical solutions and advantages of the present invention, the present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification.

[0068] The present invention may also be implemented or applied through other different specific examples, and the details in this specification may also be modified and changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention.

[0069] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0070] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. Secondly, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0071] Preferably, the enhanced large language model processing method based on Monte Carlo tree search of the present invention is applied in one or more terminals or servers. The terminal is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (APPlication Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0072] The terminal can be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal can interact with the client through a keyboard, a mouse, a remote control, a touch pad, or a voice control device.

[0073] The present invention is to realize an enhanced large language model processing method, system, platform and storage medium based on Monte Carlo tree search.

[0074] like Figure 1 , which is a flow chart of an enhanced large language model processing method based on Monte Carlo tree search provided by an embodiment of the present invention.

[0075] In this embodiment, the enhanced large language model processing method based on Monte Carlo tree search can be applied to a terminal with a display function or a fixed terminal. The terminal is not limited to a personal computer, a smart phone, a tablet computer, a desktop computer or an all-in-one computer equipped with a camera, etc.

[0076] The enhanced large language model processing method based on Monte Carlo tree search can also be applied to a hardware environment consisting of a terminal and a server connected to the terminal through a network. The network includes but is not limited to: a wide area network, a metropolitan area network or a local area network. The enhanced large language model processing method based on Monte Carlo tree search in the embodiment of the present invention can be executed by a server, can be executed by a terminal, or can be executed by both a server and a terminal.

[0077] For example, for a terminal that needs to perform enhanced large language model processing based on Monte Carlo tree search, the enhanced large language model processing function based on Monte Carlo tree search provided by the method of the present invention can be directly integrated on the terminal, or a client for implementing the method of the present invention can be installed. For another example, the method provided by the present invention can also be run on a server or other device in the form of a software development kit (SDK), and an interface for the enhanced large language model processing function based on Monte Carlo tree search is provided in the form of SDK. The terminal or other device can implement the enhanced large language model processing function based on Monte Carlo tree search through the provided interface. The present invention is further described below in conjunction with the accompanying drawings.

[0078] like Figure 1-Figure 2 As shown, the present invention provides an enhanced large language model processing method based on Monte Carlo tree search, wherein the method comprises the following steps:

[0079] S1. Create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model;

[0080] S2, dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent subtree expansion, and generating at least two candidate outputs in each subtree based on the first model, and dynamically adjusting the corresponding beam width in combination with confidence data;

[0081] S3. Independently simulate and process the expansion path of each subtree, and generate corresponding first data. Backtrack and update the confidence data of each node based on the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; the second data is the inference response data after enhanced optimization processing.

[0082] The datasets include math problem solving, chain thinking reasoning, context-related question-answering reasoning, long-span logical reasoning, and domain-specific knowledge question-answering reasoning.

[0083] The method further comprises: dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent subtree expansion, and generating at least two candidate outputs in each subtree based on the first model, and dynamically adjusting the corresponding beam width in combination with confidence data.

[0084] S21. Calculate and generate confidence data corresponding to the path nodes; wherein the calculation formula is as follows:

[0085]

[0086] In the formula, R n is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients;

[0087] S22. Based on the search tree, the paths with high confidence are preferentially expanded, and the beam width is dynamically adjusted in combination with the score difference of the process reward model; wherein, the adjustment calculation formula is as follows:

[0088]

[0089] Where M min is the minimum beam width, M max is the maximum beam width, and variance(PRM_scores) is the variance of the current process reward model score.

[0090] The independent simulation processes the expansion path of each subtree and generates corresponding first data, and the confidence data of each node is backtracked and updated according to the first data to generate corresponding second data, and also includes:

[0091] S31, creating a second model corresponding to the subtree extension path; wherein the second model is a process reward model;

[0092] S32. Based on the second model and in combination with the ε-greedy strategy, corresponding third data are generated; wherein the third data are path quality evaluation data; wherein the calculation formula of the exploration ratio ε is:

[0093]

[0094] In the formula, ∈ min is the minimum value of the exploration probability, ∈0 is the initial exploration probability, and t is the current number of iterations;

[0095] S33, based on the path cumulative reward, combined with the process reward and path length penalty, automatically adjust the corresponding path selection during the reasoning process to balance the breadth and depth of the search; wherein the cumulative reward value calculation formula is:

[0096]

[0097] Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty.

[0098] The independent simulation processes the expansion path of each subtree and generates corresponding first data, and the confidence data of each node is backtracked and updated according to the first data to generate corresponding second data, and also includes:

[0099] S34. After the path simulation is completed, the node information is updated and processed based on the simulation results, and subsequent decisions are optimized; the accumulated rewards of the nodes are updated in real time based on the current simulation rewards; the update formula is:

[0100]

[0101] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It is a simulated reward;

[0102] S35. After the backtracking path update simulation is completed, backtrack from the leaf node to the root node to update the number of visits and historical rewards of all visited nodes.

[0103] Specifically, in an embodiment of the present invention, an enhanced large language model reasoning method based on Monte Carlo tree search is provided, and the algorithm includes the following steps:

[0104] S01, data preparation and model fine-tuning, use a dataset containing logical reasoning process to fine-tune all parameters of the large language model;

[0105] S02, initial subtree generation. At the beginning of the search, the search tree is divided into N initial paths based on the task requirements. Each path is assigned M nodes for independent subtree expansion, where N is the number of initial paths and M is the number of nodes assigned to each path.

[0106] S03, Dynamic Path Expansion: In each subtree, multiple candidate outputs are generated based on the fine-tuning model. Based on the confidence calculation of each node, the search tree will prioritize the expansion of paths with high confidence and dynamically adjust the beam width in combination with the score difference of the process reward model, thereby balancing the search depth and breadth;

[0107] S04, path simulation: The expansion path of each subtree is simulated independently. Based on its confidence score, more paths with higher confidence may be simulated selectively and the ε-greedy strategy is used to combine the process reward model score and the logical consistency of the path for path simulation;

[0108] S05, backtracking update: based on the simulation results and process reward model feedback, backtrack and update the confidence score, cumulative reward and number of visits of each node;

[0109] S06. Reflection mechanism to summarize and optimize the strategies and methods in the reasoning process, so as to continuously improve the efficiency and accuracy of reasoning.

[0110] The data set in step S01 includes but is not limited to math problem solving, chain thinking reasoning data set, context-related question-answering reasoning data set, long-span logical reasoning data set, and domain-specific knowledge question-answering reasoning data set.

[0111] The search tree in step S02 is used to represent different reasoning steps and their subsequent decision spaces. Each node represents an intermediate state of the problem, and each edge represents the transition from one intermediate step to another. Root node: represents the initial state of the problem. Child nodes represent possible subsequent steps extending from the current node. Leaf nodes represent the final candidate solutions. The initial subtree generates N and M values, which are adjusted according to the task complexity and computing resources to optimize the reasoning efficiency.

[0112] The confidence calculation formula in step S03 is:

[0113]

[0114] Where, is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients used to balance the contribution of different factors to the confidence.

[0115] In step S03, the beam width is dynamically adjusted by using the variance based on the process reward model score. The specific calculation formula is:

[0116]

[0117] Among them, M min is the minimum beam width, M max is the maximum value of the beam width, and variance(PRM_scores) is the variance of the current process reward model score. After each step is completed, the beam width is adjusted dynamically. If the paths of some subtrees are relatively certain, the beam width can be reduced to reduce the calculation of low-confidence paths; if the paths of some subtrees are uncertain, the beam width can be increased for further exploration.

[0118] In step S04, the expansion path of each subtree is simulated independently to predict the quality of the reasoning path. The simulation adopts the ε-greedy strategy, which combines the process reward model score with the logical consistency of the path to evaluate and select the path. The calculation formula of its exploration ratio ε is:

[0119]

[0120] where ∈ minis the minimum value of the exploration probability to prevent it from becoming a completely greedy choice (for example, set to 0.1), ∈0 is the initial exploration probability (for example, 0.8), and t is the current number of iterations. As the number of steps increases, the exploration probability gradually decreases, enhancing the utilization of high-confidence paths.

[0121] The path simulation in step S04 includes generating multiple reasoning steps until a termination mark is generated or a preset maximum depth is reached, and the path cumulative reward value includes a path length penalty item to avoid the generation of an excessively long path.

[0122] The cumulative reward of the path combines the process reward and the path length penalty. The model can automatically adjust the path selection during the reasoning process to balance the breadth and depth of the search. The cumulative reward value calculation formula is:

[0123]

[0124] Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty, which determines the degree of influence of the path length on the cumulative reward.

[0125] In the backtracking update mechanism in step S05, each time a path is simulated, a node of the tree is visited once, and the number of visits to the node increases by 1. Each time a node is visited, the accumulated reward of the node is updated according to the current simulation reward, and the update formula is:

[0126]

[0127] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It's a simulated reward.

[0128] The reflection mechanism in step S06 includes the following sub-steps:

[0129] Error analysis: Analyze the reasoning trace of the generated path to identify and record wrong steps or low-quality paths;

[0130] Strategy optimization: adjust the weight parameters, subtree allocation strategy or path generation rules of the process reward model according to the error analysis results to optimize the reasoning efficiency;

[0131] Iterative learning: Convert analysis results into new fine-tuned data or model rules, and feed them back into the model to form a closed-loop optimization process to further improve the quality of reasoning.

[0132] That is to say, the scheme of the present invention proposes an enhanced large language model reasoning method based on Monte Carlo tree search, which aims to optimize reasoning efficiency and accuracy through reasonable path expansion, dynamic adjustment strategy and simulation feedback mechanism. This method combines the fine-tuning technology of the large language model and the path selection algorithm of Monte Carlo tree search, and can dynamically select and optimize the reasoning path according to the requirements of different reasoning tasks, thereby improving the performance of the large language model in complex reasoning tasks.

[0133] Among them, an enhanced large language model reasoning method based on Monte Carlo tree search includes:

[0134] S010, Data preparation and model fine-tuning: Use a dataset containing logical reasoning processes to fine-tune all parameters of the large language model to ensure that the model has strong reasoning ability and logical consistency in reasoning tasks. The dataset includes but is not limited to math problem solving, chained reasoning, context-related question answering, long-span logical reasoning, and domain-specific knowledge question answering.

[0135] S020, initial subtree generation: At the beginning of reasoning, the search tree is divided into N initial paths based on task requirements, and each path is assigned M nodes for independent subtree expansion, where N is the number of initial paths and M is the number of nodes assigned to each path. This step is dynamically adjusted based on task complexity and computing resources to optimize reasoning efficiency.

[0136] S030, Dynamic Path Extension: In each subtree, multiple candidate outputs are generated based on the fine-tuned model, the confidence of each node is calculated, and the beam width is adjusted dynamically. The beam width is adjusted based on the score difference of the process reward model, and the paths with high confidence are preferentially expanded, and the reasoning process is optimized by balancing the search depth and breadth.

[0137] S040, Path Simulation: The expansion path of each subtree is simulated independently, and the ε-greedy strategy is used to combine the process reward model score and the logical consistency of the path to evaluate the path quality. The exploration ratio ε will gradually decrease according to the current iteration step, thereby enhancing the utilization of high-confidence paths.

[0138] S050, Backtracking Update: Based on the simulation results and process reward model feedback, backtrack and update the confidence score, cumulative reward and number of visits of each node. Each time a path is simulated, the number of node visits increases by 1, and the node reward value is updated.

[0139] S060, Reflection mechanism: Summarize the lessons learned in the reasoning process, conduct error analysis, identify low-quality paths and optimize strategies. According to the error analysis results, adjust the parameters of the process reward model and the path generation rules, and feed the results back to the model to form a closed-loop optimization process to continuously improve the efficiency and accuracy of reasoning.

[0140] The search tree structure is used to represent the decision space in the reasoning process. The search tree structure is designed as follows:

[0141] (1) Node representation: Each node represents an intermediate state in the reasoning task. The node contains historical rewards, number of visits, logical consistency score, confidence, parent node and child node references. The root node represents the initial state of the reasoning task. The child nodes are generated by the expansion of the current node and represent possible subsequent steps. The leaf nodes represent possible terminal states or final solutions.

[0142] The confidence calculation formula is:

[0143]

[0144] Where, is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients used to balance the contribution of different factors to the confidence.

[0145] (2) Edge representation: Each edge represents the transition from one intermediate step to another, recording the operations required or the output generated along the path. Connecting two nodes represents the transition or derivation step from one intermediate state to the next state.

[0146] Edges represent reasoning steps from one node to the next:

[0147] Transfer process: The inference steps generated by the large language model determine the generation of edges, such as from "factorization step one" to "factorization completed".

[0148] Score update: Each edge records the reward value passed, affecting the node’s cumulative reward.

[0149] Pruning strategy: If the cumulative reward of an edge is lower than the threshold for a long time, it can be pruned while dynamically adjusting the bundle width.

[0150] The reasoning action strategy is the core part of implementing enhanced large language model reasoning based on Monte Carlo Tree Search (MCTS), which is used to guide the node expansion, path selection and final reasoning decision of the search tree.

[0151] (1) Confidence-driven expansion: Path expansion gives priority to nodes with higher confidence. The confidence calculation formula is:

[0152]

[0153] (2) Dynamically adjust the cluster width: In order to optimize resource allocation, the path expansion stage dynamically controls the search width:

[0154]

[0155] Among them, M min is the minimum beam width, M max is the maximum value of the beam width, and variance(PRM_scores) is the variance of the current process reward model score. After each step is completed, the beam width is adjusted dynamically. If the paths of some subtrees are relatively certain, the beam width can be reduced to reduce the calculation of low-confidence paths; if the paths of some subtrees are uncertain, the beam width can be increased for further exploration.

[0156] (3) Beam search: When expanding, a beam search strategy is used to retain only the M with the highest confidence. new Nodes with low confidence are removed to avoid wasting resources.

[0157] Path simulation is used to evaluate the potential quality of the current path, combining the ε-greedy strategy and path logic consistency for simulation.

[0158] (1) ε-greedy strategy: In the simulation stage, path selection combines exploration and utilization, and the calculation formula of its exploration ratio ε is:

[0159]

[0160] where ∈ min is the minimum value of the exploration probability to prevent it from becoming a completely greedy choice (for example, set to 0.1), ∈0 is the initial exploration probability (for example, 0.8), and t is the current number of iterations. As the number of steps increases, the exploration probability gradually decreases, enhancing the utilization of high-confidence paths. Exploration phase: randomly select paths to discover potential high-quality solutions. Utilization phase: select paths with higher confidence to optimize reasoning efficiency.

[0161] (2) Calculation of cumulative reward value: During the simulation process, the cumulative reward value of the path is used to measure the overall quality of the path. The cumulative reward of the path combines the process reward and the path length penalty. The model can automatically adjust the path selection during the reasoning process to balance the breadth and depth of the search. The cumulative reward value calculation formula is:

[0162]

[0163] Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty, which determines the degree of influence of the path length on the cumulative reward.

[0164] The cumulative reward balances the depth and quality of the path, giving priority to solutions with high rewards and short paths, and penalizing overly long paths to save computing resources.

[0165] Backtracking update strategy: After the path simulation is completed, the node information is backtracked and updated according to the simulation results to optimize subsequent decisions. The node's cumulative reward will be updated according to the current simulation reward. The update formula is:

[0166]

[0167] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It's a simulated reward.

[0168] As the number of visits increases, the confidence of the node gradually stabilizes, avoiding over-reliance on a single simulation result.

[0169] After the backtracking path update simulation is completed, backtrack from the leaf node to the root node and update the visit count T of all visited nodes n and historical reward R n , the updated confidence C n Will affect the next round of expansion.

[0170] Dynamic adjustment of reasoning action strategy

[0171] (1) Balance between path exploration and utilization: Dynamically adjust according to the complexity of the task and the characteristics of the search space:

[0172] Exploration ratio: Gradually reduce randomness and enhance utilization through the ε-greedy strategy.

[0173] Beam width: Adjust exploration depth and breadth via variance control.

[0174] (2) Balance between rewards and penalties: Paths with high rewards are prioritized, and paths with poor logical consistency are gradually eliminated.

[0175] Long paths are penalized for their length, avoiding lengthy reasoning.

[0176] (3) Reflection mechanism guides optimization: adjust the weight parameters of the process reward model through error analysis and optimize the path selection rules.

[0177] Feed the features of low-quality paths back into model fine-tuning to continuously improve the inference strategy.

[0178] The reflection mechanism includes the following sub-steps: Error analysis: Analyze the reasoning trajectory of the generated path, identify and record the wrong steps or low-quality paths; Strategy optimization: Adjust the weight parameters of the process reward model, subtree allocation strategy or path generation rules according to the error analysis results to optimize the reasoning efficiency; Iterative learning: Convert the analysis results into new fine-tuning data or model rules, and feed them back to the model to form a closed-loop optimization process to further improve the reasoning quality.

[0179] Example 1

[0180] Taking the review of medical device standards (taking the review of technical requirements for medical protective masks as an example) as the background, the preferred embodiment of the present invention is described in detail, with reference to Figure 1 and Figure 2 , elaborate on the specific implementation methods of this patent:

[0181] Assume that a certain medical protective mask needs to be reviewed for the following technical requirements:

[0182] Filtration efficiency: Is the particle filtration efficiency (PFE) greater than or equal to 95%.

[0183] Pressure difference: whether the breathing resistance meets the standard (such as ≤49Pa).

[0184] Microbiological indicators: whether the bioburden (number of bacteria or fungi) is below a set threshold.

[0185] Material safety: whether it meets biocompatibility requirements (such as sensitization and cytotoxicity tests).

[0186] Comprehensive evaluation: Based on the above test results, comprehensively evaluate whether the mask complies with national standards (such as GB19083-2010).

[0187] Taking the review of technical requirements for medical protective masks as an example, the present invention specifically includes the following steps:

[0188] S0100, Data preparation and model fine-tuning: Dataset: Use standard documents related to medical device audits (such as "Technical Requirements for Medical Protective Masks" GB19083-2010) and specific audit cases to call the GPT-4o model to generate a fine-tuning dataset.

[0189] The dataset includes:

[0190] Specific terms of the technical standards (such as filtration efficiency ≥ 95%, pressure difference ≤ 49Pa, etc.).

[0191] Historical audit records (including cases of compliance and non-compliance with standards).

[0192] The logical deduction relationship between various indicators and audit conclusions.

[0193] Model fine-tuning:

[0194] All parameters are fine-tuned based on the Qwen2-7B-Instruct model to enable it to understand the professional terminology and logical reasoning process of medical device standards.

[0195] Ensure that the model can generate step-by-step reasoning paths and output logical conclusions.

[0196] S0200, initial subtree generation

[0197] Initialize the audit task:

[0198] Model the auditing task as a Monte Carlo Tree Search (MCTS) problem:

[0199] Root node: represents the initial state of the audit task.

[0200] Subnode: represents the intermediate status of each step in the audit process, such as "filtration efficiency qualified" or "filtration efficiency unqualified".

[0201] Leaf node: indicates the final status of the audit process (such as "passed the audit" or "failed the audit").

[0202] Generate the initial path:

[0203] According to technical requirements, the audit task is divided into multiple initial paths, each path corresponds to an audit indicator:

[0204] Path 1: Audit filter efficiency.

[0205] Path 2: Audit pressure difference.

[0206] Path 3: Review microbiological indicators.

[0207] Path 4: Review material safety.

[0208] Each path is assigned several nodes, for example:

[0209] Initial node of path 1: particle filtration efficiency test value (such as 97%, 94%, 90%).

[0210] Initial node of path 2: pressure difference test result (such as 40Pa, 50Pa).

[0211] S0300, Dynamic Path Extension

[0212] (1) Node confidence calculation

[0213] The confidence of each node C n The calculation formula is:

[0214]

[0215] Where, is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients used to balance the contribution of different factors to the confidence.

[0216] (2) Expansion strategy

[0217] Prioritize expanding nodes with higher confidence, for example:

[0218] If the filtration efficiency test value is 97% and the logic consistency is high, the path is expanded first.

[0219] If the pressure difference test value is 50Pa, the logic consistency is low, and the path is extended with a delay.

[0220] (3) Dynamically adjust the beam width

[0221] Bunch width M new Dynamic Adjustment:

[0222]

[0223] Among them, M min is the minimum beam width, M max is the maximum beam width, and variance(PRM_scores) is the variance of the current process reward model score.

[0224] If the variance of the PRM scores is high (high path uncertainty), increase the beam width to perform a wider search.

[0225] If the score variance is low (path quality is relatively certain), reduce the beam width to save computing resources.

[0226] S0400, Path Simulation

[0227] (1) Simulation target

[0228] Simulate the audit results for each pathway and assess the quality and potential conclusions of the pathway.

[0229] Combined with the ε-greedy strategy, a dynamic balance is achieved between exploration and exploitation:

[0230] Exploration phase: Randomly select paths (such as trying to expand the microbial indicator path).

[0231] Utilization phase: Select the path with the highest confidence (such as preferentially expanding the filtration efficiency path).

[0232] (2) Calculation of cumulative reward value

[0233] The cumulative reward value R of the path total Comprehensively consider process rewards and path length penalties.

[0234] The calculation formula for the cumulative reward value is:

[0235]

[0236] Among them, R total is the cumulative reward value of the path (for example, nodes with filtering efficiency that meets the standard will receive higher rewards), r t is the reward value of the tth step in the path, L is the length of the path (to avoid lengthy reasoning), and λ is the coefficient of the path length penalty, which determines the degree of influence of the path length on the cumulative reward.

[0237] S05, Retroactive Update

[0238] (1) Node information update

[0239] According to the simulation results, the number of visits to each node T is updated backtrackingly. n 、Historical Reward R n and confidence C n .

[0240] The node's cumulative reward will be updated according to the current simulation reward. The update formula is:

[0241]

[0242] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It's a simulated reward.

[0243] (2) Path priority optimization

[0244] The accumulated reward value of high-quality paths (such as paths with qualified filtering efficiency) increases and their priority is improved.

[0245] Low-quality pathways (such as pathways with unqualified microbiological indicators) are given lower priority.

[0246] S06. Reflection Mechanism

[0247] (1) Error analysis

[0248] Analyze low-quality paths in the audit process, such as:

[0249] If the logic of the filtration efficiency path is inconsistent (eg, the test value is lower than 95%, but is mistakenly judged as qualified), an error feature is recorded.

[0250] If there is a major problem with the microbial indicator path (such as excessive bioburden), it is flagged as a critical error.

[0251] (2) Strategy Optimization

[0252] Adjust the parameters of the process reward model (such as increasing the weight of the logical consistency score).

[0253] Optimize path generation rules (such as giving priority to expanding filtration efficiency paths).

[0254] (3) Closed-loop optimization

[0255] Feed the error analysis results back into model fine-tuning to form new training data.

[0256] Avoid similar mistakes in the next round of review and improve reasoning efficiency and accuracy.

[0257] To achieve the above object, the present invention also provides an enhanced large language model processing system based on Monte Carlo tree search, wherein the system is used to implement the enhanced large language model processing method based on Monte Carlo tree search; Figure 3-Figure 4 As shown, the system specifically includes:

[0258] A first data processing unit, configured to create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model;

[0259] A second data processing unit is used to divide the search tree into at least two initial paths based on task requirements, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding beam width is dynamically adjusted in combination with confidence data;

[0260] The third data processing unit is used to independently simulate and process the expansion path of each subtree, and generate corresponding first data, backtrack and update the confidence data of each node according to the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; and the second data is the inference response data after enhanced optimization processing.

[0261] The datasets include math problem solving, chain thinking reasoning, context-related question answering, long-span logical reasoning, and domain-specific knowledge question answering.

[0262] The second data processing unit further includes:

[0263] The first calculation module is used to calculate and generate confidence data corresponding to the path node; wherein the calculation formula is as follows:

[0264]

[0265] In the formula, R n is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients;

[0266] The first processing module is used to preferentially expand and process high-confidence paths based on the search tree, and dynamically adjust the bundle width in combination with the score difference of the process reward model; wherein the adjustment calculation formula is as follows:

[0267]

[0268] Where M min is the minimum beam width, M max is the maximum value of the beam width, variance(PRM_scores) is the variance of the current process reward model score;

[0269] And / or, the third data processing unit further includes:

[0270] A first creation module is used to create a second model corresponding to the subtree expansion path; wherein the second model is a process reward model;

[0271] The first generating module is used to generate corresponding third data based on the second model and in combination with the ε-greedy strategy; wherein the third data are path quality evaluation data; wherein the calculation formula of the exploration ratio ε is:

[0272]

[0273] In the formula, ∈ min is the minimum value of the exploration probability, ∈0 is the initial exploration probability, and t is the current number of iterations;

[0274] The second processing module is used to automatically adjust the corresponding path selection during the reasoning process based on the path cumulative reward, combined with the process reward and the path length penalty item, to balance the breadth and depth of the search; wherein the cumulative reward value calculation formula is:

[0275]

[0276] Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty.

[0277] The third data generating unit further includes:

[0278] The third processing module is used to update the node information according to the simulation results after the path simulation is completed, and optimize the subsequent decision-making; the accumulated rewards of the node are updated in real time according to the current simulation rewards; the update formula is:

[0279]

[0280] Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It is a simulated reward;

[0281] The fourth processing module is used to trace back from the leaf node to the root node after the backtracking path update simulation is completed, and update the number of visits and historical rewards of all visited nodes.

[0282] That is to say, the solution of the present invention provides an enhanced large language model reasoning system based on Monte Carlo tree search, comprising:

[0283] Model fine-tuning module: customizes the language model to ensure that it generates high-quality reasoning output in the task;

[0284] Monte Carlo Tree Search Module: Uses the Monte Carlo Tree Search algorithm to search and expand paths, ensuring the breadth and depth of the exploration space while optimizing path selection;

[0285] Process reward module: provides a reward score for each reasoning path and calculates the confidence of the path to guide the path expansion and search process;

[0286] Reflection mechanism module: summarizes the lessons learned in the reasoning process, continuously optimizes system strategies and reasoning capabilities, and promotes the iterative evolution of the system.

[0287] Specifically, an embodiment of the present invention provides an enhanced large language model reasoning system based on Monte Carlo tree search, the system is used to execute an enhanced large language model reasoning method based on Monte Carlo tree search described in the first aspect embodiment, the system comprising:

[0288] Model fine-tuning module: This module is responsible for customizing the language model to ensure that it can generate high-quality reasoning output based on specific tasks.

[0289] Monte Carlo Tree Search Module: This module uses the Monte Carlo tree search algorithm to expand and select paths, ensuring the breadth and depth of the reasoning space while optimizing path selection.

[0290] Process reward module: provides a reward score for each reasoning path, calculates the confidence of the path, and guides the path expansion and search process.

[0291] Reflection mechanism module: This module summarizes the lessons learned in the reasoning process, continuously optimizes system strategies, promotes the iterative evolution of the system, and improves reasoning capabilities.

[0292] In the system solution embodiment of the present invention, the specific details of the method steps involved in the enhanced large language model processing based on Monte Carlo tree search have been explained above, that is, the functional modules in the system are used to implement the steps or sub-steps in the above method embodiment, which will not be repeated here.

[0293] To achieve the above objectives, the present invention also provides an enhanced large language model processing platform based on Monte Carlo tree search, such as Figure 5 As shown, it includes a processor, a memory, and an enhanced large language model processing platform control program based on Monte Carlo tree search; wherein, the processor executes the enhanced large language model processing platform control program based on Monte Carlo tree search, the enhanced large language model processing platform control program based on Monte Carlo tree search is stored in the memory, and the enhanced large language model processing platform control program based on Monte Carlo tree search implements the steps of the enhanced large language model processing method based on Monte Carlo tree search. For example:

[0294] S1. Create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model;

[0295] S2, dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent subtree expansion, and generating at least two candidate outputs in each subtree based on the first model, and dynamically adjusting the corresponding beam width in combination with confidence data;

[0296] S3. Independently simulate and process the expansion path of each subtree, and generate corresponding first data. Backtrack and update the confidence data of each node based on the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; the second data is the inference response data after enhanced optimization processing.

[0297] The specific details of the steps have been explained above and will not be repeated here.

[0298] In the embodiment of the present invention, the built-in processor of the enhanced large language model processing platform based on Monte Carlo tree search can be composed of integrated circuits, for example, a single packaged integrated circuit, or a plurality of integrated circuits with the same or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor uses various interfaces and lines to connect various components, and executes or executes programs or units stored in the memory, and calls data stored in the memory to perform various functions and process data of the enhanced large language model based on Monte Carlo tree search;

[0299] The memory is used to store program codes and various data, and is installed in an enhanced large language model processing platform based on Monte Carlo tree search, and realizes high-speed and automatic access to programs or data during operation.

[0300] The memory includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electronically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0301] To achieve the above object, the present invention also provides a computer readable storage medium, such as Figure 6 As shown, the computer-readable storage medium stores an enhanced large language model processing platform control program based on Monte Carlo tree search, and the enhanced large language model processing platform control program based on Monte Carlo tree search implements the steps of the enhanced large language model processing method based on Monte Carlo tree search, for example:

[0302] S1. Create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model;

[0303] S2, dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent subtree expansion, and generating at least two candidate outputs in each subtree based on the first model, and dynamically adjusting the corresponding beam width in combination with confidence data;

[0304] S3. Independently simulate and process the expansion path of each subtree, and generate corresponding first data. Backtrack and update the confidence data of each node based on the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; the second data is the inference response data after enhanced optimization processing.

[0305] The specific details of the steps have been explained above and will not be repeated here.

[0306] In the description of the embodiments of the present invention, it should be noted that any process or method description in the flowchart or otherwise described herein may be understood as representing a module, fragment or portion of a code comprising one or more executable instructions for implementing steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations, in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present invention belong.

[0307] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processing module, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM).

[0308] In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0309] In an embodiment of the present invention, in order to achieve the above-mentioned purpose, the present invention further provides a chip system, wherein the chip system includes at least one processor, and when the program instructions are executed in the at least one processor, the chip system executes the steps of the enhanced large language model processing method based on Monte Carlo tree search, for example:

[0310] S1. Create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model;

[0311] S2, dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent subtree expansion, and generating at least two candidate outputs in each subtree based on the first model, and dynamically adjusting the corresponding beam width in combination with confidence data;

[0312] S3. Independently simulate and process the expansion path of each subtree, and generate corresponding first data. Backtrack and update the confidence data of each node based on the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; the second data is the inference response data after enhanced optimization processing.

[0313] The specific details of the steps have been explained above and will not be repeated here.

[0314] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0315] The present invention creates a first model corresponding to a large language through a method, generates and obtains a data set corresponding to a logical reasoning process, and fine-tunes the first model according to the data set; wherein the first model is a large language model; based on task requirements, the search tree is divided into at least two initial paths, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding cluster width is dynamically adjusted in combination with confidence data; the expansion path of each subtree is independently simulated and processed, and the corresponding first data is generated, and the confidence data of each node is backtracked and updated according to the first data, and the corresponding second data is generated at the same time; wherein the first data is the simulation result data; the second data is the reasoning response data after enhanced optimization processing, as well as the system, platform and storage medium corresponding to the method, so as to improve the efficiency and accuracy of reasoning processing.

[0316] That is to say, the present invention proposes an enhanced large language model push processing method based on Monte Carlo tree search, which optimizes the reasoning efficiency and accuracy through reasonable path expansion, dynamic adjustment strategy and simulation feedback mechanism. The scheme of the present invention combines the fine-tuning technology of the large language model and the path selection algorithm of Monte Carlo tree search, and can dynamically select and optimize the reasoning path according to the requirements of different reasoning tasks, thereby improving the performance of the large language model in complex reasoning tasks.

[0317] In other words, the enhanced large language model reasoning method based on Monte Carlo tree search based on a large language model provided by the present invention can significantly improve the performance of LLM in complex reasoning tasks by combining Monte Carlo tree search and a large language model, especially in applications such as mathematical problem solving, chain thinking reasoning, and context-related question and answer reasoning.

[0318] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. An enhanced large language model processing method based on Monte Carlo tree search, characterized in that: The method comprises: Creating a first model corresponding to the large language, generating and acquiring a data set corresponding to the logical reasoning process, and fine-tuning the first model according to the data set; wherein the first model is a large language model; Based on the task requirements, the search tree is divided into at least two initial paths, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding beam width is dynamically adjusted in combination with the confidence data; Independently simulate and process the expansion path of each subtree and generate corresponding first data. Backtrack and update the confidence data of each node based on the first data and generate corresponding second data at the same time; wherein the first data is the simulation result data; the second data is the inference response data after enhanced optimization processing.

2. The enhanced large language model processing method based on Monte Carlo tree search according to claim 1, characterized in that: The datasets include math problem solving, chain thinking reasoning, context-related question-answering reasoning, long-span logical reasoning, and domain-specific knowledge question-answering reasoning.

3. The enhanced large language model processing method based on Monte Carlo tree search according to claim 1, characterized in that: The method further comprises: dividing the search tree into at least two initial paths based on task requirements, allocating at least two nodes to each path for independent subtree expansion, and generating at least two candidate outputs in each subtree based on the first model, and dynamically adjusting the corresponding beam width in combination with confidence data. Calculate and generate confidence data corresponding to the path nodes; the calculation formula is as follows: In the formula, R n is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients; Based on the search tree, the paths with high confidence are preferentially expanded, and the beam width is dynamically adjusted in combination with the score difference of the process reward model; the adjustment calculation formula is as follows: Where M min is the minimum beam width, M max is the maximum beam width, and variance(PRM_scores) is the variance of the current process reward model score.

4. The enhanced large language model processing method based on Monte Carlo tree search according to claim 1, characterized in that: The independent simulation processes the expansion path of each subtree and generates corresponding first data, and the confidence data of each node is backtracked and updated according to the first data to generate corresponding second data, and also includes: Creating a second model corresponding to the subtree expansion path; wherein the second model is a process reward model; Based on the second model and in combination with the ε-greedy strategy, corresponding third data are generated; wherein the third data are path quality evaluation data; wherein the calculation formula of the exploration ratio ε is: In the formula, ∈ min is the minimum value of the exploration probability, ∈0 is the initial exploration probability, and t is the current number of iterations; Based on the path cumulative reward, combined with the process reward and path length penalty, the corresponding path selection is automatically adjusted during the reasoning process to balance the breadth and depth of the search; the cumulative reward value calculation formula is: Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty.

5. The enhanced large language model processing method based on Monte Carlo tree search according to claim 1 or 4, characterized in that: The independent simulation processes the expansion path of each subtree and generates corresponding first data, and the confidence data of each node is backtracked and updated according to the first data to generate corresponding second data, and also includes: After the path simulation is completed, the node information is updated and processed based on the simulation results, and subsequent decisions are optimized; the node's cumulative reward is updated in real time based on the current simulation reward; the update formula is: Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It is a simulated reward; After the backtracking path update simulation is completed, backtrack from the leaf node to the root node to update the visit count and historical rewards of all visited nodes.

6. An enhanced large language model processing system based on Monte Carlo tree search, characterized in that: The system is applied to an enhanced large language model processing method based on Monte Carlo tree search as claimed in any one of claims 1 to 5; the system comprises: A first data processing unit, configured to create a first model corresponding to a large language, generate and obtain a data set corresponding to a logical reasoning process, and fine-tune the first model according to the data set; wherein the first model is a large language model; A second data processing unit is used to divide the search tree into at least two initial paths based on task requirements, each path is assigned at least two nodes for independent subtree expansion, and at the same time, in each subtree, at least two candidate outputs are generated based on the first model, and the corresponding beam width is dynamically adjusted in combination with confidence data; The third data processing unit is used to independently simulate and process the expansion path of each subtree, and generate corresponding first data, backtrack and update the confidence data of each node according to the first data, and generate corresponding second data at the same time; wherein, the first data is the simulation result data; and the second data is the inference response data after enhanced optimization processing.

7. The enhanced large language model processing system based on Monte Carlo tree search according to claim 6, characterized in that: The datasets include math problem solving, chain thinking reasoning, context-related question answering, long-span logical reasoning, and domain-specific knowledge question answering. The second data processing unit further includes: The first calculation module is used to calculate and generate confidence data corresponding to the path node; wherein the calculation formula is as follows: In the formula, R n is the historical reward of node n, T n is the number of visits to node n, L n The logical consistency score of the node, α, β, γ are weight coefficients; The first processing module is used to preferentially expand and process high-confidence paths based on the search tree, and dynamically adjust the bundle width in combination with the score difference of the process reward model; wherein the adjustment calculation formula is as follows: Where M min is the minimum beam width, M max is the maximum value of the beam width, variance(PRM_scores) is the variance of the current process reward model score; And / or, the third data processing unit further includes: A first creation module is used to create a second model corresponding to the subtree expansion path; wherein the second model is a process reward model; The first generating module is used to generate corresponding third data based on the second model and in combination with the ε-greedy strategy; wherein the third data are path quality evaluation data; wherein the calculation formula of the exploration ratio ε is: In the formula, ∈ min is the minimum value of the exploration probability, ∈0 is the initial exploration probability, and t is the current number of iterations; The second processing module is used to automatically adjust the corresponding path selection during the reasoning process based on the path cumulative reward, combined with the process reward and the path length penalty item, to balance the breadth and depth of the search; wherein the cumulative reward value calculation formula is: Among them, R total is the cumulative reward value of the path, r t is the reward value of the tth step in the path, L is the length of the path, and λ is the coefficient of the path length penalty.

8. The enhanced large language model processing system based on Monte Carlo tree search according to claim 6 or 7, characterized in that: The third data generating unit further includes: The third processing module is used to update the node information according to the simulation results after the path simulation is completed, and optimize the subsequent decision-making; the accumulated rewards of the node are updated in real time according to the current simulation rewards; the update formula is: Among them, Q(s) is the current cumulative reward of node s, N(s) is the number of visits to node s, r t It is a simulated reward; The fourth processing module is used to trace back from the leaf node to the root node after the backtracking path update simulation is completed, and update the number of visits and historical rewards of all visited nodes.

9. An enhanced large language model processing platform based on Monte Carlo tree search, characterized in that: The invention comprises a processor, a memory and an enhanced large language model processing platform control program based on Monte Carlo tree search; wherein the enhanced large language model processing platform control program based on Monte Carlo tree search is executed by the processor, the enhanced large language model processing platform control program based on Monte Carlo tree search is stored in the memory, and the enhanced large language model processing platform control program based on Monte Carlo tree search implements the enhanced large language model processing method based on Monte Carlo tree search as described in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an enhanced large language model processing platform control program based on Monte Carlo tree search, and the enhanced large language model processing platform control program based on Monte Carlo tree search implements the enhanced large language model processing method based on Monte Carlo tree search as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Large model reasoning system, method and equipment based on Monte Carlo tree search

    CN120354953A

  • Database logic error discovery-oriented large model enhanced query rewriting synthesis method

    CN120492487A

  • Large-Model Enhanced Query Rewriting Synthesis Method for Database Logical Error Detection

    CN120492487B

  • Multi-path legal inference engine based on MCP protocol

    CN120764667A

  • Intelligent application interaction method and system driven by multi-mode end side model

    CN121188100A