Question answering method and device, equipment, medium and program product

By introducing a Monte Carlo tree search method based on decision tree in the large language model, the problem of low efficiency and accuracy of multi-step questions is solved, and a more efficient and accurate answering process is achieved.

CN119990314APending Publication Date: 2025-05-13IFLYTEK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510056814.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When answering multi-step questions, existing large language models have low efficiency and accuracy, and lack of generalization and high-cost tree search limits the search depth.

Method used

The Monte Carlo tree search method based on decision tree is adopted to improve the efficiency and accuracy of the large language model in multi-step questions through the steps of selection, expansion, simulation and backtracking. The specific steps include selecting the next child node based on the root node of the decision tree, expanding the child node of the leaf node, deleting unnecessary child nodes, and performing simulation and backtracking to determine the target child node.

Benefits of technology

By reducing the number of simulations and backtracking, the efficiency and accuracy of multi-step questions are improved, making it suitable for handling complex multi-step questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990314A_ABST
    Figure CN119990314A_ABST
Patent Text Reader

Abstract

The invention provides a question answering method, device and equipment, a medium and a program product, and the question answering method comprises the steps: selecting a next child node based on a root node of a decision tree until a leaf node is reached; nodes of the decision tree comprise answering content composed of at least one answering step of the question to be answered; under the condition that the leaf node is not the terminal node, expanding each child node of the leaf node; determining features corresponding to each child node of the leaf node, and based on each feature, deleting a part of child nodes of the leaf node to obtain reserved child nodes of the leaf node; and performing simulation and backtracking based on the reserved child nodes of the leaf nodes, determining a target child node of the root node, determining the target child node as the root node of the decision tree, and returning to execute the step of selecting the next child node based on the root node of the decision tree until the complete answer content of the question to be answered is generated. According to the invention, the efficiency and accuracy of multi-step question answering can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text processing technology, and in particular to a question answering method, device, equipment, medium and program product. Background Art

[0002] In recent years, large language models have attracted widespread attention and have shown great potential for answering multi-step questions. Unlike conventional text generation tasks, answering multi-step questions has the characteristics of many steps and strong logic. Therefore, using large language models to answer multi-step questions is still a big challenge.

[0003] In order to improve the reasoning and planning capabilities of large language models, existing research has introduced various technologies, such as multi-step reasoning and tree search, so that large language models can think in steps when generating answers, thereby improving their reasoning accuracy. Through these technologies, large language models can think in multiple steps like humans and find more accurate solutions. This method is in line with human thinking habits and has strong interpretability, so it has received more and more attention.

[0004] However, the efficiency and accuracy of using large language models to answer multi-step questions are currently low. Summary of the invention

[0005] The present application provides a question-solving method, apparatus, device, medium, and program product for improving the efficiency and accuracy of multi-step question-solving.

[0006] According to a first aspect of an embodiment of the present application, a method for answering a question is provided, comprising:

[0007] Selecting a next child node based on the root node of the decision tree until a leaf node is reached; wherein the node of the decision tree includes a solution content consisting of at least one solution step of the question to be solved; and the edge of the decision tree includes the solution step;

[0008] When the leaf node is not a terminal node, expanding each child node of the leaf node;

[0009] Determine the features corresponding to each child node of the leaf node, and based on each of the features, delete some child nodes of the leaf node to obtain the retained child nodes of the leaf node;

[0010] Based on the retained child nodes of the leaf node, simulation and backtracking are performed to determine the target child node of the root node, the target child node is determined as the root node of the decision tree, and the step of selecting the next child node based on the root node of the decision tree is returned to execute until the complete answer content of the question to be answered is generated.

[0011] According to a second aspect of an embodiment of the present application, a question answering device is provided, comprising:

[0012] A selection unit, configured to select a next child node based on a root node of a decision tree until a leaf node is reached; wherein the node of the decision tree includes a solution content consisting of at least one solution step of a question to be solved; and the edge of the decision tree includes a solution step;

[0013] An expansion unit, used for expanding each child node of the leaf node when the leaf node is not a terminal node;

[0014] A deleting unit, used to determine the features corresponding to the respective child nodes of the leaf node, and based on the respective features, delete some of the child nodes of the leaf node to obtain the retained child nodes of the leaf node;

[0015] The processing unit is used to simulate and backtrack based on the retained child nodes of the leaf node, determine the target child node of the root node, determine the target child node as the root node of the decision tree, and return to execute the step of selecting the next child node based on the root node of the decision tree until the complete answer content of the question to be answered is generated.

[0016] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including a memory and a processor;

[0017] The memory is connected to the processor and is used to store programs;

[0018] The processor is used to implement the question-solving method as described in the first aspect by running the program in the memory.

[0019] According to a fourth aspect of an embodiment of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for solving a problem as described in the first aspect is implemented.

[0020] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer program instructions, which, when executed by a processor, cause the processor to execute the question-solving method as described in the first aspect.

[0021] In the present application, the next child node is selected based on the root node of the decision tree until a leaf node is reached, wherein the node of the decision tree includes a solution content composed of at least one solution step of the question to be answered, and the edge of the decision tree includes a solution step. When the leaf node is not a terminal node, the various child nodes of the leaf node are expanded, the corresponding features of the various child nodes of the leaf node are determined, and based on the various features, some child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node. Based on the retained child nodes of the leaf node, simulation and backtracking are performed to determine the target child node of the root node, and the target child node is determined as the root node of the decision tree. The step of selecting the next child node based on the root node of the decision tree is returned to execute until the complete solution content of the question to be answered is generated. In the present application, the Monte Carlo tree search method of selection, expansion, simulation and backtracking is used to find the next solution step faster and better in the process of solving the question, until the complete solution content of the question to be solved is generated, thereby improving the efficiency and accuracy of solving multi-step questions. Moreover, after expanding the child nodes of the leaf node, based on the corresponding features of the child nodes of the leaf node, some of the child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node, and simulation and backtracking are performed based on the retained child nodes of the leaf node. Since some of the child nodes of the leaf node are deleted, the number of simulations and backtracking is reduced, which further improves the efficiency of solving multi-step problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0023] Figure 1 A flowchart of a method for answering a question provided in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of a process of selecting the next child node from a root node of a decision tree provided in an embodiment of the present application;

[0025] Figure 3 A flowchart of a training process of a process reward model provided in an embodiment of the present application;

[0026] Figure 4 A flowchart of a training process of a correction model provided in an embodiment of the present application;

[0027] Figure 5 A schematic diagram of a process for obtaining a retained child node of a leaf node provided in an embodiment of the present application;

[0028] Figure 6 A flowchart of a method for answering a question provided in an embodiment of the present application;

[0029] Figure 7 A schematic diagram of a process of step 602 provided in an embodiment of the present application;

[0030] Figure 8 A schematic diagram of a process of step 603 provided in an embodiment of the present application;

[0031] Fig. 9 A schematic diagram of the structure of a question-solving device provided in an embodiment of the present application;

[0032] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] In recent years, large language models have attracted widespread attention and have shown great potential for answering multi-step questions. Unlike conventional text generation tasks, answering multi-step questions has the characteristics of many steps and strong logic. Therefore, using large language models to answer multi-step questions is still a big challenge.

[0034] In order to improve the reasoning and planning capabilities of large language models, existing research has introduced various technologies, such as multi-step reasoning and tree search, so that large language models can think in steps when generating answers, thereby improving their reasoning accuracy. Through these technologies, large language models can think in multiple steps like humans and find more accurate solutions. This method is in line with human thinking habits and has strong interpretability, so it has received more and more attention.

[0035] However, existing solutions often rely on the large language model itself to evaluate the current state, which lacks generalization. In addition, the high cost of tree search limits the depth of the tree, which is not conducive to exploring multi-step questions. Therefore, how to design an efficient search method to improve the answering performance of the large language model has become an urgent problem to be solved.

[0036] In order to enhance the question-answering ability of large language models, existing search solutions are mainly divided into the following two categories:

[0037] Category I

[0038] Chain of Thought (CoT) requires the large language model to first decompose the problem and gradually reason about the results of each step, rather than directly giving the answer. This reasoning method is linear and cannot try multiple different ideas at the same time, which limits its reasoning ability and is not suitable for handling complex multi-step problems that require multiple different ideas.

[0039] Category II

[0040] Tree search, such as Tree-of-Thought (ToT) and Reasoning-via-Planning (RAP), converts the reasoning process into a tree structure, allowing large language models to select and evaluate multiple possible paths.

[0041] The tree search method has the following disadvantages: (1) Multi-step questions have many steps and strong logic, but they use Monte Carlo tree search or depth / breadth first search, which has high search cost and limits the search depth, and is not suitable for multi-step questions. (2) The large language model itself is used to score the intermediate steps, which lacks generalization and robustness. (3) Training is often only performed once, and the improvement effect is limited.

[0042] In summary, the efficiency and accuracy of using large language models to answer multi-step questions are currently low.

[0043] In order to improve the efficiency and accuracy of multi-step question answering, the present application provides a question answering method, apparatus, device, medium and program product.

[0044] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0045] Exemplary Implementation Environment

[0046] The method for answering questions according to the embodiment of the present application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user device, a mobile device, a computing device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be implemented by a processor calling a computer-readable program instruction stored in a memory.

[0047] Exemplary Methods

[0048] See also Figure 1 In an exemplary embodiment, a method for solving a problem is provided. Figure 1 As shown, the process of the question-solving method mainly includes:

[0049] Step 101, select the next child node based on the root node of the decision tree until a leaf node is reached.

[0050] The nodes of the decision tree include solution content consisting of at least one solution step of the problem to be solved; and the edges of the decision tree include solution steps.

[0051] In an exemplary embodiment, each node of the decision tree is generated by the large language model for answering questions based on the questions to be answered. The question-answering process of the large language model for answering questions is modeled as a multi-step Markov decision process, in which generating a section of answering steps is defined as an action, and the answer content generated by the large language model for answering questions is defined as a state. By applying the idea of ​​Monte Carlo tree search to the large language model for answering questions, exploration and selection are performed among multiple possible paths, helping the large language model for answering questions to make higher-quality decisions when answering questions, and improving the multi-step question-answering ability of the large language model for answering questions. A node of a decision tree is actually a state, and an edge of a decision tree is actually an action.

[0052] In an exemplary embodiment, an action may be defined as generating a solution step of a fixed length (eg, 80 in length), or may be defined as generating a solution step, which may not be a solution step of a fixed length, and the present application does not impose any limitation on this.

[0053] In an exemplary embodiment, the Monte Carlo tree search refers to the Monte Carlo tree search of AlphaZero - α.

[0054] For example, the root node of the decision tree is the solution content consisting of step A and step B, and there are three selectable edges after the root node, namely step C, step D and step E. The child nodes of the root node include child node 1, child node 2 and child node 3. Child node 1 is the solution content consisting of step A, step B and step C, child node 2 is the solution content consisting of step A, step B and step D, and child node 3 is the solution content consisting of step A, step B and step E. After selecting child node 1 based on certain node selection rules, select another child node from each child node of child node 1, and so on, until a leaf node is reached.

[0055] In an exemplary embodiment, a leaf node refers to a node without child nodes.

[0056] In an exemplary embodiment, the question to be answered can be determined as the initial state, that is, the first root node of the decision tree; or the first solution step of the question to be answered can be determined as the first root node of the decision tree, and the present application does not limit this.

[0057] In an exemplary embodiment, the problem to be solved may be a multi-step problem, such as a math problem, a physics problem, and the like.

[0058] In some embodiments, selecting the next child node based on the root node of the decision tree in step 101 includes: selecting the next child node based on the root node of the decision tree by using a UCB (Upper Confidence Bounds) method.

[0059] In other embodiments, Figure 2 As shown, in step 101, selecting the next child node based on the root node of the decision tree includes:

[0060] Step 201 , performing the following operations for any candidate child node of the root node: determining a prediction score corresponding to any candidate child node.

[0061] In an exemplary embodiment, the prediction score corresponding to any candidate sub-node is used to predict the correctness of any candidate sub-node.

[0062] In an exemplary embodiment, during the selection phase, the Monte Carlo tree search uses the current state as the root node.

[0063] In some embodiments, determining the predicted score corresponding to any candidate sub-node in step 201 includes: determining the scores of each step in the standard answer to the question to be answered; and determining the predicted score corresponding to any candidate sub-node based on the scores of each step in the standard answer to the question to be answered.

[0064] In other embodiments, determining the prediction score corresponding to any candidate subnode in step 201 includes: inputting the question to be answered and any candidate subnode into a process reward model to obtain the prediction score corresponding to any candidate subnode. The process reward model is used to predict the correctness of any candidate subnode.

[0065] In an exemplary embodiment, the preset score may be a value between 0 and 1, or may be a specific score, for example, 5 points, and the present application is not limited to this.

[0066] In an exemplary embodiment, the process reward model is used to predict the correctness of the intermediate answer content of the question to be answered based on the question to be answered and the intermediate answer content of the question to be answered. The intermediate answer content can be a complete answer content or a partial answer content, and the present application does not limit this. For example, the complete answer content of the question to be answered is the answer content composed of step A, step B and step D. The process reward model can predict the correctness of step A, the correctness of the answer content composed of step A and step B, and the correctness of the answer content composed of step A, step B and step D.

[0067] By predicting the correctness of the intermediate solutions to the questions to be solved through the process reward model, the tree search can be guided more accurately, the efficiency of the tree search can be improved, and the efficiency of solving multi-step problems can be improved.

[0068] In some embodiments, each node of the decision tree is generated by the large language model for answering questions based on the questions to be answered.

[0069] like Figure 3 As shown in Figure 1, the training process of the process reward model includes:

[0070] Step 301: input a first sample question into a large question-answering language model to obtain a first sample answer content for the first sample question.

[0071] In an exemplary embodiment, the first sample answer content may be a complete answer content of the first sample question, or may be a partial answer content of the first sample question, that is, an intermediate step of the first sample question.

[0072] Step 302: input the first sample question, the standard answer to the first sample question and the first sample solution content into the grading model to obtain the first sample score corresponding to the first sample solution content.

[0073] The first sample score is used to predict the correctness of the first sample answer content.

[0074] In an exemplary embodiment, the first sample score may be a value between 0 and 1, or may be a specific score, for example, 5 points, and the present application is not limited to this.

[0075] In an exemplary embodiment, the grading model is used to predict the correctness of the first sample answer content based on the first sample question, the standard answer to the first sample question, and the first sample answer content of the first sample question.

[0076] Step 303: training the process reward model based on the first sample question, the first sample answer content and the first sample score.

[0077] In some embodiments, step 301 includes: inputting the first sample question once into the large language model for answering questions, and obtaining a first sample answer content of the first sample question. Step 302 includes: inputting the first sample question, the standard answer of the first sample question, and a first sample answer content into the grading model, and obtaining a first sample score corresponding to the first sample answer content. Step 303 includes: training the process reward model based on the first sample question, a first sample answer content of the first sample question, and a first sample score corresponding to the first sample answer content of the first sample question.

[0078] Among them, the number of the first sample questions can be set freely.

[0079] In other embodiments, step 301 includes: repeatedly inputting the first sample question into the large language model for answering a preset number of times to obtain a preset number of first sample answer contents for the first sample question. Step 302 includes: inputting the first sample question, the standard answer to the first sample question, and a preset number of first sample answer contents into the grading model to obtain scores corresponding to the preset number of first sample answer contents; selecting a first sample answer content for the first sample question from the preset number of first sample answer contents for the first sample question; and determining the average of the scores corresponding to the preset number of first sample answer contents as a first sample score corresponding to a first sample answer content. Step 303 includes: training the process reward model based on the first sample question, a first sample answer content for the first sample question, and a first sample score corresponding to the first sample answer content.

[0080] Among them, the number of the first sample questions can be set freely.

[0081] For example, the preset number of times is 20, and the large language model for answering questions is used to repeatedly sample 20 times to generate 20 first sample answer contents for the first sample questions; the first sample questions, the standard answers to the first sample questions, and the 20 first sample answer contents are input into the grading model to obtain the scores corresponding to the 20 first sample answer contents; one first sample answer content is selected from the 20 first sample answer contents, and the average value of the scores corresponding to the 20 first sample answer contents is determined as the first sample score corresponding to the selected first sample answer content; based on the first sample question, the selected first sample answer content of the first sample question, and the first sample score corresponding to the first sample answer content, the process reward model is trained.

[0082] Repeat inputting the first sample question a preset number of times into the large language model for answering the question, obtaining a preset number of first sample answer contents for the first sample question, determining the average of the scores corresponding to the preset number of first sample answer contents as a first sample score corresponding to a first sample answer content, and training the process reward model can improve the accuracy of the first sample score corresponding to the first sample answer content, thereby more accurately predicting the correctness of the first sample answer content, and training the process reward model can further improve the accuracy of the process reward model in predicting the correctness of the intermediate answer content of the question to be answered.

[0083] The first sample question is input into the large language model for answering questions to obtain the first sample answer content of the first sample question, the first sample question, the standard answer of the first sample question and the first sample answer content are input into the correction model to obtain the first sample score corresponding to the first sample answer content, wherein the first sample score is used to predict the correctness of the first sample answer content, and the process reward model is trained based on the first sample question, the first sample answer content and the first sample score. By inputting the first sample answer content, the first sample question and the standard answer of the first sample question generated by the large language model for answering questions into the correction model, the correctness of the first sample answer content generated by the large language model for answering questions can be accurately predicted, and the process reward model is trained based on the first sample question, the first sample answer content and the first sample score, which can further improve the accuracy of the process reward model in predicting the correctness of the intermediate answer content generated by the large language model for answering questions to be answered.

[0084] In some embodiments, Figure 4 As shown in Figure 1, the training process of the correction model includes:

[0085] Step 401: input the second sample question into the question-answering language model to obtain a second sample answer content for the second sample question.

[0086] Step 402 , training the grading model based on the second sample questions, the standard answers to the second sample questions, the second sample solution content, and the real sample scores of the second sample solution content.

[0087] The true sample score is used to indicate the actual correctness of the second sample answer content.

[0088] In an exemplary embodiment, the real sample score of the second sample answer content can be obtained by manually correcting the second sample answer content.

[0089] In an exemplary embodiment, the true sample score of the second sample answer content may include 0 or 1. When the second sample answer content is correct, the true sample score of the second sample answer content is 1; when the second sample answer content is wrong, the true sample score of the second sample answer content is 0.

[0090] Training the correction model based on the second sample questions, the standard answers to the second sample questions, the second sample answer contents generated by the question-answering large language model for the second sample questions, and the actual sample scores of the second sample answer contents can improve the performance of the correction model. In the subsequent process, the correction model can provide supervision signals for the training of the process reward model and the multiple iterative optimization of the question-answering large language model and the process reward model.

[0091] Step 202: Determine the probability of selecting any candidate child node when the root node has been selected.

[0092] In an exemplary embodiment, step 202 includes: determining the probability of the large question-answering language model selecting any candidate child node when the root node has been selected.

[0093] In an exemplary embodiment, selecting any candidate child node when the root node has been selected is the same as selecting the action corresponding to any candidate child node when the root node has been selected. The action corresponding to any candidate child node refers to the edge connecting the root node and any candidate child node. When the root node has been selected, any candidate child node can be reached from the root node through the action corresponding to any candidate child node.

[0094] In an exemplary embodiment, the probability of selecting any one of the candidate child nodes when the root node has been selected is the same as the probability of selecting an action corresponding to any one of the candidate child nodes when the root node has been selected.

[0095] In an exemplary embodiment, step 202 may include: determining the probability of each token (word) in the action corresponding to any candidate child node generated by the question-answering language model when the answer content of the root node has been generated, and determining the average value of the probabilities of each token in the action corresponding to any candidate child node as the probability of selecting any candidate child node when the root node has been selected.

[0096] Step 203: Determine the number of visits to any candidate child node when the root node has been selected.

[0097] In an exemplary embodiment, the number of visits to select any one of the candidate child nodes when the root node has been selected is the same as the number of visits corresponding to the action of selecting any one of the candidate child nodes when the root node has been selected.

[0098] Step 204 , based on the prediction scores, probabilities, and access times corresponding to the candidate child nodes of the root node, and the total access times of the root node, determine the next child node from the candidate child nodes of the root node.

[0099] In the exemplary embodiment, the root node is represented by state s, and the action corresponding to any candidate child node is represented by action a. Any candidate child node is the result of the concatenation of state s and action a. Determining the next child node from each candidate child node of the root node can be regarded as selecting action a under state s. The specific formula for selecting action a under state s is as follows:

[0100]

[0101] Among them, Q(s,a) represents the reward for taking action a in state s, that is, the prediction score corresponding to any candidate child node; P(s,a) represents the probability of the large language model taking action a in state s, that is, the probability of selecting any candidate child node when the root node has been selected; N(s,a) represents the number of visits to take action a in state s, that is, the number of visits to select any candidate child node when the root node has been selected; ∑ b N(s,b) represents the sum of the number of visits of all access actions b in state s, that is, the total number of visits to the current state s, that is, the total number of visits to the root node; c puct It is a hyperparameter used for exploration and utilization in the balanced tree search process; argmax is the set of independent variables that corresponds to the maximum value of the dependent variable, that is, the set of independent variable points with the maximum value.

[0102] The following operations are performed for any candidate child node of the root node: determining the prediction score corresponding to any candidate child node, determining the probability of selecting any candidate child node when the root node has been selected, determining the number of visits to select any candidate child node when the root node has been selected, and determining the next child node from each candidate child node of the root node based on the prediction score, probability, and number of visits corresponding to each candidate child node of the root node, as well as the total number of visits to the root node. In the process of determining the next child node from each candidate child node of the root node, both the prediction score corresponding to any candidate child node, that is, the correctness of any candidate child node, and the probability of selecting any candidate child node by the large language model for answering questions when the root node has been selected are considered, which can guide tree search and improve the efficiency of tree search, thereby improving the efficiency and accuracy of multi-step question answering.

[0103] Step 102: When the leaf node is not a terminal node, expand each child node of the leaf node.

[0104] In an exemplary embodiment, a terminal node means that when the terminal node is executed, the task of answering the unanswered question has been completed and a complete answer content has been generated.

[0105] When the leaf node is not a terminal node, it is necessary to continue the tree search and expand the child nodes of the leaf node.

[0106] In an exemplary embodiment, the expansion of each child node of a leaf node is achieved through a large language model for answering questions.

[0107] Step 103, determining the features corresponding to each child node of the leaf node, and based on the features, deleting some child nodes of the leaf node to obtain the retained child nodes of the leaf node.

[0108] In an exemplary embodiment, the features corresponding to each sub-node may refer to the semantic features corresponding to each sub-node.

[0109] In an exemplary embodiment, a trained language model may be used to extract features corresponding to each sub-node.

[0110] In some embodiments, Figure 5 As shown, in step 103, based on various features, some child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node, including:

[0111] Step 501, determining the similarity between each feature.

[0112] In an exemplary embodiment, step 501 may include: determining the semantic similarity between each feature, for example, determining the cosine similarity between each feature, or determining the Euclidean distance between each feature, or other similarity calculation methods, which are not limited in this application.

[0113] Step 502: determine each target feature from each feature based on each similarity.

[0114] Among them, the similarity between each target feature is greater than the similarity threshold.

[0115] In an exemplary embodiment, the cosine similarity between each target feature may be greater than the similarity threshold, or the Euclidean distance between each target feature may be less than the distance threshold. The smaller the distance, the greater the similarity, and there is a negative correlation between the distance and the similarity.

[0116] Step 503: Delete some sub-nodes from the sub-nodes corresponding to each target feature to obtain the retained sub-nodes of the leaf node.

[0117] In an exemplary embodiment, some sub-nodes are deleted from the sub-nodes corresponding to each target feature to obtain the retained sub-nodes of the leaf node. It can be that only one sub-node is retained as the retained sub-node of the leaf node among the sub-nodes corresponding to each target feature, or a preset number of sub-nodes is retained as the retained sub-nodes of the leaf node among the sub-nodes corresponding to each target feature, or a preset proportion of sub-nodes is retained as the retained sub-nodes of the leaf node among the sub-nodes corresponding to each target feature. The present application is not limited to this.

[0118] For example, if only two characters in the two child nodes expanded from a leaf node are different and the other contents are the same, only one child node can be retained.

[0119] Determine the similarity between each feature, and based on each similarity, determine each target feature from each feature, wherein the similarity between each target feature is greater than the similarity threshold, delete some sub-nodes from the sub-nodes corresponding to each target feature, and obtain the retained sub-nodes of the leaf node. Delete some sub-nodes from the sub-nodes whose similarity is greater than the similarity threshold, and improve the efficiency of tree search without affecting the accuracy of tree search, thereby improving the efficiency and accuracy of multi-step problem solving.

[0120] In other embodiments, in step 103, based on various features, some child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node, including: based on various features, the child nodes of the leaf node are clustered, and only one child node is retained in each cluster set as the retained child node of the leaf node.

[0121] Step 104 , based on the retained child nodes of the leaf node, simulation and backtracking are performed to determine the target child node of the root node.

[0122] In an exemplary embodiment, simulation means starting from the retained child node of the leaf node, running a simulated output until the task of answering the unanswered question is completed, and obtaining the score and the number of visits of the retained child node of the leaf node. The score of the retained child node of the leaf node can be the correct degree of predicting the retained child node of the leaf node using the process reward model.

[0123] In an exemplary embodiment, backtracking refers to updating the simulation results to all passed nodes of the tree, and the updated content includes information such as the number of visits and scores.

[0124] In an exemplary embodiment, after completing multiple searches, based on the number of visits to each of the child nodes of the root node, the child node with the largest number of visits is selected as the target child node of the root node.

[0125] Step 105 , determine whether a complete answer to the question to be answered has been generated, if so, execute step 106 , otherwise, execute step 107 .

[0126] In an exemplary embodiment, determining whether a complete answer to a question to be answered has been generated may be performed by determining whether the number of simulated searches has reached a preset number. One round of selection, expansion, simulation, and backtracking is called a simulated search. Of course, determining whether a complete answer to a question to be answered has been generated may also be performed in other ways, and the present application does not limit this.

[0127] Step 106, end the process.

[0128] Step 107 , determining the target child node as the root node of the decision tree, and returning to execute step 101 .

[0129] In the present application, the Monte Carlo tree search method of selection, expansion, simulation and backtracking is used to find the next solution step faster and better in the process of solving the problem, until the complete solution content of the problem to be solved is generated, thereby improving the efficiency and accuracy of solving multi-step problems. Moreover, after expanding each child node of the leaf node, based on the corresponding features of each child node of the leaf node, some child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node, and simulation and backtracking are performed based on the retained child nodes of the leaf node. Since some child nodes of the leaf node are deleted, the number of simulations and backtracking is reduced, further improving the efficiency of solving multi-step problems.

[0130] In some embodiments, Figure 6 As shown, the method for solving the question also includes:

[0131] Step 601, input each target sub-node corresponding to the question to be answered, the standard answer to the question to be answered and the complete answer content of the question to be answered into the correction model to obtain the second sample score corresponding to each target sub-node.

[0132] The second sample score is used to predict the correctness of the target subnode.

[0133] In an exemplary embodiment, a question to be answered may have one complete answer or at least two complete answers, and this application is not limited to this.

[0134] In an exemplary embodiment, a complete answer content is generated, which corresponds to a search path during the tree search process, and a search path is composed of multiple target sub-nodes.

[0135] In an exemplary embodiment, before step 601, the question answering method further includes: deleting erroneous complete answer contents in each complete answer content of the question to be answered, and retaining only correct complete answer contents.

[0136] Step 602: training the large language model for answering questions based on the target sub-nodes corresponding to the questions to be answered and the complete answers to the questions to be answered.

[0137] In some embodiments, Figure 7 As shown, step 602 includes:

[0138] Step 701 : determining the strategy distribution corresponding to each target sub-node in the tree search process of the decision tree.

[0139] In an exemplary embodiment, during the tree search process of the decision tree, the kth target child node s k The strategy distribution is π k express.

[0140] In an exemplary embodiment, ~ refers to equivalence; N(s k ,a) refers to the target child node s k The number of visits to take action a when the target child node s k Select the next target child node s (k+1) Number of visits; ∑ b N(s k ,b) refers to the state s k The sum of the number of visits to all access actions b, that is, the target child node s k The total number of visits.

[0141] Step 702, determining the prediction probability corresponding to each target sub-node generated by the question answering language model.

[0142] In an exemplary embodiment, the large language model for answering questions generates the kth target child node s k The predicted probability is expressed as p k express.

[0143] Step 703, determining the posterior probability of the complete answer content of the question to be answered generated by the large language model for answering the question.

[0144] In the exemplary embodiment, the complete answer content of the question to be answered is represented by s select Indicates that P θ represents the posterior probability output by the large language model for answering questions, P θ (s select ) represents the posterior probability that the large language model for answering questions generates a complete answer to the question to be answered.

[0145] Step 704: Determine the supervised fine-tuning loss of the large language model for answering questions based on the questions to be answered and the complete answers to the questions to be answered.

[0146] In an exemplary embodiment, the supervised fine-tuning loss is express.

[0147] Step 705, based on each strategy distribution, each prediction probability, posterior probability and supervised fine-tuning loss, determine the first target loss of the question answering language model.

[0148] In an exemplary embodiment, the specific formula of the first target loss of the large language model for answering questions is as follows:

[0149]

[0150] Among them, in the tree search process of the decision tree, the kth target child node s k The strategy distribution is π k Indicates; the large language model for answering questions generates the kth target child node s k The predicted probability is expressed as p k Indicates; the complete answer to the question to be answered is indicated by s select Indicates that P θ represents the posterior probability output by the large language model for answering questions, P θ (s select ) represents the posterior probability of the complete answer content of the question to be answered by the large language model; the supervised fine-tuning loss is Represents; γ and λ are hyperparameters that control the weights of each loss.

[0151] Among them, ∑ k π k logp k The steps used to make the generation of the large language model for answering questions close to the optimal steps obtained by Monte Carlo tree search; θ (s select ) is used to make the probability of the path obtained during the Monte Carlo tree search as high as possible.

[0152] Step 706: Train the question-answering language model based on the first target loss.

[0153] Determine the strategy distribution corresponding to each target sub-node in the tree search process of the decision tree, determine the prediction probability corresponding to each target sub-node generated by the large language model for answering questions, determine the posterior probability of the large language model for answering questions generating the complete answer content of the question to be answered, determine the supervised fine-tuning loss of the large language model for answering questions based on the question to be answered and the complete answer content of the question to be answered, determine the first target loss of the large language model for answering questions based on each strategy distribution, each prediction probability, posterior probability and supervised fine-tuning loss, and train the large language model for answering questions based on the first target loss. The steps generated by the large language model for answering questions can be made close to the optimal steps obtained by Monte Carlo tree search, and the probability of the path obtained in the Monte Carlo tree search process can be as high as possible, guiding the large language model for answering questions to generate complete answer content faster and better, and improving the efficiency and accuracy of multi-step question answering.

[0154] Step 603: Training the process reward model based on the unanswered questions, the target sub-nodes corresponding to the complete answers to the unanswered questions, and the second sample scores.

[0155] In an exemplary embodiment, each target sub-node and each second sample score corresponding to the question to be answered and the complete answer content of the question to be answered can be determined as an enhanced data set. Based on the enhanced data set, the large language model for answering questions and the process reward model are trained.

[0156] Input the target sub-nodes corresponding to the questions to be answered, the standard answers to the questions to be answered, and the complete answers to the questions to be answered into the grading model, obtain the second sample scores corresponding to the target sub-nodes, train the large language model for answering questions based on the target sub-nodes corresponding to the questions to be answered and the complete answers to the questions to be answered, and train the process reward model based on the target sub-nodes and the second sample scores corresponding to the questions to be answered and the complete answers to the questions to be answered. Use Monte Carlo tree search to assist the large language model for answering questions to be answered. After generating the complete answers to the questions to be answered, further train the large language model for answering questions and the process reward model based on the target sub-nodes and the second sample scores corresponding to the questions to be answered and the complete answers to the questions to be answered. Fine-tune the parameters of the large language model for answering questions and the process reward model. After the training of the large language model for answering questions and the process reward model is completed, use Monte Carlo tree search to assist the large language model for answering questions, implement multiple iterative optimizations of the large language model for answering questions and the process reward model, and simultaneously improve the performance of the large language model for answering questions and the process reward model.

[0157] In some embodiments, after training the large language model for answering questions and the process reward model, the trained large language model for answering questions and the process reward model are used to continue to perform the Monte Carlo tree search of steps 101 to 107, which can achieve better answering results in the Monte Carlo tree search process. The process of training the large language model for answering questions and the process reward model and the process of Monte Carlo tree search can be performed alternately, and multiple iterative optimizations are performed to continuously optimize the large language model for answering questions and the process reward model until the large language model for answering questions and the process reward model converge.

[0158] In some embodiments, Figure 8 As shown, step 603 includes:

[0159] Step 801, for any target sub-node corresponding to the complete answer content of the question to be answered, perform the following operations: based on the question to be answered and any target sub-node, determine the third sample scores corresponding to each candidate sub-node output by the process reward model for any target sub-node.

[0160] The third sample score is used to predict the correctness of the candidate child nodes of any target child node.

[0161] Step 802: sort candidate child nodes of any target child node according to the number of visits to obtain a sorting result.

[0162] In the exemplary embodiment, the question to be answered is represented by q, and any target subnode is represented by s. t Indicates that t There are l candidate actions. Sort these l candidate actions in descending order according to the number of visits, and get (a1,…,a l ), the candidate child nodes of any target child node are sorted according to the number of visits, and (s t +a1,…,s t +a l ). i and r j Respectively represent the given question q to be answered and any target child node s t In the case of i and a j The third sample score output, that is, the process reward model for the candidate child node (s t +a i ) and (s t +a j ) outputs the third sample score.

[0163] Step 803: Determine the sorting loss corresponding to any target sub-node based on the sorting result and each third sample score.

[0164] In an exemplary embodiment, ∑ l≥i>j≥1 log(σ(r i ―r j )) is any target child node s t The corresponding sorting loss, where σ refers to the sigmoid function, which is the activation function in the neural network.

[0165] Because representations with high visit counts are more likely to generate the correct answer process, the process reward model can be further optimized by combining the ranking loss.

[0166] Step 804: input the to-be-answered question and each target sub-node into the process reward model to obtain a fourth sample score corresponding to each target sub-node.

[0167] The fourth sample score is used to predict the correctness of the target child node.

[0168] In an exemplary embodiment, V θ represents the reward output by the process reward model, V θ (s t ) refers to any target child node s tThe corresponding fourth sample score.

[0169] Step 805 : Determine the second target loss of the process reward model based on each second sample score, each fourth sample score, and the ranking loss corresponding to each target sub-node.

[0170] In an exemplary embodiment, z t refers to any target child node s t The score of the second sample predicted by the revised model.

[0171] In an exemplary embodiment, the specific formula of the second target loss of the process reward model is as follows:

[0172]

[0173] The question to be answered is represented by q, and any target child node is represented by s. t Indicates that t There are l candidate actions. Sort these l candidate actions in descending order according to the number of visits, and get (a1,…,a l ), the candidate child nodes of any target child node are sorted according to the number of visits, and (s t +a1,…,s t +a l ). i and r j Respectively represent the given question q to be answered and any target child node s t In the case of i and a j The third sample score output, that is, the process reward model for the candidate child node (s t +a i ) and (s t +a j )The third sample score output. ∑ l≥i>j≥1 log(σ(r i ―r j )) is any target child node s t The corresponding sorting loss, where σ refers to the sigmoid function, which is the activation function in the neural network. V θ represents the reward output by the process reward model, V θ (s t ) refers to any target child node s t The corresponding fourth sample score. t refers to any target child node s t The score of the second sample predicted by the corrected model. η is a hyperparameter that controls the weight of each loss.

[0174] Step 806: training the process reward model based on the second target loss.

[0175] Based on the second sample scores, the fourth sample scores, and the ranking losses corresponding to the target sub-nodes, the second target loss of the process reward model is determined. Based on the second target loss, the process reward model is trained, which can make the process reward model's ability to predict the accuracy of the intermediate answer content closer to the correction model. Moreover, a high number of accesses indicates a greater hope of generating a correct answer process. Combined with the ranking loss, the process reward model can more accurately assist the answering language model in generating the correct answering process, guide the answering language model to generate complete answer content faster and better, and improve the efficiency and accuracy of answering multi-step questions.

[0176] After optimizing the large language model and process reward model for answering questions, better answering results can be achieved during the tree search process.

[0177] In an exemplary embodiment, the grading model is trained first, and after the grading model training is completed, the process reward model is trained. In the process of training the grading model, the parameters of the large language model for answering questions are frozen, and the parameters of the grading model are fine-tuned. In the process of training the process reward model, the parameters of the large language model for answering questions are frozen, the parameters of the grading model are frozen, and the parameters of the process reward model are fine-tuned. After the process reward model training is completed, the Monte Carlo tree search is used to assist the large language model for answering questions to be answered. After the complete answer content of the question to be answered is generated, the large language model for answering questions and the process reward model are further trained based on the question to be answered and the complete answer content of the question to be answered. Each target sub-node and each second sample score corresponding to the complete answer content of the question to be answered are used to further train the large language model for answering questions and the process reward model. The parameters of the large language model for answering questions and the process reward model are fine-tuned. After the training of the large language model for answering questions and the process reward model is completed, the Monte Carlo tree search is used to assist the large language model for answering questions, and multiple iterative optimizations of the large language model for answering questions and the process reward model are achieved, and the performance of the large language model for answering questions and the process reward model are simultaneously improved. The performance of the large language model for answering questions and the process reward model is continuously optimized.

[0178] In summary, in the present application, the next child node is selected based on the root node of the decision tree until a leaf node is reached, wherein the node of the decision tree includes a solution content consisting of at least one solution step of the question to be answered, and the edge of the decision tree includes a solution step. When the leaf node is not a terminal node, the various child nodes of the leaf node are expanded, the corresponding features of the various child nodes of the leaf node are determined, and based on the various features, some child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node. Based on the retained child nodes of the leaf node, simulation and backtracking are performed to determine the target child node of the root node, and the target child node is determined as the root node of the decision tree. The step of selecting the next child node based on the root node of the decision tree is returned to execute until the complete solution content of the question to be answered is generated. In the present application, the Monte Carlo tree search method of selection, expansion, simulation and backtracking is used to find the next solution step faster and better in the process of answering the question until the complete solution content of the question to be answered is generated, thereby improving the efficiency and accuracy of answering multi-step questions. Moreover, after expanding the child nodes of the leaf node, based on the corresponding features of the child nodes of the leaf node, some of the child nodes of the leaf node are deleted to obtain the retained child nodes of the leaf node, and simulation and backtracking are performed based on the retained child nodes of the leaf node. Since some of the child nodes of the leaf node are deleted, the number of simulations and backtracking is reduced, which further improves the efficiency of solving multi-step problems.

[0179] Exemplary Devices

[0180] Accordingly, the present application embodiment also provides a question answering device, such as Fig. 9 As shown, the question answering device comprises:

[0181] A selection unit 901 is used to select a next child node based on a root node of a decision tree until a leaf node is reached; wherein the node of the decision tree includes a solution content consisting of at least one solution step of the question to be solved; and the edge of the decision tree includes a solution step;

[0182] An expansion unit 902, configured to expand each child node of the leaf node when the leaf node is not a terminal node;

[0183] A deleting unit 903 is used to determine the features corresponding to the respective child nodes of the leaf node, and based on the respective features, delete some of the child nodes of the leaf node to obtain the retained child nodes of the leaf node;

[0184] Processing unit 904 is used to perform simulation and backtracking based on the retained child nodes of the leaf node, determine the target child node of the root node, determine the target child node as the root node of the decision tree, and return to execute the step of selecting the next child node based on the root node of the decision tree until the complete answer content of the question to be answered is generated.

[0185] Optionally, the deletion unit 903 is specifically configured to:

[0186] determining similarities between each of the features;

[0187] Based on each of the similarities, determining each target feature from each of the features; wherein the similarity between each of the target features is greater than a similarity threshold;

[0188] Some sub-nodes are deleted from the sub-nodes corresponding to the respective target features to obtain the retained sub-nodes of the leaf node.

[0189] Optionally, the selection unit 901 includes:

[0190] A prediction score determination subunit is used to perform the following operations on any candidate child node of the root node: determine a prediction score corresponding to the any candidate child node;

[0191] A probability determination subunit, used to determine the probability of selecting any one of the candidate child nodes when the root node has been selected;

[0192] An access count determination subunit, used to determine the access count of selecting any one of the candidate child nodes when the root node has been selected;

[0193] The child node determination subunit is used to determine the next child node from each candidate child node of the root node based on the prediction score, the probability, the number of visits corresponding to each candidate child node of the root node, and the total number of visits to the root node.

[0194] Optionally, the prediction score determination subunit is specifically configured to:

[0195] Inputting the to-be-answered question and any one of the candidate subnodes into a process reward model to obtain a prediction score corresponding to any one of the candidate subnodes;

[0196] The process reward model is used to predict the correctness of any candidate child node.

[0197] Optionally, each node of the decision tree is generated by the large language model for answering questions based on the questions to be answered;

[0198] The question-solving device also includes a first training unit of a process reward model;

[0199] The first training unit of the process reward model is used to:

[0200] Inputting a first sample question into the large question-answering language model to obtain a first sample answer content for the first sample question;

[0201] Inputting the first sample question, the standard answer to the first sample question and the first sample answer content into the grading model to obtain a first sample score corresponding to the first sample answer content; wherein the first sample score is used to predict the correctness of the first sample answer content;

[0202] The process reward model is trained based on the first sample question, the first sample answer content and the first sample score.

[0203] Optionally, the question answering device further comprises a correction model training unit;

[0204] Correction model training unit, used for:

[0205] Inputting the second sample question into the question-answering large language model to obtain a second sample answer content for the second sample question;

[0206] The correction model is trained based on the second sample questions, the standard answers to the second sample questions, the second sample solution content and the real sample scores of the second sample solution content; wherein the real sample scores are used to indicate the actual correctness of the second sample solution content.

[0207] Optionally, the question answering device further includes:

[0208] A sample score determination unit is used to input each of the target sub-nodes corresponding to the question to be answered, the standard answer to the question to be answered and the complete answer content of the question to be answered into the correction model to obtain a second sample score corresponding to each of the target sub-nodes; wherein the second sample score is used to predict the correctness of the target sub-node;

[0209] A question answering language model training unit, configured to train the question answering language model based on the target sub-nodes corresponding to the question to be answered and the complete answer content of the question to be answered;

[0210] The second training unit of the process reward model is used to train the process reward model based on the questions to be answered, the target sub-nodes corresponding to the complete answers to the questions to be answered, and the second sample scores.

[0211] Optionally, the question answering large language model training unit is specifically used for:

[0212] Determine the strategy distribution corresponding to each of the target sub-nodes during the tree search process of the decision tree;

[0213] Determine the prediction probability corresponding to each of the target sub-nodes generated by the question-answering large language model;

[0214] Determine the posterior probability that the large language model for answering the question generates a complete answer to the question to be answered;

[0215] Determining a supervised fine-tuning loss of the large question-answering language model based on the question to be answered and the complete answer content of the question to be answered;

[0216] Determine a first target loss of the question-answering large language model based on each of the strategy distributions, each of the predicted probabilities, the posterior probability, and the supervised fine-tuning loss;

[0217] The question-answering large language model is trained based on the first target loss.

[0218] Optionally, the second training unit of the process reward model is specifically used for:

[0219] For any target sub-node corresponding to the complete answer content of the question to be answered, perform the following operations:

[0220] Based on the to-be-answered question and the arbitrary target subnode, determine the third sample scores corresponding to the outputs of the process reward model for each candidate subnode of the arbitrary target subnode; wherein the third sample scores are used to predict the correctness of the candidate subnodes of the arbitrary target subnode;

[0221] Sorting each candidate child node of any target child node according to the number of visits to obtain a sorting result;

[0222] Determine a sorting loss corresponding to any one of the target subnodes based on the sorting result and each of the third sample scores;

[0223] Inputting the to-be-answered questions and each target sub-node into the process reward model to obtain a fourth sample score corresponding to each target sub-node; wherein the fourth sample score is used to predict the correctness of the target sub-node;

[0224] Determine a second target loss of the process reward model based on each of the second sample scores, each of the fourth sample scores, and the ranking loss corresponding to each of the target subnodes;

[0225] The process reward model is trained based on the second objective loss.

[0226] The question-solving device provided in this embodiment belongs to the same application concept as the question-solving method provided in the above-mentioned embodiment of this application, and can execute the question-solving method provided in any of the above-mentioned embodiments of this application, and has the corresponding functional modules and beneficial effects of executing the question-solving method. For technical details not fully described in this embodiment, please refer to the specific processing content of the question-solving method provided in the above-mentioned embodiment of this application, and will not be repeated here.

[0227] The functions implemented by the above selection unit 901, expansion unit 902, deletion unit 903 and processing unit 904 can be implemented by the same or different processors respectively, and the embodiment of the present application is not limited thereto.

[0228] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by the configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of hardware circuits, or in part by a processor calling software, and the remaining part is implemented in the form of hardware circuits.

[0229] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.

[0230] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0231] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0232] Exemplary Electronic Devices

[0233] An embodiment of the present application provides an electronic device, see Fig.10 As shown, the device includes:

[0234] Memory 200 and processor 210;

[0235] The memory 200 is connected to the processor 210 and is used to store programs;

[0236] The processor 210 is used to implement the question-solving method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0237] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .

[0238] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected to each other via a bus.

[0239] A bus may include a pathway that transfers information between components of a computer system.

[0240] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0241] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0242] The memory 200 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0243] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0244] Output device 240 may include a device that allows information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0245] The communication interface 220 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0246] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any one of the problem-solving methods provided in the above embodiments of the present application.

[0247] Exemplary computer program products and storage media

[0248] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the question-solving method according to various embodiments of the present application described in any of the above-mentioned embodiments of this specification.

[0249] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0250] In addition, the embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to execute the steps of the method for answering questions according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically the following steps may be implemented:

[0251] Step 101, select the next child node based on the root node of the decision tree until a leaf node is reached.

[0252] The nodes of the decision tree include solution content consisting of at least one solution step of the problem to be solved; and the edges of the decision tree include solution steps.

[0253] Step 102: When the leaf node is not a terminal node, expand each child node of the leaf node.

[0254] Step 103, determining the features corresponding to each child node of the leaf node, and based on the features, deleting some child nodes of the leaf node to obtain the retained child nodes of the leaf node.

[0255] Step 104 , based on the retained child nodes of the leaf node, simulation and backtracking are performed to determine the target child node of the root node.

[0256] Step 105 , determine whether a complete answer to the question to be answered has been generated, if so, execute step 106 , otherwise, execute step 107 .

[0257] Step 106, end the process.

[0258] Step 107 , determining the target child node as the root node of the decision tree, and returning to execute step 101 .

[0259] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0260] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0261] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0262] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be combined, divided and deleted according to actual needs.

[0263] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0264] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0265] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0266] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0267] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0268] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0269] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for solving a problem, characterized in that: include: Selecting a next child node based on the root node of the decision tree until a leaf node is reached; wherein the node of the decision tree includes a solution content consisting of at least one solution step of the question to be solved; and the edge of the decision tree includes the solution step; When the leaf node is not a terminal node, expanding each child node of the leaf node; Determine the features corresponding to each child node of the leaf node, and based on each of the features, delete some child nodes of the leaf node to obtain the retained child nodes of the leaf node; Based on the retained child nodes of the leaf node, simulation and backtracking are performed to determine the target child node of the root node, the target child node is determined as the root node of the decision tree, and the step of selecting the next child node based on the root node of the decision tree is returned to execute until the complete answer content of the question to be answered is generated.

2. The method for answering a question according to claim 1, characterized in that: The deleting some child nodes of the leaf node based on each of the features to obtain the retained child nodes of the leaf node includes: determining similarities between each of the features; Based on each of the similarities, determining each target feature from each of the features; wherein the similarity between each of the target features is greater than a similarity threshold; Some sub-nodes are deleted from the sub-nodes corresponding to the respective target features to obtain the retained sub-nodes of the leaf node.

3. The method for answering a question according to claim 1, wherein: The step of selecting a next child node based on the root node of the decision tree comprises: For any candidate child node of the root node, perform the following operations: Determine a prediction score corresponding to any one of the candidate child nodes; Determining a probability of selecting any one of the candidate child nodes when the root node has been selected; Determine the number of visits to select any one of the candidate child nodes when the root node has been selected; Based on the prediction scores, the probabilities, the access times corresponding to the candidate child nodes of the root node, and the total access times of the root node, the next child node is determined from the candidate child nodes of the root node.

4. The method for answering a question according to claim 3, characterized in that: The determining the prediction score corresponding to any one of the candidate child nodes includes: Inputting the to-be-answered question and any one of the candidate subnodes into a process reward model to obtain a prediction score corresponding to any one of the candidate subnodes; The process reward model is used to predict the correctness of any candidate child node.

5. The method for answering a question according to claim 4, characterized in that: Each node of the decision tree is generated by the large language model for answering questions based on the questions to be answered; The training process of the process reward model includes: Inputting a first sample question into the large question-answering language model to obtain a first sample answer content for the first sample question; Inputting the first sample question, the standard answer to the first sample question and the first sample answer content into the grading model to obtain a first sample score corresponding to the first sample answer content; wherein the first sample score is used to predict the correctness of the first sample answer content; The process reward model is trained based on the first sample question, the first sample answer content and the first sample score.

6. The method for answering a question according to claim 5, characterized in that: The training process of the correction model includes: Inputting the second sample question into the question-answering large language model to obtain a second sample answer content for the second sample question; The correction model is trained based on the second sample questions, the standard answers to the second sample questions, the second sample solution content and the real sample scores of the second sample solution content; wherein the real sample scores are used to indicate the actual correctness of the second sample solution content.

7. The method for answering a question according to claim 6, characterized in that: The method further comprises: Inputting each of the target sub-nodes corresponding to the question to be answered, the standard answer to the question to be answered, and the complete answer content of the question to be answered into the grading model to obtain a second sample score corresponding to each of the target sub-nodes; wherein the second sample score is used to predict the correctness of the target sub-node; Based on the target sub-nodes corresponding to the questions to be answered and the complete answers to the questions to be answered, the large language model for answering the questions is trained; The process reward model is trained based on the questions to be answered, the target sub-nodes corresponding to the complete answers to the questions to be answered, and the second sample scores.

8. The method for answering a question according to claim 7, characterized in that: The step of training the large language model for answering questions based on the target sub-nodes corresponding to the questions to be answered and the complete answers to the questions to be answered includes: Determine the strategy distribution corresponding to each of the target sub-nodes during the tree search process of the decision tree; Determine the prediction probability corresponding to each of the target sub-nodes generated by the question-answering large language model; Determine the posterior probability that the large language model for answering the question generates a complete answer to the question to be answered; Determining a supervised fine-tuning loss of the large question-answering language model based on the question to be answered and the complete answer content of the question to be answered; Determine a first target loss of the question-answering large language model based on each of the strategy distributions, each of the predicted probabilities, the posterior probability, and the supervised fine-tuning loss; The question-answering large language model is trained based on the first target loss.

9. The method for answering a question according to claim 7, characterized in that: The training of the process reward model based on the to-be-answered question, each of the target sub-nodes corresponding to the complete answer content of the to-be-answered question, and each of the second sample scores includes: For any target sub-node corresponding to the complete answer content of the question to be answered, perform the following operations: Based on the to-be-answered question and the arbitrary target subnode, determine the third sample scores corresponding to the outputs of the process reward model for each candidate subnode of the arbitrary target subnode; wherein the third sample scores are used to predict the correctness of the candidate subnodes of the arbitrary target subnode; Sorting each candidate child node of any target child node according to the number of visits to obtain a sorting result; Determine a sorting loss corresponding to any one of the target subnodes based on the sorting result and each of the third sample scores; Inputting the to-be-answered questions and each target sub-node into the process reward model to obtain a fourth sample score corresponding to each target sub-node; wherein the fourth sample score is used to predict the correctness of the target sub-node; Determine a second target loss of the process reward model based on each of the second sample scores, each of the fourth sample scores, and the ranking loss corresponding to each of the target subnodes; The process reward model is trained based on the second objective loss.

10. A question-solving device, characterized in that: include: A selection unit, configured to select a next child node based on a root node of a decision tree until a leaf node is reached; wherein the node of the decision tree includes a solution content consisting of at least one solution step of a question to be solved; and the edge of the decision tree includes a solution step; An expansion unit, used for expanding each child node of the leaf node when the leaf node is not a terminal node; A deleting unit, used to determine the features corresponding to the respective child nodes of the leaf node, and based on the respective features, delete some of the child nodes of the leaf node to obtain the retained child nodes of the leaf node; The processing unit is used to simulate and backtrack based on the retained child nodes of the leaf node, determine the target child node of the root node, determine the target child node as the root node of the decision tree, and return to execute the step of selecting the next child node based on the root node of the decision tree until the complete answer content of the question to be answered is generated.

11. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the problem-solving method according to any one of claims 1 to 9 by running the program in the memory.

12. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for solving a problem according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, enable the processor to execute the method for solving a problem as claimed in any one of claims 1 to 9.

Citation Information

Cited By

  • Mathematical problem reasoning trajectory generation method, system and equipment based on large model

    CN120354952A