Search task processing method and target scoring model training method

By introducing the target scoring model in Monte Carlo tree search, the problem of Monte Carlo tree search taking time to deal with complex problems and inaccurate scoring is solved, and efficient and accurate search task processing is achieved.

CN120218173AActive Publication Date: 2025-06-27ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510679123.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Monte Carlo tree search takes time and inaccurate ratings when dealing with complex problems, which limits its practical application.

Method used

The target scoring model is introduced to replace the LLM for scoring, and the target scoring model obtained by training is used to score the target nodes, reducing the scoring time and improving the stability of scoring.

Benefits of technology

It significantly reduces the time-consuming rating, improves the stability and accuracy of ratings, and achieves efficient and accurate search task processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218173A_ABST
    Figure CN120218173A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a search task processing method and a target scoring model training method.The search task processing method comprises the steps that a target node of a target search task is determined, and the target node is an answer needing iterative optimization; a scoring result of the target node is determined through a target scoring model, the target scoring model is obtained based on training of at least one sample node and a corresponding label score, and the label score is determined based on a correctness reward or an execution efficiency reward of the sample node; and performing node expansion based on the scoring result of the target node to determine a search result of the target search task. According to the method, the correctness reward or the execution efficiency reward of the sample node is introduced to determine the label score, the target scoring model is obtained through specific label score training, the special target scoring model is introduced to score the target node, the quality of the target node can be evaluated quickly, stably and accurately, the scoring efficiency is improved, and the scoring accuracy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and particularly to a method for processing search tasks and a method for training a target scoring model. Background Art

[0002] With the rapid development of computer technology and artificial intelligence technology, generation tools based on large language models (LLMs) have become increasingly important. When LLMs reason about complex problems, it is often difficult to directly generate optimal answers. Therefore, Monte Carlo Tree Search (MCTS) can be used to optimize the answers generated by LLMs. However, problems such as long time consumption and inaccurate scoring in Monte Carlo Tree Search limit its practical applications. Therefore, there is an urgent need for a more efficient and accurate search task processing solution. Summary of the Invention

[0003] In view of this, the embodiments of this specification provide a method for processing search tasks. One or more embodiments of this specification also relate to a method for training a target scoring model, a search task processing device, a target scoring model training device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0004] According to the first aspect of the embodiments of this specification, a method for processing search tasks is provided, including: Determine the target node of the target search task, where the target node is the answer that needs to be iteratively optimized; Determine the scoring result of the target node through a target scoring model, where the target scoring model is trained based on at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; Perform node expansion based on the scoring result of the target node to determine the search result of the target search task.

[0005] According to the second aspect of the embodiments of this specification, a method for training a target scoring model is provided, including: Obtain a training sample set under a target search task, where the training sample set includes at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; Based on the sample nodes in the training sample set, obtain the corresponding predicted scores through an initial scoring model; Based on the label score of the sample node and the prediction score, determine the target loss of the initial scoring model, and train the initial scoring model based on the target loss to obtain a trained target scoring model, where the target scoring model is used to score the target node in the above search task processing method.

[0006] According to the third aspect of the embodiments of the present specification, a search task processing device is provided, including: A first determination module configured to determine a target node of a target search task, where the target node is an answer that needs to be iteratively optimized; A scoring module configured to determine a scoring result of the target node through a target scoring model, where the target scoring model is trained based on at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; A second determination module configured to perform node expansion based on the scoring result of the target node to determine the search result of the target search task.

[0007] According to the fourth aspect of the embodiments of the present specification, a training device for a target scoring model is provided, including: An acquisition module configured to acquire a training sample set under a target search task, where the training sample set includes at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; A prediction module configured to obtain a corresponding prediction score through an initial scoring model based on the sample nodes in the training sample set; A training module configured to determine the target loss of the initial scoring model based on the label score of the sample node and the prediction score, and train the initial scoring model based on the target loss to obtain a trained target scoring model, where the target scoring model is used to score the target node in the above search task processing method.

[0008] According to the fifth aspect of the embodiments of the present specification, a computing device is provided, including: A memory and a processor; Wherein, the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the above search task processing method or the training method of the target scoring model are implemented.

[0009] According to the sixth aspect of the embodiments of the present specification, an electronic device is provided, including: A memory and a processor, and the memory and the processor are connected through a bus; Among them, the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned search task processing method or the training method of the target scoring model are implemented.

[0010] According to the seventh aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned search task processing method or the training method of the target scoring model are implemented.

[0011] According to the eighth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned search task processing method or the training method of the target scoring model are implemented.

[0012] An embodiment of the present specification provides a search task processing method, which determines a target node of a target search task, where the target node is an answer that needs to be iteratively optimized; determines a scoring result of the target node through a target scoring model, where the target scoring model is trained based on at least one sample node and a corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; and performs node expansion based on the scoring result of the target node to determine a search result of the target search task.

[0013] An embodiment of the present specification realizes introducing the correctness reward or execution efficiency reward of the sample node, determining the label score of the sample node, obtaining the target scoring model through training with specific label scores, and scoring the target node through the introduced proprietary target scoring model, enabling it to quickly, stably, and accurately evaluate the quality of the target node, improving the efficiency of scoring and ensuring the accuracy of scoring, so as to efficiently and accurately obtain the search result of the target search task. Description of the Drawings

[0014] Figure 1 is a schematic diagram of the process of Monte Carlo tree search provided by an embodiment of the present specification; Figure 2 is a schematic diagram of single-node expansion in the optimization of Monte Carlo tree search provided by an embodiment of the present specification; Figure 3 is a flowchart of a search task processing method provided by an embodiment of the present specification; Figure 4 is another schematic diagram of single-node expansion in the optimization of Monte Carlo tree search provided by an embodiment of the present specification; Figure 5It is a flowchart of a method for training a target scoring model provided by an embodiment of this specification; Figure 6 It is a flowchart of a processing procedure of a search task processing method provided by an embodiment of this specification; Figure 7 It is a schematic structural diagram of a search task processing device provided by an embodiment of this specification; Figure 8 It is a schematic structural diagram of a training device for a target scoring model provided by an embodiment of this specification; Figure 9 It is a structural block diagram of a computing device provided by an embodiment of this specification; Figure 10 It is a structural block diagram of an electronic device provided by an embodiment of this specification. Detailed implementation manners

[0015] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0016] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any or all possible combinations of one or more of the associated listed items.

[0017] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".

[0018] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0019] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with over hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.

[0020] When a large model is actually applied, it only needs to be fine-tuned with a small number of samples for the pre-trained model to be applied to different tasks. Large models can be widely applied in the fields of natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0021] First, the noun terms involved in one or more embodiments of this specification are explained.

[0022] Monte Carlo Tree Search (MCTS): It is a heuristic tree search algorithm based on a random sampling strategy, widely used in decision-making processes and planning problems, such as game AI (Artificial Intelligence), path planning, etc. Its core idea is to gradually build an asymmetric search tree by repeatedly simulating and evaluating possible action paths and select a better strategy.

[0023] Monte Carlo Tree Search Refine (MCTS Refine): It represents the integration of Monte Carlo tree search and large models. It abstracts the iterative refinement process of the answers to solve problems into a search tree structure. Nodes on this tree represent different versions of the answers, and edges represent improvement attempts.

[0024] Reward Model: It is a key component in reinforcement learning, especially reinforcement learning from human feedback (RLHF), used to quantify the quality of generated content (such as text, code) and guide the model optimization direction. The reward model can score the quality of the output generated by the model (such as a piece of code, an answer) and convert it into an optimizable scalar reward signal.

[0025] Refine: Improve and optimize existing answers to improve their quality and accuracy.

[0026] MBPP (Measuring Big - Model Programming Proficiency) dataset: It is a dataset used to evaluate the programming ability of large models. It can evaluate the code generation ability of large models in multi - language programming scenarios. It can be composed of code snippets from multiple open - source code repositories, covering multiple programming languages, aiming to simulate programming tasks and scenarios in the real world. MBPP helps to more comprehensively evaluate the performance of the model on programming tasks with different languages and different difficulty levels by providing diverse programming problems and corresponding solutions (code), so as to promote the development and optimization of code - generation models in the field of multi - language programming.

[0027] It should be noted that with the rapid development of artificial intelligence technology, code - generation tools based on LLM play an increasingly important role in software development. However, when dealing with complex problems, it is often difficult for LLM to directly obtain satisfactory answers. Therefore, the method of MCTS is used to improve the reasoning ability of the model. Among them, the MCTS Refine method has achieved good results, that is, it uses MCTS to search and optimize the code generated by LLM.

[0028] In one implementation, the operation workflow of MCTS Refine follows the general pattern of the MCTS algorithm. MCTSRefine uses self - reflection - driven self - improvement to refine answers, samples the rewards of different versions of answers through self - reward ability, and realizes the iterative refinement of answers.

[0029] Figure 1 It is a schematic diagram of the process of a Monte Carlo tree search provided by an embodiment of this specification, as Figure 1As shown, Monte Carlo Tree Search is divided into four stages: Selection, Expansion, Evaluation, and Back-Propagation. In the Selection stage, starting from the root node (the current chessboard state), the most promising child node is recursively selected until an unexpanded node is encountered. In the Expansion stage, if the selected node has not been fully explored (i.e., there are untried actions), a new child node is expanded. In the Evaluation stage, starting from the newly expanded node, the game is randomly simulated until the end (a win, loss, or draw is determined). In the Back-Propagation stage, the simulation result (win / loss / reward) is propagated backward to all nodes on the path, updating their statistical information. The above steps are repeated multiple times (e.g., 1000 simulations), and finally, the child node with the most visits or the highest win rate is selected as the actual move.

[0030] Figure 2 It is a schematic diagram of single-node expansion in the optimization of Monte Carlo Tree Search provided by an embodiment of this specification. For any node S0 to be expanded, first, the LLM is called to score the node. Based on the scoring result, the LLM is called to evaluate the node. Then, the LLM is called to generate a new answer based on the evaluation result, which is used as the expanded node S1.

[0031] In the above process of optimizing Monte Carlo Tree Search, for each node expansion, the large model needs to be called three times (scoring, evaluation, generation), resulting in an overly long overall inference time, especially when dealing with complex problems, the time consumption increases significantly. Moreover, due to directly relying on the large model for scoring, the scoring results may have large fluctuations and inconsistencies, affecting the selection strategy of MCTS and the quality of the final answer. Furthermore, a fixed number of iterative refinement rounds is usually preset, regardless of the complexity of the problem, resulting in excessive time consumption for simple problems and insufficient optimization for complex problems.

[0032] Therefore, an embodiment of this specification provides a search task processing solution. By introducing a reward model (i.e., the target scoring model) to replace the LLM for scoring, the scoring time consumption is significantly reduced and the scoring stability is improved. By replacing the LLM evaluation process with a fixed text, one large model call is reduced for each expansion. Moreover, a specific loss formula can be designed to enable the reward model to set a threshold to adaptively adjust the search depth, achieving the goal of quickly answering simple problems and deeply optimizing complex problems. By introducing the reward model to optimize the MCTS Refine process, not only can the efficiency and accuracy of scoring be significantly improved, but also a threshold judgment mechanism can be added, realizing efficient and accurate processing of problems with different complexities, achieving one-pass for simple problems and in-depth thinking for complex problems.

[0033] To solve the above technical problems, in this specification, a method for processing a search task is provided. One or more embodiments of this specification are also related to a method for training a target scoring model, a search task processing apparatus, a target scoring model training apparatus, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.

[0034] See Figure 3 , Figure 3 which shows a flowchart of a method for processing a search task according to an embodiment of this specification, specifically including the following steps.

[0035] Step 302: Determine the target node of the target search task, where the target node is the answer that needs to be iteratively optimized.

[0036] Embodiments of this specification are applied to an application program, a website, or a mini-program with a search task processing function. On the application program, website, or mini-program, the search task processing function is implemented. For example, on a website where a target scoring model is deployed, the processing functions of code search optimization, text search optimization, etc. are implemented. Also, for example, a third-party application can call the deployed target scoring model through an Application Programming Interface (API) to implement the corresponding search task processing function.

[0037] It should be noted that the target search task refers to a task waiting for search optimization, such as a code search optimization task or a text search optimization task. The target search task can carry initial task data, which is the data that needs to be iteratively searched and optimized, such as initially generated code or initially generated text.

[0038] In actual implementation, the Monte Carlo tree search method can be used to take the initial task data as the initial node in the Monte Carlo tree, and continuously iteratively search to obtain optimized task data as the search result of the target search task.

[0039] For example, taking the target search task as a code search optimization task, assume the initial code is Code 1. Code 1 is expanded to obtain Code 2, and Code 2 is expanded to obtain Code 3. When Code 3 meets the stop condition, Code 3 is used as the search result of the target search task, that is, the initial Code 1 is iteratively optimized to the final Code 3.

[0040] In actual implementation, the target node of the target search task can be determined to facilitate subsequent iterative expansion of the target node. The target node is the node to be expanded currently, and the target node can be the answer that needs to be iteratively optimized. For example, the target node can be answers such as code, text, etc. Initially, the initial task data carried by the target search task can be used as the target node. Subsequently, based on the upper confidence bound (UCB) of each child node of the current node, the target node to be expanded currently can be selected from each child node to balance exploration (attempting actions that have not been fully evaluated) and exploitation (selecting the current optimal action). Herein, each child node of the current node refers to the node expanded based on the current node.

[0041] Specifically, the UCB value of each child node of the current node can be calculated through the following formula (1), and the child node with the largest UCB value is selected as the target node to be expanded currently: (1) where refers to the cumulative reward of the node, refers to the number of times the node is visited, refers to the number of times the parent node is visited, refers to the exploration coefficient.

[0042] Exemplarily, taking the code search optimization task as the target search task as an example, assuming the initial code is Code 1, Code 1 is used as the target node for expansion to obtain Code 2 and Code 3. Calculate the UCB values of Code 2 and Code 3, and select Code 2 with a higher UCB value as the target node, and then expand it, and so on until it stops.

[0043] In an optional implementation manner of this embodiment, before determining the target node of the target search task, it further includes: For the target generation task of the target scenario, obtain the target generation content corresponding to the target generation task through the target generation model; Create a corresponding target search task based on the target generation content, where the target search task is used to update the target generation content.

[0044] Herein, the target scenario refers to the scenario for content generation, such as the code generation scenario, the answer text generation scenario, etc.; the target generation model is a generative large model that can generate corresponding answers for the input question. For example, in the code generation scenario, for the input question, the target generation model can generate the corresponding code; in the answer text generation scenario, for the input question, the target generation model can generate the corresponding answer text.

[0045] In actual implementation, the task problem carried by the target generation task can be input into the target generation model, and the target generation model can output the corresponding target generation content. Subsequently, a corresponding target search task can be created based on the target generation content to iteratively search for the target generation content generated by the target generation model, achieving further optimization.

[0046] For example, in the code generation scenario, the task problem carried by the target generation task is "Please generate a piece of code to sort multiple data from largest to smallest". The target generation model can generate Code 1, which can sort multiple data from largest to smallest. At this time, Code 1 can be used as the initial task data waiting to be searched and optimized, and a corresponding target search task can be created. Thus, through the Monte Carlo tree search method, Code 1 can be iteratively searched to obtain the optimized code. That is to say, the target generation content generated by the target generation model is the initial task data, and subsequent iterative optimization is performed.

[0047] In the embodiments of this specification, after obtaining the target generation content corresponding to the target generation task through the target generation model, a corresponding target search task can be created based on the target generation content, so as to iteratively search for the content generated by the target generation model, realizing the update and optimization of the target generation content of the large model, and improving the reliability of the model generation result.

[0048] In an optional manner, the Monte Carlo tree search method can be used to iteratively search for the content generated by the target generation model, and the Monte Carlo tree search method is used to update and optimize the target generation content of the large model, improving the reliability of the model generation result.

[0049] Step 304: Determine the scoring result of the target node through the target scoring model, where the target scoring model is trained based on at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node.

[0050] Specifically, the target scoring model is a reward model pre-trained based on at least one sample node and the corresponding label score, and is used to score the quality of the target node. The label score is determined based on the correctness reward and / or execution efficiency reward of the sample node. The correctness reward is used to evaluate whether the content of the sample node is correct, and the execution efficiency reward is used to evaluate the high or low processing efficiency of the sample node. For example, the execution efficiency reward can include a time reward and / or a content reward. The time reward refers to the processing time reward of the sample node for the test case, and the content reward refers to the content reward of the sample node. For example, the shorter the processing time, the higher the execution efficiency and the higher the time reward score. The shorter the content length of the sample node, the more concise the sample node and the more efficient the execution, and the higher the content reward score.

[0051] In actual implementation, MCTS Refine directly relies on a large model for scoring, and the scoring results may have significant fluctuations and inconsistencies, affecting the selection strategy of MCTS and the quality of the final solution. By designing a scientific reward formula, it can stably and accurately evaluate the quality of answers, and the reward model outputs the specific scores of each answer instead of just performing relative ranking to ensure the reliability and consistency of scoring.

[0052] In the implementation of this specification, by introducing a target scoring model to replace the LLM for node scoring, the scoring time is significantly reduced and the scoring stability is improved. Moreover, by comprehensively considering the correctness reward and / or execution efficiency reward of sample nodes, the label score of sample nodes is determined as the sample label, and the target scoring model is trained to improve the accuracy and applicability of scoring.

[0053] It should be noted that the training process of the target scoring model can refer to the content shown below Figure 5 and will not be elaborated in this embodiment of this specification.

[0054] Step 306: Expand nodes based on the scoring results of target nodes to determine the search results of the target search task.

[0055] In actual implementation, based on the scoring results of target nodes, the parts with poor quality in the target nodes can be optimized to obtain expanded nodes. Based on the current nodes, target nodes are continuously selected, and then the process of scoring and expanding the target nodes is returned until the stop condition is met, and the finally iteratively searched nodes are obtained as the search results of the target search task.

[0056] In one implementation, the scoring results and target nodes can be input into a generative large model to guide the generative large model to optimize the target nodes according to the scoring results and output the optimized expanded nodes, and the process of continuously selecting target nodes, scoring, and expanding is continued. In another implementation, it can be determined whether to continue expanding based on the scoring results. If expansion is needed, the optimized expanded nodes are output through the generative large model; if no further expansion is required, the search can be stopped to obtain the search results of the target search task.

[0057] In an optional implementation of this embodiment, expanding nodes based on the scoring results of target nodes to determine the search results of the target search task includes: Determine whether the scoring results of the target nodes meet the stop search constraint; If the stop search constraint is met, the target node is used as the search result of the target search task; If the stop search constraint is not satisfied, expand the target node to obtain an expanded node, and determine the search result of the target search task based on each current node of the target search task, where each current node of the target search task includes the target node and the expanded node.

[0058] The stop search constraint is a constraint condition configured based on the scoring result of the target node and is used to control the timing of stopping the search.

[0059] In actual implementation, if the scoring result of the target node satisfies the stop search constraint, there is no need to perform iterative search anymore, and the target node is directly used as the search result of the target search task. If the scoring result of the target node does not satisfy the stop search constraint, further expansion is required. At this time, the target node can be expanded to obtain an expanded node, and then the search optimization is iteratively performed based on each current node of the target search task to determine the search result of the target search task.

[0060] It should be noted that for a simple target search task, a high-quality target node can be obtained through fewer iteration times, such that the score of the target node satisfies the stop search constraint; for a complex target search task, more iteration times are required to obtain a high-quality target node, such that the score of the target node satisfies the stop search constraint. That is to say, the simple target search task has fewer search times, and the complex target search task has more search times. In other words, by introducing the verification of whether the scoring result of the target node satisfies the stop search constraint, the search times of target search tasks with different complexities can be controlled, avoiding using a fixed number of search times for target search tasks with different complexities.

[0061] In the embodiments of this specification, by introducing the verification of whether the scoring result of the target node satisfies the stop search constraint, it is realized that the search depths of target search tasks with different complexities are different. A simple target search task may generate high-quality results and stop searching after one or two searches. Only complex tasks will search for a higher number of times, improving the reasoning efficiency and having better adaptability.

[0062] In an optional implementation manner of this embodiment, determining whether the scoring result of the target node satisfies the stop search constraint includes: When the scoring result of the target node is greater than the set scoring threshold, it is determined that the stop search constraint is satisfied; When the scoring result of the target node is less than or equal to the set scoring threshold, it is determined that the stop search constraint is not satisfied.

[0063] It should be noted that the stop search constraint may be that the scoring result of the target node is greater than the set scoring threshold. The set scoring threshold is a preset value and is used to determine whether the quality of the target node is already relatively good and can meet the requirements.

[0064] In actual implementation, if the scoring result of the target node is greater than the set scoring threshold, it indicates that the quality of the target node is already relatively good, meeting the stop search constraint. Then, stop the search and use the target node as the search result of the target search task. If the scoring result of the target node is less than or equal to the set scoring threshold, it indicates that the quality of the target node is still poor and does not meet the stop search constraint. In this case, it is necessary to continue iterative search optimization until the search result of the target search task is obtained.

[0065] In the embodiments of this specification, regardless of the complexity of the problem, the MCTS Refine method can preset a fixed number of iteration rounds, resulting in excessive time consumption for simple problems and insufficient optimization for complex problems. By adding a threshold judgment mechanism and based on the score output by the target scoring model, when the score reaches the preset threshold, the search is immediately terminated, adapting to the complexity of the problem, enabling the model to distinguish between simple and complex problems, achieving one-pass for simple problems and in-depth search for complex problems.

[0066] In an optional implementation manner of this embodiment, the target node is the target generated content; expanding the target node to obtain an expanded node includes: Based on the set evaluation prompt words, through the target generation model, generate the expanded content corresponding to the target generated content, and use the expanded content as the expanded node.

[0067] Among them, the set evaluation prompt words are a pre-configured fixed text used to prompt the target generation model to generate the corresponding expanded content, which is the optimized content. For example, the set evaluation prompt words can be "Wait a moment, there seems to be some error here".

[0068] In actual implementation, input the target generated content and the set evaluation prompt words into the target generation model. The target generation model modifies the input target generated content and outputs the corresponding expanded content as the expanded node of the target node.

[0069] It should be noted that during the processing of the target generation model, the target generation model will conduct a detailed analysis and modification of the target generated content based on the guidance of the set evaluation prompt words. For example, if the set evaluation prompt words indicate an error, the target generation model can try to identify and correct the error part in the target generated content. Under the indication of the set evaluation prompt words, the target generation model can generate more accurate and optimized expanded content without calling the large model to provide specific evaluation details, thus effectively improving the processing effect of the search task.

[0070] In the embodiments of this specification, by replacing the LLM evaluation process with the use of fixed text, one large model call is reduced each time an expansion is performed, further improving the inference efficiency of the search task.

[0071] Figure 4 It is another schematic diagram of single - node expansion in Monte Carlo tree search optimization provided by an embodiment of this specification. As Figure 4 shown, for the target node S0, the LLM model can be called for scoring. The scoring result includes a score and evaluation content. Based on the evaluation content, the LLM is called again for generation to obtain the expanded node S1. This process still requires calling the large model 2 times, and the efficiency of extended reasoning is relatively low.

[0072] Taking the target scoring model as the reward model as an example, for the target node S0, the RM (reward model) can be called for scoring. The scoring result includes a score. Based on this score, the LLM is called for evaluation to obtain the evaluation content. Based on the evaluation content, the LLM is called again for generation to obtain the expanded node S1. This process still requires calling the large model 2 times, and the efficiency of extended reasoning is relatively low.

[0073] For the target node S0, the RM (reward model) can be called for scoring. The scoring result includes a score. Directly based on the set evaluation prompt words, the LLM is called for generation to obtain the expanded node S1. This process only needs to call the large model 1 time, which improves the efficiency of extended reasoning. And a threshold can also be set. If the score of the target node S0 reaches the threshold, the search stops.

[0074] In an optional implementation manner of this embodiment, based on each current node of the target search task, the search result of the target search task is determined, including: Determine the comprehensive scores of each node respectively. Based on the comprehensive scores, determine the updated nodes to be expanded from each node. Take the updated nodes to be expanded as the updated target nodes, and return to execute the step of determining the scoring result of the target node through the target scoring model until the search result of the target search task is obtained.

[0075] Among them, the comprehensive score of each node refers to the score calculated by comprehensively considering factors such as the historical scores and the number of visits of each node, which can indicate the potential value of each node. For example, this comprehensive score can be the UCB value of each node.

[0076] In actual implementation, the UCB value of each node can be calculated, and the node with the highest UCB value is selected as the target node to continue iterative expansion. Return to iteratively execute the scoring and expansion process until the search result of the target search task is obtained.

[0077] It should be noted that by comprehensively considering the comprehensive scores of each node, it is ensured that nodes with higher potential value are selected as the target nodes for the next round of iteration, and continue to iteratively perform scoring and expansion until the final search result is obtained, which improves the accuracy of node expansion and search.

[0078] An embodiment of this specification provides a search task processing method, which realizes introducing the correctness reward and / or execution efficiency reward of sample nodes, determining the label scores of sample nodes, obtaining a target scoring model through training with specific label scores, and scoring target nodes through introducing a proprietary target scoring model, enabling it to quickly, stably, and accurately evaluate the quality of target nodes, improving the efficiency of scoring and ensuring the accuracy of scoring, so as to efficiently and accurately obtain the search results of target search tasks.

[0079] See Figure 5 , Figure 5 shows a flowchart of a method for training a target scoring model provided by an embodiment of this specification, which specifically includes the following steps.

[0080] Step 502: Obtain a training sample set under a target search task, where the training sample set includes at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node.

[0081] Among them, the training sample set under the target search task includes at least one sample node and the corresponding label score for each sample node. The target search task refers to a task waiting for search optimization. The sample node refers to multiple different versions of content, carrying the corresponding label score. For example, in a code search optimization task, the sample node is multiple different versions of code, and in a text search optimization task, the sample node is multiple different versions of content. The label score is determined based on the correctness reward or execution efficiency reward of the sample node, and is used as a sample label to train the target scoring model.

[0082] In actual implementation, the training sample set can be obtained from a database or other third-party application platforms, or the training sample set can also be constructed based on the target search task.

[0083] In an optional implementation manner of this embodiment, the sample node is a sample answer; obtaining the training sample set under the target search task includes: Obtain test cases under the target search task; Generate at least one sample answer corresponding to each sample question under the target search task, test each sample answer based on the test cases, and obtain the test information of each test case and each sample answer; Determine the label score of each sample answer based on the correctness reward or execution efficiency reward of each sample answer, where the correctness reward or execution efficiency reward of the sample answer is determined based on the test information of the sample answer; Obtain the training sample set under the target search task based on each sample answer and the corresponding label score.

[0084] Specifically, a test case refers to data used to test the quality of a sample answer. For example, if the sample answer is a sorting code, the test case can be a set of values {5, 4, 8, 1, 2, 6, 9}. Inputting the test case into the sample answer, the sample answer can sort the test case and output the sorting result. Through the test case, it can be tested whether the sample answer can sort the test case and whether the sorting result is correct, etc., thereby reflecting the quality of the sample answer.

[0085] In actual implementation, the test case can be obtained from a database or other third - party platforms, or can also be constructed using other large models or generation tools.

[0086] It should be noted that at least one sample answer corresponding to each sample question under the target search task can be generated. The sample answer can be generated using other large models or generation tools, or directly obtained from a specific database. For example, taking the code generation scenario as an example, the MBPP dataset can be used as a basis to obtain the sample questions therein. Using an open - source large model to generate a batch of codes for each sample question in the MBPP dataset and removing duplicates, at least one sample answer corresponding to each sample question can be obtained, thereby ensuring the richness of the sample answers and further ensuring the model training effect.

[0087] After generating the sample answers, each test case can be used to test each sample answer to obtain the test information of each test case and each sample answer, so as to determine the test case of the sample answer based on the test information to test each sample answer, obtain the test information of each test case and each sample answer, and further determine the label score of each sample answer. Among them, the test information refers to the relevant information of the sample answer testing the test case, such as time, whether it passes, the length of the test case, etc.

[0088] In the embodiments of this specification, the test cases under the target search task can be obtained, sample answers can be generated, each sample answer can be tested based on the test cases, and the test information of each test case and each sample answer can be obtained. Thus, based on the test information, the correctness reward or execution efficiency reward of each sample answer can be determined, and then the label score of each sample answer can be determined, realizing the automatic annotation of the label score of each sample answer as a sample label, automatically constructing a training sample set with label scores, that is, constructing a dedicated dataset to train the target scoring model, improving the training efficiency and training effect of the target scoring model.

[0089] In an optional implementation manner of this embodiment, obtaining the test cases under the target search task includes: Obtaining at least one sample question under the target search task, and the reference answers of each sample question; Generating at least one set of test cases corresponding to each sample question; The test samples are verified based on the reference answers to the sample questions to obtain the verified test samples.

[0090] It should be noted that in addition to generating sample answers to sample questions as training data, reference answers to sample questions can also be obtained. The reference answers can be high-quality answers that are used to evaluate the quality of test samples, thereby screening out high-quality test samples whose sample answers can be successfully identified and processed, thereby ensuring the quality of the test samples.

[0091] In actual implementation, taking the code generation scenario as an example, the MBPP dataset can be used as a basis to obtain sample questions and corresponding codes as reference answers. Then, multiple open source big models are used to generate multiple groups of test samples for each sample question. The richness of the generated test samples can be improved by using multiple open source big models. Then, the generated test samples are verified using the MBPP built-in code (i.e., the reference answer), and the passed test samples are removed and retained to obtain the test sample set.

[0092] In specific implementation, a passed test sample refers to a test sample for which the reference answer can successfully output a result, and is not necessarily a correct processing result. If there are multiple reference codes for the same sample problem, the test results of the majority of reference answers in the multiple reference codes can be used to determine whether the test sample has passed the test.

[0093] As an example, the sample question is "Please generate code that can sort multiple numerical values", and the reference answer is code A. Code A can sort multiple input numerical values. The multiple groups of test samples generated for this sample question using a variety of open source large models are test sample 1 {5, 4, 8, 1, 2, 6, 9}, test sample 2 {2, 1, 3, 4, 7, 5, 6}, test sample 3 {5, 4, a, 1, 2, 6, 9}, test sample 4 {5, 4, 8, 1, 2, b, 9}, and test sample 5 {5, 4, 8, 1, 2, 6, 9}. Test samples 1-5 are input to the code A respectively. For test sample 1, code A outputs {1, 2, 3, 4, 5, 7, 6} for test sample 2, code A outputs "an error occurred" for test sample 3, code A outputs "an error occurred" for test sample 4, and code A outputs {1, 2, 4, 6, 5, 8, 9} for test sample 5. As can be seen from the above, test samples 1, 2, and 5 all output results normally. Even if the sorting results of test samples 2 and 5 are wrong, they can output results normally. Test samples 1, 2, and 5 all pass the test. In addition, since test sample 5 is repeated with test sample 1, the final test samples obtained after deduplication are test sample 1 and test sample 2.

[0094] In the embodiments of this specification, the reference answers to sample questions can be obtained to test the generated test cases, deduplicate the test cases that pass the test, and obtain the final test cases for testing the sample answers, thereby determining the quality of the sample answers, screening out the incorrect sample data in the generated test cases, ensuring that the retained test cases are correct test cases that can enable the sample answers to be executed smoothly, and ensuring the accuracy of subsequent testing of the sample answers.

[0095] In the embodiments of this specification, the target scoring mode can be a reward model. By constructing a dedicated dataset to train the reward model in the above manner, the training efficiency and training effect of the reward model are improved.

[0096] In an optional implementation manner of this embodiment, based on the correctness reward or execution efficiency reward of each sample answer, the label score of each sample answer is determined, including: Based on the test information, it is determined whether there is a test case that fails to pass the test for the first sample answer, where the first sample answer is any one of the sample answers; If so, based on the test information, the correctness reward of the first sample answer is determined, and the correctness reward of the first sample answer is mapped to the first threshold range to obtain the label score of the first sample answer; If not, based on the test information, the execution efficiency reward of the first sample answer is determined, and the execution efficiency reward of the first sample answer is mapped to the second threshold range to obtain the label score of the first sample answer.

[0097] Among them, the first sample answer is any one of the sample answers. For each sample answer, the above judgment is made to determine whether there is a test case that fails to pass the test, so as to calculate the corresponding label score in different ways.

[0098] Specifically, the correctness reward is used to evaluate whether the content of the sample answer is correct, and the execution efficiency reward is used to evaluate the high or low processing efficiency of the sample answer. The first threshold range is a numerically limited range of the label score configured in advance in the case of having a test case that fails to pass the test; the second threshold range is a numerically limited range of the label score configured in advance in the case of having no test case that fails to pass the test. For example, the first threshold range can be , and the second threshold range can be , where λ can be configured based on actual requirements.

[0099] It should be noted that when the sample answer fails to pass all the test cases, it means that for some test cases, the sample answer cannot successfully output the result. In this case, only the correctness reward is considered, and the score is mapped to the first threshold range. When the sample answer passes all the test cases, it means that for each test case, the sample answer can successfully output the result. In this case, the execution efficiency reward of the sample answer is comprehensively considered. For example, the execution efficiency reward can include time reward and / or content reward, and the score is mapped to the second threshold range.

[0100] In the embodiments of this specification, by designing a proprietary reward formula, it can stably and accurately evaluate the answer quality, so that the obtained target scoring model can output the specific scores of each sample answer, rather than just performing relative ranking, ensuring the reliability and consistency of the scores output by the target scoring model.

[0101] In an optional implementation manner of this embodiment, determining the correctness reward of the first sample answer based on the test information includes: Obtaining the test results of whether the first sample answer passes the test on each test case; Determining the example scores of each test case, where the example scores of the test case are determined based on the answer scores of the passed sample answers, the total scores of each sample answer, and the test results; Determining the correctness reward of the first sample answer based on the example scores of the test cases passed by the first sample answer, the total scores of each test case, and the test results.

[0102] In actual implementation, the test results of whether each sample answer passes the test on each test case are a [0,1] matrix. The (i,j)-th element in this matrix represents whether the i-th sample answer passes the j-th test case. If the (i,j)-th element is 0, it means that the i-th sample answer fails the j-th test case; if the (i,j)-th element is 1, it means that the i-th sample answer passes the j-th test case.

[0103] Specifically, the correctness reward of the first sample answer can be calculated through the following formulas (2)-(4): (2) (3) (4) Among them, represents the code score, represents the test case score, M is a 0,1 matrix of whether the code passes the test case, and are damping factors. The formula means that for a code, the more correct test cases it has, the higher the correct example score, and the code score The higher; for a test set, the more test error codes, the higher the error code score, and the higher the test set score The higher.

[0104] In specific implementation, both the code score and the test case score are initialized to 1, and then the above formulas (3) and (4) are continuously iterated until the code score and the test case score are stable. A fixed number of iteration rounds can be set, such as 100 times.

[0105] In addition, M is a 0, 1 matrix indicating whether the code passes the test cases. M' is used to calculate the test case score of the above formula (4). For the code, passing will increase the score, but for the test case, passing the code will decrease the score. Therefore, the elements "0" and "1" in the 0, 1 matrix indicating whether the code passes the test cases need to be reversed, that is, the above formula (2).

[0106] It should be noted that when the sample answer does not pass all the test cases, only the correctness reward is considered, and the score is mapped to the first threshold range, such as , the label score can be calculated by the following formula (5): (5) Among them, the above "λ" can be flexibly configured according to actual needs.

[0107] In the embodiments of this specification, when the sample answer does not pass all the test cases, only the correctness reward is considered. By comprehensively considering the test results of whether the first sample answer passes the tests on each test case and the example scores of each test case, the correctness reward is comprehensively determined, and a proprietary correctness reward formula is designed, which improves the accuracy of the label score of the labeled sample answer and ensures the model training effect.

[0108] In an optional implementation manner of this embodiment, the execution efficiency reward includes a time reward and / or a content reward; determining the execution efficiency reward of the first sample answer based on the test information and mapping the execution efficiency reward of the first sample answer to the second threshold range to obtain the label score of the first sample answer includes: Based on the test information, determining the maximum execution time and the minimum execution time of the first sample answer on each test case, and determining the time reward score of the first sample answer based on the maximum execution time and the minimum execution time; Based on the test information, determining the maximum content length and the minimum content length of each sample answer, and determining the content reward score of the first sample answer; Based on the time reward score and / or the content reward score, and combining the threshold of the second mapping range, determining the label score of the first sample answer.

[0109] In actual implementation, the test information may include execution time, content length, etc. When the sample answer passes all test cases, an execution efficiency reward can be considered. This execution efficiency reward can include a time reward and / or a content reward, that is, consider the execution time of the sample answer passing the test samples, the content length of the sample answer, etc. to analyze the time reward / content reward of the sample answer, map the time reward score / content reward score to the second threshold range, and obtain the label score of the sample answer.

[0110] Specifically, the time reward score (i.e., the time reward) is calculated by the following formula (6): (6) where t max refers to the maximum execution time among the execution times of each test case for the first sample answer; tmin refers to the minimum execution time among the execution times of each test case for the first sample answer, and t refers to the current time.

[0111] The content reward score (i.e., the length reward) is calculated by the following formula (7): (7) where I max refers to the longest content length among each sample answer, I min refers to the shortest content length among each sample answer, and I refers to the current content length. For example, taking the sample answer as code, this content length can refer to the number of tokens of the code, or it can be the number of words of the code task. For example, the content length of "import pandas" is two words.

[0112] It should be noted that when the sample answer passes all test cases, comprehensively consider the time reward and / or content reward of the code, map the score to the second threshold range, such as , and calculate the label score by the following formula (8): (8) where the above "λ" can be flexibly configured according to actual needs. refers to the time reward score, refers to the content reward score.

[0113] In the embodiments of this specification, when the sample answer passes all test cases, the time reward and content reward can be comprehensively considered. Based on the maximum execution time and minimum execution time of the sample answer on each test case, the time reward score is determined. Based on the maximum content length and minimum content length of each sample answer, the content reward score is comprehensively determined, and a proprietary execution efficiency reward formula is designed, which improves the accuracy of the label score of the annotated sample answer and ensures the model training effect.

[0114] Step 504: Based on the sample nodes in the training sample set, obtain the corresponding predicted scores through the initial scoring model.

[0115] In actual implementation, the first sample node in the training sample set can be input into the initial scoring model to obtain the predicted score corresponding to the first sample node. This predicted score is the scoring result output by the initial scoring model after analyzing the sample node, and this scoring result is the model prediction value, which can be used for subsequent calculation of the model loss and further training of the initial scoring model.

[0116] In an optional implementation manner of this embodiment, the method further includes: Obtain a candidate generation model; Replace the language model layer in the candidate generation model with a linear layer and a mapping layer to obtain the initial scoring model.

[0117] It should be noted that any generative model can be used to construct the initial scoring model, such as the Qwen2.5-Coder-7B-Instruct model.

[0118] Among them, Qwen2.5-Coder-7B-Instruct is an open-source large language model for code, belonging to the Qwen2.5-Coder series, with 7 billion parameters. There are also other models with different parameter scales in this series, which can meet the needs of different users. The training data scale is large, and it performs well in code generation, reasoning, and repair. The training tokens are extended to 5.5 trillion, and the training data includes source code, text code base, synthetic data, etc.; the coding ability is strong, supporting long contexts, which is beneficial to processing large code libraries.

[0119] Specifically, since the generative model needs to generate corresponding content, the model needs to include a language model layer (lm_head) for content generation. In the embodiment of this specification, the initial scoring model does not need to generate content and only needs to output scores. Therefore, the language model layer of the generative model can be removed. And since the initial scoring model needs to output scores, two linear layers and a mapping layer (i.e., the Tanh activation function) can be connected to obtain the initial scoring model. Through the linear layer, relevant features can be refined, and through the mapping layer, the result can be mapped to a specific numerical range (such as mapped to 0-1), transforming the generative model from a generation task (such as code completion) into a scoring task, so that by training the initial scoring model, a target scoring model capable of scoring nodes can be obtained.

[0120] In the embodiment of this specification, any generative model can be used to construct the initial scoring model, without building a completely new scoring model, which improves the model building efficiency.

[0121] Step 506: Determine the target loss of the initial scoring model based on the label score and prediction score of the sample node, and train the initial scoring model based on the target loss to obtain the trained target scoring model, where the target scoring model is used to score the target node in the above search task processing method.

[0122] In actual implementation, the model parameters of the initial scoring model can be adjusted backward based on the target loss. When the training stop condition is not reached, return to continue training the model based on the next sample answer until the training stop condition is reached to obtain the trained target scoring model.

[0123] Among them, the label score refers to the result that the initial scoring model actually wants to output, that is, the label score is the real result. When the sample answer is input into the initial scoring model, the output prediction score is the prediction result. When the difference between the prediction result and the real result is small enough, it means that the prediction result is close enough to the real result. At this time, the initial scoring model is trained and it is determined that the training stop condition is met to obtain the target scoring model.

[0124] In addition, in addition to judging whether the training stop condition is reached based on the loss value, the iteration times can also be combined to determine whether the training stop condition is reached. Specifically, if the target loss is greater than the loss value threshold, it can be further judged whether the current iteration times reach the preset iteration times. If the current iteration times do not reach the preset iteration times, it can be determined that the training stop condition is not reached, and the model parameters of the initial scoring model can be continuously adjusted, and return to continue training the model based on the next sample answer until the preset iteration times are reached, and it is determined that the training stop condition is reached, stop the iteration, and obtain the trained target scoring model.

[0125] Among them, the loss value threshold and the preset iteration times are set according to the actual situation, and the embodiments of the present application do not make any limitations on this. When the number of training times reaches the preset iteration times, it means that the number of training times of the initial scoring model is sufficient. At this time, the prediction result of the initial scoring model is extremely close to the real result, and the training can be stopped.

[0126] It should be noted that through the target loss, the difference between the prediction result and the real result of the model can be intuitively shown. Then, targeted training is carried out on the initial scoring model and the parameters are adjusted, which can effectively improve the training rate and training effect of the model.

[0127] In an optional implementation manner of this embodiment, determining the target loss of the initial scoring model includes: Determine the recall loss of the sample node based on the label score and prediction score of the sample node; Determine the target loss of the initial scoring model based on the recall loss and the set threshold penalty factor.

[0128] It should be noted that, since it is necessary to determine the specific scores of sample nodes, the ranking loss commonly used in the reward model is discarded and the regression loss is adopted. Moreover, in order to ensure the reliability of the threshold, the threshold penalty factor can be further increased. The objective loss of the initial scoring model is specifically determined by the following formula (9): (9) where is the predicted score, is the label score, is the penalty coefficient, λ refers to the threshold pre-configured based on actual requirements, is the threshold penalty factor, ensuring that the model can be appropriately penalized when the predicted score exceeds the threshold.

[0129] In the embodiments of this specification, by designing a scientific reward formula and a dedicated data set, designing a specific loss formula, and training the initial scoring model, the reward model can adaptively adjust the search depth based on the set threshold, achieving the goal of quickly answering simple questions and deeply optimizing complex questions, and improving the inference efficiency and accuracy.

[0130] The following, in combination with the attached Figure 6 , taking the application of the search task processing method provided in this specification in the code generation scenario as an example, further illustrates the search task processing method. Among them, Figure 6 shows the processing procedure flowchart of a search task processing method provided by an embodiment of this specification, specifically including the following steps.

[0131] The task problem carried by the target generation task is "Please generate a piece of code to sort multiple data from largest to smallest". The target generation model can generate Code 1, and this Code 1 can sort multiple data from largest to smallest.

[0132] Call the reward model to score Code 1.

[0133] If the score of Code 1 exceeds the set threshold, then Code 1 is used as the final search result.

[0134] If the score of Code 1 does not exceed the set threshold, then based on the set evaluation prompt words (fixed text), through the target generation model, generate Code 2 corresponding to Code 1, and Code 2 optimizes some errors in Code 1. Calculate the UCB values of Code 1 and Code 2, select the code with the higher UCB value among Code 1 and Code 2 as the target node, and then iterate for scoring and expansion, and so on until it stops.

[0135] Next, an experimental test is carried out on the search task processing method of the embodiments of this specification, and the experimental results are as follows: Test set: HumanEval Test model: Qwen2.5-Coder-32B-Instruct The accuracy verification of the target scoring model (i.e., the reward model) is as follows: Experimental setup: Use Qwen2.5-Coder-32B-Instruct to sample 5 answers for each question in HumanEval, and then sort them using different models. The results are shown in Table 1 and Table 2 below. Among them, qwen2.5-coder-7b-instruct-Generate-Train means generative training with the same data and labels of yes / no, and then taking the probability of yes as the score. qwen2.5-coder-7b-instruct-RM is the reward model trained using the embodiments of this specification.

[0136] Table 1 Effect table of Qwen2.5-Coder-32B-Instruct generating code

[0137] Table 2 Correct rate of the final answer after each model sorts each answer and scores

[0138] The performance and overall accuracy analysis of the target scoring model (i.e., the reward model) are as follows: Experimental setup: For the Qwen2.5-Coder-32B-Instruct model, the MCTS Refine method sets the search depth to 5, RM+MCTS Refine sets the maximum depth to 5, and the score threshold to 0.3. Applying the solution of the embodiments of this specification, 85% of the scenarios terminate the search in advance, and the termination accuracy rate reaches 95.7%. The overall accuracy rate is 93.2%, an increase of 3.6%. The time consumption is reduced by 10 times. The average number of calls is optimized from 14 LLM calls to 1.64 RM calls plus 1.64 LLM calls, and the average time consumption per question is optimized from 71.1s to 6.4s.

[0139] Table 3 Performance and accuracy table

[0140] Corresponding to the above method embodiments, this specification also provides embodiments of a search task processing device, Figure 7 showing a schematic structural diagram of a search task processing device provided by an embodiment of this specification. As Figure 7 shown, the device includes: The first determination module 702 is configured to determine the target node of the target search task, where the target node is the answer that needs to be iteratively optimized; A scoring module 704, configured to determine a scoring result of a target node through a target scoring model, where the target scoring model is trained based on at least one sample node and a corresponding label score, and the label score is determined based on a correctness reward or an execution efficiency reward of the sample node; A second determination module 706, configured to perform node expansion based on the scoring result of the target node to determine a search result of the target search task.

[0141] Optionally, the second determination module 706 is further configured to: Determine whether the scoring result of the target node satisfies a stop search constraint; If the stop search constraint is satisfied, use the target node as the search result of the target search task; If the stop search constraint is not satisfied, expand the target node to obtain an expanded node, and determine the search result of the target search task based on each current node of the target search task, where each current node of the target search task includes the target node and the expanded node.

[0142] Optionally, the target node is target generated content; the second determination module 706 is further configured to: Generate extended content corresponding to the target generated content through a target generation model based on a set evaluation prompt, and use the extended content as the expanded node.

[0143] Optionally, the second determination module 706 is further configured to: Determine the comprehensive score of each node respectively, determine the node to be expanded from each node based on the comprehensive score, use the node to be expanded as the updated target node, and return to execute the step of determining the scoring result of the target node through the target scoring model until the search result of the target search task is obtained.

[0144] Optionally, the second determination module 706 is further configured to: Determine that the stop search constraint is satisfied when the scoring result of the target node is greater than a set scoring threshold; Determine that the stop search constraint is not satisfied when the scoring result of the target node is less than or equal to the set scoring threshold.

[0145] Optionally, the apparatus further includes a creation module, configured to: For a target generation task of a target scenario, obtain target generated content corresponding to the target generation task through a target generation model; Create a corresponding target search task based on the target generated content, where the target search task is used to update the target generated content.

[0146] An embodiment of this specification provides a search task processing device, which realizes introducing the correctness reward and / or execution efficiency reward of sample nodes, determining the label score of sample nodes, obtaining a target scoring model through training with specific label scores, and scoring target nodes through introducing a proprietary target scoring model, so that it can quickly, stably and accurately evaluate the quality of target nodes, improve the scoring efficiency and ensure the scoring accuracy, thereby efficiently and accurately obtaining the search results of the target search task.

[0147] The above is a schematic solution of a search task processing device according to this embodiment. It should be noted that the technical solution of this search task processing device and the technical solution of the above search task processing method belong to the same concept. For the details not described in detail in the technical solution of the search task processing device, reference can be made to the description of the technical solution of the above search task processing method.

[0148] Corresponding to the above method embodiment, this specification also provides an embodiment of a training device for a target scoring model. Figure 8 The structure diagram of a training device for a target scoring model provided by an embodiment of this specification is shown. As Figure 8 shown, the device includes: An acquisition module 802, configured to acquire a training sample set under a target search task, where the training sample set includes at least one sample node and a corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; A prediction module 804, configured to obtain a corresponding predicted score through an initial scoring model based on the sample nodes in the training sample set; A training module 806, configured to determine the target loss of the initial scoring model based on the label score and predicted score of the sample node, and train the initial scoring model based on the target loss to obtain a trained target scoring model, where the target scoring model is used to score target nodes in the above search task processing method.

[0149] Optionally, the training module 806 is further configured to: Determine the recall loss of the sample node based on the label score and predicted score of the sample node; Determine the target loss of the initial scoring model based on the recall loss and a set threshold penalty factor.

[0150] Optionally, the sample node is a sample answer; the acquisition module 802 is further configured to: Acquire test examples under the target search task; Generate at least one sample answer corresponding to each sample question under the target search task, test each sample answer based on the test examples, and obtain the test information of each test example and each sample answer; Determine the label score of each sample answer based on the correctness reward or execution efficiency reward of each sample answer, where the correctness reward or execution efficiency reward of the sample answer is determined based on the test information of the sample answer; Obtain the training sample set under the target search task based on each sample answer and the corresponding label score.

[0151] Optionally, the obtaining module 802 is further configured to: Obtain at least one sample question under the target search task and the reference answer of each sample question; Generate at least one set of test examples corresponding to each sample question; Verify the test examples based on the reference answers of each sample question, and obtain the test examples that pass the verification.

[0152] Optionally, the obtaining module 802 is further configured to: Determine whether there is a test example for which the first sample answer fails the test based on the test information, where the first sample answer is any one of each sample answer; If so, determine the correctness reward of the first sample answer based on the test information, and map the correctness reward of the first sample answer to the first threshold range to obtain the label score of the first sample answer; If not, determine the execution efficiency reward of the first sample answer based on the test information, and map the execution efficiency reward of the first sample answer to the second threshold range to obtain the label score of the first sample answer.

[0153] Optionally, the obtaining module 802 is further configured to: Obtain the test results of whether the first sample answer passes the test on each test example; Determine the example score of each test example, where the example score of the test example is determined based on the answer score of the passed sample answer, the total score of each sample answer, and the test results; Determine the correctness reward of the first sample answer based on the example scores of the test examples passed by the first sample answer, the total scores of each test example, and the test results.

[0154] Optionally, the execution efficiency reward includes a time reward and / or a content reward; the obtaining module 802 is further configured to: Based on the test information, determine the maximum execution time and the minimum execution time of the first sample answer on each test example, and determine the time reward score of the first sample answer based on the maximum execution time and the minimum execution time; Based on the test information, determine the maximum content length and minimum content length of each sample answer, and determine the content reward score of the first sample answer; Based on the time reward score and / or content reward score, and in combination with the threshold of the second mapping range, determine the label score of the first sample answer.

[0155] Optionally, the device further includes an obtaining module, configured to: Obtain a candidate generation model; Replace the language model layer in the candidate generation model with a linear layer and a mapping layer to obtain an initial scoring model.

[0156] In the embodiments of the present specification, by designing a scientific reward formula and a dedicated data set, designing a specific loss formula, and training the initial scoring model, the reward model can adaptively adjust the search depth based on the set threshold, achieving the goal of quickly answering simple questions and deeply optimizing complex questions, and improving the inference efficiency and accuracy.

[0157] Figure 9 The structural block diagram of a computing device provided by an embodiment of the present specification is shown.

[0158] The computing device 900 includes: A memory 910 and a processor 920; The memory 910 is used to store computer programs / instructions, and the processor 920 is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 920, the steps of the above search task processing method or the training method of the target scoring model are implemented.

[0159] In one or more embodiments of the present specification, the computing device 900 may be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), a model all-in-one machine, a mobile phone, a tablet computer, or other portable intelligent terminals, etc. And, the model in the above embodiments of the present application may be pre-installed in the computing device.

[0160] Specifically, the computing device 900 can pre-install various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi-modal task processing, etc., so as to provide diverse model selections. In different product forms, the computing device 900 can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the computing device 900 also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on model evaluation tools), etc. In other product forms, the computing device 900 can also create applications based on models, provide API (Application Programming Interface) invocation capabilities, and can call the model into the created application through the API interface, while providing application management tools to realize the management and monitoring of applications.

[0161] Furthermore, the computing device 900 can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn artificial intelligence technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated artificial intelligence development, training, deployment, and application device is provided.

[0162] Figure 10 The block diagram of an electronic device provided by an embodiment of this specification is shown.

[0163] A memory 1010 and a processor 1020, the memory 1010 and the processor 1020 are connected through a bus 1030; The memory 1010 is used to store computer programs / instructions, the processor 1020 is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor 1020, the steps of the above search task processing method or the training method of the target scoring model are implemented.

[0164] Specifically, the components of the electronic device 1000 include but are not limited to the memory 1010 and the processor 1020. The processor 1020 and the memory 1010 can be connected through a bus 1030.

[0165] The electronic device 1000 may further include an access device 1040, which enables the electronic device 1000 to communicate with a database 1050 storing data via one or more networks 1060. Examples of these networks 1060 include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).

[0166] In one embodiment of the present specification, the above components of the electronic device 1000 and Figure 10 other components not shown may also be connected to each other, for example, via a bus 1030. It should be understood that Figure 10 the block diagram of the electronic device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0167] The electronic device 1000 may be any type of stationary or mobile electronic device, including a mobile computer or mobile electronic device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable electronic device (e.g., a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary electronic device such as a desktop computer or a Personal Computer (PC). The electronic device 1000 may also be a mobile or stationary server.

[0168] The above is a schematic solution of an electronic device according to this embodiment. It should be noted that the technical solution of this electronic device and the technical solutions of the above search task processing method or the training method of the target scoring model belong to the same concept. For the details not described in detail in the technical solution of the electronic device, reference can be made to the descriptions of the technical solutions of the above search task processing method or the training method of the target scoring model.

[0169] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above search task processing method or the training method of the target scoring model are implemented.

[0170] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above search task processing method or the training method of the target scoring model belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above search task processing method or the training method of the target scoring model.

[0171] An embodiment of this specification also provides a computer program product including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above search task processing method or the training method of the target scoring model are implemented.

[0172] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above search task processing method or the training method of the target scoring model belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above search task processing method or the training method of the target scoring model.

[0173] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0174] A computer program / instructions includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. A computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0175] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.

[0176] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0177] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A search task processing method, comprising: Determine the target node of the target search task, where the target node is the answer that needs to be iteratively optimized; Determine the scoring result of the target node through a target scoring model, where the target scoring model is trained based on at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; Perform node expansion based on the scoring result of the target node to determine the search result of the target search task.

2. The method according to claim 1, wherein performing node expansion based on the scoring result of the target node to determine the search result of the target search task comprises: Determine whether the scoring result of the target node meets the stop search constraint; If the stop search constraint is met, use the target node as the search result of the target search task; If the stop search constraint is not met, expand the target node to obtain an expanded node, and determine the search result of the target search task based on each current node of the target search task, where each current node of the target search task includes the target node and the expanded node.

3. The method according to claim 2, wherein the target node is target generated content; expanding the target node to obtain an expanded node comprises: Based on a set evaluation prompt, use a target generation model to generate extended content corresponding to the target generated content, and use the extended content as the expanded node.

4. The method according to claim 2, wherein determining the search result of the target search task based on each current node of the target search task comprises: Respectively determine the comprehensive scores of each node, determine the node to be expanded from each node based on the comprehensive scores, use the node to be expanded as the updated target node, and return to execute the step of determining the scoring result of the target node through the target scoring model until the search result of the target search task is obtained.

5. The method according to claim 2, wherein determining whether the scoring result of the target node meets the stop search constraint comprises: In the case where the scoring result of the target node is greater than a set scoring threshold, determine that the stop search constraint is met; In the case where the scoring result of the target node is less than or equal to the set scoring threshold, determine that the stop search constraint is not met.

6. The method according to claim 1, before determining the target node of the target search task, further comprising: For the target generation task of the target scenario, obtain the target generated content corresponding to the target generation task through a target generation model; Create a corresponding target search task based on the target generated content, where the target search task is used to update the target generated content.

7. A training method for a target scoring model, comprising: Obtain a training sample set under a target search task, where the training sample set includes at least one sample node and the corresponding label score, and the label score is determined based on the correctness reward or execution efficiency reward of the sample node; Based on the sample nodes in the training sample set, obtain the corresponding predicted scores through the initial scoring model; Based on the label scores and the predicted scores of the sample nodes, determine the target loss of the initial scoring model, and train the initial scoring model based on the target loss to obtain the trained target scoring model, where the target scoring model is used to score the target nodes in the search task processing method described in any one of claims 1-5 above.

8. The method according to claim 7, wherein the determining the target loss of the initial scoring model based on the label scores and the predicted scores of the sample nodes includes: Based on the label scores and the predicted scores of the sample nodes, determine the recall loss of the sample nodes; Based on the recall loss and the set threshold penalty factor, determine the target loss of the initial scoring model.

9. The method according to claim 7, wherein the sample nodes are sample answers; the obtaining the training sample set under the target search task includes: Obtain the test examples under the target search task; Generate at least one sample answer corresponding to each sample question under the target search task, and test each sample answer based on the test examples to obtain the test information of each test example and each sample answer; Based on the correctness reward or execution efficiency reward of each sample answer, determine the label score of each sample answer, where the correctness reward or execution efficiency reward of the sample answer is determined based on the test information of the sample answer; Based on each sample answer and the corresponding label score, obtain the training sample set under the target search task.

10. The method according to claim 9, wherein the obtaining the test examples under the target search task includes: Obtain at least one sample question under the target search task and the reference answers of each sample question; Generate at least one set of test examples corresponding to each sample question; Verify the test examples based on the reference answers of each sample question to obtain the test examples that pass the verification.

11. The method according to claim 9, wherein the determining the label score of each sample answer based on the correctness reward or execution efficiency reward of each sample answer includes: Based on the test information, determine whether there is a test example in which the first sample answer fails the test, where the first sample answer is any one of each sample answer; If so, determine the correctness reward of the first sample answer based on the test information, and map the correctness reward of the first sample answer to the first threshold range to obtain the label score of the first sample answer; If not, determine the execution efficiency reward of the first sample answer based on the test information, and map the execution efficiency reward of the first sample answer to the second threshold range to obtain the label score of the first sample answer.

12. The method according to claim 11, wherein the determining the correctness reward of the first sample answer based on the test information includes: Obtain the test results of whether the first sample answer passes the test on each test example; Determine the sample scores of each test sample, where the sample scores of the test samples are determined based on the answer scores of the passed sample answers, the total scores of each sample answer, and the test results; Determine the correctness reward of the first sample answer based on the sample scores of the test samples passed by the first sample answer, the total scores of each test sample, and the test results.

13. The method according to claim 11, wherein the execution efficiency reward includes a time reward and / or a content reward; the determining the execution efficiency reward of the first sample answer based on the test information and mapping the execution efficiency reward of the first sample answer to a second threshold range to obtain the label score of the first sample answer includes: Based on the test information, determine the maximum execution time and the minimum execution time of the first sample answer on each test sample, and determine the time reward score of the first sample answer based on the maximum execution time and the minimum execution time; Based on the test information, determine the maximum content length and the minimum content length of each sample answer, and determine the content reward score of the first sample answer; Based on the time reward score and / or the content reward score, and in combination with the thresholds of the second mapping range, determine the label score of the first sample answer.

14. The method according to claim 7, wherein the method further includes: Obtain a candidate generation model; Replace the language model layer in the candidate generation model with a linear layer and a mapping layer to obtain the initial scoring model.

15. A computing device, comprising: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1-14 are implemented.

16. An electronic device, comprising: A memory and a processor, the memory and the processor are connected by a bus; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1-14 are implemented.

17. A computer-readable storage medium, which stores computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1-14 are implemented.

18. A computer program product, comprising computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1-14 are implemented.

Citation Information

Patent Citations

  • Method and device for evaluating large code model

    CN117194258A

  • Large code model self-evolution method based on Monte Carlo tree search

    CN119398173A

  • Code generation model training method and device, electronic equipment and storage medium

    CN119621029A

  • Dynamic retrieval decision scheme determination method and system based on Monte Carlo tree search

    CN120011413A

  • Test Case Generation Through Reinforcement Learning With Static Code Quality Rewards

    US20250111239A1

Cited By

  • Job task extraction model training method and device, equipment, medium and product

    CN121071087A

  • Search model training method, search task processing method and webpage search task processing method

    CN121456105A