Method and device for realizing model reasoning search based on computing power of intelligent computing center
By combining large language models, Monte Carlo tree search and thinking chain technology on the intelligent computing center, dynamic generation and evaluation substeps are solved, and the problem of how to effectively support large language model inference in the intelligent computing center is achieved, achieving high-quality solution generation and efficiency optimization.
Patent Information
- Application Number
- CN202510096019.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-21
AI Technical Summary
How to effectively utilize the computing power resources of the intelligent computing center to support the reasoning of large language models and improve the quality of solutions generated by large language models in complex scenarios.
By combining large language models, Monte Carlo tree search algorithms and thinking chain technology, the computing power resources of the intelligent computing center are used to dynamically generate, execute and evaluate substeps until the score result is greater than the preset threshold or the solution is effective.
It realizes the generation of high-quality solutions in complex scenarios, optimizes the generation efficiency of solutions, and improves the quality of solutions.
Smart Images

Figure CN120030197A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of computing power infrastructure, and in particular, to a method and device for implementing model reasoning search based on the computing power of an intelligent computing center. Background Art
[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0003] "Large language model" refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0004] "Scaling Law" describes the power-law relationship between model performance and model parameters, data volume, and computing resources. Studies have shown that as resources increase, the improvement in model performance shows a phenomenon of diminishing returns. This means that when the model is small, increasing resources can significantly improve performance; while in large-scale models, the impact of increasing resources on performance improvement gradually weakens. Therefore, Scaling Law not only reveals the balance between resource input and performance improvement, but also provides researchers with a theoretical framework for effective performance optimization under conditions of limited data and computing power.
[0005] As Internet data resources are depleted, the traditional Scaling Law paradigm has also changed, gradually shifting from a pre-training-based model to a path of reasoning expansion. The new paradigm emphasizes the integration of inference models, synthetic data generation, and reinforcement learning (RL), indicating that the reasoning process will play an increasingly important role in future AI (Artificial Intelligence) research. This shift not only reflects the need for more efficient data utilization, but also promotes in-depth exploration of reasoning services.
[0006] In this context, the o1 model released by OpenAI on September 13, 2024, uses reinforcement learning technology to break through the limits of LLM reasoning and create a new reasoning expansion paradigm. The o1 model places special emphasis on the computational requirements of the training and reasoning stages, and proposes two new RL post-training scaling laws, "train-time compute" and "test-time compute". By showing that the performance of the model improves significantly with the increase of reinforcement learning time and reasoning thinking time, the o1 model provides strong empirical support for the importance of reasoning services.
[0007] Under the new paradigm, the computing power requirements have not decreased, but have increased significantly due to the introduction of inference models. Traditionally, the computing requirements for pre-training and post-training are in a 1:1 ratio, but in the framework of the o1 model, the computing requirements of the inference phase account for a larger proportion. Components such as generators, reward models, policy models, and verifiers generate a large amount of "berry training data" (inference path data), which greatly increases the overall computing power requirements. During the inference process, the amount of forward pass computation increases significantly, while the proportion of gradient operations and weight corrections decreases, which supports the potential of large distributed clusters.
[0008] In summary, as AI systems develop, computing overhead will be more concentrated on inference services rather than just pre-training calculations. Therefore, how to effectively use the computing resources of intelligent computing centers to support the inference of large language models and improve the quality of solutions generated by large language models under inference in complex scenarios has become a technical problem that needs to be solved urgently. Summary of the invention
[0009] The embodiments of the present application provide a method and device for implementing model reasoning search based on the computing power of an intelligent computing center, so as to solve the technical problem of how to effectively utilize the computing power resources of the intelligent computing center to support the reasoning of a large language model and improve the quality of the solutions generated by the large language model under the reasoning of complex scenarios.
[0010] In order to solve the above technical problems, this application is implemented as follows:
[0011] In a first aspect, the present application provides a method for implementing model reasoning search based on the computing power of an intelligent computing center, the method comprising:
[0012] Step S1: obtaining a request input by a user, inputting the request into a large language model, and obtaining a plurality of sub-steps output by the large language model, wherein the large language model is used to use the request as a root node of a Monte Carlo tree, and based on a thought chain technique, inferring the plurality of sub-steps equivalent to leaf nodes at the same level of the Monte Carlo tree;
[0013] Step S2: sending the plurality of sub-steps to the executor for execution respectively, and scoring the plurality of sub-steps respectively based on the large language model to obtain a score result corresponding to each sub-step;
[0014] Step S3: selecting the sub-step with the highest score based on the score result, and inputting the sub-step with the highest score and the request into the large language model to obtain multiple new sub-steps output by the large language model, wherein the large language model is used to infer the multiple new sub-steps equivalent to the leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the thought chain technology, and the leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score;
[0015] Step S4: sending the multiple new sub-steps to the executor for execution respectively, and scoring the multiple new sub-steps respectively based on the large language model and the score of the sub-step with the highest score, to obtain a score result corresponding to each new sub-step;
[0016] Step S5: Repeat the steps S3 and S4 until the score result of the new sub-step is greater than a preset threshold, and then record the solution formed by the sub-step with the highest score in the leaf nodes of each level; or, until the solution formed by the sub-step with the highest score in the leaf nodes of each level is determined to be valid by the large language model based on the request and the large language model;
[0017] Step S6: taking the solution as the target solution for the request.
[0018] Optionally, the sub-steps include: the large language model infers a calling method and input parameters corresponding to the calling method based on the request and the thought chain technology.
[0019] Optionally, the calling method includes at least one of the following: a basic calling function, and a calling method generated by the large language model based on the basic calling function and the request.
[0020] Optionally, step S2 includes:
[0021] Step S21: sending the plurality of sub-steps to an executor for execution respectively, and obtaining an execution result, wherein the execution result includes any one of the following: execution success, execution failure;
[0022] Step S22: if the execution result is the execution failure, then based on the large language model, a first score result corresponding to each sub-step in which the execution failed is obtained, and the first score result is 0 points;
[0023] Step S23: If the execution result is that the execution is successful, then based on the large language model, a second score result corresponding to each successfully executed sub-step is obtained, and the second score result is a positive number.
[0024] Optionally, step S6 includes:
[0025] Step S61: inputting the solution into the actuator for verification;
[0026] Step S62: If the verification is successful, the solution is used as the target solution for the request.
[0027] In a second aspect, an embodiment of the present application provides a device for implementing model reasoning search based on the computing power of an intelligent computing center, the device comprising:
[0028] An acquisition module, used to execute step S1: acquire a request input by a user, input the request into a large language model, and obtain multiple sub-steps output by the large language model, wherein the large language model is used to use the request as a root node of a Monte Carlo tree, and based on a thought chain technology, infer the multiple sub-steps equivalent to leaf nodes at the same level of the Monte Carlo tree;
[0029] An execution module, used to execute step S2: sending the multiple sub-steps to the executor for execution respectively, and scoring the multiple sub-steps respectively based on the large language model to obtain a score result corresponding to each sub-step;
[0030] Used to execute step S3: select the sub-step with the highest score based on the score result, and input the sub-step with the highest score and the request into the large language model to obtain multiple new sub-steps output by the large language model, wherein the large language model is used to infer the multiple new sub-steps equivalent to the leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the thought chain technology, and the leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score;
[0031] Used to execute step S4: send the multiple new sub-steps to the executor for execution respectively, and score the multiple new sub-steps respectively based on the large language model and the score of the sub-step with the highest score, to obtain a score result corresponding to each new sub-step;
[0032] for executing step S5: repeatedly executing step S3 and step S4 until the score result of the new sub-step is greater than a preset threshold, then recording the solution formed by the sub-step with the highest score in the leaf nodes of each level; or, until, based on the request and the large language model, the large language model determines that the solution formed by the sub-step with the highest score in the leaf nodes of each level is valid;
[0033] Used to execute step S6: taking the solution as the target solution of the request.
[0034] Optionally, the sub-steps include: the large language model infers a calling method and input parameters corresponding to the calling method based on the request and the thought chain technology.
[0035] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for realizing model reasoning search based on the computing power of an intelligent computing center as described in the first aspect.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of a method for realizing model reasoning search based on the computing power of an intelligent computing center as described in the first aspect are implemented.
[0037] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of a method for realizing model reasoning search based on the computing power of an intelligent computing center as described in the first aspect.
[0038] In the method provided in the embodiment of the present application, the reasoning ability of the large language model is combined with the Monte Carlo tree search and the thought chain technology to fully utilize the advantages of each. For example, the large language model can generate solutions described in natural language, and the Monte Carlo tree as the search skeleton can expand the solution space of the original large language model and solve the reasoning problem of complex problems, and the thought chain technology can provide support for logical reasoning.
[0039] Specifically, the Monte Carlo tree search algorithm allows for extensive exploration in the solution space. The algorithm can generate multiple possible solutions by randomly sampling and evaluating different decision paths, thus avoiding falling into local optimal solutions. The mind chain technology emphasizes solving complex problems through step-by-step reasoning. Combined with the reasoning ability of large language models, detailed sub-steps can be generated in each step and analyzed and evaluated. This step-by-step optimization method ensures that each sub-step is carefully considered, thereby improving the quality of the final solution. In the dynamic generation, execution, and evaluation cycle, the generated sub-steps can be continuously adjusted and optimized based on the execution results and feedback information. By evaluating the effectiveness of each sub-step in real time, low-quality options can be quickly identified and eliminated, thereby improving the quality of the overall solution.
[0040] The Intelligent Computing Center can provide powerful computing resources, enabling the Monte Carlo tree search algorithm to perform efficient random sampling and evaluation in large-scale state spaces; at the same time, the reasoning process of the thinking chain technology and large language models can also be accelerated in the parallel computing environment provided by the Intelligent Computing Center. The Intelligent Computing Center can also dynamically adjust the allocation of computing resources according to the complexity of the task and real-time requirements. For example, in the Monte Carlo tree search stage, more computing power can be allocated for extensive exploration, while in the thinking chain reasoning stage, resources can be concentrated for in-depth analysis.
[0041] In summary, by combining the computing power resources of the intelligent computing center with the Monte Carlo tree search algorithm, thought chain technology and reasoning of the large language model, high-quality solutions can be obtained in complex scenarios. The dynamic generation, execution and evaluation cycle mechanism not only optimizes the solution generation efficiency, but also improves the quality of the solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0043] Figure 1 A flowchart of a method for implementing model reasoning search based on the computing power of an intelligent computing center provided in an embodiment of the present application;
[0044] Figure 2 A flowchart of a method for implementing model reasoning search based on the computing power of an intelligent computing center provided in an embodiment of the present application;
[0045] Figure 3A structural block diagram of a device for implementing model reasoning search based on the computing power of an intelligent computing center provided in an embodiment of the present application;
[0046] Figure 4 A schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0048] The following is a brief description of the technical terms involved in this application.
[0049] The "computing power" mentioned in this application is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0050] The "computing power" (Computational Power, CP) described in this application is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is about the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP = CP general + CP intelligent + CP super.
[0051] The "Network Power" (NP) described in this application is a manifestation of the data transmission capability of computing facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. The carrying capacity involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities. In the embodiment of this application, the carrying capacity uses the memory bandwidth.
[0052] The "Storage Power" (SP) described in this application refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.
[0053] The "computing power infrastructure" described in this application is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, and green and low-carbon.
[0054] The “computing power” described in this application includes general computing power, intelligent computing power and super computing power.
[0055] The "general computing power" described in this application refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0056] The "intelligent computing power" described in this application is a computing platform for large-scale deployment of special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), etc. for various innovative applications of artificial intelligence, such as natural language processing and machine vision.
[0057] The "supercomputing power" mentioned in this application is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.
[0058] The "intelligent computing center" mentioned in this application refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning scenarios) by using large-scale heterogeneous computing resources, including general computing power (CPU: Central Processing Unit) and intelligent computing power (GPU: Graphics Processing Unit, FPGA: Field Programmable Gate Array, ASIC: Application Specific Integrated Circuit, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0059] The "computing resources" mentioned in this application refer to technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as supporting and guarantee resources such as wind, fire, water and electricity.
[0060] The "large language model" mentioned in this application refers to a large-scale language model, which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0061] The “Post-training key technologies” described in this application are divided into the following aspects:
[0062] "Self-play RL": Self-play is a reinforcement learning framework in which an agent collects training data by playing multiple games against itself. This method can effectively reduce the reliance on complex environments. The idea of self-play originally came from AlphaGo Zero, which generated data by playing against itself and used Monte Carlo Tree Search (MCTS) to guide each step of the decision. This method is also applicable to other task scenarios that require strategy planning.
[0063] "Chain of Thought (CoT)": Chain of Thought is the key to reasoning about complex problems for large models. In 2022, Wei et al. added a paragraph of "step-by-step reasoning" text to the dataset to stimulate the thinking ability of large models. This method allows the model to decompose multi-step reasoning problems into intermediate steps, thereby improving the accuracy and interpretability of the model when dealing with complex problems. The chain of thought technology simulates the thinking process of humans when solving problems, enabling large language models to analyze problems step by step and generate intermediate reasoning steps to finally obtain accurate answers. This method not only improves the performance of the model in mathematical, common sense, and symbolic reasoning tasks, but also has a significant impact on the model size, with large models benefiting more. In addition, the chain of thought prompts also make large language models more interpretable and provide an opportunity to debug reasoning path errors.
[0064] "Monte Carlo Tree Search": is a heuristic search algorithm used in certain decision-making processes. Its core idea is to put resources on branches that are more worthy of search, that is, to concentrate computing power on more valuable places. In response to the Inference scaling law caused by the o1 model, the search method is used to allow LLM to explore a larger solution space and find a complete solution. In essence, it uses the CoT method to optimize the path of LLM sampling to the correct solution.
[0065] Each cycle of Monte Carlo Tree Search consists of four steps:
[0066] Selection: Starting from the root node R, select child nodes continuously downward to the leaf node L;
[0067] Expansion: Unless the game ends at L due to a win or loss by either party, create one or more child nodes and select one of them, C.
[0068] Simulation: Start from node C again and play the game with a random strategy, also known as playout or rollout;
[0069] Backpropagation: Use the results of the random game to update the node information on the path from C to R.
[0070] Figure 1 A method for implementing model reasoning search based on the computing power of an intelligent computing center according to an embodiment of the present application is shown. Figure 1 As shown, the method includes:
[0071] Step S1: Obtain a request input by a user, input the request into a large language model, and obtain multiple sub-steps of the large language model output;
[0072] The large language model is used to take the request as the root node of the Monte Carlo tree, and based on the thought chain technology, it infers multiple sub-steps equivalent to the leaf nodes at the same level of the Monte Carlo tree;
[0073] Step S2: Send the multiple sub-steps to the executor for execution respectively, and score the multiple sub-steps respectively based on the large language model to obtain the score results corresponding to each sub-step;
[0074] Step S3: Select the sub-step with the highest score based on the score result, and input the sub-step with the highest score and the request into the large language model to obtain multiple new sub-steps output by the large language model;
[0075] Among them, the large language model is used to infer multiple new sub-steps equivalent to the leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the thought chain technology, and the leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score;
[0076] Step S4: Send the multiple new sub-steps to the executor for execution respectively, and score the multiple new sub-steps respectively based on the large language model and the score of the sub-step with the highest score, so as to obtain the score result corresponding to each new sub-step;
[0077] Step S5: Repeat step S3 and step S4 until the score result of the new sub-step is greater than the preset threshold, and then record the solution formed by the sub-step with the highest score in the leaf nodes of each level; or, until the solution formed by the sub-step with the highest score in the leaf nodes of each level is determined to be valid by the large language model based on the request and the large language model;
[0078] Step S6: The solution is taken as the target solution of the request.
[0079] In step S1, the request (query) input by the user can be a question, task or requirement to accurately convey the user's intention to the large language model (such as GPT, etc.), and the large language model is used to use the request as the root node of the Monte Carlo tree, and to infer multiple sub-steps based on the thinking chain technology. These sub-steps are equivalent to the leaf nodes at the same level of the Monte Carlo tree, indicating multiple preliminary actions that can be taken from the current request. The sub-steps include: the calling method and the input parameters corresponding to the calling method inferred by the large language model based on the request and the thinking chain technology. The calling method includes at least one of the following: a basic calling function, a calling method generated by the large language model based on the basic calling function and the request. In other words, when the defined basic calling function cannot be processed, the large language model can adaptively generate code based on the basic calling function and embed it into the workflow to ensure the complete execution of the solution.
[0080] In step S2, the multiple sub-steps generated in step S1 can be input into the executor. The executor is responsible for actually executing these sub-steps, and can execute the sub-steps by calling external services, running algorithms, or performing other operations, and the large language model can score each sub-step to obtain the score results corresponding to each sub-step. Therefore, by executing and scoring the sub-steps, the actual effect of each sub-step can be evaluated, providing a basis for selecting the best path for the next step.
[0081] It should be noted that, unlike typical mathematics and code fields, for the generation of solutions, each sub-step outputs a standard JSON dictionary, which contains the method to be called (such as the function to be called), input parameters, etc., and uses an external executor to execute the calling method to provide a complete solution. For example: If the user's request is to develop a climbing recognition algorithm to automatically detect and send real-time alerts to whether the person entering the detection area has raised his hands and his feet are above the set threshold to ensure the safety of smart parks and building real estate. Then sub-step 1 can be as follows:
[0082] Step 1:
[0083] "current_step":"Use the posture estimation function to process the video and detect the key points of the people in the video, especially the positions of hands and feet, to provide basic data for subsequent climbing recognition.",
[0084] "function":"pose",
[0085] "input":{
[0086] "video_path":"Input video path"
[0087] }
[0088] In one possible implementation, Figure 2 As shown, step S2 includes:
[0089] Step S21: Sending the multiple sub-steps to the executor for execution respectively, and obtaining the execution result, which includes any of the following: execution success, execution failure;
[0090] Step S22: If the execution result is execution failure, then based on the large language model, a first score result corresponding to each sub-step that failed to execute is obtained, and the first score result is 0 points;
[0091] Step S23: If the execution result is successful, then based on the large language model, a second score result corresponding to each successfully executed sub-step is obtained, and the second score result is a positive number.
[0092] Therefore, by setting different scoring results, the performance of each sub-step can be clearly evaluated to select the sub-step with the best performance to participate in the generation of the solution, and serve as the basis for the generation of subsequent sub-steps, which can further improve the quality of the solution.
[0093] In step S3, the Monte Carlo tree search algorithm and the chain of thought technology are combined. The large language model can infer multiple new sub-steps equivalent to the leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the chain of thought technology. The leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score. Specifically, after obtaining the score results of each sub-step, the sub-step with the highest score will be selected to ensure that the selected sub-step is the most potential, and the sub-step with the highest score will be input into the large language model together with the original request to generate new sub-steps, which are equivalent to the leaf nodes of the next level of the Monte Carlo tree. And this process can be based on the chain of thought technology to ensure that the generated new sub-steps are logically reasonable and request-related extensions. As a result, the large language model can gradually deepen reasoning and explore the possibility of solutions, ensuring that each step is based on the extension of existing successful steps to improve the quality of the solution.
[0094] In step S4, the multiple new sub-steps need to be sent to the executor for execution, and the multiple new sub-steps are scored based on the large language model and the score of the sub-step with the highest score, and the score results corresponding to each new sub-step are obtained. In other words, the multiple new sub-steps generated in step S3 are input to the executor for execution and scored again. The difference between this scoring and the previous one is that the score of the sub-step with the highest score will be combined to further evaluate the value of the newly generated sub-step.
[0095] In step S5, steps S3 and S4 need to be repeated until the score result of the new sub-step is greater than the preset threshold, and the solution formed by the sub-step with the highest score in the leaf nodes of each level is recorded; or, based on the request and the large language model, the large language model determines that the solution formed by the sub-step with the highest score in the leaf nodes of each level is valid. In other words, steps S3 and S4 are repeated to form a loop until the score result of the new sub-step exceeds the preset threshold, or when the large language model confirms that the currently generated solution is valid. This process ensures that the solution is optimized in continuous iterations and gradually approaches the optimal result.
[0096] Step S6 may further include: Step S61: input the solution to the executor for verification; Step S62: if the verification is successful, the solution is used as the target solution for the request. Therefore, after the solution is obtained, it can be input to the executor for unified verification to ensure the successful execution of the solution. If the verification fails, detailed feedback information about the verification failure can be collected from the executor, such as error information, execution logs, performance indicators, etc., and the specific reasons for the failure can be analyzed. According to the analysis results, the original solution can be adjusted as necessary and verified again. In the adjustment of the solution, steps S1 to S6 can be re-executed, and the overall verification of the solution can be re-performed until the verification is successful and the target solution is obtained.
[0097] In summary, in the method provided in the embodiment of the present application, the reasoning ability of the large language model is combined with the Monte Carlo tree search and the thought chain technology to fully utilize their respective advantages. For example, the large language model can generate solutions described in natural language, and the Monte Carlo tree, as the search skeleton, can expand the solution space of the original large language model and solve the reasoning problem of complex problems, and the thought chain technology can provide support for logical reasoning.
[0098] Specifically, the Monte Carlo tree search algorithm allows for extensive exploration in the solution space. The algorithm can generate multiple possible solutions by randomly sampling and evaluating different decision paths, thus avoiding falling into local optimal solutions. The mind chain technology emphasizes solving complex problems through step-by-step reasoning. Combined with the reasoning ability of large language models, detailed sub-steps can be generated in each step and analyzed and evaluated. This step-by-step optimization method ensures that each sub-step is carefully considered, thereby improving the quality of the final solution. In the dynamic generation, execution, and evaluation cycle, the generated sub-steps can be continuously adjusted and optimized based on the execution results and feedback information. By evaluating the effectiveness of each sub-step in real time, low-quality options can be quickly identified and eliminated, thereby improving the quality of the overall solution.
[0099] The Intelligent Computing Center can provide powerful computing resources, enabling the Monte Carlo tree search algorithm to perform efficient random sampling and evaluation in large-scale state spaces. At the same time, the reasoning process of the thinking chain technology and large language models can also be accelerated in the parallel computing environment provided by the Intelligent Computing Center. The Intelligent Computing Center can also dynamically adjust the allocation of computing resources according to the complexity of the task and real-time requirements. For example, in the Monte Carlo tree search stage, more computing power can be allocated for extensive exploration, while in the thinking chain reasoning stage, resources can be concentrated for in-depth analysis.
[0100] In summary, by combining the computing power resources of the intelligent computing center with the Monte Carlo tree search algorithm, thought chain technology and reasoning of the large language model, high-quality solutions can be obtained in complex scenarios. The dynamic generation, execution and evaluation cycle mechanism not only optimizes the solution generation efficiency, but also improves the quality of the solution.
[0101] In addition, based on the method provided in the embodiment of the present application, each search can generate usable data, which can provide support for subsequent basic model fine-tuning, such as SFT (Supervised Fine-Tuning) and fine-tuning of the reward model. Both positive and negative data can play an important role. Positive examples help enhance the model's understanding of the target task, while negative examples can help the model identify and avoid errors, thereby improving overall performance. By making rational use of synthetic data, the model can be continuously optimized to perform better in practical applications.
[0102] The method for implementing model reasoning search based on the computing power of an intelligent computing center provided in an embodiment of the present application is now explained from the perspective of specific application scenarios and codes.
[0103] Taking Titan (a platform built by Agent) as an example, the root node of the Monte Carlo tree is the user's query (equivalent to the request mentioned above); the leaf nodes below it are each different step (equivalent to sub-steps), and step is just one of the steps to solve the query; and the number of tree branches (that is, the number of extensions of sub-steps at the same level) can be pre-defined, for example, branch = 3, then for the same query, there will be 3 next steps with different contents but at the same level; simulation is to select one of the steps and use the root node (query) to the selected step as the existing solution, based on which, the simulation continues to look for the solution step downward; backtracking is when the simulation step is reached or the final solution is found, the value of all nodes on this path is changed.
[0104] It should be noted that "value" is an important concept in the search process, especially in the reward model in reinforcement learning, which is used to evaluate the impact of each step on the final solution, thereby generating a reward signal, usually a decimal between 0 and 1. These reward signals will be accumulated to guide the subsequent search direction and provide support for the selection of the final solution.
[0105] With the above introduction as the background, focusing on the CV (Computer Vision) scenario, a method for implementing model reasoning search based on the computing power of an intelligent computing center as shown in an embodiment of the present application is introduced.
[0106] The query is to develop a climbing recognition algorithm to automatically detect and send real-time alerts to people entering the detection area if their hands are raised and their feet are off the ground beyond the set threshold, to ensure the safety of smart parks and building real estate.
[0107] At this time, the existing steps are empty, and the large language model will give guidance based on the query, that is, based on the current relevant information, what method is best to call next. For example, standardized opinions: First, you need to use the "pose function" to estimate the posture of the video and detect the key points of the person in the video, especially the position of the hands and feet, which is the basis for realizing climbing recognition.
[0108] Then, the large language model will expand the nodes based on the opinions and context. For example, three next steps may be generated, one of which may be:
[0109] Step 1:
[0110] "current_step":"Use the posture estimation function to process the video and detect the key points of the people in the video, especially the positions of hands and feet, to provide basic data for subsequent climbing recognition.",
[0111] "function":"pose",
[0112] "input":{
[0113] "video_path":"Input video path"
[0114] }
[0115] It should be noted that the large language model returns relevant information about the calling method. After obtaining the calling method and input parameters, it can be sent to the executor for execution and obtain the execution result, which is success or failure. After obtaining the execution result of the step, the large language model will judge the reward score of the step based on the current query and the execution result of the step, that is, how much help this step can provide to the final solution. It should be noted that after the existing steps are generated, the large language model will judge the reward score of the current step based on the current query, the execution result of the step and the reward score of the existing step in the scoring.
[0116] After getting the reward score, you can give a basis for the next node (sub-step, step), that is, select the next node to be expanded based on the size of the reward. For example: among the three next steps mentioned above, the reward score of step 1 is 0.3, and the scores of the other two next steps are 0.2 and 0.1, then you can continue to expand the node based on step 1, and repeat the process until the score of the final step is greater than the preset threshold, or the large language model determines that the solution is valid, then the search ends.
[0117] Exemplary:
[0118] The final search results are as follows:
[0119]
[0120]
[0121] In this way, any query raised by the user can be processed, and a solution that can be successfully executed can be found through multi-step search.
[0122] In summary, by combining the computing power resources of the intelligent computing center with the Monte Carlo tree search algorithm, thought chain technology and reasoning of the large language model, high-quality solutions can be obtained in complex scenarios. The dynamic generation, execution and evaluation cycle mechanism not only optimizes the solution generation efficiency, but also improves the quality of the solution.
[0123] Figure 3 A device for implementing model reasoning search based on the computing power of an intelligent computing center according to an embodiment of the present application is shown, and the device 30 includes:
[0124] The acquisition module 301 is used to execute step S1: obtain a request input by a user, input the request into a large language model, and obtain multiple sub-steps output by the large language model, wherein the large language model is used to use the request as a root node of a Monte Carlo tree, and based on the thought chain technology, infer multiple sub-steps equivalent to leaf nodes at the same level of the Monte Carlo tree;
[0125] The execution module 302 is used to execute step S2: send the multiple sub-steps to the executor for execution respectively, and score the multiple sub-steps respectively based on the large language model to obtain the score results corresponding to each sub-step;
[0126] Used to execute step S3: select the sub-step with the highest score based on the score result, and input the sub-step with the highest score and the request into the large language model to obtain multiple new sub-steps output by the large language model, wherein the large language model is used to infer multiple new sub-steps equivalent to leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the thought chain technology, and the leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score;
[0127] Used to execute step S4: send the multiple new sub-steps to the executor for execution respectively, and score the multiple new sub-steps respectively based on the large language model and the score of the sub-step with the highest score, to obtain the score result corresponding to each new sub-step;
[0128] Used to execute step S5: repeatedly execute step S3 and step S4 until the score result of the new sub-step is greater than the preset threshold, and then record the solution formed by the sub-step with the highest score in the leaf nodes of each level; or, until the solution formed by the sub-step with the highest score in the leaf nodes of each level is determined to be valid by the large language model based on the request and the large language model;
[0129] Used to execute step S6: taking the solution as the target solution of the request.
[0130] In a possible implementation, the sub-steps include: the large language model infers the calling method and the input parameters corresponding to the calling method based on the request and the thought chain technology.
[0131] In a possible implementation, the calling method includes at least one of the following: a basic calling function, and a calling method generated by a large language model based on the basic calling function and a request.
[0132] In a possible implementation, the execution module 302 is further used to execute step S21: sending the multiple sub-steps to the executor for execution respectively, and obtaining the execution result, which includes any of the following: execution success, execution failure;
[0133] Execute step S22: if the execution result is execution failure, then based on the large language model, obtain the first score result corresponding to each sub-step in which execution failed, and the first score result is 0 points;
[0134] Execute step S23: If the execution result is successful, then based on the large language model, obtain the second score result corresponding to each successfully executed sub-step, and the second score result is a positive number.
[0135] In a possible implementation, the execution module 302 is further configured to execute step S61: inputting the solution to the actuator for verification;
[0136] Execute step S62: If the verification is successful, the solution is used as the target solution of the request.
[0137] In summary, by combining the computing power resources of the intelligent computing center with the Monte Carlo tree search algorithm, thought chain technology and reasoning of the large language model, high-quality solutions can be obtained in complex scenarios. The dynamic generation, execution and evaluation cycle mechanism not only optimizes the solution generation efficiency, but also improves the quality of the solution.
[0138] Please refer to Figure 4 The embodiment of the present application also provides an electronic device 40, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the computer program is executed by the processor 401, each process of the above-mentioned method embodiment for realizing model reasoning search based on the computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0139] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned method embodiment for realizing model reasoning search based on the computing power of an intelligent computing center can be implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0140] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of the above-mentioned method embodiment for realizing model reasoning search based on the computing power of an intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0141] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0142] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0143] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for realizing model reasoning search based on the computing power of an intelligent computing center, characterized in that: The method comprises: Step S1: obtaining a request input by a user, inputting the request into a large language model, and obtaining a plurality of sub-steps output by the large language model, wherein the large language model is used to use the request as a root node of a Monte Carlo tree, and based on a thought chain technique, inferring the plurality of sub-steps equivalent to leaf nodes at the same level of the Monte Carlo tree; Step S2: sending the plurality of sub-steps to the executor for execution respectively, and scoring the plurality of sub-steps respectively based on the large language model to obtain a score result corresponding to each sub-step; Step S3: selecting the sub-step with the highest score based on the score result, and inputting the sub-step with the highest score and the request into the large language model to obtain multiple new sub-steps output by the large language model, wherein the large language model is used to infer the multiple new sub-steps equivalent to the leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the thought chain technology, and the leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score; Step S4: sending the multiple new sub-steps to the executor for execution respectively, and scoring the multiple new sub-steps respectively based on the large language model and the score of the sub-step with the highest score, to obtain a score result corresponding to each new sub-step; Step S5: Repeat the steps S3 and S4 until the score result of the new sub-step is greater than a preset threshold, and then record the solution formed by the sub-step with the highest score in the leaf nodes of each level; or, until the solution formed by the sub-step with the highest score in the leaf nodes of each level is determined to be valid by the large language model based on the request and the large language model; Step S6: taking the solution as the target solution for the request.
2. The method according to claim 1, characterized in that The sub-steps include: the large language model infers a calling method and input parameters corresponding to the calling method based on the request and the thought chain technology.
3. The method according to claim 2, characterized in that The calling method includes at least one of the following: a basic calling function, and a calling method generated by the large language model based on the basic calling function and the request.
4. The method according to claim 1, characterized in that: The step S2 comprises: Step S21: sending the plurality of sub-steps to an executor for execution respectively, and obtaining an execution result, wherein the execution result includes any one of the following: execution success, execution failure; Step S22: if the execution result is the execution failure, then based on the large language model, a first score result corresponding to each sub-step in which the execution failed is obtained, and the first score result is 0 points; Step S23: If the execution result is that the execution is successful, then based on the large language model, a second score result corresponding to each successfully executed sub-step is obtained, and the second score result is a positive number.
5. The method according to any one of claims 1 to 4, characterized in that The step S6 comprises: Step S61: inputting the solution into the actuator for verification; Step S62: If the verification is successful, the solution is used as the target solution for the request.
6. A device for implementing model reasoning search based on the computing power of an intelligent computing center, characterized in that: The device comprises: An acquisition module, used to execute step S1: acquire a request input by a user, input the request into a large language model, and obtain multiple sub-steps output by the large language model, wherein the large language model is used to use the request as a root node of a Monte Carlo tree, and based on a thought chain technology, infer the multiple sub-steps equivalent to leaf nodes at the same level of the Monte Carlo tree; An execution module, used to execute step S2: sending the multiple sub-steps to the executor for execution respectively, and scoring the multiple sub-steps respectively based on the large language model to obtain a score result corresponding to each sub-step; Used to execute step S3: select the sub-step with the highest score based on the score result, and input the sub-step with the highest score and the request into the large language model to obtain multiple new sub-steps output by the large language model, wherein the large language model is used to infer the multiple new sub-steps equivalent to the leaf nodes of the next level of the Monte Carlo tree based on the request, the sub-step with the highest score and the thought chain technology, and the leaf nodes of the next level are the leaf nodes of the next level expanded based on the sub-step with the highest score; Used to execute step S4: send the multiple new sub-steps to the executor for execution respectively, and score the multiple new sub-steps respectively based on the large language model and the score of the sub-step with the highest score, to obtain a score result corresponding to each new sub-step; for executing step S5: repeatedly executing step S3 and step S4 until the score result of the new sub-step is greater than a preset threshold, then recording the solution formed by the sub-step with the highest score in the leaf nodes of each level; or, until, based on the request and the large language model, the large language model determines that the solution formed by the sub-step with the highest score in the leaf nodes of each level is valid; Used to execute step S6: taking the solution as the target solution of the request.
7. The device according to claim 6, characterized in that The sub-steps include: the large language model infers a calling method and input parameters corresponding to the calling method based on the request and the thought chain technology.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for realizing model reasoning search based on the computing power of an intelligent computing center as described in any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for implementing model reasoning search based on the computing power of an intelligent computing center as described in any one of claims 1 to 5.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for realizing model reasoning search based on the computing power of an intelligent computing center as described in any one of claims 1-5.
Citation Information
Patent Citations
Intelligent customer service language model optimization method and system for instruction enhancement based on Monte Carlo tree search
CN118170884A
Thinking tree-based big language model reasoning method and device
CN118709786A
Automatic analysis method for foreign-related legal information text based on retrieval enhanced thinking chain
CN118761865A
Large model agent interactive question and answer task decision-making method, device and equipment and medium
CN119166778A
Reward-model based reinforcement learning for performing reasoning tasks
US20240104391A1