A model processing method, apparatus, storage medium, and program product

CN122797702APending Publication Date: 2026-09-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338216.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

但是,这种“长思维链慢思考”的推理大幅提升了推理所需的计算量,如此导致推理设备需要较多的推理算力开销,推理设备处理速度较慢

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797702A_ABST
    Figure CN122797702A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model processing method and device, a storage medium and a program product. The method comprises: performing a search tree algorithm based on an input question, and performing the i th round of search, i being greater than or equal to 1; determining a similarity index corresponding to the i th round of search, the similarity index corresponding to the i th round of search being used to represent the similarity of inference content generated by the model in the i th round of search; in the case where the similarity index corresponding to the i th round of search satisfies a first condition, determining the number of expansion nodes corresponding to the (i+1) th round of search, the number of expansion nodes corresponding to the (i+1) th round of search being less than the number of expansion nodes corresponding to the i th round of search. In this way, the number of expansion nodes corresponding to the next round of search can be reduced, the computing power consumed on homogeneous content can be saved without affecting the inference result, the amount of calculation required for inference can be reduced, and the processing speed of the inference device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, specifically to a model processing method, apparatus, storage medium, and program product. Background Technology

[0002] With the development of artificial intelligence (AI), model-based search solutions have become a hot research topic in academia and industry.

[0003] Currently, by incorporating search schemes during model training, models can learn multi-step reasoning capabilities from synthetic data. Furthermore, by integrating search schemes during inference, multi-round, multi-path reasoning can be achieved, providing more computation to significantly enhance the model's reasoning ability. However, this "long thought chain, slow thinking" reasoning greatly increases the computational load required for inference, resulting in higher computational costs for inference devices and slower processing speeds. Summary of the Invention

[0004] This application provides a model processing method, apparatus, storage medium, and program product that can reduce the amount of computation required for inference and improve the processing speed of inference devices.

[0005] In view of this, firstly, embodiments of this application provide a model processing method, which can be executed by an inference device. Unless otherwise specified, the "inference device" in embodiments of this application can refer to the inference device itself, a component within the inference device (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the functions of the inference device. The method includes: executing a search tree algorithm based on the input question, performing an i-th round of search, where i is greater than or equal to 1; determining a similarity index corresponding to the i-th round of search, which characterizes the similarity of the inference content generated by the model in the i-th round of search; and, if the similarity index corresponding to the i-th round of search satisfies a first condition, determining the number of extended nodes corresponding to the (i+1)-th round of search, where the number of extended nodes corresponding to the (i+1)-th round of search is less than the number of extended nodes corresponding to the i-th round of search.

[0006] The method provided in this application determines whether the reasoning content generated in each round of search is highly homogeneous and has low information content by determining the similarity of the reasoning content generated in that round of search. Combined with a first condition, it determines whether the data of the extended nodes corresponding to the next round of search needs to be adjusted. If the first condition is met, the number of extended nodes corresponding to the next round of search is reduced. This saves computational power consumed on homogeneous content without affecting the reasoning result, reducing the computational load required for reasoning and improving the processing speed of the reasoning device. Furthermore, in the slow-thinking reasoning of a long thought chain model, the dynamic similarity index corresponding to each round of search is used for conditional judgment. This enables proactive perception of the model's reasoning state, reducing the total number of generated sequences (tokens) while keeping the accuracy reduction controllable, achieving reasonable compression of the search tree size. This improves the processing speed of the reasoning device while ensuring the accuracy of the reasoning result.

[0007] In conjunction with the first aspect, in one possible implementation of the first aspect, determining the similarity index corresponding to the i-th round of search includes: obtaining j child nodes of the search tree corresponding to the i-th round of search; scoring the j child nodes of the search tree to obtain j scoring results, where j is greater than or equal to 1; and determining the similarity index corresponding to the i-th round of search based on the j scoring results. By scoring the child nodes of the search tree, a corresponding similarity index can be obtained, reflecting the similarity of the reasoning content and improving the feasibility of the solution.

[0008] In conjunction with the first aspect, in one possible implementation of the first aspect, j search tree child nodes are scored to obtain j scoring results, including: scoring the j search tree child nodes using a process reward model (PRM) to obtain j scoring results. By dynamically using PRM to score the search tree child nodes, the corresponding scoring results can be dynamically obtained, and the corresponding similarity index can be dynamically acquired. This enables proactive awareness of the model's inference state, reduces the total number of tokens generated while keeping accuracy reduction controllable, and achieves reasonable compression of the search tree size. This, in turn, improves the processing speed of the inference device while ensuring the accuracy of the inference results.

[0009] In conjunction with the first aspect, in one possible implementation of the first aspect, determining the similarity index corresponding to the i-th round of search based on j scoring results includes: calculating the standard deviation based on the j scoring results to obtain the similarity index corresponding to the i-th round of search. Thus, by calculating the standard deviation, the corresponding similarity index can be obtained, reflecting the similarity of the reasoning content and improving the feasibility of the solution.

[0010] In conjunction with the first aspect, in one possible implementation of the first aspect, the first condition includes that the M similarity indicators corresponding to M consecutive rounds of searching continuously increase, and the i-th round of searching is the last round of searching in the M consecutive rounds of searching, and M is greater than or equal to 2. By setting a specific first condition, the size of the search tree can be compressed in a timely manner when the model continuously generates homogeneous content, saving computational power consumed on homogeneous content, reducing the computational load required for inference, and improving the processing speed of the inference device.

[0011] In conjunction with the first aspect, in one possible implementation of the first aspect, determining the number of expanded nodes corresponding to the (i+1)th round of search includes: determining that the number of expanded nodes corresponding to the (i+1)th round of search is k1 times the number of expanded nodes corresponding to the i-th round of search, where k1 is less than 1. By setting k1 in this way, a reasonable compression of the number of expanded nodes corresponding to the next round of search can be achieved, improving the feasibility of the solution.

[0012] In conjunction with the first aspect, in one possible implementation of the first aspect, the method further includes: determining a similarity index corresponding to the (i+n)th round of search, where the similarity index is used to characterize the similarity of the inference content generated by the model in the (i+n)th round of search, and n is greater than or equal to 1; and, if the similarity index corresponding to the (i+n)th round of search satisfies the second condition, determining that the number of expanded nodes corresponding to the (i+n+1)th round of search is k2 times the number of expanded nodes corresponding to the (i+n)th round of search, where k2 is less than k1. Thus, after compressing the search tree size once, if a large amount of homogeneous content continues to be generated, the search tree size will be further compressed, thereby further saving the computational power consumed on homogeneous content, further reducing the computational load required for inference, and further improving the processing speed of the inference device.

[0013] In conjunction with the first aspect, in one possible implementation of the first aspect, the second condition includes the continuous increase of (M+n) similarity indicators corresponding to (M+n) consecutive rounds of search, and the (i+n)th round of search being the last round of search in the (M+n) consecutive rounds of search, where M is the number of consecutive search rounds included in the first condition, and M is greater than or equal to 2. By setting a specific second condition, the size of the search tree can be further compressed in a timely manner while the model continues to generate a large amount of homogeneous content, further saving computational power consumed on homogeneous content, further reducing the computational load required for inference, and further improving the processing speed of the inference device.

[0014] Secondly, embodiments of this application provide a model processing apparatus, which includes a module for performing the model processing method in the first aspect or any optional embodiment of the first aspect.

[0015] Thirdly, embodiments of this application provide a model processing apparatus, the apparatus comprising:

[0016] The search tree expansion control module is used to execute the search tree algorithm based on the input question and perform the i-th round of search.

[0017] The similarity index calculation and perception module is used to determine the similarity index corresponding to the i-th round of search. The similarity index corresponding to the i-th round of search is used to characterize the similarity of the reasoning content generated by the model in the i-th round of search, where i is greater than or equal to 1.

[0018] The search tree expansion control module is also used to determine the number of expansion nodes corresponding to the (i+1)th round of search when the similarity index corresponding to the i-th round of search meets the first condition. The number of expansion nodes corresponding to the (i+1)th round of search is less than the number of expansion nodes corresponding to the i-th round of search.

[0019] Fourthly, embodiments of this application provide a model processing apparatus that can be applied to an inference device or circuits, chips, chip systems, etc. within the inference device. The apparatus may include at least one processor, which is used to call computer instructions in memory to cause the model processing apparatus to execute the model processing method in the first aspect or any optional embodiment of the first aspect.

[0020] In conjunction with the fourth aspect, in one possible implementation of the fourth aspect, the model processing apparatus may further include a memory.

[0021] Fifthly, embodiments of this application provide a computer-readable storage medium that may include instructions that, when executed on a computer, cause the computer to perform the model processing method in the first aspect or any optional implementation thereof.

[0022] In a sixth aspect, this application provides a computer program product containing instructions, which may include instructions that, when executed on a computer, cause the computer to perform the model processing method in the first aspect or any optional embodiment of the first aspect.

[0023] Seventhly, embodiments of this application provide a chip system including a processor for supporting a device in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing necessary program instructions and data for the device. This chip system may be composed of chips or may include chips and other discrete devices.

[0024] Eighthly, embodiments of this application provide a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, it causes the electronic device to perform the model processing method in the first aspect or any optional embodiment of the first aspect. Attached Figure Description

[0025] Figure 1 A structural diagram illustrating the main framework of artificial intelligence;

[0026] Figure 2a A schematic diagram of the structure of the question-answering system provided in the embodiments of this application;

[0027] Figure 2b Another schematic diagram of the question-and-answer system provided in the embodiments of this application;

[0028] Figure 2c A schematic diagram of the related equipment for model processing provided in the embodiments of this application;

[0029] Figure 3 A schematic diagram of the architecture of the question-answering system 100 provided in this application embodiment;

[0030] Figure 4 A schematic flowchart of the model processing method provided in the embodiments of this application;

[0031] Figure 5 Another flowchart illustrating the model processing method provided in the embodiments of this application;

[0032] Figure 6 A schematic diagram of a search tree provided in an embodiment of this application;

[0033] Figure 7 A schematic diagram illustrating an application example of the model processing method provided in this application embodiment;

[0034] Figure 8 A schematic diagram illustrating another application example of the model processing method provided in the embodiments of this application.

[0035] Figure 9 A schematic diagram of the model processing apparatus provided in the embodiments of this application;

[0036] Figure 10 This is a schematic diagram of the structure of an inference device provided in an embodiment of this application. Detailed Implementation

[0037] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0038] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. The term "at least one" should be understood as one or more, and "at least one" should be understood as one or more. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0039] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0040] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 This is a structural diagram illustrating the main framework of artificial intelligence. The following explanation of the AI ​​framework is based on two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.

[0041] (1) Infrastructure

[0042] Infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. This communication occurs through sensors; computing power is provided by intelligent chips (hardware acceleration chips such as CPUs, NPUs, GPUs, ASICs, and FPGAs); and the basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.

[0043] (2) Data

[0044] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0045] (3) Data processing

[0046] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.

[0047] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training of data by symbolizing and formalizing it.

[0048] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.

[0049] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.

[0050] (4) General ability

[0051] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0052] (5) Smart Products and Industry Applications

[0053] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.

[0054] The following sections introduce several application scenarios of embodiments of this application.

[0055] Figure 2a This is a schematic diagram of a question-and-answer system provided in an embodiment of this application. The system includes a user device and a data processing device. The user device includes smart terminals such as mobile phones, personal computers, or information processing centers. The user device is the input end for questions; as the initiator of the question-and-answer process, the user typically inputs a question through the user device to initiate a response request.

[0056] The aforementioned data processing equipment can be cloud servers, network servers, application servers, management servers, or other devices or servers with data processing capabilities. The data processing equipment receives response requests from smart terminals through an interactive interface, and then performs processing such as machine learning, deep learning, search, reasoning, and decision-making through a storage device and a data processing processor. The storage device can be a general term encompassing local storage and a database storing historical data; the database can reside on the data processing equipment or on other network servers.

[0057] exist Figure 2a In the question-and-answer system shown, the user device can receive instructions from the user. For example, the user device can obtain the question input by the user and then send a response request to the data processing device. This allows the data processing device to perform search, reasoning, and other processing on the question obtained by the user device, thereby obtaining the corresponding response result for the question.

[0058] exist Figure 2a In this context, the data processing device can execute the model processing method of the embodiments of this application.

[0059] Figure 2b This is another schematic diagram of the question-answering system provided in the embodiments of this application. Figure 2b In this context, the user equipment (UE) directly functions as a data processing device. This UE can directly acquire input from the user and process it directly through its own hardware. The specific process is similar to... Figure 2a Similar to the description above, it will not be repeated here.

[0060] exist Figure 2b In the question-and-answer system shown, the user device can receive instructions from the user. For example, the user device can obtain the question entered by the user and then perform search, reasoning and other processing on the question to obtain the corresponding answer.

[0061] exist Figure 2b In this context, the user equipment itself can execute the model processing method of the embodiments of this application.

[0062] Figure 2c This is a schematic diagram of a model processing device provided in an embodiment of this application.

[0063] The above Figure 2a and Figure 2b The user equipment in the context can specifically be Figure 2c Local device 301 or local device 302 in the system. Figure 2a The data processing equipment in the middle can specifically be Figure 2c The execution device 210 in the process includes a data storage system 250 that can store the data to be processed by the execution device 210. The data storage system 250 can be integrated into the execution device 210 or set up in the cloud or on other network servers.

[0064] Figure 2a and Figure 2b The processor in the system can perform data training / machine learning / deep learning through the model, and use the data to finally train or learn the model to perform search, reasoning and other processing on the problem, thereby obtaining the corresponding answer result.

[0065] Furthermore, this application also provides another question-answering system, which includes a data processing device (such as a cloud server, network server, application server, and management server, etc., devices or servers with data processing functions) and a user device. The data processing device trains a model and deploys the trained model on the user device, enabling the user device to execute the model processing method of this application embodiment.

[0066] Figure 3 A schematic diagram of the architecture of the question-answering system 100 provided in this application embodiment. Figure 3 In this embodiment, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140. The input data is a question in this application embodiment, and the question can be at least one of the following forms: text, voice, video, image, etc.

[0067] During the preprocessing of input data by the execution device 110, or during the calculation module 111 of the execution device 110 performing calculations and other related processing (such as implementing the functions of the model in this application), the execution device 110 may call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.

[0068] Finally, I / O interface 112 returns the response to client device 140, thereby providing it to the user.

[0069] It is worth noting that the training device 120 can generate corresponding target models / rules based on different training data for different objectives or tasks. These target models / rules can then be used to achieve the aforementioned objectives or complete the aforementioned tasks, thereby providing the user with the required results. The training data can be stored in the database 130 and originates from training samples collected by the data acquisition device 160.

[0070] exist Figure 3 In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 140 can automatically send input data to I / O interface 112. If user authorization is required for the client device 140 to automatically send input data, the user can set the corresponding permissions in the client device 140. The user can view the output results of the execution device 110 on the client device 140, which can be presented in various forms such as display, sound, or animation. The client device 140 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 140, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130.

[0071] It is worth noting that, Figure 3 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 3 In this context, the data storage system 150 is an external memory relative to the execution device 110. However, in other cases, the data storage system 150 can also be placed within the execution device 110. For example... Figure 3 As shown, the model can be trained using training device 120.

[0072] This application also provides a chip including a neural network processor (NPU). This chip can be configured as follows: Figure 3 The execution device 110 shown is used to perform the calculations of the calculation module 111. This chip can also be located in, for example... Figure 3 The training device 120 shown is used to complete the training work of the training device 120 and output the target model.

[0073] The Neural Processing Unit (NPU) is a coprocessor mounted on the main central processing unit (CPU) (host CPU), where tasks are assigned by the CPU. The core of the NPU is the computation circuitry, which is controlled by a controller to retrieve data from memory (weight memory or input memory) and perform calculations.

[0074] In some implementations, the arithmetic circuitry includes multiple process engines (PEs). In some implementations, the arithmetic circuitry is a two-dimensional pulsating array. The arithmetic circuitry can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuitry is a general-purpose matrix processor.

[0075] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory and caches it in each PE (Process Equipment) of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory and performs matrix operations with matrix B. The partial or final result of the obtained matrix is ​​stored in the accumulator.

[0076] Vector computation units can further process the output of computational circuits, such as vector multiplication, vector addition, exponentiation, logarithmic operations, size comparisons, etc. For example, vector computation units can be used for computation in non-convolutional / non-FC layers of neural networks, such as pooling, batch normalization, and local response normalization.

[0077] In some implementations, the vector computation unit can store the processed output vector into a unified buffer. For example, the vector computation unit can apply a nonlinear function to the output of the arithmetic circuit, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as activation input to the arithmetic circuit, for example, for use in subsequent layers of a neural network.

[0078] The unified memory is used to store input data and output data.

[0079] The weight data is directly transferred from the external memory to the input memory and / or unified memory, stored in the weight memory, and stored in the unified memory to the external memory through the direct memory access controller (DMAC).

[0080] The bus interface unit (BIU) is used to enable interaction between the main CPU, DMAC, and instruction fetch memory via a bus.

[0081] The instruction fetch buffer, connected to the controller, is used to store the instructions used by the controller.

[0082] The controller is used to invoke instructions cached in the memory to control the operation of the computing accelerator.

[0083] Generally, the unified memory, input memory, weight memory, and instruction fetch memory are all on-chip memories, while external memory is memory outside the NPU. This external memory can be double data rate synchronous dynamic random access memory (DDRSDRAM), high bandwidth memory (HBM), or other readable and writable memory.

[0084] To facilitate understanding, the relevant terms involved in the embodiments of this application will be introduced below.

[0085] (1) Search tree algorithm is a type of algorithm based on tree structure for searching. It is widely used in various scenarios where the optimal solution needs to be found in a complex search space.

[0086] Tree search algorithms can include Monte Carlo Tree Search (MCTS), beam search, and others. MCTS is a heuristic search algorithm based on stochastic simulation, primarily used for complex decision-making problems such as board games and path planning. MCTS combines classic tree search methods with Monte Carlo methods, optimizing decisions through random sampling and statistical analysis. Beam search is a heuristic graph search algorithm, typically used in systems with large solution spaces. Beam search reduces space consumption and improves time efficiency by pruning some low-quality nodes at each depth expansion step, retaining high-quality nodes.

[0087] (2) The process reward model (PRM) is a reward model used to evaluate and optimize the reasoning process. The core idea of ​​PRM is to provide fine-grained scoring for each step in the generation or reasoning process, rather than just evaluating the final result. PRM trains the model using labeled datasets at the step level, predicts the correctness of each intermediate step in the solution, identifies error locations in the reasoning process, provides more accurate feedback, and thus optimizes the entire reasoning process.

[0088] (3) The reward model is a core concept in reinforcement learning. It is used to evaluate the behavior of an agent in a specific state. By scoring the input questions and answers, it guides the model to generate outputs that are more in line with expectations and safety standards.

[0089] (4) The outcome reward model (ORM) is a reward model that focuses on the final quality of the generated result rather than the intermediate steps in the generation process. It guides the model to learn to generate more expected outputs by assigning reward scores to the final generated result.

[0090] (5) Large language model (LLM) refers to a deep learning model with a large number of parameters. It is usually built on a complex neural network architecture, capable of processing and generating large-scale data, and has strong data generalization ability and multi-task learning ability.

[0091] With the development of artificial intelligence, combining models with search schemes has become a hot research topic in academia and industry. By incorporating search schemes during model training, models can learn multi-step reasoning capabilities from synthetic data; by combining search schemes during inference, multi-round, multi-path reasoning can be achieved, providing more computation to significantly improve the model's reasoning ability. However, this "long thought chain, slow thinking" reasoning significantly increases the computational load required for reasoning, resulting in higher computational costs for inference devices and slower processing speeds. Moreover, multi-step reasoning in models suffers from problems such as repetitive thinking and redundant responses.

[0092] Currently, by introducing a formula into each round of MCTS to calculate nodes with higher success probabilities and allocating computing resources to these nodes, inference computational costs can be reduced to some extent. However, this method cannot leverage the characteristics of slow thinking to reduce computational costs, and it can only be optimized for the MCTS tree. Therefore, how to reasonably compress the search tree while maintaining model accuracy to reduce inference computational costs and improve the processing speed of inference devices is a pressing technical problem that needs to be solved.

[0093] To address this, this application discloses a model processing method, apparatus, storage medium, and program product. By determining the similarity of the inference content generated in each round of search, it determines whether the inference content generated in that round of search has a high degree of homogeneity and low information content. Combined with a first condition, it determines whether the data of the expansion nodes corresponding to the next round of search needs to be adjusted. If the first condition is met, the number of expansion nodes corresponding to the next round of search is reduced. This saves computing power consumed on homogeneous content without affecting the inference result, reducing the amount of computation required for inference and improving the processing speed of the inference device. Moreover, in the long-thinking, slow-thinking inference of the model, the dynamic similarity index corresponding to each round of search is used for condition judgment. This enables proactive perception of the model's inference state, reducing the total number of tokens generated while keeping the accuracy reduction controllable, and achieving reasonable compression of the search tree size. This can improve the processing speed of the inference device while ensuring the accuracy of the inference result.

[0094] For details, please refer to Figure 4 , Figure 4 This is a flowchart illustrating a model processing method provided for implementation of this application. This model processing method can be implemented using an inference device on which a model is deployed. Unless otherwise specified, the "inference device" in this application can refer to the inference device itself or a component within the inference device (e.g., a processor, chip, or chip system). The model processing method provided in this application embodiment may include:

[0095] S401. Execute the search tree algorithm based on the input question to perform the i-th round of search.

[0096] Where i is greater than or equal to 1, and i is a positive integer.

[0097] In this embodiment of the application, after the user inputs a question, the inference device can obtain the question accordingly, and based on the question, execute the search tree algorithm, call the model to perform the i-th round of search, generate the search tree child node corresponding to the i-th round of search, and thus obtain the inference content corresponding to the i-th round of search.

[0098] It should be noted that, in the embodiments of this application, "searching" refers to the process of exploring the nodes and paths of the search tree to find the optimal path or solution from the initial state to the target state.

[0099] The input question in this application embodiment may include at least one of the following: text, voice, image, video, etc. For example, inputting an image A and inputting "Please analyze the content in the image" together constitutes a question. It is understood that the above is only an illustrative example and should not be construed as a limitation on the embodiments of this application.

[0100] The search tree algorithm in this application embodiment may include: MCTS, beam search, etc., and this application embodiment does not limit it.

[0101] The models in this application embodiment may include: LLM, multimodal generative models, etc. Among them, LLM may include O1 series models, O3 series models, O1-like large models, O3-like large models, multimodal generative models, etc., and multimodal generative models may include autoregressive generative and diffusion models, etc., which are not limited in this application embodiment.

[0102] S402. Determine the similarity index corresponding to the i-th round of search.

[0103] The similarity index corresponding to the i-th round of search is used to characterize the similarity of the reasoning content generated by the model in the i-th round of search.

[0104] In this embodiment, the similarity index corresponding to the i-th round of search can be determined based on the child nodes of the search tree corresponding to the i-th round of search.

[0105] S403. If the similarity index corresponding to the i-th round of search meets the first condition, determine the number of expanded nodes corresponding to the (i+1)-th round of search.

[0106] The number of expanded nodes corresponding to the (i+1)th round of search is less than the number of expanded nodes corresponding to the i-th round of search.

[0107] For example, the number of expanded nodes corresponding to the second round of search is 2, and the number of expanded nodes corresponding to the third round of search is 1; or, the number of expanded nodes corresponding to the second round of search is 3, and the number of expanded nodes corresponding to the third round of search is 2, and so on. It is understood that the above is merely an illustrative example and should not be construed as a limitation on the embodiments of this application.

[0108] In this embodiment of the application, after determining the number of expanded nodes corresponding to the (i+1)th round of search, the number of expanded nodes corresponding to the (i+1)th round of search will be adjusted to reduce the number of expanded nodes generated in the (i+1)th round of search, thereby compressing the size of the search tree.

[0109] As can be seen, in this embodiment, by determining the similarity of the inference content generated in each round of search, it is determined whether the inference content generated in that round of search has a high degree of homogeneity and low information content. Combined with the first condition, it is determined whether the data of the expansion nodes corresponding to the next round of search needs to be adjusted. If the first condition is met, the number of expansion nodes corresponding to the next round of search is reduced. In this way, without affecting the inference result, the computing power consumed on homogeneous content is saved, which can reduce the amount of computation required for inference and improve the processing speed of the inference device. Moreover, in the long thinking chain and slow thinking inference of the model, the similarity index corresponding to each round of search is dynamically determined for condition judgment. In this way, the model's inference state can be actively perceived, and the total number of tokens generated can be reduced while the accuracy reduction is controllable. This achieves reasonable compression of the search tree size, which can improve the processing speed of the inference device while ensuring the accuracy of the inference result.

[0110] Please see Figure 5 , Figure 5 A flowchart illustrating another model processing method for implementing this application is provided. Figure 5 exist Figure 4 Based on the provided model processing methods, a more detailed explanation of the model processing methods will be given. Figure 5 The model processing methods can also be executed by the inference device. Figure 5 The definition of "inference device" in the illustrated embodiment is the same as described above. Figure 4 The definition of "inference device" in the illustrated embodiments is the same and will not be repeated here. The model processing method provided in the embodiments of this application may include:

[0111] S501. Problem of obtaining input.

[0112] In this embodiment, the user can input a question through the display interface, the inference device will obtain the input question, and the subsequent steps will be executed based on the input question.

[0113] S502. Execute the search tree algorithm based on the problem to perform the i-th round of search.

[0114] It is understood that S502 is the same as S401 in the above embodiments, so it will not be described again.

[0115] S503. Determine the similarity index corresponding to the i-th round of search.

[0116] It is understood that S503 is similar to S402 in the above embodiments, and the same parts will not be described again here.

[0117] In one possible implementation, in this embodiment, the i-th round of search generates j search tree child nodes, where j is greater than or equal to 1. Accordingly, S503 may include:

[0118] a1. Obtain the j child nodes of the search tree corresponding to the i-th round of search.

[0119] a2. Score the j child nodes of the search tree to obtain j scoring results.

[0120] In this embodiment, the PRM can be used to score j child nodes of the search tree, resulting in j scoring results. By using the PRM, the child nodes of the search tree generated in each round can be dynamically scored, allowing for the dynamic acquisition of corresponding scoring results and the dynamic generation of similarity metrics for each round of search. This enables proactive awareness of the model's inference state, reducing the total number of tokens generated while keeping accuracy reduction controllable, achieving reasonable compression of the search tree size, and thus improving the processing speed of the inference device while ensuring the accuracy of the inference results.

[0121] It is understood that other scoring models, scoring rules or methods may also be used to score the child nodes of the search tree in the embodiments of this application, such as using ORM to score the child nodes of the search tree, or any reward model to score the child nodes of the search tree, etc., and the embodiments of this application do not limit this.

[0122] It should be noted that, in the embodiments of this application, scoring j search tree child nodes to obtain j scoring results may include: scoring the j search child nodes generated in the i-th round of search to obtain j scoring results; or, scoring the paths between the j search child nodes and the nodes preceding the j search child nodes to obtain j scoring results.

[0123] Please see Figure 6 This is a schematic diagram of a search tree provided in an embodiment of this application, combined with... Figure 6 The scoring method of the embodiments of this application will be explained. Figure 6The child nodes of the search tree generated in the i-th round of search include c1, c2, c3, and c4. In this embodiment, c1, c2, c3, and c4 can be scored to obtain j scoring results; the paths b2-c1, b2-c2, b4-c3, and b4-c4 can also be scored to obtain j scoring results; the paths a1-b2-c1, a1-b2-c2, a1-b4-c3, and a1-b4-c4 can also be scored to obtain j scoring results; c1 combined with the path b2-c1, c2 combined with the path b2-c2, c3 combined with the path b4-c3, and c4 combined with the path b4-c4 can also be scored to obtain j scoring results; c1 combined with the path a1-b2-c1, c2 combined with the path a1-b2-c2, c3 combined with the path a1-b4-c3, and c4 combined with the path a1-b4-c4 can also be scored to obtain j scoring results, and so on. It is understood that the above is merely an illustrative example and should not be construed as a limitation on the embodiments of this application.

[0124] a3. Based on j scoring results, determine the similarity index corresponding to the i-th round of search.

[0125] In this embodiment, the standard deviation can be calculated based on j scoring results to obtain the similarity index corresponding to the i-th round of search. For example, in the second round of search, search tree child nodes 1, 2, and 3 are generated. Using PRM, the three search tree child nodes are scored, resulting in a score of 0.945 for child node 1, 0.942 for child node 2, and 0.951 for child node 3. The standard deviation of 0.945, 0.942, and 0.951 is calculated, yielding 0.0037. Based on the negative value of 0.0037, the similarity index corresponding to the second round of search is determined to be -0.0037, or 0.0037. It is understood that the above is merely an illustrative example and should not be construed as a limitation on the embodiments of this application. This standard deviation calculation more accurately reflects the similarity of the reasoning content, yields a more accurate similarity index, and improves the feasibility of the solution.

[0126] It is understood that, in the embodiments of this application, the standard deviation of a negative number can be used as the similarity index, or the standard deviation of a non-negative number can be used as the similarity index. Other methods can also be used to determine the similarity index, such as calculating the average of j scoring results to determine the similarity index corresponding to the i-th round of search; or comparing the size of j scoring results to determine the similarity index corresponding to the i-th round of search, etc. The embodiments of this application do not limit this.

[0127] S504. If the similarity index corresponding to the i-th round of search meets the first condition, determine that the number of expanded nodes corresponding to the (i+1)-th round of search is k1 times the number of expanded nodes corresponding to the i-th round of search.

[0128] It is understood that S504 in this embodiment is similar to S403 in the above embodiment, and the same parts will not be described again here.

[0129] In one possible implementation, the first condition may include M similarity metrics corresponding to M consecutive rounds of searching that continuously increase, and the i-th round of searching being the last round of searching in the M consecutive rounds, where M is greater than or equal to 2. Under this first condition, the standard deviation of a negative number can be used as the similarity metric. By setting a specific first condition, the size of the search tree can be compressed in a timely manner when the model continuously generates homogeneous content, saving computational power consumed on homogeneous content, reducing the amount of computation required for inference, and improving the processing speed of the inference device.

[0130] In another possible implementation, the first condition may include a continuous decrease in the M similarity indices corresponding to M consecutive search rounds, where the i-th search round is the last of the M consecutive search rounds, and M is greater than or equal to 2. Under this first condition, the standard deviation (not negative) can be used as the similarity index. By setting a specific first condition, the size of the search tree can be compressed in a timely manner when the model continuously generates homogeneous content, saving computational power consumed on homogeneous content, reducing the computational load required for inference, and improving the processing speed of the inference device.

[0131] It is understood that the first condition in this application embodiment can also be other. For example, the first condition can also be that the P1 similarity indicators corresponding to the first P1 rounds of search in the continuous P rounds of search continuously increase, and the P2 similarity indicators corresponding to the last P2 rounds of search in the continuous P rounds of search continuously decrease, and the i-th round of search is the last round of search in the continuous P rounds of search, where P is greater than or equal to 3, the sum of P1 and P2 equals P, P1 is greater than or equal to 1, and P2 is greater than or equal to 1, that is, the similarity indicators first increase and then decrease; or, the first condition can also be that the similarity indicator corresponding to the i-th round of search is lower than the first score threshold; or, the first condition can also be that the similarity indicator corresponding to the i-th round of search is higher than the second score threshold, etc. This application embodiment does not limit this, as long as it is ensured that the similarity indicators correspond to the first condition, and the first condition is satisfied under the condition of continuously generating homogeneous content.

[0132] In this embodiment, k1 is less than 1, and the number of expanded nodes corresponding to the (i+1)th round of search is less than the number of expanded nodes corresponding to the i-th round of search. By setting k1 in this way, the number of expanded nodes corresponding to the next round of search can be reasonably compressed, improving the feasibility of the solution.

[0133] In this embodiment of the application, after determining the number of expanded nodes corresponding to the (i+1)th round of search, the number of expanded nodes corresponding to the (i+1)th round of search will be adjusted to k1 times the number of expanded nodes corresponding to the i-th round of search, thereby reducing the number of expanded nodes generated in the (i+1)th round of search and compressing the size of the search tree.

[0134] S505. Determine the similarity index corresponding to the (i+n)th round of search.

[0135] The similarity index corresponding to the (i+n)th search round is used to characterize the similarity of the reasoning content generated by the model in the (i+n)th search round, where n is greater than or equal to 1.

[0136] It is understood that the method for determining the similarity index corresponding to the (i+n)th round of search in this embodiment is similar to the method for determining the similarity index corresponding to the i-th round of search in S503, and will not be repeated here.

[0137] S506. If the similarity index corresponding to the (i+n)th round of search meets the second condition, determine that the number of expanded nodes corresponding to the (i+n+1)th round of search is k2 times the number of expanded nodes corresponding to the (i+n)th round of search.

[0138] The second condition includes a continuous increase in (M+n) similarity metrics across (M+n) consecutive search rounds, with the (i+n)th search round being the last of the (M+n) consecutive search rounds. M represents the number of consecutive search rounds included in the first condition, and M is greater than or equal to 2. It can be understood that M in the second condition is the same as M in the first condition of S504. By setting this specific condition, the size of the search tree can be further compressed in a timely manner when the model continues to generate a large amount of homogeneous content, further saving computational resources consumed on homogeneous content, further reducing the computational load required for inference, and further improving the processing speed of the inference device.

[0139] In the embodiments of this application, k2 is less than k1. For example, k1 is 0.5 and k2 is 0.4; or k1 is 0.6 and k2 is 0.5, etc. It should be understood that the above is only an exemplary description and should not be construed as a limitation on the embodiments of this application.

[0140] It is understood that the setting method of the second condition in the embodiments of this application is similar to the setting method of the first condition, and will not be described again here.

[0141] In this embodiment of the application, after determining the number of expanded nodes corresponding to the (i+n+1)th round of search, the number of expanded nodes corresponding to the (i+n+1)th round of search will be adjusted to k2 times the number of expanded nodes corresponding to the (i+n)th round of search, thereby reducing the number of expanded nodes generated in the (i+n+1)th round of search and compressing the size of the search tree.

[0142] S507. If the termination condition is met, output the answer to the question.

[0143] The termination conditions in the embodiments of this application may include at least one of the following: completing a predetermined search round, the reasoning results of several consecutive search rounds not showing significant improvement, or the reasoning results starting to repeat, etc.

[0144] The response can be at least one of the following: text, voice, image, video, etc. For example, if the question is "Please output an image of a spring field," the response will be an image. It is understood that the above is merely illustrative and should not be construed as limiting the embodiments of this application.

[0145] As can be seen, in this embodiment, by determining the similarity of the inference content generated in each round of search, it is determined whether the inference content generated in that round of search has a high degree of homogeneity and low information content. Combined with the first condition or the second condition, it is determined whether the data of the expansion nodes corresponding to the next round of search needs to be adjusted. If the first condition or the second condition is met, the number of expansion nodes corresponding to the next round of search is reduced. In this way, without affecting the inference result, the computing power consumed on homogeneous content is saved, which can reduce the amount of computation required for inference and improve the processing speed of the inference device. Moreover, in the long thinking chain and slow thinking inference of the model, the similarity index corresponding to each round of search is dynamically determined for condition judgment. In this way, the model's inference state can be actively perceived, and the total number of tokens generated can be reduced while the reduction in accuracy is controllable. This achieves reasonable compression of the search tree size, which can improve the processing speed of the inference device while ensuring the accuracy of the inference result.

[0146] Please see Figure 7 , Figure 8 This is a schematic diagram illustrating an application example of the model processing method provided in the embodiments of this application. Figure 7 , Figure 8 exist Figure 4 , Figure 5 Based on the provided model processing methods, examples are given to illustrate these methods. The following section combines... Figure 7 and Figure 8The method of using the similarity index calculation perception module, search tree trend judgment module and search tree expansion control module in the inference device to execute model processing is explained. In this application example, the standard deviation of the negative number is used as the similarity index. The first condition includes that the M similarity indices corresponding to the M consecutive rounds of search continue to rise, and the i-th round of search is the last round of search in the M consecutive rounds of search, where M is greater than or equal to 2.

[0147] S1: In the model inference task, given an LLM, a question is input and the search tree algorithm is executed; in each round of search, the child nodes of the generated search tree are scored by PRM to obtain the corresponding scoring results, that is, the corresponding PRM score; the similarity index calculation perception module receives the scoring results and calculates the similarity index of the content generated in the current search round of the search tree.

[0148] It should be noted that statistical analysis revealed that the standard deviation of the PRM scores of all child nodes of the search tree generated in the current search round during the slow thinking process can reflect the overall similarity of the inference content generated by the model in the current search round. Based on this finding, this application's embodiments summarize and propose a similarity index calculation perception module, which calculates the negative of the standard deviation of the PRM score corresponding to each search round, and uses this as a similarity index. The larger the similarity index, the more homogeneous or similar the inferred content is.

[0149] Combination Figure 8 To explain, Figure 8 In the current search round, four search tree child nodes are generated. The PRM scores corresponding to the four search tree child nodes are 0.945, 0.942, 0.951, and 0.948, respectively, and the similarity index is -0.00335. It should be understood that the above is only an illustrative example and should not be construed as a limitation on the embodiments of this application.

[0150] S2: During the search tree reasoning process, the search tree trend judgment module continuously obtains similarity indicators from the similarity indicator calculation and perception module, and judges the trend of similarity indicator changes. When the similarity indicator of the most recent M rounds of search continues to rise, the search tree expansion control module actively initiates the adjustment of the number of expansion nodes in the next round.

[0151] It should be noted that statistical analysis revealed that the similarity index exhibits a trend of high initial value, continuous decline in the middle stage, and continuous increase in the later stage during the slow thinking process. If the similarity index declines for M consecutive rounds, it can generally be considered to be in the early stage of long thinking chain search; if the similarity index rises for M consecutive rounds, it is highly likely to be in the middle to late stage. Based on this finding, this application's embodiments summarize and propose a search tree trend judgment module.

[0152] S3: When the similarity index of the search tree trend judgment module continues to rise in the most recent M rounds of search, the search tree expansion control module calculates the new search tree expansion width by combining hyperparameter settings, obtains the number of expansion nodes, and compresses the number of search tree node expansions in the next round of search, thus completing the optimization and compression of the search tree.

[0153] It should be noted that statistical analysis revealed that when the similarity index continues to rise, the model tends to generate relatively homogeneous content, characterized by high similarity and low information content. If the search tree trend judgment module indicates a continuous increase in the similarity index, it instructs the search tree algorithm to reduce the number of expansion nodes in the next round to k1 times the original number, thus saving computational resources consumed on homogeneous content without affecting the inference results.

[0154] Combination Figure 7 To explain, Figure 7 The number of expanded nodes in the current search round is 2, meaning that one child node in the search tree from the previous round will be expanded into two child nodes in the search tree. Figure 7 The expansion node for the next round of search is 1, meaning that a child node in the search tree in the current search round will not be expanded; instead, only one child node will be generated. It is understood that the above is merely an illustrative description and should not be construed as a limitation on the embodiments of this application.

[0155] Combination Figure 7 To explain, Figure 7 In this model, X is 2, Y is 2, and k1 is 0.5. For example, using a search tree algorithm, X nodes are selected in each round. Based on the hyperparameter settings, each node calls the large model to infer one step forward, resulting in Y*X search tree child nodes, where Y is the number of expanded nodes, i.e., the number of search tree child nodes expanded by each node. The search tree algorithm may infer multiple steps forward for new nodes, but this embodiment only considers the similarity of the newly generated search tree child nodes, not the simulated inference after the newly generated search tree child nodes. The PRM is called to score the newly generated search tree child nodes, obtaining the score [s_1,s_2,…,s_Y*X] of the currently newly generated search tree child nodes, and the negative of the standard deviation d of the sequence [s_1,s_2,…,s_Y*X] is calculated. The list is initialized, and the similarity index d of each round is recorded, obtaining the list [d_1,d_2,…,d_i], where i is the current inference tree depth, i.e., the number of inference rounds. If [d_(i-M+1), d_(i-M+2), ..., d_i] is an increasing sequence, the number of expanded nodes is modified. The expansion amount Y in the current round is obtained, and the expansion amount in the next round is modified to Y*k1. It should be understood that the above is only an exemplary illustration and should not be construed as a limitation on the embodiments of this application.

[0156] It should be noted that, compared with the model processing method without the present application embodiment, when tested on the MATH500 math problem dataset, with M=2 and k1=0.5, the total inference generation amount (equivalent inference computing power overhead) can be compressed by about 14% without loss of accuracy, while maintaining the same accuracy. Moreover, the large model can maintain the ability to accurately answer user questions during multi-step inference, and at the same time optimize the amount of tokens generated for inference. The number of characters generated by the large model in the later stages of inference is significantly compressed.

[0157] The model processing method provided in this application can be used for all models and systems that perform thought chain reasoning based on search trees, and also for all computing power systems that support such models and systems for reasoning, such as graphics cards and CPUs. It is understood that the model processing method provided in this application is based on current large model designs and can also be applied to similar generative scenarios, such as multimodal large models, multimodal generative models, etc. The similarity-based pruning scheme can be applied to any scenario where the similarity of nodes in the search tree or other metrics can be evaluated. Furthermore, the model processing method provided in this application can be used as a reasoning acceleration plugin within a large model reasoning framework, allowing users to achieve accelerated reasoning functionality for large model search trees through this software plugin.

[0158] It should be noted that the similarity index calculation formula involved in the embodiments of this application, the trend judgment parameter such as M, and the adjustment parameter for the number of expanded nodes such as k1, can be adjusted by implementing the model processing method of this application. The performance of the large model in multi-step response can be observed from statistical analysis. Alternatively, the same task can be input to observe the pattern of the large model's generated results, such as the efficiency of each round of reasoning in the generation process, the length of the generated sequence, and whether there are obvious characteristics in the pattern of the large model's generated results, such as the correlation between the reduction of similarity and the adjustment of the number of expanded nodes in the search tree. When the reasoning tree structure can be returned, it can also be observed whether the reasoning tree structure has been pruned and compressed, and whether there is a correlation with the similarity index.

[0159] In this embodiment of the application, during the large-scale model inference process, statistical analysis experience is combined to dynamically use score evaluation models, such as PRM and ORM, to quantify the similarity of the content generated by multiple nodes, rather than using a fixed static threshold. Compared to using a static threshold, the model processing method of this embodiment combines the characteristics of similarity during the large-scale model inference process for conditional judgment, which can proactively perceive the model's inference state and ensure model accuracy. Moreover, the calculation of similarity indicators can be directly based on the process scores required by the search algorithm, which not only reduces computational overhead but also has a good ability to reflect the model's inference state in statistical analysis, ensuring model accuracy. In addition, the model processing method of this embodiment does not require training and fine-tuning, the original model performance remains unchanged, and the implementation cost is low.

[0160] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0161] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.

[0162] Please see Figure 9 This is a schematic diagram of the structure of a model processing device provided in an embodiment of this application. The model processing device 900 includes: a similarity index calculation and perception module 901, a search tree trend judgment module 902, and a search tree expansion control module 903.

[0163] The search tree expansion control module 903 is used to execute the search tree algorithm based on the input question and perform the i-th round of search.

[0164] The similarity index calculation perception module 901 is used to determine the similarity index corresponding to the i-th round of search. The similarity index corresponding to the i-th round of search is used to characterize the similarity of the reasoning content generated by the model in the i-th round of search, where i is greater than or equal to 1.

[0165] The search tree trend judgment module 902 is used to determine whether the similarity index corresponding to the i-th round of search meets the first condition;

[0166] The search tree expansion control module 903 is also used to determine the number of expansion nodes corresponding to the (i+1)th round of search when the similarity index corresponding to the i-th round of search meets the first condition, wherein the number of expansion nodes corresponding to the (i+1)th round of search is less than the number of expansion nodes corresponding to the i-th round of search.

[0167] In some possible implementations, in the model processing device 900 provided in this application embodiment, the similarity index calculation and perception module 901 is specifically used to obtain the j search tree child nodes corresponding to the i-th round of search; score the j search tree child nodes to obtain j scoring results, where j is greater than or equal to 1; and determine the similarity index corresponding to the i-th round of search based on the j scoring results.

[0168] In some possible implementations, the similarity index calculation and perception module 901 in the model processing device 900 provided in this application embodiment is specifically used to score j search tree child nodes through a process reward model to obtain j scoring results.

[0169] In some possible implementations, the similarity index calculation and perception module 901 in the model processing device 900 provided in this application embodiment is specifically used to calculate the standard deviation based on j scoring results to obtain the similarity index corresponding to the i-th round of search.

[0170] In some possible implementations, in the model processing device 900 provided in this application embodiment, the first condition includes that the M similarity indicators corresponding to the M consecutive rounds of search continuously increase, and the i-th round of search is the last round of search in the M consecutive rounds of search, and M is greater than or equal to 2.

[0171] In some possible implementations, the search tree expansion control module 903 in the model processing device 900 provided in this application embodiment is further used to determine that the number of expansion nodes corresponding to the (i+1)th round of search is k1 times the number of expansion nodes corresponding to the i-th round of search, where k1 is less than 1.

[0172] In some possible implementations, in the model processing device 900 provided in this application embodiment, the similarity index calculation and perception module 901 is further used to determine the similarity index corresponding to the (i+n)th round of search. The similarity index corresponding to the (i+n)th round of search is used to characterize the similarity of the reasoning content generated by the model in the (i+n)th round of search, where n is greater than or equal to 1.

[0173] The search tree trend judgment module 902 is used to determine whether the similarity index corresponding to the (i+n)th round of search meets the second condition;

[0174] The search tree expansion control module 903 is also used to determine, when the similarity index corresponding to the (i+n)th round of search meets the second condition, that the number of expansion nodes corresponding to the (i+n+1)th round of search is k2 times the number of expansion nodes corresponding to the (i+n)th round of search, where k2 is less than k1.

[0175] In some possible implementations, in the model processing apparatus 900 provided in this application embodiment, the second condition includes the continuous increase of (M+n) similarity indicators corresponding to (M+n) consecutive rounds of search, and the (i+n)th round of search is the last round of search in the (M+n) consecutive rounds of search, where M is the number of consecutive search rounds included in the first condition, and M is greater than or equal to 2.

[0176] It should be noted that the information interaction and execution process between the modules of the above-mentioned device are based on the same concept as the method embodiment of this application, and the resulting technical effects are the same as those of the method embodiment of this application. For details, please refer to the description in the method embodiment shown above in this application, and it will not be repeated here.

[0177] The following describes an inference device provided in an embodiment of this application. Please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram of the structure of an inference device provided in an embodiment of this application. The inference device 1000 may include one or more processors 1001 and memory 1005.

[0178] Memory 1005 may include read-only memory and random access memory, and provides instructions and data to processor 1001. A portion of memory 1005 may also include non-volatile random access memory (NVRAM). Memory 1005 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0179] The processor 1001 controls the operation of the device. In specific applications, the various components of the inference device 1000 are coupled together through a bus system, which may include not only the data bus but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.

[0180] The model processing method disclosed in the above embodiments of this application can be applied to processor 1001, or implemented by processor 1001. Processor 1001 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 1001 or by instructions in the form of software. The processor 1001 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1001 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1005. Processor 1001 reads the information in memory 1005 and, in conjunction with its hardware, completes the aforementioned knowledge-based decoding steps.

[0181] The inference device 1000 may also include one or more power supplies 1002, one or more wired or wireless network interfaces 1003, one or more input / output interfaces 1004, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0182] In this embodiment of the application, the processor 1001 is used to execute... Figure 4 or Figure 5 The corresponding embodiment describes the model processing method. It should be noted that the specific manner in which the processor 1001 executes the aforementioned steps differs from that in this application. Figures 4 to 8 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 4 to 8 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0183] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the aforementioned actions. Figure 4 or Figure 5 The model processing method described in the illustrated embodiment.

[0184] This application also provides a computer program product, which includes a program that, when run on a computer, causes the computer to perform the aforementioned actions. Figure 4 or Figure 5 The model processing method described in the illustrated embodiment.

[0185] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.

[0186] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CLUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0188] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0189] The aforementioned computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A model processing method, characterized in that, The method includes: Based on the input question, a search tree algorithm is executed to perform the i-th round of search, where i is greater than or equal to 1; Determine the similarity index corresponding to the i-th round of search, which is used to characterize the similarity of the reasoning content generated by the model in the i-th round of search; If the similarity index corresponding to the i-th round of search meets the first condition, the number of expanded nodes corresponding to the (i+1)-th round of search is determined, wherein the number of expanded nodes corresponding to the (i+1)-th round of search is less than the number of expanded nodes corresponding to the i-th round of search.

2. The method according to claim 1, characterized in that, The determination of the similarity index corresponding to the i-th round of search includes: Obtain the j child nodes of the search tree corresponding to the i-th round of search; The j child nodes of the search tree are scored to obtain j scoring results, where j is greater than or equal to 1; Based on the j scoring results, the similarity index corresponding to the i-th round of search is determined.

3. The method according to claim 2, characterized in that, The scoring of the j child nodes of the search tree yields j scoring results, including: The process reward model is used to score the j child nodes of the search tree, resulting in j scoring results.

4. The method according to any one of claims 2 to 3, characterized in that, The step of determining the similarity index corresponding to the i-th round of search based on the j scoring results includes: The standard deviation is calculated based on the j scoring results to obtain the similarity index corresponding to the i-th round of search.

5. The method according to any one of claims 1 to 4, characterized in that, The first condition includes M similarity indicators corresponding to M consecutive rounds of search continuously increasing, and the i-th round of search being the last round of search in the M consecutive rounds of search, where M is greater than or equal to 2.

6. The method according to any one of claims 1 to 5, characterized in that, Determining the number of expanded nodes corresponding to the (i+1)th round of search includes: The number of expanded nodes corresponding to the (i+1)th round of search is determined to be k1 times the number of expanded nodes corresponding to the i-th round of search, where k1 is less than 1.

7. The method according to claim 6, characterized in that, The method further includes: Determine the similarity index corresponding to the (i+n)th search round, wherein the similarity index corresponding to the (i+n)th search round is used to characterize the similarity of the reasoning content generated by the model in the (i+n)th search round, wherein n is greater than or equal to 1; If the similarity index corresponding to the (i+n)th round of search meets the second condition, the number of expanded nodes corresponding to the (i+n+1)th round of search is determined to be k2 times the number of expanded nodes corresponding to the (i+n)th round of search, where k2 is less than k1.

8. The method according to claim 7, characterized in that, The second condition includes the continuous increase of (M+n) similarity indicators corresponding to (M+n) consecutive rounds of search, and the (i+n)th round of search is the last round of search in the (M+n) consecutive rounds of search, where M is the number of consecutive search rounds included in the first condition, and M is greater than or equal to 2.

9. A model processing device, characterized in that, The apparatus includes a module for performing the model processing method as described in any one of claims 1 to 8.

10. A model processing apparatus comprising at least one processor, the processor being configured to invoke computer instructions in memory to cause the model processing apparatus to perform the model processing method as claimed in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the model processing method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the model processing method as described in any one of claims 1 to 8.

13. A chip, characterized in that, The device includes one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of the electronic device and send the signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the model processing method as described in any one of claims 1 to 8.