Hybrid bonding and speculative sampling based large model inference system and method
Patent Information
- Application Number
- CN202611019046.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-25
AI Technical Summary
然而,上述方式虽然能够从算法上加速推理,但是并未考虑大模型推理系统的实际硬件情况,导致大模型推理速度依旧较慢
[0018]本发明提供的基于混合键合与投机采样的大模型推理系统及方法,草稿模型执行的解码过程是内存带宽密集型任务,目标模型执行的预填充过程、验证过程是算力密集型及高内存容量任务;将草稿模型部署在混合键合堆叠处理单元上,能够充分利用混合键合堆叠处理单元的高内存带宽的特性;将目标模型部署在异构加速器处理单元上,能够充分利用异构加速器处理单元的高内存容量及高算力的特性,从而将模型算法的特性与硬件能力特性相结合,能够提高大模型的推理速度。
Smart Images

Figure CN122819477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of online reasoning services for large language models, and in particular to a large model reasoning system and method based on hybrid bonding and speculative sampling. Background Technology
[0002] LLM (Large Language Model), often simply called a large model, is widely used in latency-sensitive scenarios such as dialogue, content generation, and question answering. To improve the efficiency of large model inference, speculative sampling techniques are commonly used to accelerate inference. This involves deploying a lightweight draft model and a standard-sized target model within a single hardware setup of the large model inference system. The draft model generates tokens, and the target model validates these tokens, thus increasing inference speed. However, while this approach algorithmically speeds up inference, it doesn't consider the actual hardware limitations of the large model inference system, resulting in a still relatively slow inference speed. Summary of the Invention
[0003] This invention provides a large model inference system and method based on hybrid bonding and speculative sampling to improve the inference speed of large models.
[0004] This invention provides a large-model inference system based on hybrid bonding and speculative sampling, comprising: The system includes a hybrid bonded stack processing unit and a heterogeneous accelerator processing unit. The hybrid bonded stack processing unit employs a hybrid bonded stack structure. The memory bandwidth of the hybrid bonded stack processing unit is greater than the memory bandwidth of the heterogeneous accelerator processing unit, and the memory capacity of the heterogeneous accelerator processing unit is greater than the memory capacity of the hybrid bonded stack processing unit. A draft model runs in the hybrid bonded stack processing unit, and a target model runs in the heterogeneous accelerator processing unit. The heterogeneous accelerator processing unit is used to call the target model to pre-fill the currently acquired task request to obtain an initial processing result; generate a decoding task based on the initial processing result and send the decoding task to the hybrid bonding stack processing unit; call the target model to verify the candidate word sequence of the currently acquired verification task to obtain a sequence verification result; and send the sequence verification result to the hybrid bonding stack processing unit. The hybrid bonding stack processing unit is used to call the draft model to generate candidate lexical sequences for the currently acquired decoding task; generate a verification task based on the candidate lexical sequences, and send the verification task to the heterogeneous accelerator processing unit; and perform the next round of decoding task for the corresponding task request or end the decoding task for the corresponding task request based on the verification result of the currently acquired sequence.
[0005] According to the present invention, a large model inference system based on hybrid bonding and speculative sampling is provided, wherein the heterogeneous accelerator processing unit is provided with a first task pool and the hybrid bonding stacking processing unit is provided with a second task pool. The heterogeneous accelerator processing unit is further configured to add the received task requests and verification tasks to the first task pool; and, when the idle resources of the heterogeneous accelerator processing unit meet the first preset resource idle condition, to obtain at least one task request and / or verification task from the first task pool. The hybrid bonding stack processing unit is further configured to add the received decoding task and sequence verification result to the second task pool; and, when the idle resources of the heterogeneous accelerator processing unit meet the second preset resource idle condition, to obtain at least one decoding task and / or sequence verification result from the second task pool.
[0006] According to the present invention, in a large model inference system based on hybrid bonding and speculative sampling, the processing priority of the task request is higher than that of the verification task, and the heterogeneous accelerator processing unit processes the higher priority task first.
[0007] According to the present invention, a large model inference system based on hybrid bonding and speculative sampling is provided, wherein the first task pool includes a request task pool and a verification task pool; the request task pool is used to cache task requests, and the verification task pool is used to cache verification tasks. Specifically, when the idle resources of the heterogeneous accelerator processing unit meet the first preset resource idle condition, the task request is preferentially obtained from the request task pool according to the preset data volume; if there is no task request in the request task pool or the data volume of the task request does not meet the preset data volume requirement, the verification task is then obtained from the verification task pool.
[0008] According to the present invention, a large model inference system based on hybrid bonding and speculative sampling is provided. The heterogeneous accelerator processing unit is specifically used to split the currently acquired task request into multiple task blocks according to the preset length threshold when the length of the currently acquired task request exceeds the preset length threshold, wherein the length of each task block does not exceed the preset length threshold; and to call the target model to pre-fill each task block to obtain an initial processing result.
[0009] According to a large model inference system based on hybrid bonding and speculative sampling provided by the present invention, the heterogeneous accelerator processing unit is further configured to obtain a first proportion by acquiring the ratio of the amount of data in its own memory key-value cache to the total amount of data in memory; if the first proportion is greater than the first proportion threshold, the task acquisition request is suspended. The hybrid bonding stacking processing unit is also used to obtain a second percentage by measuring the ratio of the amount of key-value cached data in its own memory to the total amount of memory data; if the second percentage is greater than the second percentage threshold, the heterogeneous accelerator processing unit is notified to suspend the acquisition of task requests.
[0010] According to the present invention, a large model inference system based on hybrid bonding and speculative sampling is provided, wherein the hybrid bonding stacking processing unit is further configured to adjust the draft generation budget and tree width according to the first computational intensity of the hybrid bonding stacking processing unit and the second computational intensity of the heterogeneous accelerator processing unit.
[0011] According to the present invention, a large-model inference system based on hybrid bonding and speculative sampling is provided. The hybrid bonding stacking processing unit is specifically configured to: query a preset lookup table to determine the optimal tree width corresponding to the current draft generation budget, wherein the preset lookup table records the correspondence between the draft generation budget and the optimal tree width; obtain its current computational intensity to obtain a first computational intensity; if the first computational intensity is less than a first preset intensity threshold, increase the current tree width; if the first computational intensity is greater than the first preset intensity threshold, decrease the current tree width, wherein the lower limit of the tree width is a first value, and the upper limit of the tree width is the optimal tree width corresponding to the current draft generation budget; call the draft model to generate a candidate lexical sequence for the currently acquired decoding task according to the current tree width; add the current candidate lexical sequence to the draft tree; if the budget of the draft tree does not reach the current draft generation budget, return to obtain its current computational intensity to obtain the first computational intensity; if the budget of the draft tree reaches the current draft generation budget, generate a verification task based on the draft tree.
[0012] According to a large model inference system based on hybrid bonding and speculative sampling provided by the present invention, the heterogeneous accelerator processing unit is further configured to determine its current load to obtain a second load; and send the second load to the hybrid bonding stack processing unit. The hybrid bonding stack processing unit is specifically used to increase the current draft generation budget if the second load is less than the second preset load threshold, and decrease the current draft generation budget if the second load is greater than the second preset load threshold, wherein the lower limit of the draft generation budget is the second value and the upper limit of the draft generation budget is the third value.
[0013] According to the present invention, a large model inference system based on hybrid bonding and speculative sampling is provided, wherein the hybrid bonding stacked processing unit adopts a hybrid bonding stacked structure of a single-layer dynamic random access memory layer and a single-layer logic chip layer.
[0014] According to the present invention, a large model inference system based on hybrid bonding and speculative sampling is provided. The hybrid bonding stacking processing unit is specifically used to terminate the decoding task of the corresponding task request when the currently acquired sequence verification result indicates that the end symbol is correct; and to perform the next round of decoding task according to the currently acquired sequence verification result when the currently acquired sequence verification result does not have an end symbol or indicates that the end symbol is incorrect; wherein, the corresponding task request is the task request corresponding to the currently acquired sequence verification result.
[0015] This invention also provides a large model inference method based on hybrid bonding and speculative sampling, executed by any of the large model inference systems based on hybrid bonding and speculative sampling described in this invention, the method comprising: The heterogeneous accelerator processing unit calls the target model to pre-fill the currently acquired task request to obtain an initial processing result; and generates a decoding task based on the initial processing result. The hybrid bonding stacking processing unit calls the draft model to generate a candidate lexical sequence for the currently acquired decoding task; and generates a verification task based on the candidate lexical sequence. The heterogeneous accelerator processing unit calls the target model to verify the candidate word sequence of the currently acquired verification task and obtains the sequence verification result; The hybrid bonding stacking processing unit performs the next round of decoding task for the corresponding task request or terminates the decoding task for the corresponding task request based on the currently acquired sequence verification result.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the large model inference methods based on hybrid bonding and speculative sampling described in the present invention.
[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the large model inference methods based on hybrid bonding and speculative sampling described in the present invention.
[0018] The large model inference system and method based on hybrid bonding and speculative sampling provided by this invention have the following characteristics: the decoding process performed by the draft model is a memory bandwidth intensive task, while the pre-filling and verification processes performed by the target model are computationally intensive and memory-intensive tasks. Deploying the draft model on a hybrid bonding stacked processing unit can fully utilize the high memory bandwidth of the hybrid bonding stacked processing unit; deploying the target model on a heterogeneous accelerator processing unit can fully utilize the high memory capacity and high computational power of the heterogeneous accelerator processing unit. Thus, the characteristics of the model algorithm are combined with the characteristics of the hardware capabilities, which can improve the inference speed of large models. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the large model inference system based on hybrid bonding and speculative sampling provided by the present invention; Figure 2 This is a schematic diagram illustrating the reasoning speed of the large language model in different scenarios provided by this invention; Figure 3 This is a flowchart illustrating the large model inference method based on hybrid bonding and speculative sampling provided by the present invention; Figure 4 This is a flowchart illustrating the method for adjusting tree width and draft generation budget provided by the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0022] LLM autoregressive inference has two core hardware requirements: high storage bandwidth to support rapid token generation and large-capacity storage to hold massive model parameters and a continuously growing KV (Key-Value) cache. Related technologies employ speculative sampling, deploying the draft model and target model on a single hardware platform. A small, lightweight draft model quickly generates candidate tokens, which are then verified in parallel by the target model, thus accelerating inference while ensuring equivalent output distribution. However, this technology suffers from the following problems: homogeneous hardware cannot adapt to the different bandwidth and capacity requirements of the draft and target models, inevitably leading to resource mismatch; and in dynamic online service scenarios, there are no effective solutions for issues such as batch processing scheduling, resource contention, and memory overflow.
[0023] To address at least one of the aforementioned problems, this invention provides a large-model inference system and method based on hybrid bonding and speculative sampling, which will be discussed below. Figures 1 to 4 Please provide a detailed explanation.
[0024] Figure 1This is an illustration of a large-model inference system based on hybrid bonding and speculative sampling provided by the present invention, such as... Figure 1 As shown, the system includes: The system includes a hybrid bond stacking processing unit 101 and a heterogeneous accelerator processing unit 102. The hybrid bond stacking processing unit 101 employs a hybrid bond stacking structure. The memory bandwidth of the hybrid bond stacking processing unit 101 is greater than the memory bandwidth of the heterogeneous accelerator processing unit 102, and the memory capacity of the heterogeneous accelerator processing unit 102 is greater than the memory capacity of the hybrid bond stacking processing unit 101. A draft model runs in the hybrid bond stacking processing unit 101, and a target model runs in the heterogeneous accelerator processing unit 102. The heterogeneous accelerator processing unit 102 is used to call the target model to pre-fill the currently acquired task request to obtain an initial processing result; generate a decoding task based on the initial processing result and send the decoding task to the hybrid bonding stack processing unit 101; call the target model to verify the candidate word sequence of the currently acquired verification task to obtain a sequence verification result; and send the sequence verification result to the hybrid bonding stack processing unit 101. The hybrid bonding stacking processing unit 101 is used to call the draft model to generate candidate lexical sequences for the currently acquired decoding task; generate a verification task based on the candidate lexical sequences, and send the verification task to the heterogeneous accelerator processing unit 102; and perform the next round of decoding task for the corresponding task request or end the decoding task for the corresponding task request based on the verification result of the currently acquired sequence.
[0025] The hybrid bonding stacked processing unit 101 adopts an HB (Hybrid Bonding) stacked structure. HB, as an advanced 3D packaging technology, achieves vertical stacking of DRAM (Dynamic Random Access Memory) and logic chips through extremely fine-pitch interconnects. Compared to traditional HBM (High Bandwidth Memory), it provides higher interconnect density and bandwidth; compared to PIM (Processing in Memory), it better preserves DRAM capacity and computing unit capabilities. However, due to manufacturing and heat dissipation limitations, the storage capacity of a single-layer HB is limited, and multi-layer stacking significantly increases cost and difficulty. While it cannot meet the high-concurrency service scenarios of LLM with hundreds of billions of parameters, it is sufficient to support the operation of the draft model. In one example, the hybrid bonding stacked processing unit can adopt a hybrid bonding stacked structure consisting of a single-layer dynamic random access memory layer and a single-layer logic chip layer.
[0026] The heterogeneous accelerator processing unit 102 can adopt an XPU architecture, where XPU is a heterogeneous computing architecture that integrates multiple computing units. The computing core of the heterogeneous accelerator processing unit 102 adopts a mainstream GPU (Graphics Processing Unit) / TPU (Tensor Processing Unit), and the storage uses DRAM.
[0027] The memory bandwidth of the hybrid bonded stacked processing unit 101 is greater than that of the heterogeneous accelerator processing unit 102, the memory capacity of the heterogeneous accelerator processing unit 102 is greater than that of the hybrid bonded stacked processing unit 101, and the computational power of the heterogeneous accelerator processing unit 102 is greater than that of the hybrid bonded stacked processing unit 101. Draft models are typically small, with only about 1 / 10 the number of parameters of the target model, and are deployed on the hybrid bonded stacked processing unit 101. The target model, a conventional large model, is deployed on the heterogeneous accelerator processing unit 102. The task request for large model inference is divided into three parts: pre-filling, decoding, and verification. Pre-filling and verification are processes with high computational power and high memory requirements, while decoding is a task with high bandwidth requirements.
[0028] Pre-filling is a computationally intensive and memory-intensive task, executed by the heterogeneous accelerator processing unit 102. After receiving the task request, the heterogeneous accelerator processing unit 102 calls the target model to perform large-scale parallel matrix multiplication operations on all tokens in the task request, generating an initial KV cache and the first token to obtain the initial processing result, and generating a decoding task based on the initial processing result.
[0029] The decoding task requires generating a candidate token sequence, which is a memory bandwidth intensive task requiring extremely high memory bandwidth. It is executed by the hybrid bonded stack processing unit 101. After obtaining the decoding task, the hybrid bonded stack processing unit 101 calls the draft model to continue inference based on the decoding task, thereby generating a candidate token sequence, and then packages the candidate token sequence into a verification task.
[0030] The verification task is executed by the heterogeneous accelerator processing unit 102. After receiving the verification task, the heterogeneous accelerator processing unit 102 calls the target model to perform parallel forward propagation on each token in the candidate token sequence to check their correctness. After the check is completed, the sequence verification result is sent to the hybrid bonded stack processing unit 101.
[0031] The hybrid bonding stacking processing unit 101 determines whether a task request is complete based on the sequence verification result. If complete, it terminates the decoding task of the requested task and releases the corresponding hardware resources. If incomplete, it proceeds to the next round of decoding for the requested task. In one example, the hybrid bonding stacking processing unit 101 is specifically used to terminate the decoding task of the corresponding task request if the currently acquired sequence verification result indicates a correct end-of-sequence character; and to proceed to the next round of decoding according to the currently acquired sequence verification result if the currently acquired sequence verification result does not contain an end-of-sequence character or indicates an incorrect end-of-sequence character. The corresponding task request is the task request corresponding to the currently acquired sequence verification result. Specifically, if the currently acquired sequence verification result does not contain an end-of-sequence character or indicates an incorrect end-of-sequence character, it updates the context state of the next round of decoding based on the accepted tokens in the sequence verification result and proceeds to the next round of decoding.
[0032] In this embodiment of the invention, the decoding process performed by the draft model is a memory bandwidth intensive task, while the pre-filling and verification processes performed by the target model are computationally intensive and memory-intensive tasks. Deploying the draft model on a hybrid bonded stack processing unit can fully utilize the high memory bandwidth of the hybrid bonded stack processing unit. Deploying the target model on a heterogeneous accelerator processing unit can fully utilize the high memory capacity and high computational power of the heterogeneous accelerator processing unit, thereby combining the characteristics of the model algorithm with the characteristics of the hardware capabilities, which can improve the inference speed of large models.
[0033] In one possible implementation, the heterogeneous accelerator processing unit 102 is provided with a first task pool, and the hybrid bonding stacking processing unit 101 is provided with a second task pool. The heterogeneous accelerator processing unit 102 is further configured to add the received task requests and verification tasks to the first task pool; and, when its own idle resources meet the first preset resource idle condition, to obtain at least one task request and / or verification task from the first task pool. The hybrid bonding stacking processing unit 101 is further configured to add the received decoding task and sequence verification result to the second task pool; and, if its own idle resources meet the second preset resource idle condition, to obtain at least one decoding task and / or sequence verification result from the second task pool.
[0034] The first preset resource idle condition can be customized according to actual conditions. When the idle resources meet the first preset resource idle condition, it means that the heterogeneous accelerator processing unit 102 has enough resources to process new tasks. In one example, when the processor utilization and / or memory utilization of the heterogeneous accelerator processing unit 102 is less than a preset utilization threshold, it is determined that the first preset resource idle condition is met.
[0035] The second preset resource idle condition can be customized according to actual conditions. When the idle resources meet the second preset resource idle condition, it means that the hybrid bonding stack processing unit 101 has enough resources to process new tasks. In one example, when the processor utilization and / or memory utilization of the hybrid bonding stack processing unit 101 is less than a preset utilization threshold, it is determined that the second preset resource idle condition is met.
[0036] In this embodiment of the invention, tasks are cached through a task pool, which enables asynchronous parallel execution of multiple task requests. While the heterogeneous accelerator processing unit verifies the verification task of the previous task request, the hybrid bonding stack processing unit can simultaneously execute the decoding task of the current task request. The two are completely decoupled, thereby improving the inference speed of large models.
[0037] In one possible implementation, the processing priority of the task request is higher than that of the verification task, and the heterogeneous accelerator processing unit 102 processes the higher-priority task first.
[0038] In one example, the first task pool includes a request task pool and a verification task pool; the request task pool is used to cache task requests, and the verification task pool is used to cache verification tasks; the heterogeneous accelerator processing unit 102, specifically, when the idle resources of the heterogeneous accelerator processing unit meet a first preset resource idle condition, preferentially obtains task requests from the request task pool according to a preset data volume; if there are no task requests in the request task pool or the data volume of the task requests does not meet the preset data volume requirement, then continues to obtain verification tasks from the verification task pool.
[0039] Both the hybrid bonding stack processing unit 101 and the accelerator processing unit 102 can adopt a batch processing strategy. The preset data volume is set according to the actual task data volume processed by the accelerator processing unit 102 in each batch. During the process of the accelerator processing unit 102 obtaining a batch of tasks, there are four situations: (1) the task requests in the request task pool can meet the preset data volume, that is, meet the task volume required for the current batch processing; (2) there are task requests in the request task pool, but they are insufficient to meet the preset data volume, and it is necessary to obtain some verification tasks from the verification task pool; (3) there are no task requests in the request task pool, and verification tasks that meet the preset data volume are obtained directly from the verification task pool; (4) the sum of all tasks in the request task pool and the verification task pool cannot meet the preset data volume. In any case, the accelerator processing unit 102 will prioritize selecting task requests from the request task pool, and will only obtain verification tasks from the verification task pool when the data volume of the task requests in the request task pool is insufficient to meet the preset data volume.
[0040] See Figure 2 , Figure 2 This demonstrates the acceleration effect of prioritizing task requests. The vertical axis represents the normalized large model inference speed, and the horizontal axis represents the request rate, i.e., the number of task requests received per unit time. FIFO (First In First Out) represents the large model inference speed when the heterogeneous accelerator processing unit 102 processes tasks sequentially using a FIFO queue; it is normalized to 1 for ease of viewing. PFS (Prefill-First Scheduling) represents the large model inference speed when using an embodiment of the prioritization process for task requests according to this invention. From Figure 2 It can be seen that prioritizing the pre-filling process of task requests over the processing of verification tasks achieves an average speedup of 1.10 times compared to processing them in chronological order. The speedup effect of scheduling optimization increases accordingly with the increase in request rate.
[0041] In this embodiment of the invention, higher priority task requests are processed first, thereby improving the inference speed of large models.
[0042] In addition to prioritizing pre-filling, long pre-filled sequences can also be divided into blocks. In one possible implementation, the heterogeneous accelerator processing unit 102 is specifically used to split the currently acquired task request into multiple task blocks according to the preset length threshold when the length of the currently acquired task request exceeds the preset length threshold, wherein the length of each task block does not exceed the preset length threshold; and call the target model to pre-fill each task block to obtain an initial processing result.
[0043] The preset length threshold can be set according to the actual processing capacity of the heterogeneous accelerator processing unit 102. Directly pre-filling excessively long task requests would cause them to consume a large amount of hardware resources, thus affecting the processing of other tasks. Therefore, in this embodiment of the invention, task requests exceeding the preset length threshold are split into blocks according to dimensions, ensuring that the length of each task block is no greater than the preset length threshold. For example, if the preset length threshold is 1024 and the length of the task request is 2500, the task request can be divided into three task blocks with lengths of 1024, 1024, and 452. In one example, if there is also a task request or verification task with a length of 500, the 452 task block can be concatenated with the 500-length task request or verification task to form a single 952-length task block for processing.
[0044] See Figure 2 , Figure 2This demonstrates the acceleration effect of task request splitting. The vertical axis represents the normalized large model inference speed, and the horizontal axis represents the request speed, i.e., the number of task requests received per unit time. FIFO represents the large model inference speed when the heterogeneous accelerator processing unit 102 processes tasks sequentially using a first-in-first-out queue; it is normalized to 1 for ease of viewing. CHK (Chunk) represents the large model inference speed when the task request splitting and priority processing of task requests in this invention are used simultaneously. From... Figure 2 It can be seen that CHK achieves an average speedup of 1.29 times compared to FIFO. As the request rate increases, the speedup effect of scheduling optimization also increases accordingly.
[0045] In this embodiment of the invention, by splitting task requests whose length exceeds a preset length threshold, the situation where a single task request occupies a large amount of hardware resources can be reduced, thereby alleviating resource competition, improving the utilization rate of heterogeneous accelerator processing unit hardware, and ultimately improving the inference speed of large models.
[0046] In one possible implementation, a memory watermark can be set to pause new pre-filling when the KV cache percentage exceeds a threshold, thus avoiding memory overflow and request blocking.
[0047] The heterogeneous accelerator processing unit 102 is also used to obtain a first percentage by acquiring the ratio of the amount of data in its own memory key-value cache to the total amount of data in memory; if the first percentage is greater than the first percentage threshold, the task request is suspended. The hybrid bonding stacking processing unit 101 is also used to obtain a second percentage by acquiring the ratio of the amount of data in its own memory key-value cache to the total amount of data in memory; if the second percentage is greater than the second percentage threshold, the heterogeneous accelerator processing unit is notified to suspend the acquisition of task requests.
[0048] The first percentage threshold can be set according to the total memory size of the heterogeneous accelerator processing unit 102. The larger the total memory, the larger the first percentage threshold, and the smaller the total memory, the smaller the first percentage threshold. For example, the first percentage threshold can be set to 90%, 80%, or 75%, etc.
[0049] The second percentage threshold can be set according to the total memory size of the hybrid bonding stack processing unit 101. The larger the total memory, the larger the second percentage threshold, and the smaller the total memory, the smaller the second percentage threshold. For example, the first percentage threshold can be set to 85%, 75%, or 70%, etc.
[0050] In this embodiment of the invention, by monitoring the KV cache occupancy rate, it is determined whether to process new task requests, thereby improving the stability of the system. When memory is released, the service is automatically restored, which can reduce memory overflow and request avalanche.
[0051] In one possible implementation, the hybrid bond stacking processing unit 101 is further configured to adjust the draft generation budget and tree width based on the first computational intensity of the hybrid bond stacking processing unit 101 and the second computational intensity of the heterogeneous accelerator processing unit 102.
[0052] When the first computational intensity of the hybrid bond stacking processing unit 101 indicates that the load of the hybrid bond stacking processing unit 101 is too high, the tree width can be reduced to reduce the computational power consumed by the decoding task of the hybrid bond stacking processing unit 101.
[0053] In one possible implementation, the hybrid bonding stacking processing unit 101 is specifically configured to: query a preset lookup table to determine the optimal tree width corresponding to the current draft generation budget, wherein the preset lookup table records the correspondence between the draft generation budget and the optimal tree width; obtain its current computational intensity to obtain a first computational intensity; if the first computational intensity is less than a first preset intensity threshold, increase the current tree width; if the first computational intensity is greater than the first preset intensity threshold, decrease the current tree width, wherein the lower limit of the tree width is a first value, and the upper limit of the tree width is the optimal tree width corresponding to the current draft generation budget; call the draft model to generate the candidate lexical sequence of the currently acquired decoding task according to the current tree width; add the current candidate lexical sequence to the draft tree; if the budget of the draft tree does not reach the current draft generation budget, return to obtain its current computational intensity to obtain the first computational intensity; if the budget of the draft tree reaches the current draft generation budget, generate a verification task according to the draft tree.
[0054] The first computational intensity of the hybrid bond stacked processing unit 101 can be the utilization rate of the processor in the hybrid bond stacked processing unit 101, or the total data volume of queued tasks in the hybrid bond stacked processing unit 101. The first preset intensity threshold and the second preset intensity threshold can be customized according to actual conditions. When the second computational intensity of the heterogeneous accelerator processing unit 102 indicates that the load on the heterogeneous accelerator processing unit 102 is too high, the draft generation budget can be reduced, thereby reducing the computational power consumed by the verification task of the heterogeneous accelerator processing unit 102.
[0055] In one possible implementation, the heterogeneous accelerator processing unit 102 is further configured to determine its current load and obtain a second load; and send the second load to the hybrid bonding stack processing unit. The hybrid bonding stack processing unit 101 is specifically used to increase the current draft generation budget if the second load is less than the second preset load threshold, and decrease the current draft generation budget if the second load is greater than the second preset load threshold, wherein the lower limit of the draft generation budget is the second value and the upper limit of the draft generation budget is the third value.
[0056] The first computational intensity represents the computational pressure on the processor resources in the hybrid bonded stacked processing unit 101; the higher the computational pressure, the higher the first computational intensity. In one example, the first computational intensity can be the processor utilization rate in the hybrid bonded stacked processing unit 101; in another example, it can be the total data volume of queued tasks in the hybrid bonded stacked processing unit 101. The second computational intensity represents the computational pressure on the processor resources in the heterogeneous accelerator processing unit 102; the higher the computational pressure, the higher the second computational intensity. In one example, the second computational intensity can be the processor utilization rate in the heterogeneous accelerator processing unit 102; in another example, it can be the total data volume of queued tasks in the heterogeneous accelerator processing unit 102. The first preset intensity threshold, the second preset intensity threshold, the first value, the second value, and the third value can all be customized according to actual conditions.
[0057] Specifically, an example will be used to illustrate the process of adjusting the draft generation budget and tree width.
[0058] The hybrid bonding stacking processing unit 101 can be used to perform the following steps: Step A1: Query a preset lookup table to determine the optimal tree width corresponding to the current draft generation budget. The preset lookup table records the correspondence between the draft generation budget and the optimal tree width.
[0059] Step A2: Determine your current computational strength to obtain the first computational strength.
[0060] Step A3: If the first calculated intensity is less than the first preset intensity threshold, the current tree width is increased by 1, where the upper limit of the tree width is the optimal tree width; if the first calculated intensity is greater than the first preset intensity threshold, the current tree width is decreased by 1, where the lower limit of the tree width is 1.
[0061] The upper limit of the tree width is the optimal tree width corresponding to the current draft generation budget. When the current tree width is the corresponding optimal tree width, even if the first calculation intensity is less than the first preset intensity threshold, the current tree width will not increase. The lower limit of the tree width is a first value, specifically 1 in this embodiment. When the current tree width is 1, even if the first calculation intensity is greater than the first preset intensity threshold, the current tree width will not decrease. In one example, if the first calculation intensity is equal to the first preset intensity threshold, then there is no need to adjust the current tree width.
[0062] Step A4: Invoke the draft model to generate the candidate word sequence for the currently acquired decoding task according to the current tree width.
[0063] Step A5: Add the current candidate word sequence to the draft tree, return to step A2 until the draft tree's budget reaches the current draft generation budget, then proceed to step A6.
[0064] Step A6: Generate a verification task based on the draft tree and send the verification task to the heterogeneous accelerator processing unit 102.
[0065] The heterogeneous accelerator processing unit 102 can be used to perform the following steps: Step A7: Invoke the target model to verify the candidate lexical sequence of the currently acquired verification task and obtain the sequence verification result; determine the current load and obtain the second load; send the sequence verification result and the second load to the hybrid bonding stack processing unit 101.
[0066] The hybrid bonding stacking processing unit 101 can also be used to perform the following steps: Step A8: If the second load is less than the second preset load threshold, the current draft generation budget is multiplied by 2; if the second load is greater than the second preset load threshold, the current draft generation budget is divided by 2. The lower limit of the draft generation budget is the second value, and the upper limit of the draft generation budget is the third value.
[0067] The upper limit of the draft generation budget is the third value. When the current draft generation budget × 2 is greater than the third value, even if the second computational intensity is less than the second preset intensity threshold, the current draft generation budget will only be updated to the third value. The lower limit of the draft generation budget is the second value. When the current draft generation budget / 2 is less than the second value, even if the second computational intensity is greater than the second preset intensity threshold, the current draft generation budget will only be updated to the second value. In one example, if the second computational intensity is equal to the second preset intensity threshold, then there is no need to adjust the current draft generation budget.
[0068] In this embodiment of the invention, the tree width is adjusted according to the computational intensity of the hybrid bonding stacking processing unit to approximate the optimal point of the eaves model; the draft generation budget is dynamically adjusted according to the computational intensity of the heterogeneous accelerator processing unit to reduce invalid computation; and the unit computational output can be maximized while ensuring the acceptance rate.
[0069] The following example illustrates the improved inference speed of the large model inference system based on hybrid bonding and speculative sampling in this invention. The hybrid bonding stacked processing unit 101 adopts a face-to-face single-layer stacked structure, consisting of one DRAM layer and one logic chip layer interconnected via HB. The DRAM layer provides high-bandwidth, low-latency access; the logic chip layer includes input / output caches, a controller, and a distributed MAC (Multiply-Accumulate) array, supporting tensor parallelism and computation-communication overlap. The logic chip layer uses a 28nm process with an internal clock frequency of 400MHz. The substrate uses low-cost wire bonding, and due to the extremely small amount of data across units, the bandwidth fully meets the requirements. The computational core of the heterogeneous accelerator processing unit adopts a 7nm GPU / TPU architecture, and the storage uses multi-layer stacked LPDDR5X, providing large capacity (up to 512GB per node) and moderate bandwidth to meet the storage requirements of target model parameters and high-concurrency KV cache. Llama2-13B, Qwen3-32B, and OPT-66B models were used as target models, and simplified to obtain corresponding draft models for ShareGPT real-world load testing. The average of the three LLM models was taken. Compared to GPU autoregressive inference in related technologies, the large model inference system based on hybrid bonding and speculative sampling in this embodiment of the invention reduces the average latency by about one-third and improves energy efficiency by 1.96 times; under the same service objective, the request processing rate is improved by 2.97 times; a single node supports lossless service for 10–100B parameter models without quantization. Furthermore, compared to pure HBM hardware solutions, the storage and packaging costs of the large model inference system based on hybrid bonding and speculative sampling in this embodiment of the invention are reduced by approximately 50%.
[0070] The following describes the large model inference method based on hybrid bonding and speculative sampling provided by the present invention. The large model inference method based on hybrid bonding and speculative sampling described below can be referred to in correspondence with the large model inference system based on hybrid bonding and speculative sampling described above.
[0071] See Figure 3 , Figure 3 This is a flowchart illustrating a large model inference method based on hybrid bonding and speculative sampling according to an embodiment of the present invention. The method is executed by any of the aforementioned large model inference systems based on hybrid bonding and speculative sampling. The method includes: S301, the heterogeneous accelerator processing unit calls the target model to pre-fill the currently acquired task request to obtain the initial processing result; and generates a decoding task based on the initial processing result.
[0072] S302, the hybrid bonding stacking processing unit calls the draft model to generate a candidate lexical sequence for the currently acquired decoding task; and generates a verification task based on the candidate lexical sequence.
[0073] S303, the heterogeneous accelerator processing unit calls the target model to verify the candidate word sequence of the currently acquired verification task and obtains the sequence verification result.
[0074] S304, the hybrid bonding stacking processing unit performs the next round of decoding task for the corresponding task request or ends the decoding task for the corresponding task request based on the currently acquired sequence verification result.
[0075] In one possible implementation, the heterogeneous accelerator processing unit is provided with a first task pool, and the hybrid bonding stack processing unit is provided with a second task pool; the method further includes: Step B1: The heterogeneous accelerator processing unit adds the received task request and verification task to the first task pool; when the idle resources of the heterogeneous accelerator processing unit meet the first preset resource idle condition, it obtains at least one task request and / or verification task from the first task pool.
[0076] Step B2: The hybrid bonding stack processing unit adds the received decoding task and sequence verification result to the second task pool; if the idle resources of the heterogeneous accelerator processing unit meet the second preset resource idle condition, it obtains at least one decoding task and / or sequence verification result from the second task pool.
[0077] In one possible implementation, the processing priority of the task request is higher than that of the verification task, and the heterogeneous accelerator processing unit processes the higher-priority task first.
[0078] In one possible implementation, the first task pool includes a request task pool and a verification task pool; the request task pool is used to cache task requests, and the verification task pool is used to cache verification tasks; the step of obtaining at least one task request and / or verification task from the first task pool when the idle resources of the heterogeneous accelerator processing unit meet a first preset resource idle condition includes: If the idle resources of the heterogeneous accelerator processing unit meet the first preset resource idle condition, task requests are preferentially obtained from the request task pool according to the preset data volume; if there are no task requests in the request task pool or the data volume of the task requests does not meet the preset data volume requirement, then verification tasks are continued to be obtained from the verification task pool.
[0079] In one possible implementation, the heterogeneous accelerator processing unit calls the target model to pre-fill the currently acquired task request to obtain an initial processing result, including: When the length of the currently acquired task request exceeds a preset length threshold, the heterogeneous accelerator processing unit splits the currently acquired task request into multiple task blocks according to the preset length threshold, wherein the length of each task block does not exceed the preset length threshold; and calls the target model to pre-fill each task block to obtain an initial processing result.
[0080] In one possible implementation, the method further includes: The heterogeneous accelerator processing unit obtains a first percentage by measuring the ratio of the amount of key-value cached data in its own memory to the total amount of data in memory; if the first percentage is greater than the first percentage threshold, the task request is paused. The hybrid bonding stacking processing unit obtains a second percentage by measuring the ratio of the amount of data in its own memory key-value cache to the total amount of data in memory; if the second percentage is greater than the second percentage threshold, it notifies the heterogeneous accelerator processing unit to pause the acquisition of task requests.
[0081] In one possible implementation, the method further includes: The hybrid bonding stack processing unit is also used to adjust the draft generation budget and tree width according to the first computing intensity of the hybrid bonding stack processing unit and the second computing intensity of the heterogeneous accelerator processing unit.
[0082] In one possible implementation, see Figure 4 The hybrid bonding stack processing unit is further configured to adjust the draft generation budget and tree width based on the first computational intensity of the hybrid bonding stack processing unit and the second computational intensity of the heterogeneous accelerator processing unit, including: S401, the hybrid bonding stacking processing unit queries a preset lookup table to determine the optimal tree width corresponding to the current draft generation budget, wherein the preset lookup table records the correspondence between the draft generation budget and the optimal tree width.
[0083] S402, the hybrid bonding stacking processing unit obtains its current computing strength to get a first computing strength; if the first computing strength is less than a first preset strength threshold, the current tree width is increased; if the first computing strength is greater than the first preset strength threshold, the current tree width is decreased, wherein the lower limit of the tree width is a first value, and the upper limit of the tree width is the optimal tree width corresponding to the current draft generation budget.
[0084] S403, the hybrid bonding stacking processing unit calls the draft model to generate the candidate lexical sequence of the currently acquired decoding task according to the current tree width; and adds the current candidate lexical sequence to the draft tree.
[0085] S404, the hybrid bonding stacking processing unit determines whether the budget of the draft tree has reached the current draft generation budget.
[0086] If the budget of the draft tree does not reach the current draft generation budget, return to execute S402; if the budget of the draft tree reaches the current draft generation budget, execute S405.
[0087] S405, the hybrid bonding stacking processing unit generates a verification task based on the draft tree.
[0088] S406, the heterogeneous accelerator processing unit calls the target model to verify the candidate lexical sequence of the currently acquired verification task and obtains the sequence verification result; the heterogeneous accelerator processing unit determines its current load and obtains the second load; and sends the second load and the sequence verification result to the hybrid bonding stack processing unit.
[0089] S407, if the second load is less than the second preset load threshold, the hybrid bonding stacking processing unit increases the current draft generation budget; if the second load is greater than the second preset load threshold, the hybrid bonding stacking processing unit decreases the current draft generation budget, wherein the lower limit of the draft generation budget is the second value and the upper limit of the draft generation budget is the third value.
[0090] In one possible implementation, the hybrid bonding stacking processing unit performs the next round of decoding task for the corresponding task request or terminates the decoding task for the corresponding task request based on the currently acquired sequence verification result, including: The hybrid bonding stacking processing unit terminates the decoding task of the corresponding task request if the currently acquired sequence verification result indicates that the end-of-sequence symbol is correct; if the currently acquired sequence verification result does not contain an end-of-sequence symbol or indicates that the end-of-sequence symbol is incorrect, it performs the next round of decoding tasks according to the currently acquired sequence verification result; wherein, the corresponding task request is the task request corresponding to the currently acquired sequence verification result.
[0091] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the large model inference method based on hybrid bonding and speculative sampling provided by the above methods.
[0092] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the large model inference methods based on hybrid bonding and speculative sampling provided by the above methods.
[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A large-model inference system based on hybrid bonding and speculative sampling, characterized in that, include: The system includes a hybrid bonded stack processing unit and a heterogeneous accelerator processing unit. The hybrid bonded stack processing unit employs a hybrid bonded stack structure. The memory bandwidth of the hybrid bonded stack processing unit is greater than the memory bandwidth of the heterogeneous accelerator processing unit, and the memory capacity of the heterogeneous accelerator processing unit is greater than the memory capacity of the hybrid bonded stack processing unit. A draft model runs in the hybrid bonded stack processing unit, and a target model runs in the heterogeneous accelerator processing unit. The heterogeneous accelerator processing unit is used to call the target model to pre-fill the currently acquired task request to obtain an initial processing result; generate a decoding task based on the initial processing result and send the decoding task to the hybrid bonding stack processing unit; call the target model to verify the candidate word sequence of the currently acquired verification task to obtain a sequence verification result; and send the sequence verification result to the hybrid bonding stack processing unit. The hybrid bonding stack processing unit is used to call the draft model to generate candidate lexical sequences for the currently acquired decoding task; generate a verification task based on the candidate lexical sequences, and send the verification task to the heterogeneous accelerator processing unit; and perform the next round of decoding task for the corresponding task request or end the decoding task for the corresponding task request based on the verification result of the currently acquired sequence.
2. The system according to claim 1, characterized in that, The heterogeneous accelerator processing unit is provided with a first task pool, and the hybrid bonding stacking processing unit is provided with a second task pool. The heterogeneous accelerator processing unit is further configured to add the received task requests and verification tasks to the first task pool; and, when the idle resources of the heterogeneous accelerator processing unit meet the first preset resource idle condition, to obtain at least one task request and / or verification task from the first task pool. The hybrid bonding stack processing unit is further configured to add the received decoding task and sequence verification result to the second task pool; and, when the idle resources of the heterogeneous accelerator processing unit meet the second preset resource idle condition, to obtain at least one decoding task and / or sequence verification result from the second task pool.
3. The system according to claim 2, characterized in that, The processing priority of the task request is higher than that of the verification task, and the heterogeneous accelerator processing unit processes the higher-priority task first.
4. The system according to claim 3, characterized in that, The first task pool includes a request task pool and a verification task pool; the request task pool is used to cache task requests, and the verification task pool is used to cache verification tasks. Specifically, when the idle resources of the heterogeneous accelerator processing unit meet the first preset resource idle condition, the task request is preferentially obtained from the request task pool according to the preset data volume; if there is no task request in the request task pool or the data volume of the task request does not meet the preset data volume requirement, the verification task is then obtained from the verification task pool.
5. The system according to claim 1, characterized in that, The heterogeneous accelerator processing unit is specifically used to split the currently acquired task request into multiple task blocks according to the preset length threshold when the length of the currently acquired task request exceeds the preset length threshold, wherein the length of each task block does not exceed the preset length threshold; and to call the target model to pre-fill each task block to obtain an initial processing result.
6. The system according to claim 1, characterized in that, The heterogeneous accelerator processing unit is also used to obtain a first percentage by acquiring the ratio of the amount of data in its own memory key-value cache to the total amount of data in memory; if the first percentage is greater than the first percentage threshold, the task request is paused. The hybrid bonding stacking processing unit is also used to obtain a second percentage by measuring the ratio of the amount of key-value cached data in its own memory to the total amount of memory data; if the second percentage is greater than the second percentage threshold, the heterogeneous accelerator processing unit is notified to suspend the acquisition of task requests.
7. The system according to claim 1, characterized in that, The hybrid bonding stack processing unit is also used to adjust the draft generation budget and tree width according to the first computing intensity of the hybrid bonding stack processing unit and the second computing intensity of the heterogeneous accelerator processing unit.
8. The system according to claim 7, characterized in that, The hybrid bonding stacking processing unit is specifically used to: query a preset lookup table to determine the optimal tree width corresponding to the current draft generation budget, wherein the preset lookup table records the correspondence between the draft generation budget and the optimal tree width; obtain its current computing strength to obtain a first computing strength; if the first computing strength is less than a first preset strength threshold, increase the current tree width; if the first computing strength is greater than the first preset strength threshold, decrease the current tree width, wherein the lower limit of the tree width is a first value, and the upper limit of the tree width is the optimal tree width corresponding to the current draft generation budget; call the draft model to generate the candidate lexical sequence of the currently acquired decoding task according to the current tree width; add the current candidate lexical sequence to the draft tree; if the budget of the draft tree has not reached the current draft generation budget, return to obtain its current computing strength to obtain the first computing strength; if the budget of the draft tree has reached the current draft generation budget, generate a verification task according to the draft tree; The heterogeneous accelerator processing unit is further configured to determine its current load and obtain a second load; and send the second load to the hybrid bonding stack processing unit. The hybrid bonding stack processing unit is specifically used to increase the current draft generation budget if the second load is less than the second preset load threshold, and decrease the current draft generation budget if the second load is greater than the second preset load threshold, wherein the lower limit of the draft generation budget is the second value and the upper limit of the draft generation budget is the third value.
9. The system according to claim 1, characterized in that, The hybrid bonding stacking processing unit is specifically used to terminate the decoding task of the corresponding task request when the currently acquired sequence verification result indicates that the end symbol is correct; and to perform the next round of decoding task according to the currently acquired sequence verification result when the currently acquired sequence verification result does not have an end symbol or indicates that the end symbol is incorrect; wherein, the corresponding task request is the task request corresponding to the currently acquired sequence verification result.
10. A large-model inference method based on hybrid bonding and speculative sampling, characterized in that, Performed by the system according to any one of claims 1-9, the method comprises: The heterogeneous accelerator processing unit calls the target model to pre-fill the currently acquired task request to obtain an initial processing result; and generates a decoding task based on the initial processing result. The hybrid bonding stacking processing unit calls the draft model to generate a candidate lexical sequence for the currently acquired decoding task; and generates a verification task based on the candidate lexical sequence. The heterogeneous accelerator processing unit calls the target model to verify the candidate word sequence of the currently acquired verification task and obtains the sequence verification result; The hybrid bonding stacking processing unit performs the next round of decoding task for the corresponding task request or terminates the decoding task for the corresponding task request based on the currently acquired sequence verification result.